How Paradigm4 Is Helping Organizations Remove Hidden AI Bottlenecks
Tech Talks DailyJune 13, 2026
3603
22:4420.81 MB

How Paradigm4 Is Helping Organizations Remove Hidden AI Bottlenecks

What happens when a company focused on drug discovery and life sciences encounters a data problem that nobody else seems able to solve?

Recorded at the IT Press Tour in Boston, this episode explores the fascinating story behind Paradigm4 and how a challenge in large-scale biomedical research ultimately led to the creation of flexFS, a cloud-native filesystem designed to tackle some of today's biggest data infrastructure challenges.

Joining me on the podcast is David Freund from Paradigm4, who shares how the company was originally founded to help scientists work with enormous datasets in fields such as genomics, bioinformatics, and precision medicine. As researchers began working with population-scale datasets such as the UK Biobank, the team discovered that existing storage technologies either couldn't deliver the performance they needed, lacked the functionality required, or became prohibitively expensive at scale.

Our conversation explores the moment Paradigm4 realized it would need to build its own solution, why traditional approaches to cloud storage often struggle under modern analytics workloads, and how flexFS emerged from a real-world customer problem rather than a technology trend. David also explains why object storage has become such an attractive foundation for modern infrastructure, while discussing the challenges of latency, performance, and cost that still need to be addressed.

We also discuss why many organizations investing heavily in AI infrastructure may be overlooking one of the biggest constraints on performance. While much of the industry conversation focuses on GPUs and compute power, David argues that data access, movement, and management are becoming equally important considerations as AI workloads continue to grow.

Along the way, we touch on cloud independence, resilience, large-scale analytics, and why flexibility across cloud providers is becoming an increasingly important requirement for enterprise technology leaders. Whether you're working in AI, life sciences, cloud infrastructure, or enterprise data management, this episode offers an interesting perspective on how customer problems can sometimes lead to entirely new categories of technology.

Could the next major AI bottleneck be data rather than compute? And are organizations paying enough attention to the infrastructure feeding their most important workloads? I'd love to hear your thoughts.

Useful Links


Thank you to Alex Zaharov-Reutt for sharing the iTWire TV Recording of the Paradigm4/flexFS presentation at the IT Press Tour in Boston.

Visit Tech Talks Network Sponsors Below

[00:00:00] - [Speaker 0]
So a huge thanks to Denodo for supporting the Tech Talks Network, helping us produce more than 60 interviews a month. And when it comes to trusted data products, it all starts with the right foundation. And trusted data products start with Denodo because they can help you create, manage and deliver business ready data products faster with secure real time access across all of your data sources. And you can learn more by simply visiting denodo.com. Welcome back to the Tech Talks Daily Podcast.

[00:00:35] - [Speaker 0]
And today, I'm coming to you from the IT Press Tour in Boston, Massachusetts. And I've been spending the day speaking with a company whose story began not in enterprise storage, but in the world of scientific discovery. And Paradigm four was founded to help researchers tackle some of the biggest challenges in life sciences, Because they were working with vast data sets that power drug discovery, genomics research and precision medicine. But along the way, the team encountered a problem that will sound familiar to many tech and business leaders listening today. And that is the data was growing faster than the infrastructure designed to support it.

[00:01:18] - [Speaker 0]
And what followed wasn't a search for the latest technology trend or an attempt to build another storage platform. Instead, it was an effort to solve a very real customer problem. And that was existing solutions either delivered the performance researchers needed at a cost that was just too difficult to justify, Or that they delivered lower costs while introducing limitations. Limitations that would slow down scientific progress. And it was this very challenge that ultimately led to the creation of FlexFS, a technology that sits at the heart of today's conversation.

[00:01:57] - [Speaker 0]
And what makes this story particularly relevant in 2026 as I record this is that the challenges paradigm four first encountered in life sciences are now appearing across almost every industry. So whether an organisation is building AI applications, training models, managing digital twins, or running analytics platforms, or just trying to make sense of enormous volumes of operational data, many are discovering that expensive compute resources often spend their time waiting for data to arrive. So during our conversation today, we will explore how a biomedical analytics company became a storage innovator, and understand why the AI conversation is increasingly becoming a data access conversation, and what business leaders could be overlooking as they race to invest in the next generation of infrastructure. But enough scene setting for me. Let let me introduce you to David from Paradigm four that is recorded here at the IT Press Tour Boston.

[00:03:03] - [Speaker 0]
So thank you for joining me here on the IT Press Talk. Can you tell everyone listening a little about who you are and what you do?

[00:03:11] - [Speaker 1]
Well, I'm David Freund, and I'm I'm here to help with the go to market for FlexFS. Little quick thumbnail sketch of my background. I actually started off as a software engineer working on things like cluster file systems, operating system kernels, network stacks for companies like Digital Equipment Corporation. So there, just dated myself. But also branched off into product management, and I spent half a dozen years as an IT industry analyst focused on server and and storage technologies.

[00:03:40] - [Speaker 1]
Jumped back into vendor land at as CTO for infrastructure software at EMC for example. I would spend some time at NetApp and so on. And a thread that is kind of woven through this whole tapestry of my career has been being able to have conversations with customers trying to solve real problems and being able to say well this technology can help you in this way and for example if somebody comes up to me at a conference and say hey thank you personally for some great new release or something and then tell me the bizarre thing that they were doing with it that no one had thought of before. And we'll go, Hey, what if we give you this? Oh, well then I can do that and the other thing and yeah, you and a couple million of your closest friends get right back to you on that and be able to do something about it so that's why I jumped back into vendor land.

[00:04:23] - [Speaker 1]
Mean I had fun as an analyst but I wanted to be able to have that impact And so being able to have those kinds of conversations is what eventually led me here to Paradigm four. I saw something new and novel in a file system and that doesn't happen often in file systems. And so I thought this could be really something. And so I joined to kind of help evangelize this and explain it.

[00:04:46] - [Speaker 0]
You've been on an incredible journey there and of course so has Paradigm four. So just for people listening hearing about them for the first time, tell me a bit more about the story behind the company, the problem that they were originally trying to solve. And at what point did they realize that maybe existing storage technologies were just not gonna get you where you needed to be? Tell me more about that backstory of the company.

[00:05:08] - [Speaker 1]
Well, so it was founded by Turing Laureate Michael Stonebreaker and Marilyn Mats. Marilyn's still CEO of the company. And it was originally created to evangelize a new database, SciDB for science database. And eventually, the company the team discovered, hey. It's not just about the database.

[00:05:32] - [Speaker 1]
It's about solving real problems right and so they developed a whole solution around that database and so and they lean into life sciences. So doing things like drug discovery and things like that being able to manage and analyze the data better, faster, more efficiently. And so as a part of that when it came to unstructured data there is an issue around again performance, right? It's not only all not all databases fit every use. Here we've got what do we do for a file system.

[00:06:07] - [Speaker 1]
And so the reason FlexFS was born was because nothing else available in the market would either be performant or functionally complete or it was cost prohibitive. And so we decided okay we need to figure this out for ourselves. We had looked at all kinds of prior art open source solutions and so on and one of Michael Stonebraker's students, Gary Plantaber, CTO of Paradigm four Yeah. Said, you know what? We're gonna make a run at this.

[00:06:39] - [Speaker 1]
We can solve this. And instead of taking this kind of a lift and shift concept of parallel databases like Lustre, which are great for high performance compute, but have become prohibitively expensive, they're also hard to operate. And instead of in the cloud, the de facto standard has been lift and shift Lustre into the cloud. Okay? And now they're managed service offerings from all the major cloud vendors of a Lustre but it's in Alaska.

[00:07:07] - [Speaker 1]
It's hard to manage. You have to do capacity planning and so on. And rather than take that kind of approach what Gary decided to do is, hey, let's take this from the opposite end. Let's leverage the person centuries that have been invested by the hyperscalers creating performant and really scalable elastic object storage but bolsters weaknesses. Okay?

[00:07:32] - [Speaker 1]
Object storage on the hyperscalers is great for throughput, great for capacity. They're elastic on both dimensions independently. Latency, however, is a big problem. Okay? How many IOs per second can you get through that thing?

[00:07:45] - [Speaker 1]
Not many. Both for metadata and for file data. So Gary's approach is let's solve those problems but still leverage the hyperscaler. And so FlexFS was born. And and it was able to solve problems like one of the first things that had our clients had hit that we responded to was genome wide association studies in the UK Biobank.

[00:08:09] - [Speaker 1]
So that was basically customer need is what drove this. We didn't just invent this because you know somebody thought it was cool. We were solving a real problem. And then once we saw this really worked well for all our life science customers and the vast majority of them and it's like, well, wait a minute. You can do other things.

[00:08:28] - [Speaker 1]
In fact, our own clients discovered, hey, we can make use of this particular feature of this platform, this reveal thing FlexFS, this is really great. We'll also use it for this and for that and for the other thing. And so the company here realized this is both wonderfully and horribly horizontal.

[00:08:45] - [Speaker 0]
Yeah. Yeah. And what I love about your story here is it became it all began from a customer problem. There's so much hype around tech trends, etcetera. I mean, if we fast forward to present day, many people think AI is all about GPUs.

[00:08:58] - [Speaker 0]
But your presentations today suggested that the real bottleneck is often data movement. So what are tech leaders and business leaders missing when they they only focus on compute?

[00:09:10] - [Speaker 1]
The enemy in AI compute is idle CPUs or even worse GPUs. Yeah. Yeah. And why are they idle? They're waiting on data.

[00:09:19] - [Speaker 1]
Okay. And so what we're about is getting data into those GPUs as fast as possible and so everything that we're doing is around that. We still have some work to do, we're not done but that's that's the whole story. The other part of it is it's not just the compute and the storage needing to be efficient but the compute actually generates IO in patterns. Do those patterns make sense?

[00:09:47] - [Speaker 1]
Can the storage system absorb it, tolerate it or better yet compensate for that? And that's something that we actually help with. There's an example if you're doing say a Spark workload, you've got a data lake house based on and you're doing analysis based on Spark and you're even using an additional engine to accelerate it like Gluten or Comet. Comet actually is not terribly optimized for POSIX file systems and so and it tends to serialize its IO so trying to do things with a parallel file system unless there's caching it doesn't work well. Once you introduce the caching which we also have because that's what we needed to put in to address the latency issues, suddenly it flies much faster multiple times over what just a naked s three baseline would would enable.

[00:10:41] - [Speaker 0]
And what I love about the company as well is it started from that customer problem, and I think very often we all encounter these problems, do nothing about it, complain to our other halves or something, but you've gone out there and created this solution here. Before even getting to that stage though, the company evaluated many existing companies before creating FlexFS. So what limitations kept appearing again and again? And and why weren't they good enough for your customers? Just going right back to that.

[00:11:08] - [Speaker 0]
We're having a problem. We're evaluating other other solutions there. What what what was missing there?

[00:11:13] - [Speaker 1]
What was missing there was either elasticity or it was just raw performance. I mean if you the sort of an obvious choice for just a file system would be something like EFS, the elastic file system from Amazon. Great for a lot of purposes. Okay? But if you're trying to get a lot of throughput through it, you will rapidly find by the time you've got like the third node in your HPC cluster accessing it, you've hit the ceiling.

[00:11:39] - [Speaker 1]
Okay? There's just only so much throughput in aggregate that it can provide. So that's not gonna work. And even Amazon would admit, yeah, that's not a good choice. K?

[00:11:49] - [Speaker 1]
But are they pitching the object storage? No. They're pitching FSX for luster. Right? Because you need that same kind of parallelism and satisfy the POSIX need, that's the way to do it, only that came with management overhead, it came with inelasticity, it came with frankly money, okay?

[00:12:09] - [Speaker 1]
Whether you're managing it or a hyperscaler is managing it you're paying quite a bit to operate that. And what we're doing instead is giving you the same kind of parallelism, okay, leveraging the hyperscalers object stores parallel capabilities but solving the latency problem. So we can solve both of the things that something like Lustre was built to solve in a physical data center. We're doing it using cloud native technologies just making better use of those technologies and so making it that much more cost effective. So you're you're kind of getting the best of all possible worlds here.

[00:12:47] - [Speaker 0]
And for a business leader listening, hearing about Paradigm four and FlexFS for the the very first time, if I ask you to explain FlexFS to maybe a business leader, CIO, but in plain English, what business problem does it solve for Lemon and why does this matter so much? What it solves is basically giving you better bang for your IT dollar.

[00:13:09] - [Speaker 1]
Okay? And ideally you need that balanced system, okay? You need there's an axiom that's your system is only as fast as the slowest component and so a hidden bottleneck quite often is storage. Yeah. Okay?

[00:13:26] - [Speaker 1]
Becoming what's called IO bound and that costs you money and it costs you time. So especially if you're say now you're working with agents and AI and so you need to really get a response quickly to a human that's just asked a question, okay? And so that time to first token, basically the beginning of the response as well as the economics of how much did it cost me to complete that task, Okay? And so both of those aspects are really important. It's keeping the engagement with the human, being able to make it as natural interaction as possible, but be able to do that at scale because you wanna be able to do that for a few million of your closest friends.

[00:14:05] - [Speaker 1]
And it's about getting all of that value that you can from your your compute spend and your storage spend in a balanced system.

[00:14:15] - [Speaker 0]
And if we look at much of the company's early success, it has come from pharmaceutical and life science organizations. I'm curious, are the challenges that they face now appearing in other industries too? Because we're not talking about one industry here, are we, that you serve?

[00:14:31] - [Speaker 1]
No. No. That's exactly right. And while it's interesting, a lot of the problems that they solve are specific to drug discovery and clinical trials and so on but they're still solving fundamental IT problems in order to do that. And so what what we need to do in going in other verticals is to be able to look at those use cases that will resonate to senior management of those companies.

[00:14:56] - [Speaker 1]
Hey I've got manufacturing line and I've got IoT instrumentation all through it. I've got this digital twin I have to deal with in order to try to do maintenance or optimization. That's a lot of data to have to deal with. How do I do this? How do I do it in a performant way and so on?

[00:15:16] - [Speaker 1]
So technologies can be used in similar patterns in different problem spaces and so that's we just need to get better at articulating that and having those conversations with folks.

[00:15:30] - [Speaker 0]
And one of the key themes I heard throughout our conversation today and the presentation was that organizations are spending huge sums on infrastructure but still waiting for data. How much of the AI conversation is really becoming data access conversation? Is that what you're seeing more and more now?

[00:15:47] - [Speaker 1]
There's a lot of that aspect coming into the AI conversation. Probably not as much as it should in all honesty. There are some companies out there that have realized, okay, I need to actually maintain a cache of my key values and my my tokens and just maintaining context of all the stuff that I'm doing. Because now you've got people in kind of intense chatbot conversations asking questions and then details about that and details about that and details about that scaling to several million users. And so it's one term I've seen coined is tokenomics.

[00:16:22] - [Speaker 1]
And but it's not just a compute conversation. Tokenomics is all about the data. And so how how do you handle that data? How do you transmit it? How do you manage it?

[00:16:32] - [Speaker 1]
How do you do it that's both time and cost efficient because both are vital in this environment.

[00:16:38] - [Speaker 0]
And Paradigm four is seen as platform enabled clinical and Omnics evidence generation company. Yet we're also talking about today storage, analytics, AI, and infrastructure. What's really caught your attention?

[00:16:52] - [Speaker 1]
One thing I think would be about cloud vendor independence. I've seen that rising as a concern and on on a couple of dimensions. One is around just surviving catastrophic hyperscaler cloud failures. An entire region goes out or something and we've seen it on occasion with Microsoft, Amazon, Google they've all had these kinds of outages. And so the business continuity concern is well what do I do I've got my company Joules in that particular hyperscaler in North America, now what?

[00:17:28] - [Speaker 1]
Right? So it's like how do you become independent not only of a regional failure but just that vendor. And so one of the things that was important for us in designing FlexFS is to make it completely cloud agnostic. And so we run on any of the big hyperscalers also the neo clouds I mean that's going to be of rising importance I think to a lot of folks and even on premise. So for folks that are in that hybrid space and there's a lot I'm sure you know that.

[00:18:01] - [Speaker 1]
And and so but you can actually use a technology like FlexFS because it is so agnostic as a way to sort of federate your data across these various platforms and to be able to use them as a kind of sort of common lingua franca for those and be able to stay kind of independent of all those vendors whether it's for geopolitical reasons, geo specific regulations so on, sovereignty you know whatever the reasoning we see that as an increasingly important aspect of doing business in the modern era is actually being able to transit clouds at will and smoothly.

[00:18:46] - [Speaker 0]
And for anybody, I think that is a thought provoking moment to end on. And for anyone listening in wanting to find find out more about anything we talked about today, anywhere in particular you'd like me to point them to, I will include a link to your own LinkedIn but just wanted to find out more about FlexFS etcetera. Where do you like me to point them?

[00:19:03] - [Speaker 1]
Well certainly flexfs.io and all the documentation is actually right there at docs.flexfs.io and we have a community edition available for folks who want to actually kick the tires of the technology. So there's all the online documentation even will explain how to use the community edition and it's full function, full feature. The only thing missing is something we call a proxy server. It's a a separate service that provides write back caching and you're limited as to number of volumes. It's one.

[00:19:38] - [Speaker 1]
And the the amount of capacity you can consume, it's five terabytes, but it's free and you can self install and just try it out but if you want to talk to us about starting a POC or just find out more you can feel free to contact us at info@flexfs.io.

[00:19:58] - [Speaker 0]
Awesome. I will add links to absolutely everything there. Kudos to you. You've been speaking for nearly three hours straight now. He was doing a presentation and straight into a podcast with me.

[00:20:07] - [Speaker 0]
I encourage everyone listening to check out the links that I will leave. And also, as you said, kick the tires. Have a play with the technology. I think that is one of the best things that you can do. But thank you more than anything for sitting down with me today.

[00:20:18] - [Speaker 1]
Well, thank you for having me. I really appreciate it.

[00:20:20] - [Speaker 0]
One of the things I loved about today's conversation is if you strip back everything that we talked about, it wasn't about storage at all. It's about identifying a bottleneck that customers were experiencing every single day and refusing to accept that it was simply the cost of doing business. And throughout our conversation today, David returned repeatedly to a simple idea. Organisations can invest millions in infrastructure, AI initiatives and high performance computing. But if data cannot move efficiently through these environments, much of that investment just fails to deliver its full value.

[00:20:57] - [Speaker 0]
And I think it's a reminder that the future of AI might depend just as much on how organisations manage access and move information as it does on the power of the models themselves. And I think we also touched on another theme that is becoming increasingly important today. And that is, as businesses continue to spread workloads across multiple clouds, on premise environments and emerging AI platforms, infrastructure flexibility is becoming the strategic consideration rather than just a technical preference. And the ability to avoid lock in, maintain resilience and move data where it needs to be, I think this is something that could prove just as valuable as raw performance. And as my guest mentioned today, many organisations think they have a GPU problem when they actually have a data problem.

[00:21:50] - [Speaker 0]
So a big thank you to David for joining me today here at the IT Press Tour in Boston, and for sharing the story behind FlexFS and some of the challenges that inspired it. But as always, you'll find links to everything discussed in today's show notes over at techtalksnetwork.com. I'd love to hear your perspective. As AI projects scale and data volumes continue to grow, Are you and your organisation paying enough attention to the infrastructure that feeds these systems? Or are you still focusing too heavily on compute alone?

[00:22:25] - [Speaker 0]
So many big takeaways here, but let me know techtalksnetwork.com and I will return again tomorrow with another guest. Thanks for listening. Bye for now.