Can an enterprise move quickly with AI when the information it needs remains spread across business units, acquired companies, private data centers, multiple clouds, and systems governed by different privacy rules?
In this episode , I speak with Justin Borgman, co-founder and CEO of Starburst, about the data architecture decisions now affecting how quickly businesses can turn AI investment into useful results. The conversation begins with a reality many established companies will recognize. Their technology estate reflects years of applications, acquisitions, regulatory demands, regional choices, and earlier infrastructure programs.
Justin argues that this history becomes a constraint when AI teams need governed access to information quickly. A company created within the last year may design its data environment around AI from the beginning. A multinational enterprise rarely has that freedom. It must work with valuable data held across different locations while respecting security, privacy, and sovereignty requirements.
The traditional response has been to centralize everything. Justin believes the single source of truth is often an impossible target rather than a finished destination. Drawing on his earlier experience at Teradata, he says even customers using a leading database continued to retain information elsewhere. New applications, company acquisitions, regulations, and changing business demands kept creating additional systems.
His preferred approach is to accept that distributed data will remain part of enterprise life and build an architecture that can work across it. A federated data platform can query information where it lives while giving users a common point of access. This can reduce the time spent moving data before an analyst, executive, application, or AI agent can use it.
Justin illustrates the problem through a large American bank with many lines of business and inherited data silos. Senior leaders could ask commercially important questions, but answering them required analytics teams to write queries, assemble dashboards, and combine information from several systems. The process could take weeks.
He describes how connecting those sources and adding a natural-language interface can reduce that delay. Starburst calls its interface ADA. It allows a user to ask questions in conversational language while drawing on governed enterprise information and its business context. The company presents this as a way to shorten the path from question to insight without forcing every data set into one platform first.
Business value remains harder to prove than technical access. Justin recommends connecting data projects with revenue growth, cost reduction, or risk management. He also points to usage evidence within data platforms. If a data product is accessed frequently by leaders or operational teams, and the cost of producing it is known, those signals can help a company evaluate whether the investment is serving a recurring need.
Architecture economics also influence where workloads belong. Justin favors S3-compatible object storage and open formats such as Parquet and Apache Iceberg when companies assemble data in a lake architecture. His argument is that open storage can reduce cost while allowing customers to choose among query engines rather than binding the information to one provider.
That advice does not mean every workload belongs in one format or location. Some data will remain in warehouses, operational databases, regional systems, and on-premises infrastructure. Federation can provide access across those environments, while open formats create additional choice for the information that can be consolidated economically.
The human side of architecture receives equal attention. Justin says ownership, incentives, and internal politics affect data quality because centralized teams may lack the domain knowledge held by the business unit that produced the information. Treating data as a product gives an organization a way to assign responsibility for quality, maintenance, adoption, and feedback.
Named ownership changes the conversation. A successful data product can be recognized and improved because people know who created it. A weak product can receive feedback from its internal users. Distributed control can also allow teams closest to the data to apply their knowledge while the wider company accesses it through common governance.
Governance becomes especially important when AI agents can query enterprise information directly. Justin says access controls must operate beneath the agent rather than relying on the model to decide what a user should see. Row-level and column-level permissions, data masking, and query auditing can determine which records are available and provide a record of what the system retrieved.
Security and economics are also contributing to renewed interest in running some AI workloads on-premises. Justin says businesses may prefer open-weight models on owned hardware when scale improves the economics or when confidential data cannot comfortably leave the corporate firewall. He expects cloud and on-premises capabilities to coexist rather than one replacing the other.
For employees, conversational access could change the role of traditional business intelligence. Justin expects standard KPI dashboards to remain useful, but believes many custom reports and one-off dashboards could be replaced by interactive models that answer questions directly. Analysts and engineers may spend less time responding to requests and additional time improving trusted data products and governance.
The opportunity is faster access to answers. The risk is allowing speed to outrun security, quality, or accountability. A federated architecture can connect distributed systems, but it still requires companies to know who owns the data, which users can access it, how results are audited, and whether the outcome supports revenue, cost, or risk goals.
Should enterprises keep pursuing one central source of truth, or accept distributed data as a permanent condition and build governed AI around it? Listen to the episode and share your thoughts.
Useful Links
Connect with Justin Borgman
Learn More About Starburst

[00:00:00] - [Speaker 0]
Do you need AI agents that you can trust? Well, with an AI data layer providing real time connection within your data platforms, you can trust your agents to provide accurate solutions. So scale your business by trusting your agentic AI, accurately getting the work done for you. Trust its capabilities with Denodo. And you can do that by simply visiting denodo.com to learn more.
[00:00:30] - [Speaker 0]
Can an enterprise build useful AI when its data is scattered across different clouds, data centers, business units, acquired companies, and systems that let's be honest, they were never designed to work together. Well, in this episode of AI at Work, I'm gonna be speaking with Justin Borgman, cofounder and CEO of Starburst. And we're gonna talk about why the long promised single source of truth still remains elusive and what companies can start doing instead. Because he's gonna explain how federated data access can shorten the path from business question to trusted answer and why open storage formats give customers greater choice. And why governance must sit when AI agents begin querying sensitive information.
[00:01:24] - [Speaker 0]
I will also discuss how to measure the value of data products, why ownership and accountability matter every bit as much as architecture, and whether conversational AI could replace parts of that traditional business intelligence. Yeah. We've got a lot to get through today because it turns out that asking data a question may soon become much easier than simply finding a dashboard containing yesterday's answer. But enough for me. Let me introduce you to Justin now.
[00:01:56] - [Speaker 0]
So thank you for joining me on the show today. Can you tell everyone listening a little about who you are and what you do?
[00:02:03] - [Speaker 1]
Absolutely. Yeah. I'm Justin Borgman, cofounder and CEO of Starburst. We've been building this company for eight and a half years. We're a data platform, that allows you to run fast SQL analytics both in a data lake as well as federating out into the other data sources that you have.
[00:02:21] - [Speaker 1]
So you really have a single point of access for analytics and now increasingly for agentic applications, which I'm sure we'll get into at some point.
[00:02:29] - [Speaker 0]
Yeah. I mean, all things data and AI is so... Such a big topic right now. A phrase I hear a lot of tech conferences, no data, no AI. And AI teams, of course, they want access to more data and faster.
[00:02:43] - [Speaker 0]
So where does the existing data architecture become the real limit? Because much of what we use, was built for a completely different era. Right?
[00:02:52] - [Speaker 1]
That... That's exactly right. And that's precisely the challenge that I think our customers typically face. Unless you were a AI native company born in the last twelve to eighteen months, you have carried with you a lot of technical decisions that you've made over years or decades. And the reality is that does become the bottleneck to AI progress.
[00:03:16] - [Speaker 1]
And so what we find is, you know, one of the first challenges, maybe one of the most foundational challenges to any AI project, it's figuring out how to get access to the data that you need. And as I think you rightly said, AI is fundamentally a a data story at the end of the day. Data is the fuel for the AI engine. And so setting things up to be able to get that access is really important. And and among, you know, Fortune 500 customers, large multinational enterprises, this is an even more extreme challenge because they might have data in different locations.
[00:03:52] - [Speaker 1]
They might have it on prem, in the cloud, in different regions for data privacy or data sovereignty. So this is a real challenge that that people are facing today.
[00:04:01] - [Speaker 0]
And I'd love to bring to life what we're talking about here and the problems and also the solutions out there because is there a, I don't know, a customer story. You don't have to name any particular names that really show what changes when analytics move closer to operational data. I think it will it will really hammer home what we're talking about.
[00:04:20] - [Speaker 1]
Yeah. Absolutely. We have a lot of customers where this is exactly the case. Maybe one that I will use is a very large bank in The United States, one of the largest banks in The US, where they have many lines of business and as a result, many different data silos across those different lines of business. They've also grew through acquisition, like many large organizations have, and that has also inherited even more data silos.
[00:04:49] - [Speaker 1]
And so, one of the challenges that they had was even at the top of the house among leadership team, you know, even the CEO, how do they get answers to his questions? You know, when he's getting on a plane to go to Davos and he wants to know what's the latest going on, even, you know, from an economic perspective because these banks have such an incredible view on what's going on in real time in our our global economy. Those were really challenging questions to answer. You know, it required creating custom dashboards and really weeks of the analytics team to answer those questions. And, you know, by putting in a platform that connects to all those data sources directly, and then on top of that, even if you add a agentic interface...
[00:05:34] - [Speaker 1]
And we've built something we call AIDA, which is a a natural language interface to your enterprise data, that now puts, you know, the answers at your fingertips. Now you can interact with your data much the way that you would ChatGPT or Claude, but you're connected to your real enterprise data with the context of what that enterprise data means. And that can be very powerful in terms of shortening that time to value, that time to insight, that people really care about.
[00:06:02] - [Speaker 0]
And I think over the last eighteen months, the costs that surround AI have also placed a huge emphasis now on return on investment from just about any tech project. So for business leaders listening, what business result would ultimately justify the cost and complexity of a a new data platform? Because I think over the last few years, many have got distracted by the shiny next big thing, but now we're coming back to business outcomes, measurable improvements. So what would justify the cost here?
[00:06:32] - [Speaker 1]
Yeah. Absolutely. It's interesting. We were having this discussion yesterday at a... At an event that I attended with some customers.
[00:06:39] - [Speaker 1]
And, you know, measuring value can sometimes be tricky, especially when the value might seem difficult to measure. Like, for example, agility or speed of decision making or the quality of the decisions you make. Sometimes those can be a bit subjective and difficult to measure. But one of the things that we've seen be a successful proxy for value is within our platform, we, of course, measure every query that takes place and every dataset, data product that gets accessed. And, you know, we can tell you what data products are essentially accessed more frequently, and by whom, than other data products.
[00:07:24] - [Speaker 1]
We can also tell you the cost of the creation of those data products, you know, how many CPU cycles, how much compute time is being consumed when those data products are being accessed. And you can get some... At least some data to back up the claims of like, hey. You know what? Actually, assembling this data product is accessed by our business leaders every day, all the time.
[00:07:47] - [Speaker 1]
Clearly, there's a high return there relative to the investment in assembling that that that data product. So that was, one interesting discussion that we had. But I think more broadly, you know, generally, customers are looking to either increase revenue or reduce cost. And so if you can tie your projects to one of those two things, the third would be perhaps reducing risk. Again, if you're a a a bank or some regulated business, risk might actually be a very meal...
[00:08:14] - [Speaker 1]
Material value driver as well. But generally, you know, if you can tie your projects to, you know, increasing revenue, reducing cost, or managing risk, you know, those are good places to start.
[00:08:25] - [Speaker 0]
And I've been recording this podcast for, well, eleven, nearly twelve years now. And on... In that time, I've been talking around fragmented data silos. I think now AI has exacerbated this. So why is it that companies still struggle to make govern data across teams even after investing heavily in cloud infrastructure in all this time?
[00:08:45] - [Speaker 1]
Yeah. I think there are a couple reasons. I think, first of all, it is really impossible to get everything in one place. And I think this was one of the lessons that I had when I used to work at Teradata. So before starting Starburst, I was an executive at Teradata.
[00:09:01] - [Speaker 1]
Teradata had acquired my first startup fifteen years ago. And, you know, Teradata at the time was the the leader. They they were the snowflake of, you know, the past thirty, forty years. And what was interesting was despite being, you know, the leader, despite having a phenomenal database system, by the way, not one of their customers actually centralized everything there. This this single source of truth was always kind of a a myth that that people were pursuing.
[00:09:31] - [Speaker 1]
And that is really what struck me that no matter how hard you try, you're always going to have new applications, new use cases, new acquisitions of companies, new regulatory constraints, new sovereignty, you know, privacy constraints. There there are so many things in our dynamic world that are going to make it extremely difficult to consolidate in one platform. And so what we argue for our customers is you should kind of embrace that as the truth. And rather than continuing to pursue this impossible dream of having everything in one database system, build an architecture that handles that heterogeneity, that that dynamism, that change that's inherent to just operating a business today. And and I think this becomes even more important in the era of AI because speed is everything.
[00:10:22] - [Speaker 1]
You know? Speed might be the difference between your business growing and becoming an industry leader versus being disrupted and displaced. And we've seen that so many times in technology history where a new technological leap comes in, whether it's computing, you know, the Internet, mobile, cloud, and companies are disrupted. So speed is everything, and I think building an architecture allows you to get access to the data where it lives. Will always be faster than one that forces a a centralization paradigm, which is very difficult.
[00:10:53] - [Speaker 1]
We... You know, one example of this, we see a lot of people who are trying to do that, spending years on the migrations. And so, you know, they... They're waiting waiting and waiting and waiting for this, eventual reality that they think they can achieve. In the in the process, they're losing ground to their competitors.
[00:11:13] - [Speaker 0]
And I think a decision that many leaders are grappling with right now is deciding which workloads belong in a warehouse, lake house, or another architecture. And I appreciate this question on its own is an entire podcast episode. But for people listening that are struggling with this right now, any any tips or advice on how they should best decide that?
[00:11:35] - [Speaker 1]
Yeah. So first and foremost, putting myself in the customer's shoes, I I think about the economics of what you're trying to do. You know, economics drives everything at the end of the day or certainly at scale, economics really matters. The lake format, if you will, where you're storing your data... And and when I say lake, I really mean something that is s three compatible object storage.
[00:11:57] - [Speaker 1]
It doesn't have to be s three. It could be GCS and Google. It could be in Azure. It could be on prem object storage. And we're seeing a lot of investment these days in various types of s three compatible object storage on prem.
[00:12:11] - [Speaker 1]
You know, NetApp, Dell, Everpure, MinIO. There there are a number of these storage players that basically look like s three even on prem. So the key is you standardize on something that looks like s three, is s three compatible. And then you use open formats like Parquet files. And more importantly, today, Iceberg as a file format is this open source format that you own.
[00:12:33] - [Speaker 1]
You put that in object storage, and now you've created by far the lowest cost storage layer in your overall data architecture. So I always recommend that you do that for as much data as you can. Wherever you are going to assemble data in one place, you should use open formats and and s three compatible object storage. And that is essentially the the foundations of the lake architecture. And as you do that, now you have optionality as a customer to decide what query engine is going to access that data.
[00:13:04] - [Speaker 1]
Of course, we'd love for you to choose Starburst, but you have many choices. And choice is leverage for you as customers. You know, this is the best possible advice I can give you that the more choice you have at various layers of your stack, the more leverage you have to control your own costs and performance. So I think that is the sort of ideal setup. But then again, acknowledging that you're gonna have data that lives outside and being able to federate into those other data sources gives you a complete picture, a complete way of of interacting with your enterprise data holistically.
[00:13:38] - [Speaker 0]
And I've been to many tech conferences this year and seen many vendors that often present architecture choices as purely a technical debate. But I'm curious inside many organizations, where are you seeing ownership incentives and internal organizational politics or culture? How do you see those things mattering more?
[00:13:58] - [Speaker 1]
This is such a great point. And you're right. It is very often overlooked or or not discussed sufficiently. We we see that all the time. I mean, I I think any large organization is going to have an element of organizational politics.
[00:14:14] - [Speaker 1]
It's it's sort of inevitable. Any organization of size, I mean, we even see it in our own company. You know? So I think you have to think about who is, who is responsible, who is accountable for curating and creating and maintaining high quality data products. And that that is sort of at least the way that I focus the conversation because that is ultimately what you're gonna have a lot of people in the organization consume.
[00:14:43] - [Speaker 1]
So if you get that data product right, if you've curated appropriately and think about data as a product. This is one of the the key takeaways from, Jamak Degani had written a lot about something called data mesh many years ago. And and this this concept of treating data as a product was one of her her key pillars, if you will. And I think it's one that endures, even even as perhaps data mesh has has become less less popular as a paradigm, which is really to say, like, if you treat it like a product just like any other product, maybe every product that your business actually sells, you're gonna be really focused on the quality and delighting the end customer, which in this case may be your own internal data consumers. And so I think you wanna get that right.
[00:15:31] - [Speaker 1]
You wanna empower those people. You wanna give them as much flexibility and and control to create a product that delights, and then you wanna tie back accountability. You want everybody to know that that data product is created by, you know, John Smith, if you will. And and so John should get credit for its success. He should also get feedback when it's not meeting the mark, and that creates, an accountability system.
[00:15:58] - [Speaker 1]
I think also back to the topic of federated architectures versus centralized, you know, federated architectures allow for more distributed control. They allow for more autonomy within the various lines of business or within the parts of the organization that know the data the best, that the rest of the organization is set to consume. Very often when you're taking data out of different parts of the business and dumping it in a central place, the central team actually doesn't know anything about the data. And and it's no fault of their own, but they lack that domain expertise. So, you know, federated architectures can also allow you to have more domain knowledge about the data itself, which leads to better data products to be consumed.
[00:16:43] - [Speaker 0]
And for any cautious IT work... IT teams listening in today, when AI systems can increasingly query large volumes of enterprise data, are there any new security or government risk... Governance risk that that you see appearing here?
[00:16:59] - [Speaker 1]
Yeah. I I think governance and security need to be first class citizens, foundational things that you set up upfront, and they have to happen below where the agent is operating. Meaning that as the agent is querying data or asking for data, those access controls need to be applied. And, you know, I would say you you need very fine grained access access controls, you know, row level, column level, data masking. Of course, you want query auditing.
[00:17:27] - [Speaker 1]
You wanna know everything that it did pull or retrieve after the fact to understand the reasoning. So these are all very, very important. It's also a reason we see renewed interest in on prem systems, which, again, this is this is sort of counter to the trend of the last fifteen years of moving to the cloud. But we're seeing really two two factors driving more investment in the data center. One is certainly just the economics.
[00:17:53] - [Speaker 1]
You know, token maxing is a real thing. And so if you can deploy open weight models on prem, on your own hardware that you've purchased, you can very often get greater economies of scale. So economics are one. But the other is security, that there is a fear or concern that my, quote, unquote, secret data, my my confidential data, whether that's PII data, whether that's, you know, company secrets of various kind, you know, leave my firewall and leave to to go to some public LLM even if it's not supposed to be, you know, training on it or not supposed to be using that that data in any way. You know, that makes people very uncomfortable.
[00:18:34] - [Speaker 1]
And so that is another reason that we see people investing in having on prem capabilities as well. And, and again, I think there's gonna continue to be a lot of investment in open weight models. NVIDIA is obviously a huge supporter of this. And, and so I think that will be an important part of our future for both security and economic reasons.
[00:18:57] - [Speaker 0]
Now there is a huge focus on the impacts of AI in the workplace, the future of work, etcetera. So how do you see these platform decisions changing the daily work of an analyst engineer or business user listening? How are you seeing things evolve here?
[00:19:13] - [Speaker 1]
Yeah. I think it is hopefully going to supercharge the... All all three of those personas that you mentioned by allowing them to do more faster. You know, kind of back to the example I gave of the large bank where a typical workflow might have been picking up the phone, calling somebody on the analytics team saying, hey. Here's my question.
[00:19:35] - [Speaker 1]
Can you go figure out the answer? Then them diving in, writing a bunch of custom queries, building a dashboard, presenting that back to the executive. You know, that's a long feedback loop. By the way, I even have have had this over, you know, eight and a half years at Starburst where I have some question about my sales sales analytics. And I go ask my revenue operations team for data, and it takes them a couple days to get back to me.
[00:20:00] - [Speaker 1]
But leveraging agents here who have access to your government enterprise data means that I am now not bound by really any constraint other than how fast I can type and read the response from the LLM, and that's gonna allow me to move a lot faster. So I I think this really turbocharges everyone involved and allows them to get a lot more done. I do think it it probably means that traditional BI changes. Like, I do think AI will replace a lot of traditional BI. Not all of it.
[00:20:32] - [Speaker 1]
I think there's always gonna be your KPI dashboards that are really important, but a lot of the custom dashboarding or custom BI that you've done in the past, I think that will be superseded by these conversational models just because I think they're they're gonna be faster and and more interactive for end users. So that's the exciting promise of, I think, what this ultimately represents for companies.
[00:20:56] - [Speaker 0]
And as enterprises will inevitably accelerate investments in everything from AI analytics and data driven decision making, I suspect many listening will find themselves held back by fragmented data across on premise systems, and we have multiple clouds and hybrid environments. I know this is an area that you're passionate about. So for anyone listening interested in Starburst Enterprise Intelligence at scale, where should they go? Where can they find more information and connect with you or your team?
[00:21:26] - [Speaker 1]
Yeah. We'd love for you to check out our website, starburst.io, and and connect with us on LinkedIn, both both the the company on LinkedIn, and you're very welcome to connect with me as well. We produce a lot of content on both of those those channels as well as the blog on on our website itself. So we'd love to stay in touch. Please do reach out, and and we'd love to connect.
[00:21:49] - [Speaker 0]
Anyone interested in, tidying up their infrastructure, giving their organization secure, governed access across all data wherever it lives. I urge you to check out the links. They'll be in the show notes. I'd love to stay in touch with you and see how this continuously evolves too. I'm sure there'll be some big changes next year, so be look...
[00:22:09] - [Speaker 0]
Great to get you back on. But more than anything, thank you for bringing all this to life today in a language that everyone can understand. Really appreciate your time.
[00:22:16] - [Speaker 1]
My pleasure, Neil. Thanks thanks again for having me.
[00:22:18] - [Speaker 0]
Big thank you to Justin for joining me today and giving some very practical advice. You can continue waiting for every data source to arrive in one central system, or you can begin designing for the distributed reality that your business already has. And that means using open formats where consolidation makes economic sense, connecting securely to data that must remain elsewhere, and then applying access controls before an agent reaches that information. And this also means treating data products as products with named owners, clear quality expectations, usage evidence, and feedback from the people who depend on them. So a big thank you to Justin again for connecting data architecture with business speed, cost, risk, and the daily experience of employees who simply need an answer.
[00:23:12] - [Speaker 0]
So remember, you can learn more about Starburst at starburst.io. Follow the company and Justin directly on LinkedIn. And if your AI program needed reliable access to enterprise data tomorrow, Would that underlying architecture be ready, or would yet another migration stand in its way? Let me know. Techtalksnetwork.com.
[00:23:36] - [Speaker 0]
I'll be back again real soon with another guest, but thanks for listening. Bye for now.

