What does a rising cloud bill actually tell you about the value your business is creating?
Eight years after our first conversation, I welcome Kunal Agarwal, co-founder and CEO of Unravel Data, back to Tech Talks Daily. We compare the data infrastructure he was optimizing during the Hadoop era with today's enterprise stacks built around Databricks, Snowflake, BigQuery, AI pipelines, and autonomous agents.

Kunal says Unravel Data has analyzed over 10 billion workloads across hundreds of enterprises. From that work, he argues that data platforms and infrastructure can account for up to 60% of cloud spending at some global businesses, while 30% to 40% of data platform spending may produce no business value. These are company claims, but they frame a problem many technology and finance leaders will recognize. The cloud bill arrives after thousands of individual engineering decisions have already been made.
We discuss where cloud waste hides, including oversized clusters, hot storage holding cold data, abandoned pipelines, inefficient queries, duplicate datasets, and development jobs consuming production-level resources. The people creating those workloads seldom see the price attached to their decisions, leaving technology leaders with an aggregated bill that explains what was purchased but not why it was needed.
AI adds another complication. Humans create workloads at human speed, while agents can generate queries, launch infrastructure, and consume tokens around the clock. An agent is designed to complete its task, not worry about whether a single query costs $5 or $5,000. Kunal argues that machine-speed consumption cannot be governed through monthly human reviews.
We also discuss the difference between cost cutting and cost optimization, why aggressive reductions can damage performance and reliability, and how FinOps must connect cost with business outcomes. Kunal explains why leaders should measure cost per pipeline, model, agent, successful run, customer report, and business result.
Finally, we consider the benefits and risks of autonomous data platform optimization. Kunal describes autonomy as a dial, with bounded, reversible, and validated actions earning wider authority as trust develops.
Does your cloud bill show healthy growth, or is expensive waste hiding behind the headline number? Share your thoughts with me.
Useful Links
Learn more about Unravel Data

[00:00:04] - [Speaker 0]
How can a company tell whether rising cloud costs reflect healthy growth or just more expensive waste? Well, today on Tech Talks Daily, I welcome an old friend back to the podcast eight years after our first conversation. Yep. Believe it or not, I've actually been doing this eleven years now, but he is the cofounder and CEO of Unravel Data, where his team studies how data platforms consume cloud resources, a topic that is certainly much more important now. And he will explain today why AI is accelerating costs across Databricks, Snowflake, BigQuery, pipelines, storage, and compute.
[00:00:52] - [Speaker 0]
And we'll also discuss the claim that 30 to 40% of data platform spending can produce no business value. So what I wanna discuss today why a cloud bill or how a cloud bill is a receipt rather than a diagnosis and how AI agents could become a company's most expensive users. And if your dashboard show rising spend but cannot explain the business outcome behind it, I'm hopeful today's conversation could save you an awkward meeting with the finance department. But enough for me. Let me reintroduce you to my guest.
[00:01:28] - [Speaker 0]
So a massive warm welcome back to the show. It's been an incredible eight years since we last spoke. So can you remind everyone listening with a little about who you are and what you do?
[00:01:39] - [Speaker 1]
Yeah. No. Thank you, for having me here, Neil. So cofounder and CEO of Unravel Data and Arvix, founder of this company with, Shivnath Babu, our CTO. And it's been roughly twelve years, that we've been obsessed with this one problem that data platforms cost too much money and run too slow.
[00:02:01] - [Speaker 1]
And we've been on that endeavor to fix it. Eight years ago, I think we were, let's say, the Fitbit for the data stack. And now we're the actual doctor that solves and treats the problem for you, and that's really what's changed and evolved since then. Back then, the environment we were focused on were primarily Hadoop and Spark troubleshooting. Now there was a massive wave of cloud migration and the rise of Databricks, Snowflake, BigQuery, and the likes, and now we're in the AI era.
[00:02:34] - [Speaker 1]
Today, what we ultimately focus on is autonomous data platform optimization, which is really observability and FinOps, and doing it in automatic fashion using RBICs. It's not just about recommending fixes, but something that actually safely does that on behalf for you on your environment.
[00:02:57] - [Speaker 0]
Wow. That's incredible. Just how much the world has changed since we last spoke. I suspect the last time we spoke, everything would have been blockchain and crypto, maybe. Right.
[00:03:08] - [Speaker 0]
But, of course, now now AI has taken over all conversations. More recently, Agentic AI and agents, etcetera. And before, you came on today, I was reading how you're hearing from global enterprises that data platforms and infrastructure now account for as much as 60% of their cloud spend. So the question I've got to ask, I mean, the the eight years since we last spoke, did we get here, and how much of that growth is being driven by maybe AI rather than traditional analytics and data workloads?
[00:03:41] - [Speaker 1]
Right. So we've analyzed over 10,000,000,000 workloads across hundreds of enterprises, lots of exposure, and fortunate to be working with several Fortune 500 companies. How we got here really, Neil, is there was some lift and shift migrations that moved from on premises to these Cloud platforms that drove a lot of it. Democratization, far more people now, far more teams now that are running workloads on top of these platforms as well. AI, which really sits on top of everything.
[00:04:15] - [Speaker 1]
Data is the gravity, which means that every app, every dashboard, every model now inside the enterprise is running on top of this data stack. It's becoming the fastest growing line item in anybody's cloud build these days. So, is not the origin, but AI is the accelerant, right? Which means that every model needs a data pipeline that's feeding it. So, you think about the data prep, the feature engineering, the vector stores, all of that happens on top of the data stack.
[00:04:47] - [Speaker 1]
Then you have Gen AI now running on top of these different pieces, which really converts every department inside the company to a consumer of that data. So, the waste patterns actually predate AI, AI is actually compounding them at machine speed and it's happening really fast right now. So, didn't really start the fire, I would say, but it's definitely pouring jet fuel on top of it.
[00:05:14] - [Speaker 0]
Yeah. 100%. And I was also reading that you make quite an important distinction that is that the problem isn't necessarily how much companies are spending right now, but maybe how much they're actually wasting. So where does that waste typically hide, and and why is it so difficult for CIOs and engineering leaders to to see it before that cloud bill runs on their door mat?
[00:05:37] - [Speaker 1]
Yeah. Look. There's a lot of wasted spend that we're seeing in all these companies from oversized clusters and warehouses, hot storage for cold data. There's zombie pipelines that are running that nobody owns, or you're making these development jobs run at production scale. There's inefficiencies in the queries, in the pipelines, in the duplicate datasets.
[00:06:01] - [Speaker 1]
It's really all across the stack. What we see is there's about 30% to 40% wastage that is happening on these data platforms that it's really adding no business value to that company that's actually using this. The reason it's really invisible is the bill is a lagging indicator and not a leading indicator. It's really an aggregated artifact. The costs are really created by these thousands of engineering decisions that are happening in these these smaller groups inside the company.
[00:06:33] - [Speaker 1]
But the bill comes in thirty days later and it doesn't really help you understand where that money is even going. It's coming by the SKU or the account, but it's not coming by the user or the job or the query or the pipeline that actually consumed that. Nobody can act on it if you just see, Hey, my computer is $2,500,000 What do you do? The cloud build really becomes a receipt and not really a diagnosis. You can fix on the thirtieth of the month the inefficiencies that are happening on the fourth of the month.
[00:07:08] - [Speaker 1]
You really have to be real time and have the right governance in place to be able to do that. That's why it becomes after the fact scratch your head kind of a problem. And that's why a lot of people that we work with are surprised with their monthly builds, that's why they bring something that can ravel inside.
[00:07:25] - [Speaker 0]
It seems for years we've celebrate the democratization of data, giving more employees access to powerful analytics platforms, maybe one too many dashboards and removing technical barriers. But I'm curious. Has that success created this unintended consequence where thousands of people can now consume expensive compute and storage without actually understanding or caring for the economics behind their actions?
[00:07:53] - [Speaker 1]
Yeah. So look. I mean, democratization worked. Yeah. That's the problem.
[00:07:59] - [Speaker 1]
Right? We removed the friction of access, but we never added that accountability of cost and reliability. Right? Now a single SQL query can spin up and spend thousands or tens of thousands of dollars in compute, but the person who's actually spinning up the query never sees the price tag. Now you multiply that by thousands of employees that's running inside a company.
[00:08:24] - [Speaker 1]
It's like having a corporate card with no limit and no statement. Or as you may think about the analogy in The UK, think about an open bar with no menu prices. You don't know how much it's adding up and where you're actually wasting those costs as well. The answer isn't to really take the access away though because it's a good thing that a lot of people are now becoming data driven there's good science behind the decisions and the products that they're making. The answer is to really make the efficient path the default path, right?
[00:09:02] - [Speaker 1]
Is how do we run these things? How do we remove wastages from these things and make sure that we're still driving business value, but in a responsible efficient manner.
[00:09:15] - [Speaker 0]
It's so true. Although we are talking about cloud economics today and and the money that can be spent there, we I think companies are struggling because they're scanning it from all angles at the minute. We've got AI and tokenomics there. I don't know if you saw this a few maybe a few months ago. There was a fintech company, and an employee during his lunch hour was building a a game in his spare time.
[00:09:36] - [Speaker 0]
He ended up going through $80,000 in tokens. It's incredibly hard to juggle both of these big spends, I would imagine. Right?
[00:09:45] - [Speaker 1]
You're absolutely right. So just the AI stack on its own has its own inefficiencies just like the data stack does. Right? Which model should you choose? What what are you caching?
[00:09:56] - [Speaker 1]
How do you do the right prompt for efficient token usage, there's a whole science there as well. But yeah, your example of agents versus humans, humans create workloads at human speed, agents create them at machine speed. So it's happening 20 fourseven and an AI can generate queries in a minute and spin up clusters and infrastructure while you're sleeping. So it's happening all the time. Honestly, the agent doesn't care about a $5,000 query, doesn't even flinch.
[00:10:30] - [Speaker 1]
What it optimizes for is task completion, and it doesn't optimize for running this task in an efficient manner. I think that's the difference. The most expensive employee in your company, twenty twenty seven, may not be an actual employee at all. It's probably an agent that's just writing all these queries and the meter is never sleeping. This problem really exaggerates that fact, which is you cannot govern machine speed consumption with human speed reviews.
[00:11:02] - [Speaker 1]
You need to have machines running consistently measuring, governing, graduating these jobs and pipelines code and agents to make sure that all of these things are happening in an efficient manner. Again, if you go back to how you think about FinOps one point o, which is, hey. Let's look at this review thirty days later after the bill arrives. It's already too late.
[00:11:25] - [Speaker 0]
Yeah. As you say now, there's so many important points in everything you said because agents can create queries, workloads, pipelines, and infrastructure activity at machine speed. I think that's worth repeating, and it's no longer just humans consuming resources. It's AI agents that could become the most expensive users inside an enterprise as you've said there. And when cloud bills rise, though, the instinct from the CFO might be to demand a percentage reduction, but you argue organizations need to maybe move away from cost cutting to cost optimization, which again makes perfect sense.
[00:12:04] - [Speaker 0]
But what's the difference, and how do you prevent an aggressive savings target from damaging performance, reliability, or even innovation?
[00:12:13] - [Speaker 1]
Right. So it's so it's very simple. Cost cutting asks, what can we stop doing? Yeah. Optimization asks, what are we doing badly?
[00:12:23] - [Speaker 1]
It's not that we need to stop doing. We need to run it in a much more efficient manner. We've built companies where the CFO has mandated a blunt percentage handed down from the top like, Hey, we need to freeze projects. We need to shrink clusters badly. What that does is it has a negative impact, unlike you rightly said, on the outcomes of the businesses we're eventually trying to drive.
[00:12:47] - [Speaker 1]
Optimization is getting those outcomes but being cautious and responsible for how you're actually doing those things. Step one is really transparency. Everybody should know what it costs, what improvements can I actually make at the code level, at the pipeline level, at the infrastructure level? You've got to reframe the savings. To go to optimization actually could fund more innovation.
[00:13:18] - [Speaker 1]
If you're lowering the cost of work per unit, then we can get more work done. It's not about, hey, let's just reduce cost because this is what our budget demands. Because look, as much as we're talking about costs, the real talk is on the business value side, which is AI is enabling so much innovation inside the companies, improving processes inside the companies. Of course, we don't want that to stop. But the music only stops at a point where there's nothing to show for the spend that you have.
[00:13:50] - [Speaker 1]
But if you have something to show for it and you're able to show that the denominator, which is the cost was governed and this is actually an efficient cost structure, then you're never going to have the conversation with CFO, which is saying, hey, let's go and cut across the board. So we should always think about cost optimization rather than cost
[00:14:10] - [Speaker 0]
Your agentic AI might not be secure even with real time data and proper guardrails. But Denodo makes sure your business has every avenue covered. By placing all your data platforms under one AI data layer, your business can reach semantic consistency safely and securely. So get your agents on the same page by visiting denodo.com, And you can learn more about how to start trusting your agents to make business decisions. But now back to today's guest.
[00:14:46] - [Speaker 0]
FinOps traditionally gave organizations visibility into where that infrastructure money was going. But, again, I was reading before you joined me today that you believe that model now needs to incorporate performance, reliability, context, and business outcomes too. So what does that next generation of FinOps look like when the question becomes not simply, what did this workload cost, but was it worth running?
[00:15:11] - [Speaker 1]
Correct. So, I mean, I think about it as FinOps one point o and FinOps two point o. Right? So one point o was all around visibility and allocation. So it's very important.
[00:15:19] - [Speaker 1]
It's necessary, but it ultimately shows you what a team spent and it stops at the invoice. The 2.0, as we are thinking about it, is something that joins cost with performance, with reliability, and the business context. What did we actually get out from these investments as well? Then from there, you get unit economics like cost per business outcome, cost per AI model, cost per pipeline run, cost per customer report delivered, whatever that unit economics makes sense for your business. And then some questions around, hey, did this workload actually need to run at all?
[00:16:01] - [Speaker 1]
It needed to run, but did it need to run at this frequency? Does this need to get updated like every hour? Is this the right priority for running this thing, or do we need to run everything at 9AM? No, this can run at 4PM. Okay, that's fine.
[00:16:14] - [Speaker 1]
Half of the optimization is also about what not to do and what not to run, and I think that's excluded from how we think about FinOps 1.0. End state where FinOps 2.0 really leads is, let's get away from dashboards to actually autonomous actions and agentic FinOps that's continuously executing all these optimizations rather than, Here's the dashboard. Let's have a meeting. Let's send you a report. Let's figure out what to do next because that's already too late.
[00:16:48] - [Speaker 1]
So it needs to be in real time and it needs to happen at the, again, quoting that machine speed piece again, at the machine speed that these AI and data workloads are now running at.
[00:17:00] - [Speaker 0]
And at Unravel, you're now applying AgenTik AI to autonomous state platform optimization. And with early examples, including hundreds of thousands or even millions of dollars in identified savings, it feels like we've got a solution to much of what we've been talking about here today, but there's an interesting irony in using more AI to control costs partly created by AI. So what decisions are you comfortable allowing an agent to make autonomously, and and where should humans retain control?
[00:17:34] - [Speaker 1]
You've got to embrace the irony there, right? Using AI to control AI costs actually isn't ironic. It's the only thing that's fast enough, honestly, right? That's really what we've created with RBX, which is it's not an advisor, it's really an operator. It's hiring somebody inside your team that is taking care of this in a continuous manner before it even bubbles up into a problem.
[00:18:02] - [Speaker 1]
When we say autonomous control, we don't mean it as a switch, we think about it as a dial. It's something that's a comfortable or an enterprise grade governed autonomous behavior. So it's got to be reversible, it's got to be bounded, it's got to have validated actions across the spectrum of things that could be optimized from infrastructure to the code, to the data tier, to the scheduling and fixing zombie jobs, right? But you've got to make sure that every change is validated, and that's what we do before it hits production because the last thing you want is to do something bad because that breaks the trust of an enterprise really, really quickly. So, it has to be discovered, autonomous agent that's really sitting down and working on all of these problems on your behalf.
[00:18:59] - [Speaker 1]
What we usually see is people keep turning up that dial as the trust accumulates, so they start to see some things maybe in the dev environment and they say, This is a lower environment. It's good to go and act on these. Or, This is not a mission critical workload, so let's go and put it down over there. Honestly, it's really new for these enterprises as well. They don't know what they haven't seen.
[00:19:21] - [Speaker 1]
Once they start to see that, all right, it took three eighty five actions in the last week, none of them or maybe very, very few of them had a negative impact, it's mostly green. Then they start to unleash it on a bigger surface area, more protected workloads, which ultimately costs a little bit more money as well, of course, so they do get better impact out of that. I think we're in this messy middle phase where everybody wants to get autonomous across the enterprise, everybody wants to get agentic across the enterprise, and Cloud operations, AI operations is no different. They ask for it, but then the realities of their governance policies and infosec hit. I think that battle is something that we'll see clear up maybe in the next six months to nine months itself because it's moving very rapidly, where people are now saying we need to have some autonomous operators because otherwise it becomes a very, very hard and a very big problem for just humans to go and solve.
[00:20:29] - [Speaker 0]
And I think over the last few years, we've seen, a few problems around return on investment of AI and tech projects, and it has seen a return to the old IT mantra, thankfully, of you can only improve what you measure. And I always try and give people listening a valuable takeaway. So if we have a team listening who are rapidly growing Databricks, Snowflake, BigQuery, or wider data infrastructure costs, how do they determine whether that increase will represent a healthy investment or uncontrolled waste like we talked about? And most importantly, what metrics should leaders be asking for alongside that monthly cloud bill? What should they be looking for?
[00:21:09] - [Speaker 1]
Yep. So look. The monthly bill tells you the score. The unit economics tells you whether you're winning. So growth is not a signal.
[00:21:18] - [Speaker 1]
Efficiency of growth is. Right? So if spend is doubling and workloads are tripling and the SLAs that you have, those are holding, that's healthy. That's no problem at all. But if the spend doubles and the output is flat, that's leakage, right?
[00:21:34] - [Speaker 1]
So what you really need is to ask for more information alongside your bill. You really not need to start getting into the unit economics. So what's the cost of running a particular agent? What's the cost of running a particular pipeline? How are these costs correlated per business outcome?
[00:21:52] - [Speaker 1]
That trend line is something that you need to keep a very close eye on, that the business value is on par and starting to increase over those costs because the unit cost should fall as you're scaling, unit cost should not continue to keep going up as you're scaling. You need to really understand what's the waste ratio. What is your idle? What is your oversize? What is your wasted capacity?
[00:22:19] - [Speaker 1]
What's the cost per successful run of an agent? What's the spend growth versus the workload growth? My spend continues to increase, but my workloads are flat. It's a bad metric. Then what's the percentage of workloads that are actually driving business outcomes and what percentage of those workloads are actually meeting SLAs within the budget?
[00:22:41] - [Speaker 1]
There's a whole set of economics that as a company that's investing in data, in AI stacks, should be measuring right from the get go, because then you have a very good baseline and you really start to see trends and patterns over a couple of quarters time itself. So that's really how a CIO should be thinking about this and viewing this entire investment as well, which then helps them honestly go make a case to the finance side of the house around, yes, our spend is increasing and nobody is actually decreasing their spend on data and AI, especially now. But then this is the outcomes that we're getting from it, and our spend is efficient spend. That's ultimately what you're trying to answer. The music doesn't stop.
[00:23:28] - [Speaker 1]
The innovation continues, and there's no reason for having any sort of friction between, you know, the two sides inside the enterprise anymore.
[00:23:37] - [Speaker 0]
Well, thank you so much for sharing your insights today on some of the rising costs of data infrastructure. Unsurprisingly, yes, AI is adding to the problem, but global enterprises are now spending upwards of 60% of their cloud spend on data platforms and infrastructure. And it's important to remember that it used to be a real fraction of that kind of figure. So anyone listening wanting to find out more information about Unravel Data, everything we talked about today, how they can contact you or your team, where should they go?
[00:24:08] - [Speaker 1]
Yeah, Neil. So if you go to unraveldata.com, you'll get to check out, what we do. You'll see Arvix AI in action, and we run something called a health check where in a day, day and a half, we're able to return back to you and let you know where you're wasting money, where inefficiencies are hiding across Databricks, Snowflake, BigQuery, or follow us on LinkedIn. We do take some regular takes on data platforms and agenda KPIs there.
[00:24:34] - [Speaker 0]
Well, we have covered a lot today from the reasons behind what's driving the rising costs that we're talking about, the cause behind the waste in data infrastructure spend, and that mindset shift that's needed from cost cutting to cost optimization. So I will include links to everything that you mentioned now, and I urge people listening to check those links out. And, also, feedback to me, techtalksnetwork@at.com, and let me know your insights, your experiences. But more than anything, thank you for coming back on the podcast. We just need to make a promise to ourselves that we won't leave it eight years till we speak again.
[00:25:09] - [Speaker 0]
But thanks for joining me today.
[00:25:11] - [Speaker 1]
Thank you so much for having me, Neil, and thanks for the all the amazing questions.
[00:25:15] - [Speaker 0]
I think my guest left us with a useful test today. Yes. The monthly cloud bill. That tells you the score, but unit economics tell you whether you are winning. Rising spend can be perfectly healthy when workload, service levels, and business outcomes all rise faster.
[00:25:34] - [Speaker 0]
And the warning signs often appear when cost double while output remains suspiciously flat. And this means leaders should be asking for the cost per pipeline, model, agent, successful run, and customer outcome alongside waste ratios and workload growth. And another big lesson there that I think he delivered today was that autonomous optimization should almost behave like a dial with bounded, reversible, and validated actions, earning greater trust over time. So a big thank you to him for returning after eight long years. The world's a very different place now than it was in our last conversation, and you can learn more about Unravel Data and everything we talked about at unraveldata.com.
[00:26:25] - [Speaker 0]
But before I go, does your cloud bill honestly explain what the business received or merely what it has spent? Food for thought. And you can find me at techtalksnetwork.com. 4000 interviews. I'm also preparing to go back on the road with the show.
[00:26:43] - [Speaker 0]
I'm gonna be at a lot of tech conferences between October and Christmas. So, please, if you're going to any of them, check out the event page. Let me know. It'd be great to meet some of you in person. But that's it for today, so thank you for listening as always.
[00:26:56] - [Speaker 0]
Bye for now.

