Solving the $2,000 Per Hour GPU Problem With Lumen
IT Infrastructure as a ConversationAugust 14, 2026
26
00:23:2521.45 MB

Solving the $2,000 Per Hour GPU Problem With Lumen

What happens when an enterprise invests heavily in AI compute but cannot move data quickly enough to keep those expensive processors working?

In this episode of Infrastructure as a Conversation, I speak with Jim Fowler, Chief Technology and Product Officer at Lumen, about why network performance is becoming one of the defining factors in enterprise AI ROI.

Jim brings experience from both sides of the relationship. He previously led technology programs inside large enterprises and served on Lumen’s board before joining its management team. His description of Lumen as a logistics company for information provides a useful way to understand the network’s role. Models can process information, but the network must deliver the right data to the right location at the required speed.

Our conversation begins with the $2,000 per hour GPU problem. Jim recalls speaking with a Fortune 100 CIO whose company was training its own model using data stored across multiple clouds and regions. The company had GPU availability, but moving petabytes of data to those processors had become its largest bottleneck. Expensive compute was waiting for information.

We discuss why an AI workload can perform well inside a controlled development environment before slowing down when users, data sources, and cloud services become distributed. Jim says many inference workloads require response times between 10 and 20 milliseconds. Factory automation can require latency below 10 milliseconds, while natural voice conversations may need responses within five to 10 milliseconds.

Poor data movement also affects employees. Jim describes spending time with data scientists who began their mornings by moving information from three global regions. The transfers took between two and two and a half hours, leaving skilled employees waiting, attending meetings, or drinking coffee before they could begin modeling. In his words, these were “idle humans” waiting on the network.

The problem becomes harder as AI workloads span public clouds, private data centers, colocation facilities, and edge environments. Jim says bandwidth growth between clouds and data centers is rising far faster than traffic between premises and cloud environments. Enterprises are also managing multiple carriers, tools, and operating systems, making performance and predictability harder to control.

We examine data gravity and the decision to move compute closer to data rather than transporting large data sets over long distances. Security, resilience, sovereignty, compliance, and cost must be considered according to the requirements of each workload.

Jim also explains why networking is beginning to behave like a cloud service. Traditional capacity could take weeks or months to provision. A software controlled network can increase or reduce capacity as demand changes. Lumen describes this model as Cloud 2.0, with automation, APIs, and real time control replacing static capacity planning.

Our discussion closes with ownership. Network, cloud, application, data, and AI teams may each measure success differently. Jim believes a business owner should oversee the complete path from data source to AI outcome because a bottleneck can appear anywhere along the way.

Could the biggest constraint on your AI investment be the network connecting your data and compute? Listen to the episode and share your thoughts with me.

Useful Links

[00:00:00] Your agents aren't producing accurate answers because they don't have a complete semantic understanding of your data. And Denodo is solving this and solving it through semantic consistency. Through semantic consistency, your agents can start making accurate predictions in real time. So see what else Denodo can do by visiting denodo.com to learn more. But now let me introduce you to today's guest.

[00:00:30] What happens when a company spends thousands of dollars per hour on GPUs only to leave them waiting for data? It is the infrastructure equivalent of hiring a Formula One car and discovering that the fuel is arriving by bicycle. So in today's episode of IT Infrastructure as a Conversation, I'm joined by Jim Fowler. He's Chief Technology and Product Officer at a company called Lumen.

[00:00:57] And Jim is someone that's led technology inside major enterprises, but now finds himself working on the networks that they depend on. So today we're going to discuss the 2,000 per hour GPU problem. Why AI pilots can appear fast in a lab and slow down in production. And how latency has become a business metric now.

[00:01:21] And Jim will also explain where data movement breaks across clouds, data centers and edge environments. And why idle infrastructure can create idle employees. And how programmable networks could better demand, could better respond to demand in real time.

[00:01:40] So if your AI strategy begins and ends with compute, I'm hoping today's conversation might reveal the expensive piece that is hiding between the data and the decision. And with that scene set, let me officially introduce you to Jim right now. So thank you for joining me on the podcast today. Can you tell everyone listening a little about who you are and what you do?

[00:02:09] You know, my name is Jim Fowler. I am the Chief Technology and Product Officer for Lumen Technology. I've spent the past 30 years, the majority of my career as a CAO and a CTO, helping large enterprises really navigate technical transformation, digital transformation. And for the last two years, I was on the board of directors for Lumen.

[00:02:31] And I made an interesting but maybe an odd move of stepping off the board and coming into management this year, taking on tech and product. So what I think is unique about me is I used to be a customer. I've been through the other side of this, and I really understand the transformation that customers are going through. You know, for those that don't know Lumen, I really kind of describe the company as a logistics company.

[00:02:58] So just as UPS or FedEx move packages, Lumen moves information. They move packets increasingly for AI applications that are really dependent on real time access to data and information. One of the reasons I was excited to get you on the podcast today is I think you've got somewhat of a unique vantage point here. And I love digging into your origin story a little bit and quickly learn that you've, yes, led technology from inside major enterprises

[00:03:27] and now provide the infrastructure that those companies actually depend on. But I've got to ask, as someone that's been on both sides here, what did you maybe initially misunderstand about networks as a customer that became much clearer when you moved on the provider side? You know, about 25 years ago, I got my first CTO job and I had about 40,000 people that worked around the world, about 130 locations. And I really thought the network was complicated. I viewed it as plumbing.

[00:03:56] I think what's changed is what I viewed as hard there really wasn't. Because today, what I now see is the network is really a nervous system. In an AI era, data is more important, frankly, than the raw compute power. Sensing, acting, making decisions based on data in real time with AI. And with users that are physically located everywhere,

[00:04:20] AI has really turned network latency into a business metric, not just a technical one. So I think what really became clear is that enterprises aren't really struggling with compute storages as much as they're struggling with the ability to get the data efficiently where they need it to make the decisions at the speed of business. And I was reading before you came on that you've described a $2,000 an hour GPU problem, which really stood out to me.

[00:04:50] Essentially because expensive processors can sit underused, because data cannot reach them efficiently. And I think this is something we don't talk about this enough. And anyone working in networking will be applauding at that as well, I would imagine. But how common is this? And how can leaders tell whether network performance is actually reducing the return on their AI investment? You know, one of the first conversations with the CIO that I had after taking this job

[00:05:18] was a Fortune 100 CIO that was working on training their own model. They had a ton of data. They were kind of ingesting that data from multiple clouds that sat in different places around the world. And I got to a conversation. I said, well, what are you struggling with? And the answer I got back from her was, believe it or not, it's utilizing my GPUs. And she described the problem. She's like, we have petabytes upon petabytes of data, but they sit in multiple clouds in multiple regions around the world.

[00:05:45] And getting the data to the GPUs where I've got availability is my biggest bottleneck. And so it's really not surprising at this point that people are thinking about, how do I get data to utilize GPUs, one of the most expensive assets they have, in a very efficient way. And so that kind of stuck with me, that that's a problem we've got to work on solving. It's not just a latency issue, but it's also a bandwidth issue.

[00:06:12] We've got to be able to get the data where you need it as fast as you possibly can. So when an AI pilot works well, but struggles in production, what symptoms should people listening be looking out for that could suggest that the network is the constraint rather than the model, the compute capacity or application design? It's a great question, Neil. This has been kind of a fascinating one for me to watch. Kind of three things kind of come to mind. The first is latency.

[00:06:41] AI applications have to feel responsive. And many times they'll feel responsive in the lab, but slow down once users, clouds and data sources become distributed. So you've got this really controlled development environment. But once you're actually in real life, it doesn't work as well. I'll just give you an example like inference. We talk about the ability to take data and use it for inference in a large language model. You've got to be within 10 to 20 milliseconds of response time for that to actually work at the speed of business.

[00:07:11] Factory automation, under 10 milliseconds of latency for a factory to be able to continue to work the way it would expect. Tons of companies doing investment in chatbots right now where they want the feel of the human experience in an AI chatbot or from a AI driven representative. At best, you've got to be under 100 milliseconds. And if it's a real time conversation like you and I are having, it's got to be in that 5 to 10 milliseconds.

[00:07:40] So I'd say latency is the first thing that I think a lot of companies don't really take into account when they move from their development environments into production. The second we kind of talked about, which is underutilized infrastructure. You know, we have a lot of customers who say I need to move data in big chunks. But once that data is moved, I don't need the pipe that I put in place. I don't need all of the capacity you've given me. And so I think the second thing that people don't recognize is all of that infrastructure they're putting in place. They've got to pay for it when they're using it and when they're not.

[00:08:09] And so how do you get an infrastructure that you can size in real time to your needs to make sure that you don't have GPUs or network capacity that are sitting idle because those are expensive assets? The last one, and this is probably the most important one, is you see idle humans. I'm not of the camp that AI is going to replace the humans we have. I think this is a humans plus machine world. That's my view. But I'll give you an example from kind of my last job.

[00:08:38] I was talking to a team of data scientists who work with kind of large sums of data. And I said, well, tell me about your data. I was kind of doing a ride along, if you will, with them for their data, see kind of how they do their job. And they said, well, we come in in the morning and the first thing we do is we kind of set up our data moves. I said, OK, tell me about that. Well, we've got data that sits in three regions around the world. We need to pull them together to be able to run our modeling on to be able to kind of work on the next modeling that we're building. So, well, then what do you do?

[00:09:06] So, well, then we usually go either have meetings or go grab some coffee or like, wait a minute, explain that. Why the downtime? And they said, well, the network pipes and latency. So the network pipes are only sized to a certain bandwidth. And so it takes about two to two and a half hours for us to move the data. That's idle humans. That's humans that are not being enabled by machines. And so production AI is really a network system.

[00:09:30] And if the model is the brain and the network is the nervous system, connecting all of that information to humans is really what you're trying to do. So I think that's one of the last things people see is if they don't do this right, it's really not enabling the humans plus machine discussion. And it's worth highlighting here that AI workloads can span public clouds, private data centers, co-location facilities and edge environments. So I'm curious, where do the biggest data movement problems appear?

[00:10:00] And what architecture mistakes repeatedly seem to create what we're talking about here? You know, the biggest challenge right now is showing up between environments. And you can see it in the network data. Premise to cloud, premise to data center, network bandwidth growth is only growing at about 1% a year. When you look at the growth of the bandwidth and the pipes kind of data center to data center, cloud to data center, cloud to cloud,

[00:10:27] it's growing someplace between 10% to 13% a year. Moving data efficiently across multiple clouds is becoming harder and harder. And so we see complexity explode when enterprises are trying to connect multiple clouds together, multiple carriers. They've got multiple data centers using separate tools and operating systems.

[00:10:47] And really trying to figure out the predictability and the performance as those workloads become more mission critical is becoming more and more difficult for CIOs. It's the number one problem that we see. Our strategy has been threefold. One, we've been building out kind of large bandwidth, high capacity, low latency networks. But the second layer of our strategy has really been around putting a software layer in that really helps with that.

[00:11:13] And so Alkira is a company that we recently finished the acquisition of. And what it does is it provides a unified control plane across those distributed environments. What they were doing a great job of understanding is that they could actually do that cloud to cloud connectivity without having to wait to have physical pipes put in place, that there were other ways that were within those cloud providers to be able to get that work done.

[00:11:38] And so you just highlighted what is, I'd say, the number one problem of CIOs and CTOs today when it comes to their infrastructure is the sprawl of their data, their networks, their applications that's happening and how they bring all of that together to be able to run their businesses efficiently. And, of course, moving data faster can dramatically improve performance.

[00:12:02] But on the flip side of that, it also could raise questions around security, sovereignty, resilience, energy use, and, of course, cost. So how should infrastructure leaders listening be weighing some of these tradeoffs when deciding where data and compute should reside? It's the number one problem of architects right now. The enterprise architects are really struggling with this. At the starting point, it's got to be the workload.

[00:12:25] You know, different AI applications have different requirements for latency, security, and data residency. You know, starting with that workload and what you're doing with the data should be kind of job one. Then understanding that data gravity matters. You know, what we found is sometimes it's more effective to move the compute to the data than to move massive data sets across long distances. So I think sometimes you think about it from a data gravity perspective.

[00:12:54] But then, to your point, security and resiliency can't be afterthoughts. You know, with the workloads and the agentic architectures really taking off, having private networks in place, having predictable connectivity rather than best effort networking, those are going to become a bigger part of how we architect futures going forward. We're having a lot of conversations with our customers about creating the equivalent of the Internet, but for various sectors. Right. We've done a lot of this.

[00:13:24] For example, we have a product called Vivix. We've done it for the live streaming world. We've kind of created this private network that they can use to get the optimal workloads where they need around the world and be able to guarantee the bandwidth and the connectivity and the speed and the resiliency. You're going to see that happen in multiple industries. So the goal really isn't to optimize any one of those variables.

[00:13:46] It's really to kind of balance the performance, the economics that come along with the cost of the resiliency and infrastructure, the security and the compliance. But you root it all back in what's that first workload? What are its requirements? And that usually drives the rest of the answer. Yeah. And I think it also increasingly feels that networking is almost beginning to behave more like a cloud service where capacity can be provisioned and adjusted and tweaked on demand. But what does that mean in practical terms?

[00:14:14] And what must companies change in their operating models to use it more effectively? You know, it goes back to my first comment on the network was plumbing. You know, when we did cloud 1.0, I think all of us really thought about the network as being static. You ordered capacity, you waited weeks, sometimes months, and hopefully the future demand matched what you bought. And that's kind of how you thought about networking. Going forward, it's software defined.

[00:14:42] Capacity can really be provisioned, scaled and adjusted when the workload requires it. We have customers that are literally spinning up capacity and down capacity throughout the day based on their needs. We call this cloud 2.0. We think it's the next definition of how compute gets, how will we do with compute in the cloud? We do a storage of the cloud matches on the network side. Compute became programmable through the cloud. And now the network has to become more programmable as well.

[00:15:10] You know, at this point, AI is really moving too fast for a 30-day provisioning cycle. And so infrastructure needs models that are really built around automation, around APIs and real-time control of their infrastructure. The network needs to be a part of that conversation. Yeah, 100% with you. And one of the belts and braces approach to IT all over the world is that mantra of you can only improve what you measure.

[00:15:36] And, of course, network teams, cloud teams, data teams, and AI teams have so much in common, but they all measure success quite differently. So who should own end-to-end AI performance? And what shared metrics could maybe prevent some of the expensive infrastructure from being optimized in isolation almost? You know, it's a team sport now, right? No single function can optimize it. We're finding it when we go into conversations with our customers.

[00:16:02] It's usually somebody from the cloud team, the infrastructure team, the networking team, the application team, the business owner, right? It's really thinking about business performance that are coming together, and they're defining the metrics of performance. So the most successful organizations really align there around the business outcome, not the technical silos, not the technical measures, but really how quickly data becomes insight. And that might be defined in revenue. It might be defined in a cycle time metric.

[00:16:30] It might define a quality metric from a customer perspective. Sometimes we're seeing latency and data movement efficiency there, GPU utilization showing up there. Those are, though, all leading indicators of that end-to-end business performance. I think ultimately someone has to own the entire workflow from a data source to AI outcome because the bottleneck can be anywhere in the middle. And what I'm finding is more and more that should be the business leader, not necessarily the technical stack owner.

[00:17:00] They can give leading indicators, but it really is going to matter what that business process or business product owner says. And for somebody listening in a company that is preparing to move their AI workloads into production today, what should they be assessing before or should they be assessing first across connectivity, data location, latency, observability and resilience? And where is spending most often wasted?

[00:17:27] You've probably picked up more than a few war stories in your career, but what are you seeing here? Yeah, and we're going through this even inside Lumen today. And it starts with where's your data live and where's your compute live? Because odds are they're not in the same place anymore. And we see that problem getting worse, not better, right? The sprawl of SaaS providers and cloud providers and data providers is only going to get bigger and bigger. So one, just really being able to map that out.

[00:17:52] I think second, what are the latency and connectivity between the clouds, the data centers and the AI infrastructure, and how does that match up with the problems that you're trying to solve? So I guess the stats earlier on, the latency requirements for some business processes and capabilities. I think just making sure that you're rooted in understanding what those are and can your infrastructure handle that today?

[00:18:16] Third, I think you've got to assess whether your network is programmable and scalable enough to be able to respond to the changing workloads. You should demand from your networking providers the same flexibility and capability you've gotten from your cloud providers and really making sure that it is programmable enough to scale to be able to get the data where you need it, when you need it, with the latency that you need.

[00:18:37] And then the most common waste I see is that organizations overspending on infrastructure, on GPUs, on network capacity. And I think the network often drives some of that inefficiency, really understanding how that happens. You know, maybe what I leave you with is cloud 2.0 really requires the network to work like the cloud. It has to be programmable. It has to be on demand. And it has to be software controlled.

[00:19:07] It's why we did this acquisition. Now, Cura really helped us accelerate that vision of bringing software intelligence and physical infrastructure together in one single platform for AI. And I think that's what customers should expect going forward. And we've talked a lot today around that mindset change that's needed, the need to rethink how you design and architect and the infrastructure as well. So many big changes and different ways of working.

[00:19:34] But what excites you about the future if we get this stuff right and where we're heading? What excites you about that? What makes you want to jump out of bed in the morning? You know, I love that question. So much gets written right now about the pessimistic side of what's happening in the AI world. And I think right now to be a technology leader, you have to be an optimist. And I'm optimistic because there's no shortage of unsolved problems in the world. Everybody's worried about the fact that job A or job B might be automated.

[00:20:04] I'm excited because that frees up the capacity for the people that were in those jobs to go work on the unsolved problems. Think about all of the things in the world that we haven't solved yet, either from a healthcare perspective or from a monetary perspective or from a health and human services perspective. For me, all we're doing is creating capacity to go work on those problems. And that's what gets me out of bed every day. Love it. And I think that is a powerful moment to end on.

[00:20:30] And for anybody listening wanting to find out more information about you, Lumen, the work that you're doing here, or maybe carry on the conversation with you or your team. Where would you like me to point everyone listening? The best place to look is Lumen.com. It's where we share all of our kind of latest product news and information. And I'm always happy to continue the conversation. Listeners can connect with me on LinkedIn by searching Jim Fowler at Lumen. Awesome. Well, we covered a lot today from why AI performance is becoming as much a data movement problem as a compute problem.

[00:20:59] And also how networking itself is evolving from something you just provision to something that behaves so much more like a cloud service. Incredibly powerful point there. So I will include all the links to everything you mentioned there. I encourage people to check that out. But more than anything, thank you for bringing this to life today in a language everyone can understand. Really appreciate your time. Thanks for having me. Appreciate it. I loved Jim's nervous system analogy.

[00:21:25] I think it's worth carrying into all of your AI infrastructure meetings. Because a powerful model and expensive compute cannot deliver much when the data arrives late, travels through unpredictable connections or remains scattered across several clouds. And the practical starting point here is to map where your data lives, where your compute runs and how quickly information can move between them.

[00:21:53] And measuring latency against the needs of the workload and then assess whether capacity can change on demand and give one business owner responsibility for that complete path from source data to outcome feels like an incredibly sensible approach. Because without that, every technical team can hit its own target while the business sits around waiting for an answer.

[00:22:18] But of course, this wasn't a doom and gloom conversation filled with practical takeaways and particularly Jim's optimism about people and AI. Better infrastructure can release skilled employees from two-hour coffee breaks caused by data transfers. Although coffee itself should remain safely outside any automation strategy. But seriously, a massive thank you to Jim for joining me today.

[00:22:45] But over to you, where is your most expensive AI resource currently waiting for data? Love to hear your thoughts on this one. Techtalksnetwork.com. That's where you'll find all 4,000 plus interviews, ways you can leave me a message, work with me, meet me at a tech event near you. And I am, and there is a lot on there. So hopefully I can meet up with a few of you. But that is it for today. Thanks for listening as always. Bye for now.

Lumen,