What becomes possible when enterprise computer vision no longer depends on expensive GPU infrastructure?
In this episode of Tech Talks Daily, I speak with Glenn Jocher, founder and CEO of Ultralytics, about YOLO26, CPU inference, edge AI, open vocabulary vision, deployment economics, and the practical work required to move computer vision from a promising pilot into production.

Glenn's route into AI began inside the U.S. intelligence community. He worked with the National Geospatial Intelligence Agency and Defense Intelligence Agency on particle physics applications, attempting to detect and track antineutrinos.
Antineutrinos are extraordinarily difficult to detect because they pass through almost everything. Glenn describes them as the perfect spy. While searching for better detection methods, he discovered that computer vision researchers were solving similar problems with images.
His original attempt to transfer those techniques into particle physics did not succeed. However, the work introduced him to a field where the technology could create a visible effect on everyday life. That led him toward open source development and eventually the YOLO models for object detection, classification, segmentation, and tracking.
Glenn believes computer vision research has historically placed too much attention on small gains in accuracy while overlooking deployment economics. A model can perform impressively inside a laboratory and still remain unsuitable for a factory, warehouse, store, vehicle, drone, or medical environment.
Price, latency, power consumption, data privacy, and deployment speed can determine whether the technology is commercially useful. This led Glenn and Ultralytics toward smaller models capable of running close to where images and video are generated.
YOLO26 continues that approach with architectural changes designed specifically for CPU inference. Glenn says the model can process camera streams in real time at 30 frames per second and run across Intel CPUs, AMD CPUs, and lower power devices such as Raspberry Pi computers.
This matters because specialist GPUs can increase the equipment cost and power requirements of a computer vision project. Running inference on existing CPUs or edge hardware can make deployment economically possible across larger numbers of cameras and locations.
The scale already involved is difficult to comprehend. Glenn says Ultralytics models now process approximately three billion inference jobs each day, equivalent to around 30,000 every second. These jobs include images, videos, and collections of images being analyzed to detect, segment, or track objects.
He attributes the platform's maturity to thousands of mistakes and bugs corrected through a rapid feedback cycle. New models are released, users report problems and request features, and the team incorporates that information into later versions.
We also discuss the respective roles of cloud and edge infrastructure. Glenn sees cloud platforms continuing to provide the computing power required for training, while computer vision inference often belongs at the edge. Local processing can reduce latency, control operating costs, and keep sensitive video or medical information closer to where it was created.
The smallest YOLO model is approximately three megabytes, according to Glenn. That allows it to reach mobile phones, vehicles, drones, battery powered devices, and other environments where a large language model would be impractical.
Open vocabulary vision provides another development. Traditional object detection models are trained to recognize a fixed collection of objects. If a model learns to detect dogs and the user later wants it to detect cats, retraining can cause it to forget earlier knowledge unless both categories appear in the new training data.
Glenn explains how promptable models can identify common everyday objects from text or visual instructions without additional training. A user could request a person wearing a blue shirt and white shoes, for example, and the system could search an image for that description.
That flexibility could benefit businesses whose requirements change regularly. It reduces the need to create and label a new data set every time the company wants the model to recognize another common object.
The range of current applications is already extensive. Glenn describes YOLO being used across robotics, parking, industrial safety, PPE detection, warehouses, aviation, security, traffic management, food quality, and manufacturing.
Some of his favorite examples involve environmental problems. One company uses YOLO with underwater vehicles to identify and recover plastic from the ocean. Other applications detect smoke and fire early enough to support forest fire response.
For leaders considering computer vision, Glenn recommends beginning with a defined problem and measurable outcome. A manufacturing company may want to reduce defects, but it still needs labeled examples showing the model what acceptable and defective products look like.
He advises testing the idea through a limited pilot, measuring the return, and expanding only when the evidence supports further investment. Computer vision has become easier to deploy, but practical problems involving data, cameras, integration, reliability, and operating conditions still separate a demonstration from a production system.
Could CPU inference and open vocabulary models make computer vision practical for processes your organization previously considered too expensive? Listen to the episode and share your thoughts with me.
Useful Links

[00:00:03] What if one of the biggest barriers to enterprise AI isn't model accuracy, but the cost of running it? Well, computer vision can inspect factories, detect defects, improve safety, manage traffic, and even help remove plastic from our oceans. But impressive technology means very little if the business can't afford to deploy it, and deploy it at scale.
[00:00:29] Well, my guest today is Glenn Jocker, and he's the founder and CEO of Ultralytics. And today, he's going to sit down and explain why running computer vision on everyday CPUs could be about to change the very economics of AI. I will also explore how YOLO models are processing literally billions of inferences every single day,
[00:00:55] and also examine why successful AI projects should always begin with a real problem, rather than just a shiny new model. And it's on that note that it's time for me to introduce you to Glenn right now. So, thank you for joining me on the podcast today. Can you tell everyone listening a little about who you are and what you do? Yeah, hey guys, I'm Glenn Jocker. I'm the founder and CEO at Ultralytics, a computer vision startup.
[00:01:25] Super excited to be here on the podcast today, Neil. Me too. I'm glad you took the time out to sit down with me. And I also, rather than just talking about where we are now and where you're going to, I always love looking at my guest's origin story. And when I was looking at yours before you joined me today, you've got a background that spans geospatial intelligence, scientific research, and now one of the world's most widely used computer vision frameworks. Incredibly cool. But what was that journey that led to you founding Ultralytics?
[00:01:54] And what problems were you determined to solve from the outset? There's got to be a story there, right? Yeah, yeah, absolutely. The most interesting stories always have a winding path. I got my start in the intelligence community. So I was doing work for two agencies around Washington, D.C., which were the National Geospatial Intelligence Agency and the Defense Intelligence Agency. And I initially did particle physics applications for them. It's really interesting stuff.
[00:02:22] And my role in all of it was to try and detect and track anti-neutrinos, which are these really exotic, really interesting particles that we just don't know a lot about. And the team that I was in were trying to study these particles and understand more about them, how they act, and how they kind of help inform our understanding of the universe. They're one of the fundamental particles, which means that everything is built from these. But they behave very differently than a normal particle. They have flavors.
[00:02:51] They oscillate. They do strange things that we don't fully understand. And so I was trying to track these particles and detect them, which is a very difficult job because they're quite elusive. They travel through everything. They're nearly impossible to detect. They're like the perfect spy, sort of. And so anytime you catch one is incredibly valuable and you want to try and get as much information from it as you can.
[00:03:14] And I was using these reconstruction techniques, these classical techniques for identifying and tracking particles. And I wanted to do better. And so I looked around me. This is about five years ago, around 2020, to see what kind of algorithms I could see, what kind of academic papers I could read. And I realized that computer vision was starting to solve some very similar problems. Like in a picture, you could start to see, hey, this is a cat. This is a dog. Maybe we're the cats walking in a video.
[00:03:43] And I thought, wow, this is great. Like this is a cat. It's not a particle. But other than that, everything is kind of similar. And maybe I can translate some of the success over to the physics space. And that was a rough trip. I wasn't able to do that. But in the process, every time I'd go to the AI space, the grass appeared to be greener. There was more interesting things going on. There was more hype. There was more news articles. More investment, it seemed. And important for me, more everyday impact on people's lives. So anti-neutrinos are very fascinating.
[00:04:13] But at the end of the day, it just doesn't really impact us in our everyday lives. It's not something you can put in your hands to make your life better, make your day better, something like that. So that's what really attracted me to AI in the first place. I pivoted fully, then started doing open source, which I'd never done before. And all of that really picked up a lot of traction and became the YOLO models.
[00:04:35] Now that Ultralytics is most famous for, which are now the world's most popular models for detecting and tracking all kinds of objects. Particles, cats, dogs, really anything. Wow. You had me on the edge of my seat there. I didn't know where you was heading with this, what you and the government and the agencies were going to find out, what you could share with me, what you couldn't. So, well, great story. I mean, for years, the conversation around computer vision has focused on model accuracy.
[00:05:05] But I was reading how you've argued that infrastructure costs are often the real barrier to enterprise adoption of it. So why has deployment economics become that deciding factor for moving from pilot projects to production? What are you seeing there? I think if I look back at the initial work I did, I was surrounded by academics and scientists and professors. And the emphasis was always on the research. It was publishing.
[00:05:34] It was getting slightly better detection accuracies. And I was the only person in the group that really had a startup mindset. I wanted to take the tech out of the laboratory, take it into the real world, and really be able to demonstrate that impact. And I wasn't able to do that there. And I think that's my key contribution to the field initially was worrying about the real world applications, trying to connect the dots of the research to the real world. How does this get used?
[00:06:02] And there, one of the main concerns, obviously, is price. It's speed to deployment. It's things like data privacy. And I realized that a smaller, faster model that people can deploy at the edge is always going to be preferable when you're running computer vision applications in real time or at scale or dealing with private information like medical records. And so I started focusing more and more on smaller and smaller models, which is the opposite of what you see in the news these days.
[00:06:31] So these days, every model that comes out is bigger, more parameters, more resources, more GPUs, and costs more to train. And the companies that do that research go back to the investors for more billions of dollars of investment to get those GPU farms to train those models. But the inference side is different than the training side. And that's a place where if we can run models on low-power devices like your average CPU, then suddenly that opens up a lot more doors.
[00:06:59] Then you don't need a big, expensive NVIDIA GPU. You don't need a huge power budget. You're not going to be, I think, shocked by a huge bill of materials or a high running cost. And if you can do that at scale at a very economical price, then suddenly this really becomes a viable tool for many businesses, including enterprise at scale that are running millions of images per day. So it's great to see that progress.
[00:07:27] And for me, this is where things look incredibly exciting because YOLO 26 introduced the ability to run high-performance computer vision on standard CPUs rather than specialized GPU infrastructure. So tell me a little bit more about the technical advances that has made this possible and why it changes the equation for enterprise AI too. Yeah, absolutely. So YOLO 26 is our latest model that we launched this year.
[00:07:54] Every year, we're aiming to release a new generation of YOLO models for object detection, classification, segmentation. With YOLO 26, for the first time, we introduced many architectural improvements, new modules that were specifically tailored towards CPU inference. And this really aligns with our values of democratizing the technology, enabling everybody to use the technology, even at lower price points. So we don't want AI to be an expensive sport.
[00:08:23] We don't want it to be something that requires a lot of expertise, a lot of budget, and a lot of technical know-how to really run these models. We want them to be as off-the-shelf as possible. So you plug in the YOLO model, you hook it up to a camera stream, and it just starts detecting objects. And it can do it in real-time, 30 frames per second. And now you can build products and services around this type of integration that really leverages real-time AI.
[00:08:49] So again, completely different than what you typically see with a model like ChatGPT, where it's too big to run at the edge. Not only that, it's proprietary. So even if you could run it at the edge and you had a huge, huge computer, it's not available since it's not open source. YOLO models are open source. People can download them, deploy them almost everywhere. So these are Intel CPUs, AMD CPUs, and even lower-powered devices, even like Raspberry Pis, for example, that you can buy for $50.
[00:09:18] So the buried entry has never been lower if you want to get into computer vision and you want to start using AI. And your models now process something like, I think, 2.5 billion inferences every single day and across so many different industries, from manufacturing to healthcare, logistics and retail. And I've got to ask, what have these production deployments maybe taught you about that gap between impressive AI models
[00:09:46] and operating one reliably at enterprise scale? It feels like you're right in the eye of the storm here. Yeah, the numbers are just crazy and they're difficult to even put into context. So about 3 billion inference jobs per day turns out to about 30,000 inference jobs per second. And so these are images or videos or collections of images that people are passing to YOLO to understand what's inside them, to detect the objects in them, segment them, track them.
[00:10:16] And the adoption here is just incredible. It always gives me a warm, fuzzy feeling inside to see this. This is really what I've always wanted since the early days when I was in the laboratory and nobody was using what we were building. Only the government was interested in it. So now this is real, everyday technology. And to see the scale of the adoption out there and the impact that it creates in all sorts of domains, from manufacturing, healthcare, logistics, and retail is truly great.
[00:10:47] The maturity, of course, that we need to bring to get to this level is something that's taken a long time. So in 2026, we can say that we're delivering this level of capacity. In the early days, it was completely different. When I started AI, I was a newcomer. I'd never done open source before. I'd never even used Python before. And so the current capacity that we have and the maturity of the product directly relates to thousands of mistakes and bugs that we fixed since those early days.
[00:11:15] And we've done that primarily by listening to the users. So every time we would ship a new model, we would get feedback and we would rush to incorporate that as fast as possible into a new release, let everybody know that we'd updated everything, fixed their problems, incorporated their features, and then sit back and wait for more feedback. And it's this tight iterative loop. It's listening to the end user and acting as quickly as possible to incorporate their feedback into the product that I think is really driving this improvement.
[00:11:44] As Elon Musk always said, the speed of iteration is typically the limiting factor in anything. So the faster you can listen to your clients and the faster you can respond to them and get something out, then the better you're going to be. Maybe not in the first few months, but eventually in the long term, you're going to have something completely different than what you started with. And over the last two years, I think Edge AI has become increasingly attractive for organizations that are looking to reduce latency,
[00:12:11] improve their resilience, and keep sensitive data close to where it's generated. And I say that especially in Europe. But which workloads do you still see as best suited to Edge deployment? And where do you still see cloud infrastructure playing an important role too? Is it somewhat of a balancing act there? There isn't a one-size-fits-all approach? Yeah, I think in our age, it's a bit of a tug of war. And we always see, for example, perhaps cloud leading in the early 2010s,
[00:12:38] and then a response drive back to the Edge to bring down costs, to bring down governance concerns, latency, and things like that. I would say these days there's a clear delineation of AI being trained in the cloud and for computer vision applications being deployed at the Edge. The Edge sounds like a trendy word, but it's not that complicated. It could just be something as simple as your iPhone. It could be the Tesla that you're driving or maybe the DJI drone that you're flying around
[00:13:07] that is not crashing into things because of the computer vision models inside it. Depending on which company you talk to, I think they'll drive you more to cloud or more to Edge, of course. Like the bigger companies have invested tremendous amounts, and we see in the news even more than ever, more recently, into very large data centers in the U.S. for training ever and ever larger models. So there's definitely a surge in investment in data centers for training
[00:13:36] and also for large models for that inference. But when you're in a computer vision space, you can definitely take that model. A YOLO model is only the smallest one we have is about three megabytes. So basically the size of a JPEG that you might take. And you can send that almost anywhere these days, even battery-powered applications that run all day long. AI agents are only as strong as the data that they're given. When provided with an outdated data set, your agents could end up doing more harm than good.
[00:14:05] But not with Denodo. With an AI data layer built within your platform, your agents are provided with real-time data changes. So with Denodo, your agents can finally make the right business decisions. Simply visit denodo.com to learn more. And I think one of the most interesting capabilities in YOLO 26 is open vocabulary vision, which allows systems to respond to text and visual prompts without retraining.
[00:14:35] So how does that make computer vision more practical for maybe somebody listening in an organization where their business requirements are continuously changing and evolving? Yeah, this is a really good question. In the real world, when we deploy YOLO model, it's trained for specific application. In my dog-cat application, maybe it's trained to detect dogs. And maybe you deploy it, and it's working really well, and maybe you adopt a cat. And you want to know if your cat's walking around too. What do you do?
[00:15:04] And so these days, you have to train the model again on a new data set that contains cats and dogs. It's always been like this. And the reason you need to do this is because of something called catastrophic forgetting. So just like the way a person is, if you don't use something for a while, you'll forget it. And so if you train the old model on a new data set of just cats, it'll forget the dogs. If you train it on dogs, it'll forget the cats. And so you really have to train it on everything at the same time,
[00:15:30] which becomes burdensome once you're trying to expand computer vision usage within your organization. So this year, for the first time, we started working on what are called open vocabulary models or promptable models. And so YOLO has always been a traditional object detector that you train on certain objects and then deploy it, and then it detects those objects. But now we have a new version that doesn't need any training. It's been pre-trained on a massive collection
[00:15:57] of images and videos and text that accompany those to be able to identify the objects automatically, at least for common everyday objects, like bottles and cups and cars and people and so on. And so with these models, you just tell it what you want to find. If you're looking for a person in a blue t-shirt and white shoes, then it'll find that person in the picture. And this opens up a lot of new doors. And so this is much more flexible. It means that a lot of the training process
[00:16:27] that people had to do typically to deploy computer vision no longer applies if you have a common everyday use case. So this is a very innovative and new approach and I think is trending a little bit in a hybrid direction and towards what you typically see when you use ChatGPT. You tell it what you want, it does that. And now we're offering a similar capacity for computer vision at the edge with these incredibly small models.
[00:16:54] And of course, every tech project now will have to undergo several checks around ROI, measurable value that it can offer, etc. And just to bring that to life a little bit here, if we look across the organisations that you've worked with, which computer vision use cases have you seen delivering the fastest return on investment today? I don't have to mention any names, but are there any applications that you think businesses are maybe overlooking despite having the potential to solve real operational challenges?
[00:17:24] What are you seeing out there? Yeah, yeah. The interesting thing about Ultralytics is we create the models, we create the platform to train and deploy, but we don't actually create the end use cases. We hear the use cases coming back from the users and from the clients. And it's always amazing to hear the success stories. We've got so many companies now using YOLO in so many domains. When we were initially in the early days, we got a lot of advice that we should specialize in the domain. We should become very good at one thing,
[00:17:52] maybe recycling, maybe port management, maybe manufacturing. And I didn't want to do that because I saw the general potential of the technology. I thought if my eyes can do everything that I need to do, then these YOLO models should also be able to do everything that we need to do to see and detect things. And so now we have companies from robotics to parking, smart applications, industrial safety, PPE detection, forklifts, in food and quality, restaurant industries.
[00:18:21] We have warehouse automations. We have aviation applications, security, defense, smart traffic applications. There's just so many. Some of my favorite ones, though, are really kind of the feel-good applications. Like we have a company that is detecting plastic in the ocean and actually recovering it with submersibles using YOLO right there underwater. So that was the first underwater computer vision application I'd ever heard. And I really liked it. Forest fire detection.
[00:18:51] So automatic identification of like smoke and fire from centers. It's getting pretty hot. And here in Southern Europe, not in Madrid directly, but in other places in Southern Europe, there's definitely more and more forest fires that are starting from the environmental conditions. And I think anytime that we can take AI, apply it to a problem that humanity actually has, and really move the needle, not ourselves, but by empowering everyone else, by empowering companies and individuals
[00:19:19] to innovate and create these amazing applications. These aren't ideas we come up with. These are just ideas that we empower with the tools that we create. So it really makes me happy to see all this positive impact that we're driving. Yeah, 100%. Incredibly cool stories and examples there of how technology can make a real difference. And I suspect we will have many leaders listening today that are finding themselves inspired by your work and interested in computer vision, but maybe in the back of their head, they're unsure whether they're actually ready yet.
[00:19:49] So I'd love to give those people an actionable takeaway. So if you were advising a business leader like that, what practical steps would you recommend before they begin investing in Vision AI to give themselves that best chance of delivering measurable business value rather than just another promising pilot? Yeah, absolutely. Now that I've been running Ultralytics for a few years, I'm a little older and a little wiser on the business side of things. Didn't start out that way. But obviously, I think the best place that you want to start off is with the problem.
[00:20:18] So I think if you look around your organization and you think of what you want to improve, what metrics really matter to you, and what success looks like. It could be financial, it could be hitting KPIs that matter to the organization. In the computer vision space, data is also important. So if you have, say, a manufacturing facility and you want to detect defects in your manufacturing process, that's a good start. And the financial success and KPIs look pretty clear, but you need to have some label data that you can start to train.
[00:20:47] You need to show the model what good looks like and what bad looks like so it can tell the difference once it's trained and once it's deployed. Once you have that data, then I think a small prototype pilot is the right approach. Not diving in head first, but kind of putting your foot in the water and saying, let's try this on this scoped example for this time. Let's see what kind of return on investment we get here. And if it looks like this tree is bearing fruit, let's go ahead and water that and see what comes from that.
[00:21:15] So the reality for AI in the workplace, as I'm sure you've seen in the news, sometimes is not what you see in the headlines. It's really about pilots that don't meet expectations and real world problems that get in the way from theoretical deployments. It's the classical difference between kind of what you learn in school and then once you're out in the real world applying all those lessons and you realize that there's always a few more curves in the road. So the good news is Ultralytics has scaled
[00:21:43] an amazing solutions engineering team. We create models and we also help you integrate those. And so we've learned a lot of lessons. We have over 500 enterprise clients and we've been with them every step of the way, seeing the pain points, learning the lessons and helping other people not trip over those same problems. So if you are in a computer vision pilot, you're interested in deploying something, definitely come talk to us and we will give you the best advice that we can.
[00:22:11] So many big takeaways from our conversation today and I love learning more about the architectural decisions that ultimately determine whether Vision AI works in real world manufacturing floors, warehouses and retail environments and beyond. And for anyone listening that would like to come and talk with you or learn more about the work that you're doing, where would you like me to point everyone? The best place is ultralytics.com. You can find our models, documentation, enterprise solutions, community resources.
[00:22:40] If you're a developer, GitHub is the place. We've got 130,000 GitHub stars, wildly popular. Our repository is at github.com slash ultralitics. And of course, we're on LinkedIn, Twitter, YouTube, and so on. We've got a great tutorial series on YouTube for how to get started with computer vision too that takes you through all the steps. So if you are an engineer also developing with YOLO, the documentation is the best place. Just docs.ultralytics.com.
[00:23:09] Well, as we said at the beginning there, you're managing 2.5 billion daily inferences while maintaining reliability, security, and support expectations. Those numbers are just phenomenal and hard to comprehend. And what I will do, I will add links to everything you mentioned there. And anybody listening, if they go over to techtalksnetwork.com, there'll be a blog post associated to this episode. I'll also include a link to that YouTube tutorial that you mentioned there. We'll get that embedded in. And so please, everyone listening,
[00:23:37] go check that out and let me know your thoughts. But more than anything, thank you for coming on here and sitting down with me today. I'd really appreciate you. Thank you for giving me the opportunity, Neil. It's been great to be here and tell the Ultralytics story. I think today's conversation was somewhat of a fascinating reminder that some of the most exciting progress in AI really isn't about building bigger models or buying more GPUs. Glenn showed how smaller, faster, and more accessible computer vision models
[00:24:06] can bring AI into factories, warehouses, hospitals, farms, and environmental projects around the world. I loved his advice for business leaders too. Start with the problem. Understand what success looks like. Get the right data. Prove the value with the focus pilot. And only then start scaling. But I'd love to hear your thoughts. What problem in your business or your industry could computer vision help solve?
[00:24:36] Especially if the cost and complexity of deployment were no longer standing in your way. Let me know. TechTalksNetwork.com. As promised, all the links will be on the blog post associated with this episode, including that YouTube video. I'll embed that in there. But more than anything, I want to hear from you. I want to know what you're going through, what excites you, what worries you, and everything in between. So while you're thinking about that, I'm going to walk off into the sunset
[00:25:04] and find another guest for tomorrow. But thank you for listening today. Speak with you tomorrow. Bye for now.

