Preparing Unstructured Data for Enterprise AI With CTERA
Tech Talks DailyAugust 13, 2026
3681
25:0222.91 MB

Preparing Unstructured Data for Enterprise AI With CTERA

Could the real reason enterprise AI projects remain stuck in pilot mode be hidden inside the company's unstructured data?

In this episode of Tech Talks Daily, I welcome back Oded Nagel, CEO of CTERA. We discuss why enterprise AI success depends on the condition, location, permissions, and business value of the data sitting underneath models and agents.

Oded defines AI-ready data as information that is searchable, classified, and permission-aware. Many enterprises have petabytes of files distributed across offices, edge locations, legacy network-attached storage, and cloud platforms. Before introducing AI, leaders need to know what information they possess, where it resides, who can access it, and whether it remains valuable.

The cost implications are significant. Copying every available file into an AI ecosystem can create expensive ingestion and storage bills. It may also reduce answer quality when stale, duplicated, irrelevant, or personal files enter the model's source material.

Oded describes a customer classification project where approximately 80% of the data examined was stale or archival. The company also discovered personal content, including MP3 files, stored alongside enterprise information. Feeding such material into an AI system would consume resources without improving business results.

We discuss Oded's recommendation to bring AI to governed data rather than moving data outside existing controls. Keeping intelligence close to the file system can preserve access permissions, audit logs, snapshots, and recovery mechanisms. Those protections become increasingly important when autonomous agents can read, move, modify, or delete files.

Oded argues that every agent should be identifiable and its activity monitored. Businesses need to know which agent accessed which information, what action it performed, and whether the result can be reversed. Without those controls, a misunderstood instruction or malicious input could cause serious damage.

The conversation also covers CTERA InsightAI, an agentic intelligence layer built into the company's data platform. Oded says it analyzes security activity and file-system metadata, allowing users to ask questions about stale data, file types, access patterns, deleted files, and ransomware impact using natural language.

Rather than working through traditional dashboards and filters, users can question the data and request conclusions or recommended actions. Oded says some customers are piloting InsightAI while others already use it in production.

For leaders measuring enterprise AI ROI, Oded recommends concentrating on storage costs, time savings, and speed to production. AI tools should make complex information easier to understand and reduce the time required to act. A ten-page report generated instantly provides limited value if nobody knows what decision to make from it.

Does your company have enough visibility and control over its unstructured data to support production AI, or would classification uncover years of stale information and unnecessary expense? Listen to the conversation and share your thoughts with me.

Useful Links

[00:00:03] Could the biggest barrier to enterprise AI be sitting right inside of unstructured files, rather than inside the model itself that we keep seeing on our newsfeeds? Because AI models receive keynote slides, while aging file shares usually receive a support ticket and a polite request to go and behave themselves.

[00:00:23] But today, I want to welcome back the CEO of CTERA back onto the podcast and discuss why so many enterprise AI projects remain trapped in pilot mode. I've talked about this so much on this show, but I want to learn more about why. And my guest will explain why AI-ready data, what that means in practice, including searchability, classification, permission, security, and visibility across so many different distributed environments.

[00:00:53] And I also want to touch upon why copying every file into an AI platform can increase your costs and weaken existing controls. But how businesses can give agents access without surrendering governance. So you'll hear a lot about this today, and also how natural language queries can reduce investigations from days to minutes.

[00:01:17] So before your company invests in any other new model or agent, does it know what data deserves to be included? Hopefully we'll find out more right now as I officially reintroduce you to my guest, who is officially friend of the show. So a massive warm welcome back to the show. For anyone listening that's missed our previous conversation, can you tell them a little about who you are and what you do? Definitely. And again, thank you so much for hosting me again.

[00:01:47] It's a real pleasure. So my name is Oden Nagel. I'm Sitera CEO. I've been in the company since day one when we established it in 2008. Our focus is really to help enterprise customers to govern, manage, and secure their unstructured data from edge to cloud.

[00:02:11] We work mainly with governments, federal customers, healthcare providers, manufacturing, any enterprise customers that actually has a distributed environment and highly regulated about is compliance, regulations, and definitely security aware. So that's what we do. We're really focusing on protecting and unifying and activating their unstructured data across the cloud.

[00:02:39] And right now, everyone is talking about AI models and agentic AI agents. And while everyone's focusing that on the news feeds there, I was reading that you're arguing that the real constraint on enterprise AI is actually the data underneath. Tell me more about that. Yeah, I mean, as you said, I think every minute or every day we have different conversation with customers about AI strategy, AI project.

[00:03:10] Everybody is saying our management of asking us to start using and enabling AI, right? And, you know, guess what? I mean, AI is an amazing tool. It's amazing capabilities, and it's the new force of the market. But it requires some, I would say, some foundation forces in order to make sure that it will work as it should be, right?

[00:03:36] And the first thing that we are seeing in AI project is that everybody is running into a pilot phase, but they are not really being successful when they are really running into a production phase. And why is that? It's the foundation of the data. If you want to be able to have a successful project, I would say, first, make sure that your data is organized, it's classified, it's searchable,

[00:04:05] and it's also governed by the same permissions and the same security elements that you have in your file system. We see this on a daily basis. If a customer prepares his foundation correctly before he runs to start enabling this type of amazing AI capabilities, then when we see the success of the project. And that is what we are recommending to our customers, and this is really the Cetera vision is all about.

[00:04:31] And it feels like we've been talking about the problems that surround data silos for 25, 30 years now. And I think AI has finally fast-tracked that conversation. But for people listening, maybe they are having similar problems around data and data silos, etc. What does AI-ready data actually mean? And where do most enterprises discover that they're not ready? You probably have a lot of conversations there. At what moment do they suddenly realize that they have a problem like this?

[00:05:00] Yeah, it's a great question. And AI-ready, it's a term that everybody uses today. But it's really mainly from a Cetera perspective that the data is searchable, classified, and permission-aware. It's very important that when you start an AI project, you first of all make sure that the foundation of these three elements on the data is secure.

[00:05:27] And the second thing is that what we recommend to customers is not really to copy the source of the data to another location and bring it to the AI. The idea is actually to bring the intelligent and the AI into the data itself, right? And we see this is a crucial element for success in a project where you make sure, again, that first of all, you organize your data correctly. You're making it searchable and classified.

[00:05:56] And only that, based on these three elements, you can start planning your AI strategy and your AI projects. And as a company, you work with large regulated organizations as well. So I'm curious, what are you finding separates the ones that are getting the real return on investment from their AI projects, from the ones that are still seemingly stuck in these perpetual pilots and pilot purgatory that we often hear about?

[00:06:25] Again, for me, it's a little bit what I said before. The ones that actually win and make the investment worthwhile is the ones that are actually building the foundation of the data, right? What we see that the data today, and it's crossed the board, by the way. It's not like we're talking to one customer or all this.

[00:06:47] Everybody, the first problem that they are facing is where is my data, I would say, and what do I have in my data, right? Because, you know, they have petabytes of data. It's spread and siloed all over the world, right? In different locations. And by the way, you know, today 80% of the unstructured data is actually generated at the edge, right? And most of these enterprises, they don't really know which data is where, who has access to.

[00:07:16] And actually, you know, is it a stale data? Is it a real valuable data that they should actually use for AI? Because, you know, to feed everything into the AI ecosystem, right? It's going to be very costly, right? I mean, you don't really want to feed all your petabytes of data. Like, you know, we had a project with one of our customers and we ran some data classification capabilities that we have in the product. And we found out that 80% of this data, first of all, it's a stale data. It's an archive data.

[00:07:45] You don't want to feed this specifically into the AI for the first ingestion point, right? The second thing that a lot of his data was actually data that generated from a private perspective, like, you know, MP3 files or things that are not really related to enterprise data. But, you know, users are keeping this data on their legacy NAS solutions, right?

[00:08:08] And so in order really to get the right ROI, which I'm talking about, right ROI is efficiency, right? Cost, of course, as well. And by the way, security as well. Okay. I know that security is not a part of an ROI, but it's definitely important to make sure that you maintain the security and all the guardrails around AI project.

[00:08:30] So the ones that are really making the foundation correctly and not running into enabled AI, first of all, do the preparation in advance. The ones that are really getting the ROI from the projects. And if we have a CIO listening to our conversation today, nodding their head in agreement with all the issues we've raised, if they could fix just one thing before scaling AI, what should it be and why? Let's even give them a bit of advice here. I would say two things. Okay, sorry.

[00:09:01] But I know it's out. But I would say visibility and control over your unstructured data. I would say it again. Visibility and control over your unstructured data because that's the key for them to be successful. And why I'm saying that? Because we see a lot of companies bringing, I would say, and I'll say it probably a few times in this podcast, bringing the data into the AI and they don't bring the AI into the data.

[00:09:30] So they break all the security and they break all the cost model because they replicate the data itself. Right. And for me, that's the foundation. So first of all, classify, make sure that all your access control and all the data layer is organized. And then you'll make sure that you'll see that the project will be successful. And I think one topic that will divide the audience is AI agents.

[00:09:55] On one side of the coin, the opportunities that they present will excite so many people listening. But on the flip side of that coin, there'll be a lot of people with a more security conscious mindset, an IT mindset that will be terrified of some of the possibilities that could come out of this. So how do you let AI agents work across a company's file data without losing control of governance and security and all those things that keep IT teams awake at night? Yeah, that's another great question.

[00:10:24] And we can definitely talk about this topic for another hour, right? Because AI agents is everywhere, right? I believe that the future is just starting, I would say, meaning that everybody's talking about AI agent. That's an amazing tool and that's the future. And I believe that 80% of the daily basis activities and workflows and things like that in the next two to three or five years from now will be running by AI agent.

[00:10:53] But on the other hand, it's very dangerous not to be able to control these AI agents because you can just, you know, running an AI agent on your data and by mistake, right? Someone will put a malicious code inside or by mistake, the AI agent will understand something different from what you actually want to do and he will delete all your data in the file system, right?

[00:11:17] That's something we are talking and discussing with customers and partners really on a daily basis. Okay, so the first thing that we are advising the customers, and as I said before, is not to take the data outside into the AI, is actually bring the AI into the data itself. And what does it mean?

[00:11:43] It means for us making sure that the permissions and all the structure of the file system and the audit logs and all these permissions aware capabilities will remain on the file system, okay? Because that's something that it's already built in in any storage platform that is available today and definitely also in the CTERA solutions. So you need to track which AI agent is accessing. You need to be able to monitor which data is accessing.

[00:12:12] You want to provide the basic functionality of protection like snapshots and things like that. So if something goes wrong into the data, you can revert back into a single instance. So the first rule of thumb for me is bring the AI again into the file system and not moving the data outside the file system, okay? That's the most important thing for me. The second thing is about operational and the data governed, right?

[00:12:40] And how you're making sure that all the operational activities the AI agents are tracked, monitored, and audits, right? That's the second layer that is very, very important. And if you connect the AI agent into your existing file system, which provide this type of capabilities, it gives you the ability to really understand what's going on. So that's for us the foundation. Of course, there are many elements that we're working on

[00:13:06] to add more capabilities and security from, let's say, AI agent identification and authentication and a lot of other things that we're working to add into the file system. But that, for me, at least the basic step for a successful AI agent project. And before you rejoin me on the podcast today, I was doing a little research on what you've been up to since we last spoke, and you have been incredibly busy. So tell me more about the recent launch of Insight AI

[00:13:35] and how that addresses much of what we're talking about here today too. Yeah, the team is, they're working behind the club because the AI is really changing the foundation also from an engineering perspective, right? And things that you can do, like, you could do, like, a year ago, you can definitely progress and move on much faster. So the first thing that we did, we launched a new platform that we call Cetera Insight AI,

[00:14:03] which is basically an agentic intelligence layer that builds directly into the data platform, into our data fabric platform that provides the capability to the end user to monitor two things or to get data or metadata into two things, file system activities, which is more on the permissions level, on the security level, for example, which data was accessed, deleted, moved, all these security elements that you can,

[00:14:33] you want to analyze and understand on a daily basis. If you have a ransomware attack, right? I mean, actually, when it started, who actually penetrated the data? Who accessed the data? Who deleted the data? So the first data stream that we manage is the audit logs and the security element of the platform. The second thing that the Insight AI is doing is also providing the file system, the file system stream, meaning that we can identify stale data.

[00:15:02] We can identify and do data classification on a metadata level in very scalable environment, and it's a fully agentic platform, meaning that we collect all the data into a big data lake, and you don't see any more dashboards and filters, you know, and things that it's the five years ago. You talk to your data. You have a built-in agentic AI agent running into the platform, and you can start asking questions about your data,

[00:15:32] and the AI is not just going to provide you the raw data, but it's also giving you the conclusion. It's giving you some advices. It's giving you some directions, and you can talk to it because the data is changing on a daily basis. The audit logs and the permissions and things like that is changing on a daily basis, and the last thing that we also do here is the AI readiness of the data. That's what we talked, I think, in the beginning of also the podcast. This is the foundation.

[00:15:58] If you want to be successful in an AI project, first of all, build a foundation layer, and the foundation layer is your data. Make sure that it's classified. You understand which data is the AI-ready data and which data you want really to feed into the LLM because if you start feeding a lot of your data that is not AI-ready data, it's going to cost you a lot, and it's also going to confuse the users because sometimes the AI will give you answers

[00:16:28] which are not accurate as well. So that's what it's all about. The Cetera Insight AI is really, as I said, a genting interface that collects all these activities and file system activities and be able to provide a very robust, I would say, and scalable platform to give you accuracy on all the things that you have in your system. And although it is a relatively new release here,

[00:16:56] I'm curious, all the conversations you're having, what kind of feedback have you had on this? The feedbacks are amazing. You know why? Because this is, you know, sometimes when you talk about data fabric and global file system, that's things that we are doing for more than 15 years, it takes time for the customers to understand all the benefits. We need to explain what is the difference between legacy NAS and caching and all of this amazing technology.

[00:17:25] But once they start feeding the data, after five minutes, they get value. Because asking the AI just a simple question, show me all the stale data that I have in my environment five years ago, and boom, after a minute, you get the full report, everything in a single click. And then you can ask more questions about it and more questions and get into a specific folder and a specific file type. So the first impression we're getting,

[00:17:55] we need it now. How amazing this tool. Show me how I can start using it. So that's the best way from a product perspective, right? Because it's really shown them the ability to get value from this system in less than a few minutes. And we see now almost every customer of us starting to pilot it, testing it. Some of them are already in full production using it on a daily basis. And they are actually talking to their data, right? They are getting an interface

[00:18:23] which they can start talking every day. It's reducing their costs for the storage dramatically. It's reducing their IT operation significantly because think about it. If you had a ransomware attack in the past, you needed to invest days or even weeks sometimes to understand the impact on the file system. The AI agent, the agentic interface can give you the full investigation in less than a few minutes. So think about

[00:18:53] how fast and how much more efficient your organization can be by having such a platform that's plugging into your file system. Incredibly cool. And of course, the mantra in enterprise AI has always been you can only improve what you measure. I think many organizations will admit that they're guilty of maybe missing that over the last few years with the arrival of AI and every tech project is now under increasing scrutiny for return on investment. So again, to finish on that

[00:19:22] adding value note, is there anything else you can share about how it can drive new measurable efficiencies? I think you should focus on three elements storage cost and optimization because the storage cost is a lot, right? I mean, it's the unstructured data is increasing. It's very easy to feed data into the AI but you don't understand the cost behind it

[00:19:51] and then you get the bill, right? So that's one thing that I really encourage people to analyze and make sure that they do the right step in order to make sure that it will be the most efficient thing for them. The second thing is time saving, meaning that try to build a tool that you don't need to be a developer. You'll be able to ask questions on a standard natural language and get answers in a very easy way and understand

[00:20:20] and be able to digest the answers very easily, meaning that I see a lot of AI tools that are providing huge information, a lot of information but how you can actually understand the bottom line, right? I mean, you know, if you have petabytes of data and you get now a 10-page report with all this amazing information, how can you take the right decisions and the tool needs to save you time and not create more complexity. So that's the

[00:20:50] second thing piece for the efficiency and the third one is how fast you can get the value, right? How fast you can move from production, from pilots or into production, right? How fast you can classify one petabyte of data or more, of course, right? How fast you can really understand the foundation of your data before moving to production, right? And if you follow these three rules, cost, time, and access fast to the data,

[00:21:19] then I'm sure that your project ROI will be very, very well. And I think that's a great moment to end on. And for anybody listening would like to find out more information about anything that we talked about today, we did cover a lot, especially around insight AI and the solution, how that's adding value as well. Where would you like me to point everyone listening? So, first of all, you know, I encourage you to go to our website. We have some new content

[00:21:48] over there, so go to ctera.com and you'll be able to get everything from the website. If you need further information, just book a demo or access the agent and our team will happily set the time with you and explain and show a product demo. It shouldn't be more than 30 minutes call, so it's not a long demo of the product. I encourage you as well to connect to my LinkedIn account, right? I'm posting a lot of things over there. We also have our CTO, Aaron Brand,

[00:22:18] that's providing a lot of useful information to our followers, so you can follow also our CTO, Aaron Brand, that provides a lot of information about our technology. And bottom line, just, you know, reach out and we are available and the team is available all over the world and we are happy to assist. Awesome. And for everyone listening as well, if you go to techtalksnetwork.com, you go to the podcast, there will be a blog post associated with this episode. I have come across, I think, four or five different videos

[00:22:48] that will talk about everything from how to make unstructured data AI ready with Cetera and agentic action AI and customer use cases and demos and so much more. So I will embed those into the blog post as well. Everything will be there for them. Including the links that you mentioned. But as always, thank you for coming on here and being a solutions not problems kind of guy. We could talk for ages about this. Absolutely love it. Thanks for joining me again. Thank you. Pleasure at all. Thank you, Neil, for your time.

[00:23:18] I think my guest gave some practical advice today for CIOs. Most importantly, give them a practical place to begin. Gain visibility and control over unstructured data before even thinking about attempting to expand AI across your enterprise. Because classification can reveal which information carries business value, which file contains sensitive material, and which terabytes have simply just been sitting around untouched for years.

[00:23:48] So feeding everything into an AI system might increase your cost while reducing the accuracy of its responses. Makes perfect sense. And some files will deserve intelligence and others probably deserve a long conversation about retention policies. And security matters too here. Bringing AI closer to governed data can preserve permissions. auditing, snapshots, and recovery controls. An agent should have

[00:24:18] identifiable access, monitored behavior, and limited authority rather than unrestricted freedom across the file system. So if your business inspected its infrastructure data today, how much would prove useful for AI? And how much would reveal years of unnecessary cost? Yep, technical debt. As always, let me know. Tech Talks Network.com. I'd love to hear from you on this one. But I have taken up for too much of your time today.

[00:24:48] So I'll be back again tomorrow with another guest. Speak to you then. Bye for now. . . .