Commvault: Why Your Disaster Recovery Plan Could Make Ransomware Worse
Tech Talks DailyJuly 20, 2026
3650
30:0522.7 MB

Commvault: Why Your Disaster Recovery Plan Could Make Ransomware Worse

Could the disaster recovery plan designed to protect your company make a ransomware incident even worse?

In this episode, I speak with Darren Thomson, Vice President and Chief Technology Officer for EMEA at Commvault, about Resilience Operations, commonly known as ResOps, and why cyber recovery now requires security, infrastructure, identity and data teams to work from one coordinated plan.

Darren argues that many companies are accepting a difficult reality. Even with considerable investment in prevention and detection, a breach may eventually succeed. That does not make cybersecurity controls any less necessary, but it means recovery can no longer be treated as a secondary activity managed by another department.

The problem is that security operations and infrastructure teams have traditionally worked toward different objectives. Security specialists concentrate on identifying and stopping threats. Infrastructure teams protect data, maintain backups and restore systems after outages. During a cyberattack, a successful recovery requires both sets of expertise.

A backup administrator may be able to restore data quickly, but a forensic specialist must establish whether that data is clean. Without that confirmation, the company risks restoring malware and restarting the incident.

Darren explains why a conventional disaster recovery plan may be particularly dangerous during ransomware. These plans were commonly designed for physical failures such as a lost data center. Data would be copied from one location to another so operations could continue. If the source data is infected, however, fast replication can carry the malware into the recovery environment.

This is where ResOps enters the discussion. Darren describes it as an operating model rather than a product. It combines established practices from security and infrastructure management into a continuous program for testing, learning and improving recovery. Individual technology projects may come from the program, but resilience itself never reaches a final completion date.

AI adds pressure on both sides. Criminals can use it to create faster and more effective attacks, while defenders can use machine learning to inspect large volumes of information, detect patterns and identify the newest clean recovery point. Companies must also protect AI systems as they would any other business application, including the models, data repositories and identities connected with them.

Darren offers one practical starting point for CIOs and CISOs: Mean Time to Clean Recovery, or MTCR. This measures how long it takes to restore an application and its data with evidence that both are free from compromise.

Before measuring MTCR, leaders must define their minimum viable company. These are the systems and services the business cannot operate without. Once that list exists, teams can test how long a verified clean recovery would take and replace assumptions with evidence.

The initial answer may be uncomfortable. Teams may know how to restore an application without knowing whether the backup is clean. Security may know how to inspect the system but lack an established workflow with the recovery team. Darren sees those gaps as the starting point for a useful ResOps program because they provide everyone with a shared problem and a measurable objective.

If your most important systems disappeared today, how long would it take to bring the minimum viable company back using verified clean data? Listen to the episode and share your answer with me.

Useful Links

[00:00:03] What if the biggest mistake in cyber security is just assuming you'll never need your recovery plan? Well, as cyber attacks become faster, more sophisticated and increasingly powered by AI, preventing every breach is no longer a realistic strategy. But the real question is how quickly your organization can recover should the inevitable happen.

[00:00:31] Well, my guest today is Darren Thompson from Commvault and he's going to explain why resilience has become a business priority. There's a great term I want you to remember today as well called resilience operations. We're going to demystify it, understand what resilience operations or res ops really means and understand why breaking down silos could make all the difference when every minute counts.

[00:01:02] A fantastic guest Darren is and I'm looking forward to speaking with him. So enough from me. Let me introduce you to him now. So thank you for joining me on the podcast today. Can you tell everyone listening a little about who you are and what you do? Yeah, certainly. Great to be here, Neil. I'm Darren Thompson. So I work as Commvault software's field CTO across the Europe, Middle East and Africa region. So I find myself on a lot of airplanes.

[00:01:32] And between the airplane trips, what I'm doing is I'm typically advising our clients at a senior level, sometimes board level on everything around security strategy, resilience strategy, governance, regulation, all those sorts of things. Fantastic. Well, it's a pleasure to have you join me today. There's a lot I want to talk about, especially the fact that for years now, it seems that we've been talking about sec ops, backup,

[00:02:00] disaster recovery and identity management, all though as separate disciplines. So the question I've got to start with today is why is now the right time to bring them all together under the banner of a resilience operations? Well, tell me more about that. Yeah, it's a great question, Neil. And this is a sort of a byproduct really of many years of us witnessing the same thing again and again, which is, you know,

[00:02:25] we've got a threat landscape that surrounds us with increasingly dangerous actors using increasingly dangerous technology to attack us. And so, you know, I think over the past, let's say, five years or so, many, many of the CISOs and the CIOs I work with are rapidly coming to the conclusion that we're moving to a world now where the breach is inevitable. Like at some point we're going to be breached.

[00:02:50] So we move away from this, throw everything we've got into preventing a breach into, well, if one, if we assume one's going to happen, we better do something about that. That gets you to a point where actually you need to continue doing all of that good work that you were doing to try and prevent attacks and try and detect attacks as they occur to you. But you need to then join that up with the business of recovering from an attack should it happen.

[00:03:16] So so just right there, if you're going to join those things up, those traditionally be traditionally have been two different disciplines. You know, you've got the SecOps people looking after prevention and detection largely. And traditionally you had the infrastructure folk looking after disaster recovery, backup and recovery, etc. You know, for the reason that I mentioned, this has got to be joined up now. We've got to assume we've got to do what we can to stop the breach.

[00:03:42] We've got to assume one's going to happen and we've got to create plans of action now that allow us to recover very, very quickly and bounce back quickly. We'll get into the detail. But what we find is that kind of a plan needs the very skills and disciplines that exist in both of those teams. Right. So to use a simple example, you, of course, need data protection specialists to look after immutable backups and data replication and all the things that they traditionally looked after.

[00:04:12] But, you know, you also need experts in forensics who can tell you if data is clean or not. And if you keep those disciplines, just to use that example, keep those disciplines separate where you can end up with a recovery plan that doesn't check cleanliness of data. Right. That's bad. By the way, that's where we see these huge recovery times come from. People have got a backup. They try to recover it and they find that's been effective.

[00:04:39] So you need, you know, those forensic specialists experts in malware to tell you whether the backup's good or not. Right. And so these teams have got to come together. So so so we call this ResOps. This is a coming together of SecOps and infrastructure ops. And they're joined up for that reason. And the reason I wanted to start our conversation today by asking you about this is because I was reading that you describe ResOps as a new operating model rather than just another security product.

[00:05:08] It's one of the things that set off my tech spidey sensors and why I invited you on today. So what changes in a way that organizations work when they adopt this new approach? Yeah. Well, first of all, I'd go further and say not only is ResOps not a security product, it's not a product at all. Yeah. You know, it's going to leverage certain technologies in order to, you know, come alive, as it were. Commvault would like to think that their unity platform is a good start.

[00:05:35] But actually, this is this is not a product. This is just just like any other framework. This is a framework on which people can, you know, start to build their plans and their governance models and their kind of disciplines.

[00:05:48] And so what it really represents is a coming together of several best practices across things like ITIL and COBIT as well as, you know, other regulatory kind of frameworks that Commvot put together specifically designed to deal with some of those issues I mentioned earlier. So it's a continuous improvement model is, you know, I describe it importantly as a program, not a project. So for me, a project has a beginning, a middle and an end.

[00:06:18] So that's sometimes bad because you don't want it to end. You want it to learn and continue. Right. So it's a program. It continuously improves. It will probably spawn projects. So ResOps program might tell you, you need to do better at backup or you need to do better forensics or you need to do, you know, better recovery or whatever. They might turn into projects, but they're being governed by this central program. And of course, this is a tech podcast. So we need to bring AI into the conversation. Of course.

[00:06:47] AI is introducing new levels of speed and complexity into enterprise environments. So on that side of things, why are traditional cyber recovery and security strategies? Can you expand on why they're no longer enough in an AI driven world as well? Yeah, well, I mean, there's two sides to this really, aren't there? As always, I describe AI and some of the other sort of mega trends, cloud and others as double edged swords. Right. So there's always two sides to these stories. Let's start with the negative.

[00:07:18] AI, AI enabled attacks are becoming more prevalent. They're more aggressive. They're more effective for the criminal. And so that just feeds the narrative of the breach is inevitable for me. We are we are all going to be breached at some point. And if the human didn't succeed breaching us, the AI definitely will. So with that in mind, we have better be able to recover. Right.

[00:07:43] So so so all that all that that side of AI is doing is taking us to a position where this resilience thing that we're talking about becomes even more important. The our ability to bounce back quickly with data that we can trust becomes a board level mandate in an AI world just because of the threats that exist. Now, the other side of AI, of course, more positively, is that we we should be using that for good as well.

[00:08:08] And so what a what a what a what a rezox program will also necessarily demand are technologies that frankly, you wouldn't have without AI. Right. So to give you a quick example, something that we built into our platform is the ability for the system to automatically tell you where your very latest cleanest backup is.

[00:08:32] Now, it doesn't do that by guessing it does it through integration with security applications and with machine learning and AI. So there's an example of a being leveraged for good. And so on the one side, you know, I attacks mean that we've got to have results. And on the other side, res ops is enabled by I. And actually, at Convo, we look at I with three lenses in terms of our development. One is how can we help people to use AI more safely?

[00:09:01] So that's data governance and data security and masking and all that kind of stuff. We have that in our product. Secondly, how can we protect the AIs themselves? So that's just what we've always done. Right. We were protecting SAP and then we were protecting Oracle and then we were protecting virtual machines. And now we're protecting Databricks and now we're protecting Snowflake. Right. So there's that that piece. And then the final piece is leveraging AI.

[00:09:26] So what are we doing ourselves to take advantage of this trend, take advantage of what I is incredibly good at, which is basically trawling huge amount of data and looking for patterns. And we're using that to good effect in regards to cyber resilience. And of course, AI is only as trustworthy as the data that it relies on.

[00:09:48] So for people listening in around the world, business leaders and techies alike, what should they be thinking or how should they be thinking differently about protecting data quality and integrity as AI? Incredibly, it continues to become embedded inside critical business processes. It's only going to continue as the proliferation of agents, et cetera, too. Yeah, I mean, there's a number of things in there, really. I mean, choose your models very carefully.

[00:10:15] You know, I'm a little bit scared of the next generation shadow IT. We were worried about shadow IT years and years ago when laptops started to be invented. And then we were worried about shadow IT in mobile devices. And then we were worried about shadow IT in the cloud, justifiably so. You know, people building microservices in the cloud that the corporation didn't know anything about. I think AI takes that even further.

[00:10:39] So now we have people building agents, you know, building their own LLMs and building their own repositories for data. That's very dangerous and needs to be governed. My view, very simply, is that needs to be carrot and stick governed from the top. You know, technology can help you to govern that. So the Convolt platform, for example, will help you detect where sensitive data is, will mask it if necessary, will move it to where it should be, protect it in the right way. So that's very, very important.

[00:11:09] But also I think it's important to protect those AI systems themselves. They are applications like any other application. And often I think we kind of forget that they're becoming already so ubiquitous. We don't think of, you know, an AI engine as an application where we should. It needs protection like anything else does. So that's the other sort of angle on that, really.

[00:11:33] And I was also reading in a previous conversation you spoke about how assumptions can undermine cyber recovery completely with you on that. And I've got to ask, what are the most common assumptions that you've seen and heard that organizations continue to make? And why do they become so dangerous during a real incident? I'm sure you've got a few war stories throughout your career. Yeah, I've got a few. I really stand out. Yeah. I mean, assumption number one, going back to our Red Sobs thing, right?

[00:12:01] My disaster recovery plan has been around a long time and therefore is going to help me in the event of, for example, a ransomware attack. That's a huge assumption. And they are in my experience and I was a VR architect for many, many years. We designed those systems and those architectures to protect us against something quite specific, which was some sort of physical problem. Right. Data center goes away or people can't access it or something.

[00:12:30] And so the architectural principles there were, well, very expensive for one, but we would basically replicate. Right. So data center A would be mirrored with data center B and we would move data from one to the other as quickly as we could. And we would use clustering technology to move from one to the other in the event of a disaster. Well, that was fine and dandy in the world of, you know, well, once we were worried about data centers disappearing, it didn't have ransomware in mind.

[00:12:59] And so we see a lot of these DR plans that have been around for a long time, probably served their companies very well. But they're nothing in the face of a severe ransomware attack. Well, in fact, what they'll probably do is help the criminal to propagate the malware because they're, for example, synchronously replicating data from site A to site B. There's nothing in them normally that says that's bad data. I won't replicate that. They just replicate data. So we see many, many instances.

[00:13:28] I worked in the insurance industry for a few years and spoke to a lot of companies that had been breached, unfortunately. And my first question was always, where was the plan? And they would point me to a DR plan. And it was like the reason this took my only reason you were breached, but the reason it took you so long to recover is because you started bouncing from site to site. You know, you go to site B and hope that would be clean. It wasn't. You go back to site A, that's clearly not clean. And that's where all of that time goes when we're trying to recover your business. So that's the assumption number one.

[00:13:58] DR, disaster recovery is not the same as cyber recovery. They're different things. One's an addendum to the other, I would say. They all sit under business continuity management, but there are some specifics that we now need to take care of when it comes to cyber recovery. So that'd be assumption number one.

[00:14:16] Assumption number two would be the fact that your security team and your infrastructure teams are both working really hard and basically doing all of the right things based on their objectives and their measurements. But that's going to save you. Again, there's a gap very often between the two. What the SecOps folk are doing has always been right. And they're probably working really hard, investing very heavily in tools. But in the event of a breach, none of that matters anymore because you've been breached.

[00:14:45] And then on the infrastructure side, it's all very well, again, having disaster recovery and backup architectures. But if what you're recovering is not clean, that's no good either. And so the assumption that both those teams doing the right thing is enough is wrong. They've got to do the right thing together. And a lot of the work that I'm doing and the Commvault are doing is forcing these teams together in some cases for just finding an excuse to work on something together.

[00:15:13] We run something called Minutes to Meltdown, which is actually the most fun I have in my career, which is, you know, we throw these people in a room for a few hours and attack them with a virulent strain of ransomware and just help them work through it. Right. And it's amazing to see these teams start to work together and the ideas that get pinged around the room. And most of those organizations leave with, if nothing else, the conclusion that we should work together more, which is exactly what we want. So I think that's a dangerous assumption as well.

[00:15:43] It's not your fault that your agentic AI systems are acting outside of compliance. There are just too many data sets to guardrail them all. But with Denodo, you can now organize your hundreds of data sources into one layer and govern your agents with a single approach. Try it now with Denodo by visiting denodo.com to learn more.

[00:16:09] And when we're talking about these gaps here, another one I think is that many organizations are still operating with separate teams responsible for, I don't know, cybersecurity, identity, backup and recovery. So what practical steps could leaders listening take to maybe break down some of these silos without disrupting day-to-day operations? Feels like somewhat of a balancing act, but any tips? Yeah. Well, I think you've hit on a couple of interesting sort of domains there, really.

[00:16:38] So, for example, Convolt's cyber resilience platform now provides three sort of integrated domains of competence. One is what you'd expect it to be. It's data protection, essentially, right? So we're protecting workloads. We're moving data around. We're making that immutable. We're making it clean. We're doing constant scanning on it. You know, we're essentially creating backups that can work, right? So that's one.

[00:17:04] We've invested very heavily, though, in the past few years around identity protection. So identity, I guess the font of all knowledge with regards to identity is normally your Active Directory, right? Or Entra. Nobody's really protecting those effectively, in my view. And so we've invested an awful lot of money in protecting, I almost call it policing the police.

[00:17:27] So if you've got an Okta or an Active Microsoft Active Directory looking after your identities, who's looking after them? Because the criminal is going after them. I mean, that's a prime target for a criminal to steal, you know, credentials that matter to them. So identity protection is the second piece, really important. We've been investing very heavily in that. And then the third piece is good old data security. You mentioned it earlier when we talked about AI.

[00:17:53] We really need to understand what our data is, where our data is, why it matters, where the PII is, what that means in the context of regulation, how, therefore, I should be protecting it. Should I be masking it in some way? You know, in the case of AI, what am I going to allow access to and not allow access to? What is going to be masked and not masked? How am I going to tokenize things? That's essentially data security. That's the third domain.

[00:18:20] So we put those together in our platform because we think those are three core disciplines, really, or resilience operations. Now, back to your point and your question, traditionally, very separate teams, you know, running those things. You know, the backup and recovery piece would have been the infrastructure team, the identity security piece. Maybe the security team, maybe a completely separate identity team. I see that as well. You know, that's a problem.

[00:18:48] And then data security, that's, you know, another discipline. So, again, we've got to break down these barriers. This is one, ResOps is, you know, there's one goal here, and these teams have got to work together. And in terms of how we go about starting to make that happen with our customers is it's a lot of discussion and workshopping, just working through problems together. If we are going to build a cyber recovery plan, let's make sure all the right people are in the room when we build it.

[00:19:15] And then amazing things start to happen because, you know, you put these skills together and one plus one equals three all of a sudden. And so it's just I'm a big fan of just doing, you know, even if you're not 100% accurate first time out, get the right people in the room, start to build something together. So that's a lot of what takes up my time right now. And as a solutions, not problems kind of guy, I'd love for you to be able to, I don't know, bring to life what we're talking about here.

[00:19:44] And most importantly, the solution side of things. And you'd have to mention any names here, but do you have any examples of how a more coordinated approach to resilience has maybe helped an organization recover faster or reduce the impact of a cyber incident? Because it's all about ROI and improving business outcomes out there now. And it's always difficult, isn't it? ROI and risk are not really natural bedfellows. Again, I worked in the insurance industry for years. The ROI is a difficult thing to calculate.

[00:20:13] But a couple of things here. I've always been in favor of creating this kind of almost – I've always thought a good IT strategy should resemble a seesaw in some ways. At one end of that seesaw, you've got risk. You should be trying to drive that end of the seesaw down, preferably. On the other end of the seesaw, you've got efficiency. And that touches productivity of employees, how well the business is doing, how well systems are running.

[00:20:43] And so if the seesaw is operating correctly, one affects the other. You bring risk down and efficiency goes up, right? You bring efficiency down, normally risk goes up, right? And so I've always been a fan as a practicing sort of CTO. I've always been a fan of that approach. And I think that applies here. So we definitely see organizations that benefit truly from a financial standpoint by adopting ResOps and by adopting Convolt technology.

[00:21:12] We can take lots and lots and lots of money, for example, out of how people protect their cloud applications. How do we do that? Well, because we can protect them all with one system. And what we often find in cloud environments is that those environments have grown really quickly over not very much time at all. And we have lots and lots of separate teams looking after their own little applications and protecting them themselves with off-the-shelf products or native tools or whatever that might be.

[00:21:40] We strip all those native tools and those separate products out and you centralize all of that. And by the way, you centralize it on a platform that's always been looking after your on-prem stuff as well. Then all of a sudden there's a huge cost saving. So that's the obvious ROI side of things. And whether we're focusing on data protection, data security or identity protection, that becomes true. But the other side of it is we see organizations recover differently now.

[00:22:06] Our support organization do keep a log, actually, of organizations that have been breached because it was inevitable. As our customers, clearly, we want to help them through that breach. And we do that every day. We do that multiple times a week all over the world. And so we're seeing a definite trend. People that are adopting this philosophy, they are starting to recover their businesses,

[00:22:30] certainly what we call their minimal viable company, in hours rather than days or even potentially weeks or months. We are starting to see that effect. Now, that's not ROI per se because it's difficult to put a dollar value on it. But look, if your business was down for two days as opposed to three weeks, there's ROI there. It's just difficult to pin that with a number. Yeah, 100% with you.

[00:22:54] And we've seen many examples of famous retailers going offline for several weeks, not to mention automotive manufacturers. But that's a podcast episode on its own. But if there was a CIO or CISO listening today wanting to begin building a ResOps strategy, what would you recommend they prioritize first to begin improving resilience before the next major disruption occurs? Is there any advice that you would leave them there that are really embracing this ResOps concept?

[00:23:22] Yeah, I've actually got some very specific advice. So we wrote a paper about a year ago now. I started to think about a year ago about measurement. So for me, you only really get good results when you know what you're measuring and you can measure it effectively. And so, again, going back to our different teams of people across security and infrastructure, if you think about how they measure themselves, those departments, typically, again, they're very different.

[00:23:48] So the security folk are probably obsessing over meantime to root cause, meantime to getting to that root cause problem. On the infrastructure side, RPO and RTO in our world has always been a really big thing. Recovery time objective. How long does it take me to get a system back? Recovery point objective. At what point can I trust the data?

[00:24:15] And I don't think in the context of what we've been discussing in ResOps, they're both useful, but neither are enough. And so I invented something called MTCR, meantime to clean recovery, which was deliberately supposed to be a bit provocative and a bit of a synthesis of both security side of the house and infrastructure side of the house. So what mean time to clean recovery measures is exactly what it says on the tin, right? So I take an application.

[00:24:43] What is, if you were to recover this thing 10 times, what's the mean average time that it takes you to get this data, this application and data back for me? Very importantly, the C, cleanly. And I want some sort of guarantee that that's clean. I want somebody to say, probably on the security side of the house, that's clean. We've done a scan. We've done at least a scan, right? We've gone quite deep on this.

[00:25:06] So back to your question, first thing would be for a CEO to get their IT organization, the leaders in front of them, both security and infrastructure, or maybe just the CIO, if both those organizations sit under the CIO, and ask the MTCR question, right? So tell me, number one, tell me what my minimal viable company is. So in other words, what's the bit of this business that must never go away? What can't we survive without?

[00:25:35] Now, if you don't have an answer to that, by the way, that ends up being the first thing you need to do. But what is the minimal viable company? Okay, we've got that defined. Let's say, for example, the minimal viable company is Active Directory. It's never going to be that, but just to use an example, right? So Active Directory is my scope. Okay, that's good. So as a mean average, how long does it take you from complete failure? All right. So Active Directory has gone away. All of its infrastructure has gone away. There's nothing.

[00:26:05] As a mean average, how long does it take you to get me that back? Guaranteed clean. Now, the answer is not going to be a good one. And that's why it's a good first step. Okay. The answer is going to be something like, honestly, we haven't even thought about that. We know how we could get Active Directory back with a backup, but we don't know if it's clean. Or we know how to clean Active Directory, but it's not joined up with backup, right? They're not good answers.

[00:26:31] And so if you were to be honest about the actual score, the MTCR number in days, it's probably going to be a month, two months, three months, six months, as was encountered by those retailers you mentioned, right? But now you've got an objective. Now you've got something for those teams to work on together, haven't you? Because my mean time to clean recovery for my minimal viable company is two months, let's say. Well, I want it to be a week. Go at it. Make a project.

[00:27:02] Leverage the ResOps discipline and then work together, come up with us. What's the answer? And we've done this with ourselves at Convo, with our own critical applications. We do this with customers. And it's a profound question, that, because it doesn't start with technology. It's just like, look, can anybody even get close to telling me what the answer would be? Answer's probably no. That's bad. We have a problem. What do we do?

[00:27:30] You know, so that's a great place to start. Wow. So many big takeaways there. And for anybody listening that's interested in learning more about how ResOps could replace siloed approaches and with integrated workflows, enable their own teams to collaborate and respond to threats more effectively. Well, do you like me to point everyone listening to find out more about you, your work, Convo and the solutions there? Where should they go? Yeah, we'd certainly come to Commvault.com, fairly obviously.

[00:27:59] But I'm on LinkedIn. I'm fairly exposed on LinkedIn. I'm posting things pretty regularly. So by all means, reach out to me personally as well. But fundamentally, Convo.com is a very good place to start. Well, as we've said throughout our conversation today, ResOps is that new operating model that can unify cyber recovery, data security and identity resilience into one single coordinated framework.

[00:28:28] For anybody listening interested in digging a little bit deeper on that, check out the show notes in the blog post associated with this episode. All the links will be there, including your LinkedIn. So hopefully we can carry this conversation. But more than anything, just thank you for shining a light on this. Incredible what you're doing here. But thanks for sharing it with me today. Thank you, Neil. Thank you for having me.

[00:28:50] Today's conversation, I think, was a powerful reminder that resilience isn't just a product that you can buy off the shelf or a box you can tick. It's actually an ongoing discipline that brings together people, processes and technology before a crisis ever occurs. And I love how Darren challenged one of the biggest assumptions in cyber security. And that is disaster recovery and cyber recovery are the same thing.

[00:29:20] And as AI accelerates both attacks and defenses, organizations that recover the fastest will be those that have already learned to work together. But I'd love to hear your thoughts on this one. If your organization faced a major cyber attack tomorrow, how confident are you that you could recover quickly? And recover with the data that you know you can trust? Let me know how you're going to answer that one.

[00:29:49] TechTalksNetwork.com. But more than anything, just thank you for listening today. And I'll speak to you again very soon. Bye for now.