Interviews

Ian Leysen, CEO and Co-Founder of Datadobi – Interview Series

mm
Add Unite.AI to your preferred sources on Google

Ian Leysen, CEO and Co-Founder of Datadobi, is a technology executive with more than three decades of experience in software engineering, quality assurance, and enterprise data management. He co-founded Datadobi in 2009 after spending eight years at EMC as Senior Manager of Quality Assurance, preceded by leadership roles at Mediagenix and Wave Research. Throughout his career, Leysen has focused heavily on building high-quality software engineering organizations, having established three quality assurance teams from the ground up. At Datadobi, he oversees a company focused on helping large enterprises manage, govern, migrate, and protect unstructured data across on-premises, cloud, and hybrid environments. The company has evolved beyond its roots in large-scale data migration to offer StorageMAP, a vendor-neutral platform designed to give organizations greater visibility and control over complex unstructured data estates, including preparing enterprise data for AI initiatives.

Datadobi helps enterprises gain greater visibility and control over rapidly growing volumes of unstructured data. Its software can scan billions of files to identify stale data, duplicates, ownership gaps, and potential risks, while applying metadata and classification tags that support governance and automated policies for archiving, deletion, and retention. This has become increasingly important as organizations prepare enterprise data for generative AI, where poorly understood or outdated information can introduce noise, compliance issues, and security risks. Datadobi also enables companies to identify potentially valuable datasets, organize them for downstream use, and move selected information into data lakes or lakehouses while maintaining traceability and governance. The platform additionally provides insight into storage costs and carbon impact, helping organizations make more informed decisions about what data to retain and where it should reside.

You spent eight years leading quality assurance at EMC before co-founding Datadobi in 2010. What did you see in large-scale enterprise storage and data environments that convinced you there was a company to build, and how has that original vision evolved as unstructured data has become increasingly important to AI?

At EMC, I spent years watching enterprises invest heavily in storage infrastructure while having almost no visibility into what was actually sitting on it. We were exceptional at helping customers store and protect data, but nobody was asking the harder question: what is this data, who owns it, does anyone still need it, and what is it worth? That gap between infrastructure capability and data understanding was the opportunity. We started Datadobi to help organizations move and manage unstructured data intelligently, not just shift it from one array to another.

What has changed is the stakes. Fifteen years ago, an unmanaged file share was a cost and compliance problem. Today, that same unmanaged file share is a liability the moment someone points an AI model or agent at it. Unstructured data has gone from being the thing organizations stored to being the thing that determines whether their AI initiatives succeed or fail. Our original idea, that storage infrastructure alone can’t tell you what your data means to the business, hasn’t changed. It has just become urgent in a way it never was before.

You’ve argued that generative AI did not create the enterprise data problem, but instead exposed and accelerated problems that have existed for decades. What are the biggest weaknesses AI is now revealing in how organizations have historically managed their data?

Organizations have struggled to understand their enterprise data for decades. AI didn’t create that struggle, it just removed the places it used to hide. When data sat quietly on a file share or in an archive, nobody had to answer for what was in it. The moment you point a large language model or a RAG pipeline at it, every weakness becomes visible and consequential.

The biggest challenge is that most organizations have been managing storage, not managing data. They know where their volumes and buckets are, but not what’s inside them: which files are stale, which contain sensitive or regulated information, which are duplicated dozens of times across the environment, and who actually has access. AI is also exposing how fragmented ownership has become. Data accumulates across on-premises systems, multiple clouds, and SaaS repositories, and nobody owns the whole picture. Those aren’t new problems. AI has simply made the cost of ignoring them immediate and visible.

Organizations often focus their AI investments on more powerful models, GPUs, and infrastructure. Why can’t more compute or storage solve an underlying data-readiness problem, and where should enterprises be investing instead?

More compute makes a bad answer arrive faster. It doesn’t make the answer accurate, safe, or compliant. GPUs and storage infrastructure execute decisions they don’t make them. If you feed a powerful model stale, duplicated, mis-permissioned, or sensitive data, you get a powerful model producing unreliable or risky output at scale, and doing it quickly.

We believe the market has reached an important inflection point: historically, organizations optimized storage; increasingly, they need to optimize data. That means investing in the discipline that sits above the infrastructure layer, the ability to see across your entire data estate, understand what each piece of data actually is and who’s accountable for it, decide what should be retained, moved, archived, or deleted, and then execute that decision consistently. Infrastructure spend without that discipline just means organizations are able to do the wrong thing faster.

That’s precisely the problem our unstructured data management platform, was built to solve. It gives organizations a single view across on-premises, cloud, and SaaS storage, classifies data with tagging and metadata analytics so teams can see what’s redundant, obsolete, or genuinely valuable, and then executes decisions, migrating, archiving, or deleting data, through policy-driven workflows that operate continuously rather than as a one-off project. That combination of visibility, classification, and consistent execution is what turns ‘we have a lot of data’ into ‘we know exactly what we have and what to do with it.’

“AI-ready data” has become a common industry phrase. From your perspective, what actually makes unstructured data AI-ready, and what criteria should organizations use before allowing data into a generative AI, retrieval-augmented generation (RAG), or training pipeline?

AI-ready data is data an organization has already validated not just data it possesses. In practice that means the organization can answer a handful of questions with confidence before that data ever reaches a model or a pipeline: Is this data accurate and current, or has it been sitting untouched for years? Is it duplicated elsewhere in a way that will skew or contradict results? Does it contain sensitive, regulated, or personal information that shouldn’t be exposed? Who is permitted to access it, and does that still reflect who should be able to? Does it actually add business value to the use case, or is it noise?

Without answers to those questions, feeding data into a generative AI or RAG pipeline just means moving your governance problem downstream, into a system that is far better at surfacing what it finds than your file shares ever were. AI readiness is a data intelligence discipline, not a checkbox you apply once before a project kicks off.

Enterprises can have billions of files spread across on-premises infrastructure, multiple clouds, archives, and business units. How can they determine which data contains meaningful business value and which is redundant, obsolete, trivial, or simply noise that could degrade AI performance?

At that scale, nobody is going to answer that question file by file, and manual review isn’t a viable strategy. Organizations need enterprise-wide visibility first: a single, accurate view across on-premises, cloud, and SaaS repositories, because you can’t make a decision about data you can’t see. From there, it’s about applying data intelligence to classify what’s actually in the environment, so ROT (redundant, obsolete, and trivial) data, gets identified and separated from the data that genuinely carries business value.

This is where the discipline has to move beyond visibility alone. Seeing your data is necessary but not sufficient. Organizations need to progress through understanding what that data is and means, deciding what should happen to it, retain, move, archive, delete, or use it to power AI, and then executing that decision consistently across billions of objects. Skipping straight from visibility to AI ingestion is exactly how noise ends up degrading model performance and how genuinely valuable data gets buried in it.

Security and governance become especially important when AI systems can surface information that was previously difficult for employees to discover. How should organizations assess permissions, sensitive information, ownership, and regulatory risk before exposing enterprise data to AI systems?

This is one of the areas where AI has changed the risk calculus the most. A file with excessive or stale permissions used to be a theoretical exposure, because realistically, a person would have had to know it existed and go looking for it. An AI system with broad access can surface that same file to anyone who asks the right question, instantly. Obscurity was never a real control, but AI has removed the last bit of protection it accidentally provided.

Before any data is exposed to an AI system, organizations need a clear picture of who has access to it and whether that access still makes sense, what sensitive or regulated information it contains, who owns it and is accountable for it, and what regulatory obligations attach to it – data residency, retention, and privacy requirements among them. That assessment can’t be a one-time audit ahead of a launch. Enterprise data changes continuously, so permissions, ownership, and risk have to be reviewed on an ongoing basis, not just at the moment an AI project goes live.

Datadobi advocates shifting the conversation from managing storage infrastructure to managing data as a business asset. What does that transition look like in practice, and how does it change the relationship between IT teams, data teams, security leaders, and business units?

In practice, it means the conversation stops being about capacity, tiering, and uptime, and starts being about outcomes: cost reduction, risk reduction, regulatory compliance, and enabling AI. Those used to be treated as separate initiatives, each with its own tools and owners. We believe that view is increasingly outdated. They all depend on understanding the same underlying enterprise data, and  what’s needed is a new data-focused operating model that connects them, rather than treating each initiative as if it depends on a separate, isolated system. Our platform is how we put that operating model into practice.

That naturally changes who’s in the room. IT is no longer the sole owner of the conversation, because decisions about what data to keep, move, or expose to AI are business decisions, informed by data intelligence, not infrastructure decisions. Security and compliance leaders need visibility into the same data landscape IT manages. Business units need a voice in what data actually matters to their outcomes. Data management stops being a back-office IT function and becomes a shared operating discipline with IT, security, and the business making decisions off the same information.

One challenge with enterprise AI is that data is constantly changing. Is AI readiness something organizations can achieve once, or does it require an ongoing process for discovering, classifying, governing, archiving, and moving data as it evolves?

It’s ongoing, full stop. Enterprise data changes continuously, new files are created, permissions shift, employees join and leave, regulations evolve, so data management has to become a continuous operational capability rather than a sequence of independent projects. Treating AI readiness as a one-time cleanup before a project launch is a bit like declaring a building secure after a single locksmith visit and never checking the doors again.

What organizations need is an operating discipline that continuously moves through visibility, understanding, decision, and execution, discovering what data exists, classifying and understanding it, deciding what should happen to it, and then acting on that decision, on a recurring basis. The organizations that outperform their peers will be the ones that can move through that cycle continuously and at enterprise scale, not the ones that treat AI readiness as a project with an end date.

As enterprises increasingly deploy AI agents that can search across systems and take autonomous actions, does unstructured data management become even more important? What new risks emerge when an AI agent can access information scattered across an organization rather than simply responding to a user prompt?

It becomes significantly more important, because an agent changes the nature of the exposure. A chatbot answering a single prompt is limited by what one person asks and sees. An agent that can search across systems and take autonomous action can traverse far more of the environment than any individual employee typically would, and it can act on what it finds, moving, sharing, or using data, without a human necessarily reviewing each step.

That introduces risk that goes beyond simple discovery. If an agent has access to data it shouldn’t (mis-permissioned files, stale sensitive records, information that should have been archived or deleted years ago) it can act on that data at machine speed and scale, not just surface it to one curious user. The organizations deploying agents most successfully are the ones that treated data governance as a prerequisite, not an afterthought, because an agent will faithfully exploit whatever gaps exist in your data intelligence.

For an enterprise that has accumulated decades of unstructured data and wants to scale its AI initiatives, what practical steps would you recommend taking first, and what mistakes should leaders avoid as they begin getting their data estate under control?

Start with visibility. You cannot make good decisions about data you can’t see, so the first practical step is getting an accurate, enterprise-wide picture of what data exists across on-premises, cloud, and SaaS environments. From there, move into understanding, classifying that data so you know what’s valuable, what’s sensitive, and what’s simply noise, before you move into decisions about retention, migration, archiving, or deletion.

There’s also a budget reality leaders can’t ignore. Most CIOs aren’t getting a separate, unlimited AI budget, they’re working with a fixed pool of money that now has AI competing against everything else that keeps the business running. The instinct to fund AI by stripping investment out of existing infrastructure is the wrong move, because that same infrastructure, storage, data pipelines, governance, is exactly what AI depends on to succeed. The more sustainable path is creating headroom inside the existing estate: improving visibility and reducing storage waste through the kind of data optimization StorageMAP is built for frees up real budget, without touching the capacity AI initiatives will actually need.

The biggest mistake I see is organizations skipping straight to execution, pointing AI at their data estate, or launching a cleanup project, without first building that foundation of visibility and understanding. The second mistake is treating this as a one-time initiative rather than an operational capability; data keeps changing, so the discipline has to be continuous. And the third is leaving it as a purely technical exercise. The organizations that succeed treat this as a business decision, with IT, security, and business stakeholders aligned on what the data is worth and what should happen to it, not just a migration or storage project handed to IT alone.

Thank you for the great interview, readers who wish to learn more should visit Datadobi

Antoine is a visionary leader and founding partner of Unite.AI, driven by an unwavering passion for shaping and promoting the future of AI and robotics. A serial entrepreneur, he believes that AI will be as disruptive to society as electricity, and is often caught raving about the potential of disruptive technologies and AGI.

As a futurist, he is dedicated to exploring how these innovations will shape our world. In addition, he is the founder of Securities.io, a platform focused on investing in cutting-edge technologies that are redefining the future and reshaping entire sectors.