CIO Corner: The Data Hygiene Myth Most CIOs Still Believe

Nasuni’s Dalan Winbush explains why perfect data hygiene is a fantasy, and shares what CIOs should do instead to make their unstructured data agent-ready.

July 16, 2026  |  Dalan WInbush

The problem with Nora

As the CIO of an unstructured data foundation company, this is uncomfortable for me to admit: There’s a problem with one of our agents. And it’s our own unstructured data that’s causing it.

Her name is Nora, and she’s the AI assistant for our sales organization. Like any other member of the department, she has a job description, responsibilities, and a defined skillset. Nora combines internal knowledge about the stage of a deal in the pipeline with external research on the account and their market. She validates POCs, shapes demo plans, and drafts closing checklists. She’s accessible in Slack, ChatGPT, Claude, HighSpot — wherever the sales team works. She’s designed to take the grey work off their plates, helping them move faster and with more precision.

The problem is that people have stopped using her. We know why: Nora is hallucinating. And it all traces back to our data cleanliness.

Junk in, junk out, now with more junk

IT teams have been fighting a losing battle against dirty data for 50 years. Duplicate files, outdated versions, documents saved in the wrong folder by someone who left the company in 2017. None of this is new. What’s new is who’s reading it.

When a person opens Highspot and sees three versions of the same battle card, they use judgement. They check the date. They ask a colleague. An agent doesn’t do any of that. An agent takes what it has as rote and answers with conviction. So if Nora doesn’t have access to the file she needs, or if the file she does have is out of date, the sales rep gets the same output: a plausible answer that’s completely wrong.

The agents didn’t break anything. They just exposed what was already broken. Junk in, junk out has always been true. But AI is now illuminating piles of junk we didn’t know we had.

Working with imperfect data

Perfect data hygiene does not exist. Anyone who says otherwise is either a fantasist or a fool.

It would take most companies months, or years, to clean their data to the standard necessary for an AI agent to navigate faultlessly. So, we didn’t try. We built infrastructure that lets us govern the data even if the underlying estate is not perfectly clean. Infrastructure with mechanisms that can:

  • Identify the data an agent is looking at
  • Generate metadata about it
  • Tag data by sensitivity, geography, and confidentiality
  • Oversee and control who (and what) has access to it
  • Let employees, or agents acting for us, navigate around it easily

And once data is identified and tagged, automation moves it into the right folders, with the right structure, under the right permissions.

A few career stops ago, this work would have been done by hundreds of interns putting in thousands of hours on tens of thousands of files. Today, we have agents managing the permissions of folders than other agents work in. AI preparing data for AI.

Build for agents, not analysts

What I want, as a CIO, is fast access to knowledge about my data: what’s in a file, who can see it, where it lives, whether it’s current. The kind of knowledge I can act on.

I feel good about our structured data. Really good. Our data warehouse team just finished a modernization initiative with a single objective: make the data warehouse AI-ready. I saw an MVP last week and it’s impressive. They’ve built it with a semantic layer designed for AI. Not for the human mind. For agents.

Now we need to do the same thing for unstructured data.

To get there, we have to apply the same principle to unstructured data that we’ve applied to structured data: a semantic layer that has a humanistic view of the data beneath. With that layer in place, AI utility can sit on top without running an enormous cleanup operation first.

Dog food today, champagne tomorrow

MCP is fast becoming the standard for connecting enterprise environments to AI-powered utilities. I’ve already seen v1 of Nasuni’s own MCP server, AI Activate, and the efficacy is there. Once it’s live in the product, it will strengthen the two things we already do best: managing unstructured data and making it more intelligent.

In the not-too-distant future, Nasuni will generate intelligence about data that sits outside our own platform. That’s our next horizon.

Right now, we’re eating our own dog food. But when this launches to the market, I will be the first to drink the champagne.

CIO Corner is the executive lens on what’s next in IT, delivered by Nasuni CIO, Dalan Winbush. With a firsthand perspective from the frontlines of IT leadership, Dalan unpacks what it really takes to modernize infrastructure, harness AI, and lead through complexity. This series tackles the hard questions CIOs face today, such as scalability, resilience, velocity, and value. It’s a candid look at how cloud and AI are reshaping enterprise IT — from someone who’s doing it, not just talking about it.

Related resources

Learn more about the latest developments in data infrastructure

Resource CenterRequest a demo