Skip to main content

Trusted Agents Start with Trusted Data

The Agentic Enterprise Runs on Data It Can Trust, and That Starts Today, Across Your Entire Estate

Three coworkers collaborating around a mobile device in an office environment.

Share this page

Asad Khan
Asad Khan

When agentic AI stalls in an enterprise, the cause is usually not the model. It is the readiness of the data feeding it. Agents consume and act on enterprise information at far greater speed and scale than people do, which raises the requirement for accurate data, consistent permissions, and current context and memory that create knowledge to enable action.

Every enterprise wants the same thing from AI: better decisions, faster action, and measurable results from the data it already has. That is difficult to achieve when the data estate is fragmented and hard to inventory. If you don't know what data exists, where it lives, who owns it, and who can access it, you can't confidently put it to work for AI.

That is the gap the NetApp® AI Data Engine is built to close. Powering the AI Data Services layer of the NetApp Platform, it makes enterprise data discoverable, governed, and usable for AI across hybrid estates, and it begins with discovery.

That's why today we are announcing the general availability of data discovery and metadata services in the AI Data Engine across your full enterprise data estate, whether or not that data sits on NetApp. One catalog spans NetApp ONTAP®, StorageGRID, and third-party NAS and object storage over NFS, SMB, and S3, on premises and in the cloud, including ONTAP-based first-party cloud services and S3 buckets from other providers. The result is estate-wide visibility and a metadata foundation for knowing your data before AI acts on it.

The first step to agents you can trust

Enterprises have tried to get their arms around unstructured data before. The available tools worked one vendor at a time, sampled instead of scanning everything, or required copying data into a separate system first, so what came back was a partial picture that was already out of date. More than 80 enterprises told us the same thing before we built this: a tool that sees only one vendor's storage does not solve the problem.

Data discovery and metadata services are built to that requirement: one catalog across a heterogeneous estate, not a sample and not just NetApp systems. The architecture is what makes that reach possible. It deploys as software beside the storage you already own, scans in place with no data movement or copies, is designed for billion-file estates, and runs independently of the storage OS, which is why it reads third-party systems the same way it reads ONTAP. Nothing is lifted into a separate platform, so permissions, protection, and hybrid reach stay intact.

It was also built for agents, not just people. The catalog is queryable through open APIs and authenticated MCP endpoints, so an agent can ask what exists and what is relevant today. That makes it an active index rather than a classification report, and the first place agentic AI touches your estate.

Immediate value from the same foundation

The value starts with the first scan. The same estate-wide index that makes data ready for AI also helps teams reclaim capacity, plan migrations, and find data that is more widely exposed than it should be. These are not separate features but different uses of the same index, and all of it runs on metadata, so there is no content to extract, no models or GPUs to stand up, and nothing to move or copy:

  • Optimization. Identify redundant, obsolete, duplicate, and stale data across NetApp and non-NetApp storage, organized by owner and size, to right-size the footprint and inform tiering decisions.
  • Migration planning. Point the catalog at NFS, SMB, and S3 shares, NetApp or not, to understand what is active and what matters before anything moves.
  • Access hygiene and exposure. Use one inventory across ONTAP, StorageGRID, and non-NetApp storage to surface issues like open shares, stray keys, exposed logs, and over-permissioned data.
  • The agent on-ramp. Scope out files you don’t want agents to use, such as logs, binaries, duplicates, and years-old data, to create a cleaner corpus and a more trustworthy foundation for agents.

In one enterprise environment, the metadata service indexed more than three billion files in near real time without moving or copying any of it. Filtering out binaries, logs, duplicates, and files no one had opened in years narrowed that estate to roughly seventy thousand documents worth serving to an agent. The same scan grouped cold data by owner, turning a cleanup argument into a conversation with specific people, and gave the team evidence to build a migration plan for an aging platform they had been unable to retire.

Building from discovery to knowledge graph

Discovery is what is generally available today, and it is the foundation for what we are building next. The order matters: find the data, understand what is inside it, connect that content with the context and memory that turn it into knowledge, then decide who the knowledge is available to. Agents need that connection, because context and memory are what create knowledge, and knowledge is what enables action that drives results.

Understanding starts inside the file. Content extraction is designed to read what a document or object actually contains and identify what a business cares about: a customer name, a contract number, a renewal date. A knowledge graph is designed to connect those facts, linking a customer to a contract, a contract to a renewal, and a renewal to the team that owns it, so an agent can follow a relationship instead of inferring one.

Above the graph, a decision fabric is designed to hold what no file contains: what the organization decided, who decided it, and why. That organizational memory keeps agents consistent with choices the business has already made.

Knowledge that rich raises an obvious question: who should be able to use it? Existing permissions are designed to carry through, so the rules already protecting enterprise data are the rules that protect it for AI. Activation follows: semantic and visual search across text, images, and video through open protocols, and open analytics that let engines such as Starburst work with unstructured data in place. These capabilities are in development and build on the discovery foundation available today.

Getting more from the data you already own

The larger opportunity is not a better inventory. It is getting more value from data you already own while reducing the complexity, cost, and risk of preparing it for AI. Most paths to AI readiness ask you to move data or stand up separate storage: another pipeline, another set of permissions, another surface to secure. This intelligence runs where your data already sits, protected and managed as it is today, across hybrid environments, so there is less to build, less to move, and fewer places for your controls to break. For a team unable to decommission an end-of-support array because no one could say what was on it, that difference is the whole project.

The through-line is simple. The agentic enterprise is only as trustworthy as the data underneath it, and that data is already yours, wherever it lives and whoever's storage it sits on. The work is making it discoverable, understood, governed, and reachable in place so it becomes trusted knowledge AI can act on to drive business results. That first step is available today across your whole estate, and the capabilities to understand, govern, and activate your data build from it.

You can start today. Data discovery and metadata services are available to customers at no cost, across NetApp and non-NetApp storage.

We will go through it in depth at NetApp INSIGHT 2026 in Las Vegas, September 29 to October 1.

Asad Khan

Asad Khan

A defining focus of my work is shaping the future of storage for AI: building high-performance, massively scalable infrastructure that enables organizations to train, deploy, and operate the next generation of intelligent applications.

View all Posts by Asad Khan