Skip to main content

Why MCP is the missing link between AI and infrastructure

A person standing in a data center aisle while working on equipment in front of tall server racks.
Contents

Share this page

NetApp arch logo
Toby Cherasaro

I’ve experienced the evolution of the Internet, virtualization, cloud, and AI/ML. But I’ve never felt more energized about what’s happening right now.

Careful planning, deliberate change windows, and extensive manual work have defined most of my career as a storage architect and IT operations leader. Complex global IT environments carry years of operational history, accumulated technical debt, and countless “we’ll clean that up later” decisions.

Agentic AI has enormous potential, but there has always been a problem: AI cannot manage infrastructure it doesn’t understand.

That’s where Model Context Protocol (MCP) comes in.

What felt aspirational just a few months ago is now something we’re actively using inside NetApp IT. MCP has become the foundation of our AI-assisted operations, giving AI agents the infrastructure context they need to analyze issues, interrogate ONTAP environments, surface root causes, and accelerate troubleshooting in ways that weren’t possible before.

What started as a vision for the future has quickly become part of how we operate today.

The AI problem MCP solves

AI is remarkably good at reasoning, but reasoning alone isn’t enough in enterprise infrastructure.

An AI model doesn’t automatically understand your ONTAP clusters, SnapMirror relationships, storage topology, operational standards, or decades of accumulated configuration decisions. Without that context, AI is forced to guess. It can search documentation, infer intent, and make recommendations, but it lacks the operational knowledge needed to interact confidently with live infrastructure.

MCP changes the equation.

In simple terms, MCP gives AI agents a structured way to understand and interact with systems like ONTAP. It provides the commands, capabilities, and operational knowledge necessary to perform meaningful work instead of relying on assumptions.

Think of MCP as a translator between AI and infrastructure. Instead of forcing an AI model to research ONTAP every time it receives a request, MCP provides a reliable way for it to understand which actions are available, how systems are configured, and how to interact with them safely.

From “what if?” to “we’re doing it”

When I first started talking about AI-enhanced storage operations, much of it sounded theoretical.

Imagine AI agents for storage, networking, compute, cloud, and applications all working together during an outage. Each agent understands its own domain. Each one investigates its part of the infrastructure. Then those agents correlate their findings and help identify the root cause.

That was the vision. Now we’ve built it.

Inside NetApp IT, we use agentic workflows that can pick up a ServiceNow ticket, read the description, enrich it with infrastructure context, query ONTAP through MCP, cross-check related domains, and generate an RCA-style analysis in minutes.

In one example, a user reported that they could not mount a NAS path. The ticket did not say it was NFS. It did not identify ONTAP. It did not specify which cluster, SVM, volume, or region to investigate.

The agent still figured it out. Using MCP, it analyzed ServiceNow and CMDB data, inspected ONTAP, verified the UNIX side, and determined the issue was an NFS export rule. It even confirmed that other mounts from the host were working before producing a high-confidence cross-domain analysis.

Agentic AI in action

Anyone who has worked a P1 incident knows the pattern. A ticket comes in. People join a bridge. Teams begin checking their own domains. Storage looks at storage. Network looks at switches. Compute looks at hosts. Application teams look at logs.

Now imagine this instead:

  • A ServiceNow ticket triggers an AI investigation
  • A coordinator agent reads the problem statement
  • A storage agent uses ONTAP MCP to inspect relevant clusters, SVMs, volumes, exports, SnapMirror relationships, and performance data
  • A UNIX or Windows agent checks the host side
  • A network agent checks storage fabric and layer 2
  • A coordinator compares findings across domains
  • The system posts an investigation summary and recommends next steps back into the ticket or war room

Today, much of what we are doing is read-only, especially in production incident workflows. Read-only is low-stakes but high-impact. If AI can quickly tell engineers where to look, what changed, what is misconfigured, and what is most likely causing the incident, that alone can transform response time.

AI helps us attack “inhuman” problems

The other major opportunity is technical debt.

Every enterprise has problems that are not hard because they are conceptually complex. They are hard because they are too large, too tedious, or too distributed for humans to solve efficiently.

For example, we can ask AI to audit local accounts across dozens of ONTAP clusters and hundreds of nodes. That is not a simple spreadsheet exercise. Last login information may live in security logs, not directly on the user account. A human would need to collect exports, merge data, review logs, compare timestamps, and manually build a useful report.

With ONTAP MCP, we can ask questions in natural language and receive fleet-wide analysis in minutes. Tasks that require hours of manual collection, spreadsheet work, and log review can now be completed through a simple prompt.

The same idea applies to aging snapshots, stranded SnapMirror destination volumes, inconsistent configurations, missing protection relationships, lifecycle planning, and capacity hygiene.

These are the kinds of problems that can sit unresolved for years because they are too painful to chase manually. MCP-enabled AI gives us a way to fix them and keep them fixed.

If an automation gap leaves behind stranded destination volumes, the answer should not be a one-time cleanup. The answer should be an agentic workflow that regularly checks for that condition, reports it, and eventually helps remediate it through approved controls.

We are moving from manual cleanup to continuous hygiene.

Natural language changes who can automate

Automation is no longer limited to people who write code.

Earlier in this journey, I built a demo for my team that I jokingly called “Storage Hotness.” The breakthrough was connecting the model to ONTAP through MCP. I connected an AI agent to an ONTAP simulator and started typing plain-English requests:

  • Create volumes
  • Resize a volume
  • Build SnapMirror relationships
  • Show me what changed

I had not operated ONTAP directly in almost a decade, yet I was managing storage through natural language. That moment helped the team see what was possible.

Our roles are evolving. Engineers are not going away, but the work is changing. We are becoming orchestrators, reviewers, and quality controllers of AI-assisted operations. People who learn to use these tools will multiply their impact. People who remain locked into manual workflows will struggle to keep up.

Why NetApp IT is the right place to prove this

NetApp IT is in a unique position because we use the same technologies our customers use.

We run ONTAP, StorageGRID, E-Series, Amazon FSx for ONTAP, Cloud Volumes ONTAP, Data Infrastructure Insights, and other NetApp technologies across a real global enterprise environment.

We are not exploring AI for storage in a lab-only scenario. We are applying it to real infrastructure, real incidents, real technical debt, and real operational pressure.

Because we’re operating these technologies at enterprise scale, we see firsthand how MCP accelerates troubleshooting, uncovers stranded capacity, identifies configuration drift, or surfaces years of accumulated technical debt.

When we see an agent produce a useful RCA-style analysis in minutes, we understand how this can change the way IT teams respond to outages.

And when something does not work, we learn from that, too.

Storage you can talk to

We are closer than many people realize to a future in which infrastructure is managed by intent.

Executives will ask, “How much capacity do we have?”

Engineers will ask, “Which production volumes are missing protection?”

Operations teams will ask, “What changed before this incident?”

Storage teams will ask, “Where is our inactive capacity?”

AI agents will gather context, query the environment, analyze results, and recommend action. That doesn’t replace engineering judgment but rather amplifies it.

The future is not AI replacing storage engineers. The future is storage engineers commanding fleets of AI agents that can inspect, analyze, correlate, and eventually remediate at a scale humans could never match manually.

For the first time in my career, the data center's operational layer is being rewritten. MCP is one bridge connecting AI reasoning to real infrastructure action. Without it, AI remains an assistant. With it, AI becomes an operational force multiplier.

And this time, it isn’t theoretical. We’re already doing it.

NetApp arch logo

Toby Cherasaro

Toby Cherasaro is a data infrastructure leader with over two decades of experience architecting and deploying enterprise storage solutions. He leads the NetApp IT Storage Engineering team, driving innovation in unified data management and AI-driven automation to ensure the company’s infrastructure evolves with emerging technologies.

View all Posts by Toby Cherasaro
How MCP bridges AI and infrastructure for smarter operations | NetApp