跳轉至主要內容

A new data architecture for AI factories at scale

Professionals collaborating
Contents

分享本頁

Arindam Banerjee
Arindam Banerjee

Today marks the beginning of a new era in data infrastructure for gigawatt-scale AI Factories. NetApp Novus, a new class of storage architecture, exceeds 100TB/s of aggregate throughput and can keep hundreds of thousands of GPUs productive. Engineered for zettabyte-scale capacity, Novus separates metadata and data into independently scalable planes, then unifies thousands of storage nodes behind a single namespace. With these capabilities, NetApp Novus will transform efficiency, utilization, and profitability for Neocloud operators, hyperscalers, large enterprises building GPU clouds, and GPU-as-a-Service providers.

Novus is built on standards-based pNFS access, so there are no proprietary clients to deploy across GPU servers and nothing to lifecycle-manage as kernels and GPU generations change. Novus brings the resiliency, security, and multi-tenancy of NetApp ONTAP® to an architecture purpose-built for AI Factories.

The arithmetic of feeding 50,000 GPUs

AI Factory math is unforgiving. A single GPU can demand as much as 2GB/s throughput to stay busy, and sustaining this across 50,000 GPUs requires 100TB/s cumulative bandwidth. Traditional storage arrays give you 40 to 80GB/s. So reaching 100TB/s with conventional arrays means deploying more than a hundred of them, each with its own namespace, failure domain, and management overhead. Operators wind up running a fleet of storage systems and contending with complex data movement problems.

Profitability is keeping GPUs busy

For gigawatt-scale AI factories, GPUs are the largest capital item on the balance sheet and the only one that generates revenue. Every hour a GPU spends waiting on data is an hour of depreciation with no return.

That makes the difference between a well-fed GPU cluster and a starved one business-critical. Utilization in AI Factories can fall into the single digits when storage can’t deliver data fast enough to keep GPUs working. At that point, we believe the storage bottleneck is no longer a performance problem; it’s a margin problem — which only grows as operations expand.

Eliminating the metadata bottleneck

More controllers alone won’t fix the problem. Metadata is a bottleneck even before capacity or bandwidth throttles GPU activity. Every open, lookup, and layout request lands on the same controllers that are trying to serve data. A workspace import is millions of small file operations. A checkpoint is a sequential write at terabyte scale, issued by every node at once, arriving directly behind a burst of metadata traffic. When you scale the cluster, the metadata operations scale with it, until inevitably the file system slows or stalls.

It’s difficult to engineer around two access patterns that are adversarial. Small-file metadata traffic wants many independent operations addressed quickly. Checkpoint writes want the entire system dedicated to one sequential stream. Run them both on the same controllers, and each one degrades the other.

Novus decouples metadata from data, allowing each to scale on its own. To do that, we built two specialized systems, each optimized for a different job, and unified them into a single system. Data lives on enterprise grade ONTAP systems. In the first release, we will ship with high-performance NetApp AFF A90 storage. Novus Data Director moves metadata into an independently scalable, software-defined control layer that runs on standard x86 compute. A client asks the Data Director where a file lives, receives a layout, and then talks directly to the storage. Neither plane waits on the other, and data runs at line rate.

Standard clients, single namespace, many tenants

Novus is architected with standards-based pNFS/NFS access through supported Linux client environments to help operators avoid deploying, patching, and validating proprietary clients across thousands of node images, GPU generations, and tenant environments. This aims to remove a major source of recurring overhead while preserving a single namespace for shared, multi-tenant AI infrastructure.

A single namespace provides operational simplicity. When a workload’s data is spread across separate systems, it must be moved, remounted, rebalanced, and every copy accounted for. A single zettabyte-scale namespace eliminates that work, keeping datasets, checkpoints, and model artifacts continuously addressable as the environment grows. Novus is architected to allow operators to add capacity and performance without forcing application changes, disruptive remounts, or tenant interruptions. Instead of making the storage topology visible to every workload, Novus presents one consistent data environment built for AI factory scale.

At AI Factory scale, shared infrastructure only works if trust scales with performance. Novus is designed to bring the proven security, resiliency, quality-of-service, and multi-tenancy foundations of ONTAP into a single architecture built for extreme concurrency. Providers can keep tenants isolated, workloads predictable, and GPUs productive even when thousands of jobs compete for the same data plane.

Scale is a moving target

Last year, 10,000 GPUs in a cluster was a large number. This year it is 50,000. Nobody believes today’s AI factories to be the end state. The architecture you choose has to be right not just for the cluster you’re running today, but for the one you’ll be operating in three years.

Gary Grider, Senior Director for Computing Technologies at Los Alamos National Labs, put it well: “The architectures that got us to the start of the AI era won’t be able to meet the demands we place on them as we continue to accelerate innovation. As AI Factories continue to scale, we’ll see hundreds of thousands of GPUs hitting a single namespace, creating a massive backlog of metadata operations that will slow or even stall the file system. Novus is a first-of-its-kind architecture that delivers the independent scaling of both the metadata tier and the data tier so AI Factories can grow without being constrained by their ability to access data.”

What Novus enables operators to do

Maximize GPU utilization. Eliminate the data bottleneck and keep expensive GPU environments fed and monetized.

Scale Performance Independently. Pair high-performance data nodes with a disaggregated metadata tier and scale each independently.

Drive profitability. Run multi-tenant, cloud-scale operating models on an architecture built for gigawatt-scale AI factories, empowering you to raise the return on your most expensive assets.

AI Factories have rapidly outgrown the architecture they were built on, forcing operators to move to something new. Novus is designed so future growth doesn’t require a new architecture.

NetApp INSIGHT 2026 runs September 29 to October 1 in Las Vegas. Watch live or on demand. Explore more in the NetApp Novus solution brief.

Arindam Banerjee

Arindam Banerjee

Arindam is NetApp’s first Technical Fellow. He is also the Chief Architect, VP of NetApp Platforms and leads the technology vision, strategy and architecture for NetApp. He is currently spearheading the architecture and design for next generation of AI infrastructure and AI data platforms. Arindam has more than 25 years of experience in distributed storage infrastructure and data platforms. He has been in NetApp for 19 years and has championed many innovations in the areas of filesystems, distributed storage, and AI. Arindam has authored/co-authored more than 50 patents and patent publications that have received over 500 citations for reference in the field of computer data systems and technology.

查看 Arindam Banerjee 的所有文章

後續步驟

Novus: AI factory storage for GPU clouds at zettabyte scale | NetApp Blog