Operationalizing the AI Factory: Benchmarking for Predictable Outcomes
The conversation around AI has matured beyond what is possible to what is repeatable and scalable. For most enterprises, the focus has shifted to the supply chain of data: specifically, how to move from a series of successful pilots to an industrial-scale AI operating model. However, as these clusters grow, many organizations find themselves paying a “scaling tax.” This occurs when high-cost GPU compute clusters sit idle because the storage platform cannot feed them data fast enough to process.
At the Everpure PureScale Testing Center (PTC), we approach this as a first-principles problem. You cannot fix a broken foundation by simply adding shinier features or faster compute. If the “roots” of your data architecture are weak, the entire AI strategy risks collapsing under the weight of decades of data hoarding. We prioritize accountability over hype, validating our platforms through rigorous, industry-standard benchmarks and stress tests to ensure your infrastructure can sustain the high-utilization cycles required for real business ROI.
Why these benchmarks matter
AI workloads routinely push storage systems to their breaking point with massive metadata storms, exabyte-scale data sets, and relentless data movement that legacy benchmarks never anticipated. That’s why industry-standard AI benchmarks like IO500, SPECstorage AI_Image, and MLPerf Storage are now essential tools in every data architect’s toolkit. They are globally recognized and independently validate how a data platform will actually perform under real AI pressure, not just theoretical specs:
- SPECstorage Solution 2020_ai_image Benchmark: Measures real-world throughput and response times. It simulates the high-pressure environment of an active AI data center to prove the system can maintain “pass-fail” latency standards under load.
- MLCommons MLPerf Storage 2.0: Measures the system’s ability to keep GPUs fed with data at maximum utilization (typically 90% or higher), proving that the storage is not the inhibitor in training speed.
- IO500: Measures how well storage handles massive concurrency and metadata operations (such as lookups, creates, deletes) that can bottleneck large-scale training.
Navigating the architecture: Picking the ‘right tool for the job’
Storage platform selection is critical to delivering against the demands of your AI workload. Specifically, to maximize GPU utilization, you must match the storage architecture to the specific data profile, performance requirements, and scale of your installation. At Everpure, we’ve proven this with our scale-out FlashBlade® portfolio.
- FlashBlade//S™: The foundation for enterprise AI
For the vast majority of organizations, FlashBlade//S is the starting point for building out high-performance file and object-based systems. It’s engineered to deliver high I/O and throughout with low latency across high-concurrency environments where metadata management—the constant “lookup” and “tracking” of billions of small files—is the primary friction point. If you’re managing diverse training, inference, RAG, and general data analytics workloads, FlashBlade//S provides the density and metadata efficiency to keep those pipelines moving without the operational complexity of specialized “HPC” components.
- FlashBlade//EXA™: The high-throughput powerhouse for hero training at scale
FlashBlade//EXA is built for the moments when “exabyte scale” changes the nature of the problem. When you hit “hero” training jobs across thousands of GPUs, the bottleneck goes beyond metadata concurrency to predictable and sustained raw throughput. FlashBlade//EXA is engineered with a disaggregated architecture separating the metadata core from the data nodes. This ensures that even at exabyte scale, metadata lookups never compete for the same resources as the high-speed data path, effectively closing the gap between theoretical compute potential and real-world utilization.
| FlashBlade//S | FlashBlade//EXA | |
| Ideal Scale | Small to mid-size environments (up to 512 GPUs) | Large-scale AI native organizations and neoclouds (>512 GPUs) |
| Throughput (Read/Write) | 100GB/s to 1TB/s read performance in a single namespace | 1TB/s to 10+TB/s in a single namespace |
| Metadata Strategy | Best-in-class, on-demand scale data engine | Disaggregated architecture for infinite scale |
| Primary Workloads | Enterprise AI, inference, RAG, and data lake | Extreme-scale, multi-modal end-to-end training and inference workflows |
Delivering results and simplicity at scale
FlashBlade//S: Proving performance for the converged AI + HPC environment
Modern data centers no longer treat high-performance computing (HPC) and AI as separate silos. However, the operational overhead associated with traditional HPC architectures can quickly become untenable. Whether you’re running a climate simulation or an LLM training job or launching a new inference system, your storage faces the same challenge: massive concurrency. We used the IO500 benchmark to prove that FlashBlade//S can handle this convergence without the traditional HPC tradeoff in operational simplicity.
Our IO500 internal results from one uninterrupted run (with no partial component measures) for FlashBlade//S500 R2 demonstrate that a native scale-out stack outperforms niche legacy systems in the most demanding, data-intensive environments.
- Metadata scalability for AI and HPC: FlashBlade achieved 7.2 million IOPS across essential metadata operations. This is critical for RAG pipelines and HPC simulations alike. It ensures the system remains responsive even under the intense directory-level stress of billions of small files.
- A new benchmark for efficiency: With a total IO500 score of 142.32, the FlashBlade//S500 R2 system delivered approximately 2X higher performance than several competing parallel filesystems at a similar scale.
- Supercomputing on enterprise terms: The real operational win is that we achieved these numbers using standard NFS over Ethernet. You no longer need a room full of PhDs to manage a proprietary network and filesystem to get “HPC-class” results. You get the speed of a supercomputer with the operational simplicity of an enterprise array.
FlashBlade//EXA: Feeding the world’s largest GPU clusters
For training and inference workloads at exabyte scale, the goal is simple: maximize GPU utilization with predictable, linear scale. We’ve validated FlashBlade//EXA can feed the beast at any scale:
- SPECstorage: In the latest SPECstorage Solution 2020_ai_image results, FlashBlade//EXA achieved 6,300 AI_Jobs, exceeding the previous industry high of 5,000. This 26% increase represents a critical threshold. It proves that FlashBlade//EXA can handle the heaviest real-world unstructured data flows without hitting the “scaling wall” that typically causes GPU underutilization.
- MLPerf Storage v2.0: We leveraged MLPerf to prove raw GPU saturation. In our internal results (not verified by MLCommons Association), FlashBlade//EXA achieved approximately 2X the performance of the nearest competitor in 3D U-Net and took the #1 spot in ResNet-50 and CosmoFlow. These results prove that we can keep your most expensive compute resources productive, not idle.
- NVIDIA Foundation Certification: We’ve extended our existing NVIDIA certifications to FlashBlade//EXA. This isn’t just a badge; it removes any integration guesswork for your largest deployments.
- Aligned for Vera-Rubin era: As NVIDIA advances the STX modular reference architecture, Everpure is aligning FlashBlade//EXA for large-scale Vera Rubin AI Factories. We’re actively developing CMX-aligned capabilities to complement context memory acceleration and support efficient retrieval for frontier-scale models.
Beyond benchmarks and into real-world execution: STN
STN, a leading private GPU cloud provider, is currently using FlashBlade//EXA to deliver the data platform across its industrial-scale GPU One platform.
In a GPU-as-a-service model, any data stall is essentially expensive, unallocated idle time for the customer. FlashBlade//EXA architecture allows STN to sustain the massive throughput required across a 400GbE fabric and ~4 exaflops of GPU compute capacity, ensuring that dedicated NVIDIA clusters stay at peak utilization. By leveraging a disaggregated metadata core, the FlashBlade//EXA system handles the extreme concurrency of multiple AI tenants simultaneously without the scaling tax of traditional scale-out storage.
This combination allows STN’s customers to consume AI factories as a managed service, providing dedicated NVIDIA AI infrastructure with enterprise-grade performance and reliability without having to build and operate the underlying data platform themselves. For STN, this storage foundation enables a lean engineering team to scale its global GPU cloud, allowing them to shift resources from triaging infrastructure bottlenecks to accelerating customer intelligence.
The bottom line for the AI factory
Operationalizing AI requires evolving from limited-scale proof of concepts to full-scale production. The results from the Everpure PureScale Testing Center demonstrate that whether you’re consolidating enterprise-wide RAG and analytics on FlashBlade//S or driving massive-scale training on FlashBlade//EXA, the performance is predictable and the scaling is linear. The end goal is not chasing peak numbers. It’s eliminating the scaling tax and maximizing the utilization of your GPUs, allowing you to focus on the metric that actually matters: the adoption and maturity of your emerging AI applications and the people and processes they support.
Explore the full technical evidence:
Performance validation: Download the MLPerf v2.0 White Paper for GPU saturation data
Scaling verification: Review the SPECstorage 2020_ai_image White Paper for record-breaking results
Deep dive: Read our technical blog for the full breakdown of the IO500 Results