The AI Pipeline Stalls on Storage

Training, checkpointing, and inference each push storage differently — and expensive accelerators sit idle whenever data can't keep up. Teams end up staging copies in and out of object storage, running separate systems for each stage, and rewriting applications to speak object semantics.

Cloud file services deliver performance at a premium that makes large AI datasets cost-prohibitive; object storage is affordable and limitless but isn't a filesystem. The result is copy sprawl, idle GPUs, and rising cost.

MayaNAS presents a standard POSIX parallel filesystem directly on cloud object storage, so the whole pipeline reads and writes one namespace at parallel-filesystem speed — keeping accelerators highly utilized, with no staging and no app changes.

Benefits

  • One filesystem for training, checkpointing, and inference
  • Parallel-filesystem speed at cloud object-storage economics
  • Keeps accelerators highly utilized — no staging copies in or out
  • Runs in your own cloud — full data sovereignty
One Substrate for the Whole AI Pipeline
GPU Clients — Training · Checkpointing · Inference
LustrepNFS Flex FilesNFS Standard in-box clients — no proprietary software
↓ ↓ ↓ ↓
MayaNAS Parallel-FS Engine Single POSIX namespace · active-active · scales by adding HA pairs and clients
OpenZFS Data Services NVMe special vdev holds pool metadata and small blocks only
objbacker.io — bulk data direct to object storage Each OST is an OpenZFS dataset over regional, Standard-class object buckets
NVMeMetadata · small blocks
Cloud Object StorageAmazon S3 · Azure Blob · Google Cloud Storage

Proven in MLPerf® Storage v3.0

ZettaLane submitted closed-division MLPerf® Storage v3.0 results for MayaNAS and its companion NVMe-over-TCP block engine, MayaScale — run entirely on standard Google Cloud virtual machines, accessed through standard open-source clients. One approach served training, checkpointing, and inference-cache workloads across the pipeline.

WorkloadResultEngine
Checkpointing — Llama 3 8B (single client)
Entry 3.0-0137
14.43 GiB/s write · 10.40 GiB/s readMayaNAS
Checkpointing — Llama 3 70B (two clients)
Entry 3.0-0139
32.42 GiB/s write · 21.20 GiB/s readMayaNAS
Training — 3D U-Net (3 simulated B200)
Entry 3.0-0136
92.03% accelerator utilizationMayaNAS
Training — 3D U-Net (4 simulated B200)
Entry 3.0-0141
93.31% accelerator utilizationMayaScale
Training — RetinaNet (72 simulated B200, one client)
Entry 3.0-0140
86.68% accelerator utilizationMayaScale
Inference cache — KVCache, Llama 3.1 8B (storage-only)
Entry 3.0-0138
524.65 tokens/s · 9.14 GiB/s readMayaNAS
Among MLPerf® Storage v3.0 submissions, MayaNAS is the only one to run Lustre directly on a cloud object-storage tier.

Across training, checkpointing, and inference, the same POSIX parallel filesystem — with its durable tier on economical cloud object storage — kept accelerators and the client network highly utilized, without staging copies in and out. A single client fed 72 simulated B200 accelerators for RetinaNet, a high per-client accelerator density.

MLPerf® Storage v3.0 Benchmark, Closed division. ZettaLane Systems entries 3.0-0136, 3.0-0137, 3.0-0138, 3.0-0139, 3.0-0140 and 3.0-0141. Result verified by MLCommons Association. Retrieved from mlcommons.org/benchmarks/storage/ on 7 September 2026. System under test: MayaNAS and MayaScale on standard Google Cloud virtual machines. Throughput in GiB/s; utilization is measured accelerator utilization on simulated NVIDIA B200 accelerators.

The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information.

Why MayaNAS for AI

Train Directly on Object Storage — No Staging

A standard POSIX filesystem over cloud object storage means training data and checkpoints stay in place. No copy in/out, no separate cache tier, no application changes.

Keeps Accelerators Highly Utilized

Sustained ~92–93% accelerator utilization on 3D U-Net in MLPerf® Storage v3.0 — the filesystem feeds GPUs fast enough to keep them working, not waiting on I/O.

Checkpoint at Scale

Up to 32.42 GiB/s checkpoint write and 21.20 GiB/s restore on Llama 3 70B — fast saves and reloads mean less GPU idle time around every checkpoint.

High Per-Client Accelerator Density

In MLPerf® Storage v3.0, a single client fed 72 simulated B200 accelerators for RetinaNet at 86.68% utilization — fewer client nodes to feed a large accelerator fleet.

Standards-Based — No Proprietary Client

Access over Lustre, pNFS Flex Files, and NFS using the standard, in-box Linux clients. Nothing proprietary to install on GPU nodes.

Sovereign — Your Cloud, Your Buckets, Your Keys

The filesystem, buckets, and encryption keys stay in your own cloud account, meeting data-residency and sovereignty requirements for AI data.

Benchmark Configuration

AttributeDetail
BenchmarkMLPerf® Storage v3.0, closed division (ZettaLane Systems submission).
WorkloadsCheckpointing (Llama 3 8B / 70B), training (3D U-Net, RetinaNet), inference KVCache (Llama 3.1 8B).
Storage — MayaNASLustre + OpenZFS; each OST an OpenZFS dataset over regional, Standard-class object storage; NVMe special vdev for metadata.
Storage — MayaScaleCompanion NVMe-over-TCP disaggregated block engine.
Clients / accessStandard Google Cloud VMs; standard open-source Lustre / NFS clients — no proprietary client software.
Object backendRegional, Standard-class Google Cloud Storage buckets; bulk data read/written directly to object storage.
AcceleratorsSimulated NVIDIA B200, per MLPerf® Storage methodology.
DeploymentGoogle Cloud tested; portable to Microsoft Azure and AWS via Terraform (open-lustre-cloud) and cloud marketplaces.
Programs & Partnerships
Google Cloud Select Technology Partner NVIDIA Inception Program
Available on Amazon Web Services, Microsoft Azure, and Google Cloud Marketplace, and via the open-lustre-cloud project on GitHub.
Figures reflect ZettaLane's submitted MLPerf® Storage v3.0 Benchmark configuration on Google Cloud, Closed division, entries 3.0-0136 to 3.0-0141; result verified by MLCommons Association. Full attribution and trademark notice on page 2. Contact sales@zettalane.com · www.zettalane.com