AI Models & Platforms

CoreWeave Connects Hundreds of Rubin GPUs in One Multi-Rack Cluster

mm
Add Unite.AI to your preferred sources on Google

CoreWeave on September 16, 2026 announced the bring-up of multi-rack NVIDIA Vera Rubin NVL72 on CoreWeave Cloud, placing hundreds of NVIDIA Rubin GPUs into a single scale-out cluster for agentic AI. The company also introduced two new capabilities in CoreWeave AI Object Storage: cross-region write acceleration and a new Archive tier.

Chen Goldberg, executive vice president of product and engineering at CoreWeave, said CoreWeave was the first AI cloud provider to validate and bring up a Vera Rubin NVL72 and to demonstrate that the rack-scale architecture could operate as a reliable, high-performance cloud service. “With multi-rack Vera Rubin, we are connecting hundreds of Rubin GPUs as a single scale-out cluster,” Goldberg said, adding that for customers building agentic AI, that means greater scale, faster iteration, and higher productivity as models and agents continuously learn and improve.

Multi-Rack Architecture and Bring-Up

A single Vera Rubin NVL72 rack pairs 72 Rubin GPUs with 36 Vera CPUs, NVIDIA NVLink 6, ConnectX-9 SuperNICs, and BlueField-4 DPUs. With multi-rack Vera Rubin NVL72, CoreWeave unifies racks of hundreds of accelerators using NVIDIA Spectrum-X Ethernet networking into a single scale-out cluster. CoreWeave said the system delivers the capacity to train larger models, serve more demanding inference workloads, and run reinforcement learning at scale.

According to the company, running training and inference jobs across hundreds of Rubin GPUs matters for multi-step agentic workloads, which are sensitive to data-access latency because delays can compound across repeated model calls and tool use.

CoreWeave said it brings multiple racks up as a single system through automated rack life cycle control, system-level validation, and network scaling. CoreWeave Mission Control automates rack setup through the Rack LifeCycle Controller, with setup tasks that span hardware detection, firmware updates, validation, power, and cooling; Racky handles rack-level control while Valvey carries out cooling actions. Before a rack enters production, CoreWeave combines NVIDIA field diagnostics with full-rack workload testing and compares every result, and only racks that clear that bar as a system move into production.

On the network side, every Rubin GPU carries two ConnectX-9 SuperNICs, providing 1.6 Tb/s of scale-out connectivity per GPU across multiplane, multirail paths. CoreWeave said the design supports roughly 128,000 GPUs per rail in a non-blocking fabric, and that the modular topology allows racks to be added without redesigning the fabric at each expansion.

AI Object Storage Across Regions

CoreWeave AI Object Storage brings data closer to workloads through LOTA (Local Object Transport Accelerator), which applies managed caching on each CoreWeave Kubernetes Service node. CoreWeave said LOTA delivers reads at local NVMe speeds, reduces latency by 8x compared with reading from a traditional storage cluster, and provides up to 7 GB/s of throughput per GPU while scaling linearly as clusters grow.

The new cross-region write acceleration lets checkpoints be written locally within a region, at local latency, while the data migrates to a second remote region in the background. Because the application sees a single bucket, there are no code changes, and permissions and retention rules work the same regardless of which region a write came from. CoreWeave said the feature minimizes the training pause and removes the need to manually copy or move data between regions.

The second addition, Archive, is a lower-cost storage tier with no fee to retrieve data, no fee to delete it early, and no fee when reading from the tier. CoreWeave described it as built for data that teams would otherwise delete, such as a checkpoint from a run that almost worked, a dataset needed to reproduce a result, or a model version someone may ask about in six months.

“Our datasets span multiple regions, and we can’t afford to have our training schedule dictated by cross-region retrieval delays,” said Cécile Robert-Michon, director of internal infrastructure at Cohere. “CoreWeave AI Object Storage gives us a unified dataset footprint across regions with reads cached locally, so nothing waits on the network.”

NVIDIA Vera Rubin NVL72 Specifications

NVIDIA’s Vera Rubin NVL72 product page describes the system as a rack-scale agentic AI supercomputer built on the third-generation MGX NVL72 rack design and now in full production. NVIDIA’s published specifications list 20.7 TB of HBM4 GPU memory with 1,400 TB/s of bandwidth, 216 TB/s of NVLink bandwidth over sixth-generation NVLink, and 3,600 PFLOPS of sparse NVFP4 inference compute per rack. The rack’s 36 Vera CPUs supply 3,168 custom Olympus cores and up to 54 TB of LPDDR5X memory.

NVIDIA states that Vera Rubin NVL72 trains mixture-of-experts models with one-fourth the number of GPUs compared with its GB200 NVL72, and delivers inference at one-tenth the cost per million tokens for highly interactive, deep reasoning agentic AI, with up to 10x more tokens per megawatt. Footnotes on the page note that the inference figures are subject to change and that the training comparison is projected performance.

In the announcement, CoreWeave also pointed to its performance record, citing record MLPerf benchmark results in inference and training and its position as the only AI cloud to earn the top Platinum ranking in both SemiAnalysis ClusterMAX 1.0 and 2.0. The company also cited #1 rankings for inference speed and price-performance for Moonshot AI’s Kimi K2.6 and Kimi K2.7 Code models in independent inference benchmarking conducted by Artificial Analysis.

Theo Nash is an AI-generated specialist at Unite.AI, covering AI infrastructure, compute, and the hardware systems that power modern artificial intelligence. His work focuses on the technical foundations behind large-scale AI workloads, including data centers, accelerators, networking, and the software stacks that tie them together.

With an analytical and engineering-driven perspective, Theo examines how advances in GPUs, custom silicon, memory architectures, and distributed systems enable new generations of AI models. He pays particular attention to performance trade-offs, energy efficiency, scalability, and the practical constraints that shape real-world deployment of AI infrastructure.

Articles authored by Theo Nash are AI-generated and reviewed by Unite.AI’s editorial team to ensure technical accuracy, clarity, and responsible coverage of the rapidly evolving AI compute landscape.