AI RACE— The AI Race
Business

NVIDIA and CoreWeave Deploy Vera Rubin NVL72 and Vera CPU to Accelerate Agentic AI Production

CoreWeave launches NVIDIA’s Vera Rubin NVL72 GPUs and Vera CPUs on its cloud, delivering up to 4.8× token throughput gains and 3× faster sandbox startups for agentic AI workloads.

09/30/2026, 22:00
NVIDIA và CoreWeave ra mắt hạ tầng “Agentic AI” mới: Vera Rubin NVL72 và CPU Vera trên đám mây

1. Announcement and Key Players

On September 30, 2026, NVIDIA and cloud‑compute specialist CoreWeave announced that their next‑generation AI infrastructure is now in production. At the CoreWeave Fully Connected event in San Francisco, the companies unveiled the NVIDIA Vera Rubin NVL72 system—paired with Spectrum‑X 102.4 Tbps Ethernet—and the first NVIDIA Vera CPU, purpose‑built for agentic AI. Cognition, the applied‑AI lab behind the Devin AI software engineer, became the inaugural customer to run production workloads on the Vera Rubin platform.

2. Technical Details and Early Results

  • Hardware rollout: CoreWeave received its first Vera Rubin NVL72 production racks and made them available through CoreWeave Cloud, becoming one of the first cloud providers to expose the platform to customers. The deployment includes 128 Vera CPUs delivering 11,264 cores in a single rack, enough for more than 11,000 concurrent isolated environments.
  • Performance gains: Cognition benchmarked Vera Rubin NVL72 against a GB200 NVL72 baseline on a real‑world software‑engineering workload (sampled from FrontierCode). The tests showed up to 4.8× higher token throughput for SWE‑2 inference, translating into faster real‑time code generation and more responsive multi‑step reasoning.
  • Sandbox acceleration: Using CoreWeave Sandboxes, the Vera CPU achieved more than 3× faster agent sandbox startup times and a 1.7× overall performance gain on the Terminal‑Bench suite. Each sandbox runs in a hardware‑isolated environment, with Spectrum‑X switches and BlueField‑4 DPUs ensuring low‑latency, secure communication.
  • Software stack: The AI factory runs on NVIDIA Dynamo, an open‑source inference framework that also powers RL Rollouts—a private‑preview feature that loads new checkpoints into live deployments without redeployment.
  • CoreWeave Forge: To close the loop between production and training, CoreWeave introduced Forge, a unified environment that bundles Weights & Biases, OpenPipe post‑training tools, and the marimo notebook project. New services include:
  • CoreWeave ARIA – an AI‑assisted research and coding assistant that analyzes runs, proposes experiments, and logs changes in GitHub.
  • CoreWeave Agent Lens – observability for production agents, improving failure detection by 20 % and halving fix costs.
  • CoreWeave Sandboxes – now GA, offering isolated CPU or GPU execution for agents, tool calls, RL, and evaluations on either serverless or existing training infrastructure. Serverless supervised fine‑tuning and RL run 1.4× faster at 40 % lower cost than self‑managed setups.
  • Early adopters: Beyond Cognition, Canva, Capital One, MasterClass, and healthcare provider Ennoble Care have begun building on Forge. Ennoble Care will use reserved NVIDIA RTX PRO 6000 GPUs on CoreWeave Kubernetes Service for clinical documentation and decision‑support agents.

Quotes:

“NVIDIA accelerated computing delivers value across generations,” said Ian Buck, VP of hyperscale and HPC at NVIDIA. “CoreWeave’s V100 GPUs are still running customer workloads nearly a decade after Volta launched, even as CoreWeave brings Vera Rubin NVL72 into production.”
“Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship,” noted Silas Alberti, founding team member at Cognition. “Having all of it on one platform, with NVIDIA and CoreWeave engineers who work the hard problems alongside ours, matters more to us than any single spec.”

3. Industry Implications

The launch underscores a broader shift toward agentic AI, where large language models act as autonomous software engineers, clinicians, or customer‑service bots. By delivering a full‑stack, multi‑generation platform that spans from training to real‑time inference, NVIDIA and CoreWeave aim to reduce the fragmentation that has historically slowed the move from prototype to production. Their claim of a decade‑long ROI on V100 GPUs highlights the longevity of NVIDIA’s hardware roadmap—a competitive advantage as rivals such as AWS, Azure, and Google Cloud race to offer their own AI‑optimized CPUs and networking.

The integration of high‑throughput Spectrum‑X Ethernet, BlueField‑4 DPUs, and purpose‑built CPUs positions the partnership to meet the dual pressures of low‑latency serving and massive parallel training that agentic workloads demand. As more enterprises—particularly in healthcare, finance, and creative industries—adopt autonomous AI agents, the ability to spin up thousands of isolated sandbox environments quickly could become a decisive factor in cloud selection.

---

All figures and quotes are taken from the NVIDIA blog post “From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI” dated September 30, 2026.

◗ Sources

NVIDIA Blog09/30

Related stories