CoreWeave is now offering the Nvidia Vera Rubin NVL72 system on its cloud platform.

Just a couple of months after the company revealed that it had installed and was operating what it claimed was the first fully working Vera Rubin NVL72 rack, the latest generation of Nvidia hardware is now available to customers.

The company announced the news at its Fully Connected conference in San Francisco.

Vera Rubin CoreWeave
– CoreWeave

The first customer to adopt the platform is Cognition, which CoreWeave says began using the Vera Rubin systems in early September.

“Bringing up Nvidia Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU generations,” said Chen Goldberg, executive vice president of product & engineering at CoreWeave. “With customers like Cognition, that investment shows up in the ability to get production workloads running within days. When it comes to agentic tasks, long contexts, repeated model calls, and thousands of concurrent tasks put pressure on the entire platform. Our job is to make compute, networking, and software work as a single system, so customers can build increasingly complex agents without taking on the infrastructure complexity themselves.”

According to Silas Alberti, SVP research & founding team at Cognition, the company has seen up to a 4.8x increase in total token throughput for its SWE-2 inference workloads on the new generation.

The Rubin platform comprises six chips in total – the NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, and Spectrum-6 Ethernet Switch, in addition to the Vera CPU and Rubin GPU. Vera is the successor to Nvidia’s Grace CPU, while Rubin will succeed Blackwell GPUs, with Nvidia claiming Rubin will be capable of achieving 5x and 3.5x inference and training performance, respectively, compared to Blackwell.

The NVL72 system - Nvidia's rack-scale offering that comprises 36 CPUs and 72 GPUs - differs from the Vera Rubin generation in that it is 100 percent liquid-cooled and features cable‑free modular tray designs, which Nvidia claims will allow installation times to be reduced from two hours to five minutes.

In addition to the availability of Vera Rubin systems, CoreWeave is also planning to offer the Vera CPU as a standalone offering. CoreWeave has traditionally been focused on offering access to GPUs.

According to Corey Sanders, SVP of product at CoreWeave, this is still in the early stages, but the company is expecting some customers to begin testing it in the coming weeks, and a few have had early access.

Vera is designed for AI agents and to meet the needs of new workloads. The CPU will run on CoreWeave as a bare-metal offering.

Deployed at rack-scale, CoreWeave will offer 128 Vera CPUs and 11,264 cores in a single rack, with BlueField-4 DPUs and Spectrum-X Ethernet switching delivering secure, high-performance connectivity between Vera nodes.

“General-purpose infrastructure bottlenecks agentic AI; Vera is the first CPU explicitly designed to accelerate it,” said Goldberg. “Our platform natively enables Vera with products like CoreWeave Sandboxes out of the box. Teams can instantly spin up thousands of isolated environments, removing operational friction and accelerating the entire AI loop on day one.”

Also revealed at the Fully Connected conference was a new customer contract with Ennoble Care and the launch of CoreWeave Forge.

The home-based primary, palliative, and hospice care company is deploying Nvidia RTX Pro 6000 Blackwell Server Edition nodes of CoreWeave Kubernetes Service to enable the company to bring AI into everyday clinical workloads, including summarization, documentation, and clinical decision support.

“We've built a full-stack, ONC-certified EMR purpose-built for home-based primary care, and we're now developing multiple AI agents on top of it, both to augment clinical delivery and to automate back-office functions," said Jonathan Taylor, CTO, Ennoble Care. "CoreWeave is the right solution to help maintain our clinical standards. It gives us reserved capacity we can count on, a Kubernetes environment our team can move into quickly, and engineers who answer the phone. That combination allows us to scale the AI capabilities already built into our care delivery workflows and proprietary electronic medical record system, extending them to tens of thousands of additional patients as we continue to grow.”

“Clinical inference is one of the most demanding places AI can run. The workload is continuous, the latency budget is short, and the compliance requirements are absolute,” added Jon Jones, chief revenue officer at CoreWeave. “That is what The Essential Cloud for AI means in practice. It’s a cloud ready for the workloads that matter most, in the industries where precision is everything.”

CoreWeave Forge, meanwhile, is a development layer that has been added to the cloud platform. According to the company, this enables teams to build, improve, and evaluate AI models and agents in a single environment. The offering is now generally available.