Vera CPU architecture and performance claims

Nvidia detailed the specifications of its new Vera CPU, the central processor of the Vera Rubin platform built for the agentic AI era. The chip is built around Nvidia's custom Olympus core, with 88 custom cores, 176 hardware threads, up to 1.2 terabits per second of LPDDR5X memory bandwidth, 164 megabytes of unified L3 cache and up to 1.8 TB/s of coherent CPU-GPU bandwidth via NVLink-C2C. Nvidia described the design target as a "max single-threaded CPU at scale" suited to the latency-sensitive, branch-heavy behaviour of agent workloads. The company claimed twice the single-threaded performance, three times the core-to-core bandwidth and 40% lower memory latency versus competing chiplet-based designs.
Global rollout and partner ecosystem
Nvidia said the Vera Rubin NVL72 platform is in production at more than 350 factory sites across 30 countries, backed by over 300 partners. Cloud and infrastructure providers named as deployment partners include CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. Partners including CoreWeave, Google Cloud and DeepInfra reported benchmark gains in throughput, inference efficiency and AI agent orchestration on the Vera Rubin platform and Vera CPU, according to the company.
Networking, rack and cooling design
The Vera Rubin system combines seven co-designed chips and five rack trays into a single integrated platform. Nvidia said its sixth-generation NVLink delivers more than twice the throughput, three times lower latency and ten times higher packet rates than conventional Ethernet, with Spectrum-X Ethernet positioned to enhance large-scale networking. The new rack-scale design removes cables, fans and hoses from the compute tray, reducing assembly time from hours to about a minute, and uses a liquid cooling system aimed at lowering water consumption in AI factories.
Europe expansion through Microsoft and Mistral
Nvidia said Vera Rubin will serve as the computing foundation for an expanded partnership with Microsoft and Mistral in Europe. The platform is expected to power Microsoft's next-generation European AI infrastructure and Mistral Compute, supporting sovereign AI deployments across public cloud, private cloud and customer-controlled environments using tens of thousands of GPUs.
Competitive positioning against AMD and Intel
Nvidia framed Vera Rubin as a full-stack system designed to improve performance per watt and reduce token costs, intensifying its data centre push against rivals AMD and Intel. The company argued that the rise of agentic AI requires treating the CPU, network and system architecture as a single interdependent design problem rather than a collection of best-of-breed parts. Nvidia presented the platform as a response to the CPU becoming a bottleneck in the AI factory rather than a conventional server CPU refresh.
Follow-up signals
The next verifiable milestones for the rollout are production shipments of Vera Rubin NVL72 systems to named cloud and infrastructure partners and the build-out of Microsoft and Mistral's European AI infrastructure using the platform, with partner benchmark disclosures expected to continue as deployments scale.
Share this article




