
Nvidia
AI Hardware · Semiconductors · Data Center Infrastructure
Nvidia unveils Rubin GPU architecture and Vera CPU for agentic AI era
July 21, 2026
The real contest now is efficiency per watt, not raw chip speed, as Nvidia pushes deeper into the CPU market it once left to others.
- Nvidia disclosed its next-generation Rubin GPU architecture and Vera CPU, confirming the Vera Rubin platform has reached full production and setting a new MoE pre-training record on its GB300 NVL72 system.
- The Rubin GPU packs 336 billion transistors, 224 streaming multiprocessors and 288GB of HBM4 memory, purpose-built for agentic AI workloads that reason, plan and call tools across many sequential steps.
- Vera, Nvidia's custom CPU built around its Olympus core, posted 1.8x to 1.9x faster agentic performance than a 128-core AMD EPYC Turin chip, marking a deeper push into a market Nvidia previously ceded to others.
- Nvidia's GB300 NVL72 hit 1,648 teraflops per GPU pre-training the DeepSeek-V3 671B mixture-of-experts model, roughly triple the per-GPU throughput of the prior GB200 NVL72 generation.
- Partner CoreWeave measured 10x more tokens per megawatt running DeepSeek-R1 on early Vera Rubin racks compared with the previous Blackwell-based GB200 NVL72 system.
- Vera Rubin systems are already shipping to OpenAI, CoreWeave, Google Cloud, Microsoft Azure, Meta and Dell, with OpenAI expected to deploy the platform at scale in the third quarter.
- The disclosures landed just ahead of AMD's Advancing AI event in San Francisco, even as Nvidia's stock has lagged a broader chip-sector rally this year despite forecast revenue growth of 82%.
- By co-designing GPU, CPU, networking and rack-scale cooling together, Nvidia is reframing the AI infrastructure race around performance-per-watt and token cost rather than chip speed alone.