Vera Rubin NVL72 — 7-Silicon Extreme Co-Design Stack
System Architecture

Vera Rubin NVL72 — 7-Silicon Extreme Co-Design Stack

Single integrated rack · Data flow from control plane to storage
Control Plane
Vera CPU
88 Olympus cores · Spatial Multithreading · LPDDR5X 1.2 TB/s
NVLink-C2C ↓ 1.8 TB/s (7× PCIe Gen 6)
Prefill Compute
Rubin GPU ×72
TSMC 3nm · 336B transistors · 288 GB HBM4 · 22 TB/s bandwidth · 50 PFLOPS/chip
Handoff to LPU after prefill complete
Decode Accelerator
Groq 3 LPU
~500 MB on-chip SRAM · 80 TB/s memory bandwidth · Deterministic execution — no branch prediction overhead
Sequential token generation at peak throughput
Storage & Network Control
BlueField-4 DPU + ConnectX-9 SuperNIC
CMX context memory offload · NVLink 6 switch · Spectrum-6 Ethernet · NIXL pre-staging to HBM
KV cache lifecycle management
Three-Step Inference Pipeline
01 Prefill Rubin GPU processes full prompt in parallel matrix ops
02 Decode Handoff Context transferred to Groq LPU for sequential generation
03 KV Offload BlueField-4 moves cache to CMX; Grove routes next request
35×
Performance-per-watt improvement over Blackwell on trillion-parameter model inference.
Not a generational upgrade — a category redefinition.