Vera Rubin NVL72 — 7-Silicon Extreme Co-Design Stack
System Architecture
Vera Rubin NVL72 — 7-Silicon Extreme Co-Design Stack
Single integrated rack · Data flow from control plane to storage
Control Plane
Vera CPU
88 Olympus cores · Spatial Multithreading · LPDDR5X 1.2 TB/s
NVLink-C2C ↓ 1.8 TB/s (7× PCIe Gen 6)
Prefill Compute
Rubin GPU ×72
TSMC 3nm · 336B transistors · 288 GB HBM4 · 22 TB/s bandwidth · 50 PFLOPS/chip
Handoff to LPU after prefill complete
Decode Accelerator
Groq 3 LPU
~500 MB on-chip SRAM · 80 TB/s memory bandwidth · Deterministic execution — no branch prediction overhead
Sequential token generation at peak throughput
Storage & Network Control
BlueField-4 DPU + ConnectX-9 SuperNIC
CMX context memory offload · NVLink 6 switch · Spectrum-6 Ethernet · NIXL pre-staging to HBM
KV cache lifecycle management
Three-Step Inference Pipeline
01
Prefill
Rubin GPU processes full prompt in parallel matrix ops
›
02
Decode Handoff
Context transferred to Groq LPU for sequential generation
›
03
KV Offload
BlueField-4 moves cache to CMX; Grove routes next request
35×
Performance-per-watt improvement
over Blackwell on trillion-parameter model inference.
Not a generational upgrade — a category redefinition.