Nemotron 3 Ultra Architecture
NVIDIA × NAVER AI · Model Architecture
Nemotron 3 Ultra: Hybrid Architecture Breakdown
How Transformer-Mamba MoE design cuts inference cost while extending context to 1M tokens
Standard Dense LLM
Dense Transformer
Active Params / Token100% (all weights)
Sequence ComplexityO(N²) quadratic
Max Context~128K tokens
Inference CostBaseline
VS
Nemotron 3 Ultra
Hybrid Mamba-MoE
Active Params / Token10% (55B of 550B)
Sequence ComplexityO(N) linear (Mamba)
Max Context1,000,000 tokens
Inference Cost−30% vs. baseline
Nemotron 3 Ultra — Parameter Routing
550B
Total parameters — only 55B (10%) activated per inference via LatentMoE router
〰️
Mamba Layers
Recurrent sequence compression. Processes long inputs in linear time — 100× less memory overhead than full attention on 1M-token context.
🔀
LatentMoE Router
Dynamically selects the best expert weight subset for each input token. 550B total, 55B active — eliminates wasted compute on irrelevant parameters.
Inference Speed
Faster token generation vs. equivalent dense model at same parameter count
−30%
Compute Cost
Lower operational spend per token via NVFP4 precision and MoE sparsity
1M
Token Context
Processes a 2,000-page report in a single inference pass without truncation
Source: Nemotron 3 Ultra Technical Report — NVIDIA AI Research Lab · NVFP4 Precision, LatentMoE Routing Architecture