DPU automatically moves completed KV cache from HBM to G3.5 CMX storage, freeing GPU memory for next task.
3
Grove Routes Next Request
Load balancer identifies which physical node holds the user's cached context, routes the new request there.
4
NIXL Pre-stages Cache to HBM
NVIDIA Inference Transfer Library prefetches KV cache from CMX back to GPU HBM before inference begins — zero cold-start penalty.
5×
Real-time token generation throughput improvement from CMX + NIXL + Grove vs. conventional infrastructure.
Long-context agentic sessions no longer penalized by KV cache eviction.