< Back to all clusters
[TECHNOLOGY] · 3 sources

started · updated

NVIDIA revives Rubin CPX accelerator with HBM4 memory redesign

NVIDIA has reportedly revived its Rubin CPX project, a specialized AI accelerator previously placed on hold. According to supply chain analyst Ming-Chi Kuo, the project has undergone a significant architectural redesign, specifically regarding its memory configuration.

The original design planned for 128 GB of GDDR7 memory, but the revised version will utilize 168 GB of HBM4 technology. This shift to HBM4 necessitates advanced packaging solutions, likely employing TSMC’s CoWoS-S or CoWoS-L technologies to integrate the high-performance memory stacks.

The Rubin CPX is not intended as a general-purpose GPU. Instead, it is a specialized component designed to handle the prefill phase of Large Language Models (LLMs) and manage KV cache. This allows NVIDIA to separate the inference process: the Rubin CPX handles the initial context processing and prefill, while standard Rubin GPUs manage the subsequent token decoding phase. This distinction is critical for managing trillion-parameter models with extensive context windows.

A single rack tray containing eight Rubin CPX units can manage up to 1.34 TB of long-context prefill and KV cache. The hardware is expected to offer 30 PetaFLOPS of computing performance in NVFP4 format and includes integrated NVENC encoders and NVDEC decoders to facilitate heavy multimedia workloads.

Entities

Ming-Chi Kuo · Nvidia · TSMC