started · updated
AI SSDs set to become core tier in large‑language‑model inference
The rapid growth of large‑language‑model (LLM) parameters and longer context windows is hitting two limits: memory capacity and I/O bandwidth. As GPU HBM and system DRAM become insufficient for KV‑Cache and expert weights, SSDs are moving from a static storage role into the real‑time data path of inference.
Early implementations such as the Mooncake architecture separate KV‑Cache onto distributed CPU, DRAM and SSD resources, while Nvidia’s 2026 CMX platform adds an Ethernet‑connected flash layer managed by BlueField‑4. These designs treat SSDs as a pod‑level context memory between GPU memory and shared storage.
Product vendors are responding with two AI‑SSD approaches. The “AI load‑strengthening” class, exemplified by Inspur’s Dongting N3X series and Huawei’s OceanDisk line, uses low‑latency SLC NAND and XL‑Flash to cut access latency to about one‑third of traditional TLC SSDs and boost write throughput three‑fold, extending endurance by 17‑33 ×. Huawei’s OceanDisk EX 560 targets ultra‑high random write performance, while the LC 560 offers up to 245 TB per drive for massive model files.
A second class seeks to integrate SSDs more tightly with compute, adding firmware and OS‑level coordination to virtualize NAND alongside DRAM/VRAM, allowing portions of model weights and KV‑Cache to reside on flash during execution. Both routes aim to alleviate the “capacity wall” and “I/O wall” that constrain LLM inference at scale.
Entities
CMX platform · Huawei Technologies Co., Ltd. · Inspur Group · Mooncake architecture · Nvidia Corporation