NVIDIA achieves up to 25× performance‑per‑watt boost for frontier AI models
NVIDIA highlights performance‑per‑watt as the key metric for AI infrastructure, arguing that token output within a fixed power budget determines an AI factory’s revenue and profitability. The company’s Blackwell NVL72 platform, built on the Vera Rubin system, delivers the highest performance per watt at rack scale, claiming up to 25 times the efficiency of the previous Hopper generation for leading open‑source models such as DeepSeek V4 Pro, and 10–20 times for other workloads.
The efficiency gains stem from a full‑stack co‑design that integrates silicon, interconnects, cooling, and software. NVIDIA’s NVLink Switch, purpose‑built for AI workloads, and in‑network computing features like SHARP reduce GPU load. Software tools—including Dynamo, TensorRT LLM, SGLang, vLLM, and power‑steering capabilities in the DSX platform—optimize quantization, expert parallelism, and cache management, further improving token throughput while minimizing power loss to cooling. The company reports that on DeepSeek V4, performance per watt improved fivefold in a single month, illustrating rapid gains in energy‑efficient AI processing.