< Back to all clusters
[TECHNOLOGY] · 3 sources

started · updated

NVIDIA expands AI inference capabilities with Groq 3 LPX and d-Matrix partnership

NVIDIA has entered full production of the Groq 3 LPX, a specialized inference accelerator designed for the Vera Rubin platform. The technology focuses on low-latency token generation for agentic AI workloads, working alongside Vera Rubin NVL72 GPUs to separate latency-sensitive generation from context processing and prefill tasks.

Nebius is expected to be the first adopter of the Groq 3 LPX through its production inference platform. In a related strategic move, d-Matrix has announced a multi-year collaboration with NVIDIA to integrate its Raptor inference XPUs into NVIDIA MGX racks via NVLink Fusion.

The d-Matrix partnership aims to provide a heterogeneous architecture that splits workloads between GPUs and XPUs. The Raptor XPUs utilize a 3D-DRAM architecture, offering a different technical approach to inference compared to the SRAM-based Groq 3 LPX, targeting ultra-low latency for hyperscalers and AI labs.

Entities

Groq · Nebius · Nvidia · Vera Rubin · d-Matrix