started · updated
Intel unveils Crescent Island GPU for agentic AI inference
Intel has unveiled details regarding Crescent Island, a discrete data-center GPU specifically engineered for agentic AI inference. Utilizing the Xe3P architecture, the hardware prioritizes tokens per watt and high memory capacity over traditional floating-point throughput to handle the latency and memory demands of multi-step AI agents.
The GPU features up to 32 Xe cores and 256 Xe Matrix Extension engines. Notably, Intel has opted to use LPDDR5X memory instead of high-power HBM to reduce costs and power consumption. While Intel’s reference PCIe card will include 160 GB of memory, the architecture allows partners to build configurations with up to 480 GB.
Designed for both workstations and data centers, Crescent Island supports a wide range of data types from FP4 to FP64. It is intended to support frameworks such as vLLM and SGLang, targeting providers of “tokens-as-a-service” and systems running large language or multimodal models. The product is expected to be available to customers in the second half of 2026.