< Back to all clusters
[TECHNOLOGY] · United States, Taiwan · 52 sources

started · updated

OpenAI unveils Jalapeno, a custom AI inference chip

OpenAI has unveiled performance results for Jalapeno, its first custom AI inference chip co-developed with Broadcom. Designed specifically for large language model (LLM) workloads, the ASIC aims to reduce reliance on Nvidia hardware by providing a more efficient alternative for running models rather than training them.

According to benchmarks from SemiAnalysis using the InferenceX platform, Jalapeno demonstrated significant efficiency gains across several models, including GPT-OSS 120B, DeepSeek R1, and Kimi K2.5. The chip reportedly delivers 1.5 to 1.9 times more AI work per watt at peak throughput and reduces end-to-end latency by 1.7 to 3.6 times compared to existing systems like Nvidia’s GB200 and GB300. For highly interactive workloads, performance advantages reached up to 4.1 times.

Technical details indicate the chip utilizes TSMC’s N3P and N3E processes, incorporates HBM4 memory, and features a custom software stack including the Gluon programming language. OpenAI plans to begin small-scale deployment of the chip within its own infrastructure by the end of 2026, with a full production ramp expected between 2027 and 2028.

Entities

Broadcom · Nvidia · OpenAI · Richard Ho · Sam Altman · SemiAnalysis · TSMC

Claims

What the coverage asserts, and how many sources carry each claim.

Sources

about 15 hours ago
about 14 hours ago
about 14 hours ago
about 19 hours ago
about 15 hours ago
about 14 hours ago