started · updated
OpenAI unveils Jalapeno, a custom AI inference chip
OpenAI has unveiled performance results for Jalapeno, its first custom AI inference chip co-developed with Broadcom. Designed specifically for large language model (LLM) workloads, the ASIC aims to reduce reliance on Nvidia hardware by providing a more efficient alternative for running models rather than training them.
According to benchmarks from SemiAnalysis using the InferenceX platform, Jalapeno demonstrated significant efficiency gains across several models, including GPT-OSS 120B, DeepSeek R1, and Kimi K2.5. The chip reportedly delivers 1.5 to 1.9 times more AI work per watt at peak throughput and reduces end-to-end latency by 1.7 to 3.6 times compared to existing systems like Nvidia’s GB200 and GB300. For highly interactive workloads, performance advantages reached up to 4.1 times.
Technical details indicate the chip utilizes TSMC’s N3P and N3E processes, incorporates HBM4 memory, and features a custom software stack including the Gluon programming language. OpenAI plans to begin small-scale deployment of the chip within its own infrastructure by the end of 2026, with a full production ramp expected between 2027 and 2028.
Entities
Broadcom · Nvidia · OpenAI · Richard Ho · Sam Altman · SemiAnalysis · TSMC
Claims
What the coverage asserts, and how many sources carry each claim.
- [● 30 SOURCES] Jalapeno is the first of a multi-generational silicon platform being developed by OpenAI. desitalkchicago.com · fourweekmba.com · www.ithome.com · techxeber.az · www.businesstoday.in · +24 more
- [● 25 SOURCES] The Jalapeno chip was co-developed by OpenAI and Broadcom. fourweekmba.com · tech-noisy.com · diginoy.com · www.ithome.com · www.watch.impress.co.jp · +20 more
- [● 25 SOURCES] Performance results were measured using the InferenceX benchmark from SemiAnalysis. desitalkchicago.com · fourweekmba.com · tech-noisy.com · diginoy.com · www.ithome.com · +20 more
- [● 23 SOURCES] The chip achieves 1.7 to 3.6 times lower end-to-end latency than comparison systems. desitalkchicago.com · fourweekmba.com · tech-noisy.com · diginoy.com · www.ithome.com · +18 more
- [● 23 SOURCES] Jalapeno delivers 1.5 to 1.9 times more AI work per watt at peak throughput compared to existing systems. desitalkchicago.com · fourweekmba.com · tech-noisy.com · diginoy.com · www.ithome.com · +18 more
- [● 21 SOURCES] Jalapeno delivers both higher throughput and lower latency within a single architecture. desitalkchicago.com · fourweekmba.com · tech-noisy.com · diginoy.com · www.ithome.com · +16 more
- [● 19 SOURCES] OpenAI plans to begin small-scale deployment of the chip by the end of 2026, with a production ramp expected in 2027–2028. fourweekmba.com · diginoy.com · www.ithome.com · www.watch.impress.co.jp · techxeber.az · +14 more
- [● 23 SOURCES] The Jalapeno chip is an ASIC designed specifically for AI inference rather than model training. fourweekmba.com · diginoy.com · www.ithome.com · www.servethehome.com · techxeber.az · +18 more