OpenAI has released the first measured performance results for Jalapeño, its custom inference chip, alongside a detailed overview of its integrated "abundant intelligence" compute strategy. The company outlines how it combines data centers, chips, models, and software to improve efficiency and economics across its entire system.
- Jalapeño delivered more peak throughput per kilowatt and lower token latency than commercial systems on the InferenceX benchmark using GPT-OSS 120B.
- The chip performed strongly on DeepSeek R1 and Kimi K2, demonstrating gains across different model families.
- OpenAI manages a diverse hardware portfolio including Microsoft, NVIDIA, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank.
- GPT-5.6 Sol with max reasoning reached a new high on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than a leading model.
This approach aims to provide greater control over model serving economics and create a credible first-party path for silicon, ensuring the strongest performance per dollar as technology evolves.