OpenAI has released initial performance results for Jalapeño, its first custom inference chip, demonstrating significant gains in speed and power efficiency compared to existing commercial systems. The company states that the new architecture allows for higher throughput and lower latency simultaneously, addressing a common trade-off in current hardware.
- Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput than comparison systems.
- End-to-end latency was reduced by 1.7 to 3.6 times across tested models including GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T.
- For highly interactive workloads, the chip provided 2.1 to 4.1 times higher performance.
- The system achieved a sustained power draw at or below 550 watts despite a 700-watt rating.
These improvements aim to make increasingly capable AI more affordable and broadly available by lowering the cost of delivering successful results while supporting faster iteration for agentic workloads.