At the 37th Hot Chips conference, OpenAI presented benchmark details for its custom inference chip, Jalapeño, claiming superior performance per watt compared to NVIDIA's GB200 and GB300 systems. The chip delivers 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency, with deployment into OpenAI's infrastructure scheduled by year-end.
- Jalapeño achieved 2.1–4.1× higher performance for highly interactive workloads while staying at or below 550W despite a 700W rating.
- The architecture reduces the traditional throughput/latency tradeoff, performing well even without aggressive prefill/decode disaggregation or speculative decoding.
- OpenAI used GPT-Astra + Codex to optimize low-level kernels, bringing three open-weight models to high performance in two months with 1.5–1.8× speedups over human-written code.
- Other announcements included Microsoft's AutoSaddler harness improvements, Perplexity's local-first Portable Computer on NVIDIA DGX Spark, and new insights into agent memory systems.