OpenAI has introduced GPT-6 Astra as its most intelligent model to date while outlining its full-stack compute strategy, which includes the new Jalapeño inference chip. The company reports that its custom hardware delivers 1.5 to 1.9 times higher peak token throughput per watt compared to commercial systems in InferenceX tests.
- GPT-6 Astra is described as state-of-the-art in computer use, browsing, software engineering, cybersecurity, science, and professional work.
- The Jalapeño inference chip reduces end-to-end latency by 1.7 to 3.6 times compared to tested commercial systems.
- GPT-5.6 Sol improved production serving software, reducing end-to-end serving costs by 20% and increasing token-generation efficiency by more than 15 percent.
- OpenAI's research organization now utilizes 3.1 agent-workdays of effort for every workday of human labor.
These advancements aim to lower the cost and increase the speed of AI tasks, allowing customers to achieve more with fewer attempts while supporting OpenAI's growth through diversified revenue streams.