OpenAI has announced significant price reductions for its GPT-5.6 models, driven by systemic efficiency improvements and recursive self-optimization techniques. The company reports that the cost of GPT-5.4 level intelligence has dropped approximately 13x in four months, with GPT-5.6 Luna prices slashed by 80% and Terra by 20%.

  • OpenAI utilized GPT-5.6 Sol to autonomously rewrite production kernels in Triton and Gluon, reducing end-to-end serving costs by 20%.
  • Improved speculative decoding increased token-generation efficiency by over 15% through hundreds of automated experiments.
  • Agentic harness optimizations, including deferred discovery and prompt caching, reduced repetitive compute costs for tools like Codex.
  • A new "Sol Fast" tier offers 2.5x lower latency for double the standard price, while downstream agent workflows are expected to see roughly 10x lower costs.

These changes position OpenAI as highly cost-effective for non-finetuned intelligence, with observers noting that March's flagship performance is now available at a fraction of its previous token price.