Enabling two specific API settings for GPT-5.6 resulted in a tripling of its performance scores on the ARC-AGI-3 benchmark.
The improvements were achieved by retaining reasoning capabilities and enabling compaction, which also boosted efficiency.
Enabling two specific API settings for GPT-5.6 resulted in a tripling of its performance scores on the ARC-AGI-3 benchmark.
The improvements were achieved by retaining reasoning capabilities and enabling compaction, which also boosted efficiency.
OpenAI has initiated a limited preview of the GPT-5.6 model series, introducing three distinct variants: Sol as the flagship, Terra for balanced everyday work, and Luna for fast, affordable tasks. The company plans to make these models generally available in the coming weeks following this initial phase with trusted partners.
Four weeks after GPT-5.1, OpenAI has released GPT-5.2, a new model series designed for professional knowledge work that matches or exceeds Google's Gemini 3 Pro in several key benchmarks.
OpenAI has released GPT-4.5 as a research preview, describing it as the company's largest and best chat model to date. The model is initially available to ChatGPT Pro users and developers, with access expanding to Plus and Team users starting next week.
OpenAI is providing 100,000 academic researchers with free access to its most advanced AI models. This initiative aims to accelerate scientific research, collaboration, and discovery within the academic community.
A comparison of Kimi K3 and GPT-5.6 Sol on the DeepSWE benchmark reveals that while GPT-5.6 Sol leads in single-shot quality (72.7% vs 68.5%), Kimi K3 achieves higher pass@k scores for k > 1 at a significantly lower cost.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy