Two API settings tripled GPT-5.6 scores on ARC-AGI-3
Enabling two specific API settings for GPT-5.6 resulted in a tripling of its performance scores on the ARC-AGI-3 benchmark.
Enabling two specific API settings for GPT-5.6 resulted in a tripling of its performance scores on the ARC-AGI-3 benchmark.
A comparison of DeepSeek-V4 Flash 0731 and GPT-5.6 Luna on the DeepSWE benchmark reveals that while GPT-5.6 Luna is the stronger engineer, a cascade strategy using both models achieves higher accuracy at lower cost.
Alibaba has announced Qwen3.8-Max as its new flagship model, describing it as a 2.4T-parameter sparse architecture focused on long-horizon agentic work, coding, and multimodal reasoning. The company confirmed that open weights for Qwen3.8-Max will be released next week alongside the open-weight Qwen3.8-27B variant.
DeepSeek has updated its V4 Flash model, releasing a new version on 2026-07-31 that shows significant performance gains over the previous preview. The update introduces results for several new benchmarks while improving scores on existing ones.
A comparison of Kimi K3 and GPT-5.6 Sol on the DeepSWE benchmark reveals that while GPT-5.6 Sol leads in single-shot quality (72.7% vs 68.5%), Kimi K3 achieves higher pass@k scores for k > 1 at a significantly lower cost.
Anthropic has launched the Claude Opus 5 model, triggering scrutiny of its performance on coding and general capability metrics alongside renewed debate about frontier model evaluation.
Anthropic has released Claude Opus 5, a new frontier model that has triggered significant discussion regarding benchmark scores and practical performance. Independent evaluations from Epoch AI report an Epoch Capabilities Index (ECI) of 159 and a SWE-ECI of 161 for the new model.
Kimi K3, an open-weight model from Moonshot AI, achieves performance comparable to Anthropic's Claude Fable 5 on the DeepSWE benchmark while costing significantly less per task. In a July 16, 2026 evaluation, Kimi K3 reached 68.5% pass@1 compared to Fable 5's 69.9%, but surpassed it in multi-attempt scenarios and offered superior cost efficiency.
Sierra Research released τ-bench 1.0.1, which corrects the `banking_knowledge` task grading scheme to stop penalizing agents for prudent verification reads and fixes several data inconsistencies. The update ensures that re-grading leaderboard trajectories only increases scores, with no previously passing simulations failing.
On July 16th, Moonshot AI released its latest flagship model, Kimi K3, a 2.8 trillion parameter Mixture of Experts (MoE) architecture. The company has promised to release the model's weights on July 27th.
Moonshot has released Kimi K3, an open-weight model that has triggered a reassessment of how close Chinese models are to the frontier. The release is characterized by strong performance in coding, agentic tasks, and long-horizon knowledge work.
Moonshot AI has introduced Kimi K3, a new frontier-class open-weights model featuring 2.8 trillion total parameters and native multimodal input capabilities. Officially positioned as "Open Frontier Intelligence," the model supports a 1 million-token context window and utilizes novel architectural components like Kimi Delta Attention (KDA) and Attention Residuals to enhance efficiency.
Moonshot AI has launched Kimi K3, a frontier-class open-weights model with 2.8 trillion parameters and native multimodal input. The model features Kimi Delta Attention (KDA) for faster decoding in long contexts and is positioned for long-horizon agentic coding workflows.
Chinese AI lab Moonshot AI has announced Kimi K3, describing it as their most capable model to date with 2.8 trillion parameters. The model is currently available via the website and API, with an open weight release promised by July 27, 2026.
DharmaOCR achieves higher extraction quality and stability than Mistral OCR4 and Unlimited-OCR on Brazilian Portuguese through domain-specific training.
The article presents the ARC-AGI benchmark results for the DeepSeek V4 Flash 0731 model. It links to the official results page hosted by the ARC Prize organization.
DeepSeek has released DeepSeek-V4-Flash-0731, an updated version of its smaller "Flash" model that surpasses the larger DeepSeek-V4-Pro on independent benchmarks. The release utilizes a new fine-tuning process while maintaining the original Mixture-of-Experts architecture with 284 billion total parameters.
Artificial Analysis has updated its agentic index to rank Qwen 3.8 Max as the best overall model, placing it ahead of Opus 5.
The authors identify that plateauing scores on the SciCode benchmark stem from significant defects in the evaluation instrument rather than limitations in model capability. A domain-expert audit of 65 test problems uncovered 263 defects, with 192 causing correct solutions to be wrongly rejected due to issues like non-reproducible gold answers and over-tight tolerances.
Qwen 3.8 Max achieves a score of 1588 on the Debate Benchmark, improving upon Qwen 3.7 Max's score of 1462. This benchmark evaluates how well large language models hold an argument under adversarial, multi-turn opposition across various topics.