Kimi K3 has achieved the top position on AfterQuery's SpreadsheetBench 2 evaluation. This result places it ahead of Claude Fable 5 in this specific benchmark.
Kimi K3 ranks #1 on AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5
Doctorina outperforms physicians in synthetic Polish primary care diagnostics
A study compared the clinical AI system Doctorina against eight physicians and four standalone frontier language models across 150 synthetic Polish-language primary-care consultations.
Kimi K3 outperforms Opus 4.8 in oneshot HTML generation
A user evaluated Kimi K3 against Claude Opus 4.8 by running 34 oneshot prompts and assessing the resulting HTML, screenshots, and GIFs using Sonnet 4.6. The evaluation concluded that Kimi K3 produced better results than Opus 4.8.
Kimi K3 matches Claude Fable 5 on DeepSWE quality at one-third the cost
Kimi K3, an open-weight model from Moonshot AI, achieves performance comparable to Anthropic's Claude Fable 5 on the DeepSWE benchmark while costing significantly less per task. In a July 16, 2026 evaluation, Kimi K3 reached 68.5% pass@1 compared to Fable 5's 69.9%, but surpassed it in multi-attempt scenarios and offered superior cost efficiency.
Kimi K3 ranks at same level as Opus thinking on Agent Arena
According to the Agent Arena leaderboard, Kimi K3 performs at a level comparable to Opus in non-vision tasks. While some users note that Opus may have an advantage in vision capabilities, the ranking indicates parity for other use cases.
Kimi K3 (max) beats Sonnet 5 on Simple Bench
The article reports that Kimi K3 (max) outperforms Sonnet 5 on the Simple Bench benchmark.