Kimi K3 has topped the new FullStack leaderboard on arena.ai.
The model achieved the top position on this newly introduced benchmark.
Kimi K3 has topped the new FullStack leaderboard on arena.ai.
The model achieved the top position on this newly introduced benchmark.
A user evaluated Kimi K3 against Claude Opus 4.8 by running 34 oneshot prompts and assessing the resulting HTML, screenshots, and GIFs using Sonnet 4.6. The evaluation concluded that Kimi K3 produced better results than Opus 4.8.
A comparison of Kimi K3 and GPT-5.6 Sol on the DeepSWE benchmark reveals that while GPT-5.6 Sol leads in single-shot quality (72.7% vs 68.5%), Kimi K3 achieves higher pass@k scores for k > 1 at a significantly lower cost.
Kimi K3, an open-weight model from Moonshot AI, achieves performance comparable to Anthropic's Claude Fable 5 on the DeepSWE benchmark while costing significantly less per task. In a July 16, 2026 evaluation, Kimi K3 reached 68.5% pass@1 compared to Fable 5's 69.9%, but surpassed it in multi-attempt scenarios and offered superior cost efficiency.
The release of Kimi-K3 has reduced the time lag between the open-source and closed-source AI frontiers to just 1.5 months, according to Artificial Analysis.
According to the Agent Arena leaderboard, Kimi K3 performs at a level comparable to Opus in non-vision tasks. While some users note that Opus may have an advantage in vision capabilities, the ranking indicates parity for other use cases.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy