A user evaluated Kimi K3 against Claude Opus 4.8 by running 34 oneshot prompts and assessing the resulting HTML, screenshots, and GIFs using Sonnet 4.6. The evaluation concluded that Kimi K3 produced better results than Opus 4.8.
- Kimi K3 was tested on 34 oneshot prompts with outputs evaluated by Sonnet 4.6.
- The model demonstrated superior performance in generating HTML, screenshots, and GIFs compared to Claude Opus 4.8.
- Kimi K3 cost $0.44 for the full set of prompts, significantly less than Opus 4.8's $7.16.
The results suggest that Kimi K3 is not only more effective for oneshot generation tasks but also more token-efficient and cost-effective than the competing model.