An internal evaluation of open-weight models on hard agentic tasks found DeepSeek v4 flash to be the fastest and cheapest option among its peers. The test involved long-running workflows across multiple applications, scored by deterministic checks requiring full success.

  • DeepSeek v4 flash completed tasks in 164 seconds, nearly 2.5 times faster than GLM 5.2 and 1.4 times faster than Kimi K3.
  • Cost per task was approximately $0.08 for DeepSeek, compared to $0.57 for GLM 5.2 and $1.39 for Kimi K3.
  • Success rates were nearly identical across all three models: DeepSeek achieved 20/30, while GLM and Kimi each scored 21/30.
  • Despite lower success rates than frontier models like Fable 5 (24/30) and GPT-5.6 Sol (24/30), DeepSeek offered a significant performance advantage in speed and cost efficiency.

The results suggest that while state-of-the-art models remain superior for raw capability, DeepSeek v4 flash provides a highly efficient alternative for agentic workflows.