DeepSeek has updated its V4 Flash model, releasing a new version on 2026-07-31 that shows significant performance gains over the previous preview. The update introduces results for several new benchmarks while improving scores on existing ones.
- Terminal Bench score increased from 56.9 to 82.7 (note: benchmark changed from v2.0 to v2.1).
- Toolathlon score rose from 51.8 to 70.3.
- New benchmark results include NL2Repo (54.2), Cybergym (76.7), DeepSWE (54.4), Agent Last Exam (25.2), Automation Bench (25.1), DSBench-FullStack (68.7), and DSBench-Hard (59.6).
- Compared to GPT-5.6 Terra, V4 Flash leads in Terminal Bench (+4.3) and Toolathlon (+17.2), while Terra leads in DeepSWE (+15.2) and Agents' Last Exam (+25.2).
The update provides a more capable version of the model for agent-related tasks, though performance relative to competitors like GPT-5.6 Terra remains mixed depending on the specific benchmark.