DeepSeek has officially released the DeepSeek-V4.1-Flash model, which serves as the smallest member of its new architecture family and features native multimodal visual understanding. This new architecture is designed to provide a higher capability ceiling, faster inference speeds, higher throughput, and better scaling potential for larger models.
- The model achieves scores of 90.9 on GPQA Diamond, 36.8 on HLE, and 3471 on Codeforces.
- It is now available on the DeepSeek API under the name "deepseek-flash," replacing retired V4 Flash variants.
- Requests to the retired deepseek-v4-pro model will be routed to V4.1 Flash starting September 14, 2026, at the new lower price point.
- API pricing has been adjusted downward to reflect the release of this new model.
DeepSeek states that extensive testing shows V4.1 Flash outperforms the previous DeepSeek V4 Pro across performance, cost, speed, and total time.