DeepSeek has released an experimental multimodal vision understanding model named DeepSeek-V4-Flash-Vision-Exp, now available on the DeepSeek API platform. This new model is designed to enhance visual capabilities while maintaining parity with its text-only counterpart.
- The model can be accessed via the API by setting the model parameter to 'deepseek-v4-flash-vision-exp'.
- It achieves significant improvements on agent benchmarks requiring visual understanding, bringing multimodal agent capabilities close to Opus-4.8.
- In terms of pure-text capabilities such as agent, reasoning, and world knowledge, it performs on par with the official DeepSeek-V4-Flash.
- Benchmark scores include Terminal Bench 2.1 at 83.9, Chartography at 64.3, and DSBench-Hard at 59.3.
This release provides users with a new tool for multimodal tasks while retaining the robust text-based performance of the DeepSeek-V4-Flash model.