Researchers introduce Video-DeepResearch (Video-DR), a framework extending multimodal agents from static images to continuous video streams, addressing modality bias and parametric knowledge leakage. The system employs a decoupled perception-exploration pipeline with stage-wise tool unlocking to enforce cross-frame visual grounding before web retrieval.
- Training utilizes supervised fine-tuning followed by Group Relative Policy Optimization (GRPO) to enable autonomous exploration.
- A new benchmark, Video-DR-Bench, comprising 200 complex multi-hop VQA instances, is curated for evaluation.
- The Video-DeepResearch-35B-A3B model achieves a state-of-the-art 64.0% average accuracy on the benchmark.
- This performance surpasses Claude-4.5-Sonnet (59.0%), GPT-5 (52.5%), and Gemini 2.5 Pro (57.5%).
The framework demonstrates that its training paradigm is effective even at compact scales, with the 30B-A3B variant achieving 59.3% accuracy.