Researchers propose SCOUT, a framework that enhances Vision-Language Models' spatial reasoning by combining structured Chain-of-Thought with multi-objective process-reward reinforcement learning. This approach addresses credit assignment issues in existing RL methods and integrates depth perception for comprehensive 3D understanding.
- SCOUT introduces a structured CoT framework to explicitly model 3D environmental perception.
- The method utilizes a novel RL algorithm with multi-objective process rewards and tailored advantage estimation for fine-grained credit assignment.
- A new dataset, SCOUT-24k, was synthesized through a customized pipeline to support the training.
- SCOUT-3B improves upon baselines by 16.85% on general spatial benchmarks and 6.3% on complex tasks.
- The larger SCOUT-7B model outperforms GPT-4o by 4.28% despite being trained exclusively on single images.
The authors consider SCOUT a critical step toward next-generation spatially-aware VLMs, demonstrating robust out-of-domain generalization to multi-image and video scenarios.