Researchers propose SCOUT, a framework that enhances Vision-Language Models' spatial reasoning by combining structured Chain-of-Thought with multi-objective process-reward reinforcement learning. This approach addresses credit assignment issues in existing RL methods and integrates depth perception for comprehensive 3D understanding.

  • SCOUT introduces a structured CoT framework to explicitly model 3D environmental perception.
  • The method utilizes a novel RL algorithm with multi-objective process rewards and tailored advantage estimation for fine-grained credit assignment.
  • A new dataset, SCOUT-24k, was synthesized through a customized pipeline to support the training.
  • SCOUT-3B improves upon baselines by 16.85% on general spatial benchmarks and 6.3% on complex tasks.
  • The larger SCOUT-7B model outperforms GPT-4o by 4.28% despite being trained exclusively on single images.

The authors consider SCOUT a critical step toward next-generation spatially-aware VLMs, demonstrating robust out-of-domain generalization to multi-image and video scenarios.