The authors introduce EviRover, the first perception agent explicitly trained to resolve perceptual queries through interaction rather than relying on a one-shot prediction from a single image. To address the lack of data for this setting, they designed two dedicated data generation pipelines yielding EviRover-SFT-5K and EviRover-RL-12K, alongside EviLens, a human-verified benchmark with 688 instances.
- The 4B EviRover model outperforms its backbone by 30 points on average on the EviLens benchmark.
- Performance reaches levels comparable to advanced proprietary models.
- Gains transfer to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL.
The authors consider this significant because it demonstrates that agentic reinforcement learning can effectively handle perception under insufficient evidence where standard assumptions fail. All code, models, and data are released.