Researchers introduce VISTA, a visual harness designed to unlock the reasoning potential of general-purpose multimodal models in interactive environments. The system provides long-horizon vision by allowing models to directly perceive surroundings and maintain a lossless visual memory of past observations for active retrieval and reorganization.

  • On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00.
  • The model completes all 25 public games using 57.4% fewer actions than first-time human participants.
  • It substantially outperforms baselines with minimal harnesses across three additional benchmarks covering diverse visual games and puzzles.

VISTA's simple design allows for natural extension to diverse visual environments with minimal adaptation, highlighting its potential as a general-purpose tool for advancing multimodal agents in complex settings.