The authors present Show-Harness, an Embodied Harness that allows vision-language models (VLMs) to control robots through a compact semantic interface linking intent to action. This system exposes discrete semantic action units for VLM reasoning and uses embodiment-specific interpreters to ground them into local robot actions.

  • Directly unlocks closed-source frontier VLMs for zero-shot robot control.
  • Adapts small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning.
  • Introduces GUMI (GUI Manipulation Interface) to extend the semantic action space for GUI-based demonstration collection.

Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms without requiring additional model capacity or costly embodiment-specific pretraining.