Researchers introduce Mid-Harness, a method that samples and verifies candidate actions at the model-harness boundary to improve execution reliability in terminal agents. This approach allocates test-time compute by verifying multiple options before forwarding one for execution, keeping the generator and harness unchanged.

  • On TerminalBench-Lite, using a GPT-5.6 Sol verifier with 8 sampled actions raises Pass@1 from 50.00% to 68.03% compared to the base agent.
  • When TMAX-9B serves as the verifier, pairwise verification yields the best results among evaluated mechanisms.
  • Distilling responses from a stronger verifier into TMAX-9B further improves Pass@1 without changing the action generator.
  • Combining action and trajectory scaling with TMAX-9B achieves higher success at lower estimated token cost than generating more trajectories alone.

These findings identify action scaling as a promising target for test-time compute scaling in terminal agents, improving performance across additional models, benchmarks, and harnesses.