Researchers introduce Mid-Harness, a method that samples and verifies candidate actions at the model-harness boundary to improve execution reliability in terminal agents. This approach allocates test-time compute by verifying multiple options before forwarding one for execution, keeping the generator and harness unchanged.
- On TerminalBench-Lite, using a GPT-5.6 Sol verifier with 8 sampled actions raises Pass@1 from 50.00% to 68.03% compared to the base agent.
- When TMAX-9B serves as the verifier, pairwise verification yields the best results among evaluated mechanisms.
- Distilling responses from a stronger verifier into TMAX-9B further improves Pass@1 without changing the action generator.
- Combining action and trajectory scaling with TMAX-9B achieves higher success at lower estimated token cost than generating more trajectories alone.
These findings identify action scaling as a promising target for test-time compute scaling in terminal agents, improving performance across additional models, benchmarks, and harnesses.