Mid-Harness improves terminal agent action reliability via test-time compute scaling
Researchers introduce Mid-Harness, a method that samples and verifies candidate actions at the model-harness boundary to improve execution reliability in terminal agents. This approach allocates test-time compute by verifying multiple options before forwarding one for execution, keeping the generator and harness unchanged.