The authors introduce agentic meta-reasoning, an inference-time harness that structures execution choices for long-horizon tasks. A controller manages the process by consolidating progress and dispatching work based on a compact run account rather than full history.

On ProgramBench, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex, and 67.2% with Opus 4.8 against 65.5% for Claude Code. Across abstract reasoning, multi-domain long-horizon reasoning, and proof generation benchmarks, it gains between 3.6 and 4.2 points over direct control. Artifact-graph analysis shows increased reuse of earlier work and higher coverage of correct solutions.

The results indicate that spending computation on structured control becomes increasingly important as agents scale to longer runs.