Google Cloud AI Research, in collaboration with UNC-Chapel Hill, Stanford, and Washington University in St. Louis, has open-sourced RRSI (Regularized Recursive Self-Improvement). This framework allows an LLM agent to rewrite its own harness—including prompts, tools, memory, control flow, and sub-agents—without changing model weights.

  • RRSI addresses overfitting by constraining the improvement loop with regularizers such as an annealed edit budget, evidence-aware credit logging, and a leakage critic that rejects benchmark-specific logic.
  • The system uses a noise-adjusted floor and cost rule to ensure gains are real and efficient, pruning components that stop producing value.
  • On Terminal-Bench 2.1, scores rose from 74.2% to 80.2%, while SWE-bench Verified improved from 82.0% to 83.8% on held-out splits.
  • Out-of-distribution benchmarks saw gains of +4.7 points on JobBench, +3.5 on GDPval, and +3.7 on APEX-Agents.
  • The framework reduces policy token usage by approximately 30-36% compared to unregularized evolution methods.

The authors consider this important because it enables AI agents to improve their own operational structure while maintaining generalization across tasks they were not explicitly optimized against.