Researchers investigate Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure where an agent generates retrospective explanations of its own experiences and is fine-tuned using only the next-token prediction loss on those explanation tokens. This method requires no external teacher or reward-based policy update.

In software-engineering experiments with Qwen3.5-4B, ROFT reached 49.2% and 26.8% solve rates on SWE-bench Verified and Pro after 20 updates, outperforming GRPO's results after 40 updates. The approach also learns to solve tasks where all base-model attempts failed and encourages good actions by indirectly assigning credit.

These findings establish self-generated retrospections as useful training targets, demonstrating that learning to explain can improve learning to do.