Researchers introduce Eduardo, a multi-turn reinforcement learning recipe for training large language model tutors that addresses the "assistance dilemma" where models tend to give answers rather than teach. The method uses a masked near-transfer post-test and binary reward gates to discourage cognitive offloading and prevent solution handover.

  • The approach replaces continuous penalties with two binary reward gates: factual correctness of the tutor response and no solution handover.
  • A leave-one-out ablation confirms that while learning-gain rewards alone do not separate teaching from telling, the gates reduce handover and the post-test improves out-of-domain transfer.
  • Eduardo trains 4B, 9B, 14B, and 27B models from two distinct LLM architectures.
  • The post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments using 2.4-6.2x fewer thinking tokens.
  • The model more than doubles its use of the push-for-justification teacher move while training out support fading techniques.

The team open-sources the training environment, an 8,671-problem near-transfer dataset, and the trained models to facilitate further development in interactive tutoring.