TutorMoments evaluates whether AI tutors know when to help or hold back
Researchers introduce TutorMoments, a framework designed to measure if large language models can balance the pedagogical trade-off between providing support and encouraging independent reasoning. Built on real one-on-one math tutoring transcripts, the system replays decision points to evaluate model behavior against human-annotated ground truth.