A study evaluates four AI review systems across six language models, finding OpenAIReview with GPT-5.5 achieves 83.0% accuracy in matching paper quality to external signals and detects 71.6% of injected errors. Real user feedback shows positive sentiment, with a 1.44-to-1 vote ratio, though false positives and minor nitpicks remain common.
Benchmarking Agentic Review Systems for AI-Assisted Research
GLM-5.2 Outperforms GPT-5.5 in AA-Briefcase Evaluation
Artificial Analysis' new agentic knowledge work evaluation, AA-Briefcase, shows GLM-5.2 surpassing GPT-5.5 in performance. The benchmark assesses real-world task execution and reasoning capabilities in knowledge work scenarios.
TutorMoments evaluates whether AI tutors know when to help or hold back
Researchers introduce TutorMoments, a framework designed to measure if large language models can balance the pedagogical trade-off between providing support and encouraging independent reasoning. Built on real one-on-one math tutoring transcripts, the system replays decision points to evaluate model behavior against human-annotated ground truth.
OpenAI rolls out GPT-5.6 with model stratification and agent UX; Meta releases Muse Spark 1.1
OpenAI has launched GPT-5.6, introducing a new model ladder of Luna, Terra, and Sol alongside distinct effort levels for Max and Ultra modes. The release brings significant changes to the ChatGPT Work and Codex interface, though it initially faced user backlash regarding confusing navigation and rapid usage depletion.
EvoPolicyGym introduces benchmark for evaluating autonomous policy evolution
The authors introduce Autonomous Policy Evolution, a controlled evaluation setting where an agent repeatedly edits an executable policy system within a fixed interaction budget to iteratively improve explored policies.
GPT-5 outperforms humans in inducing belief states via planning
A new study evaluates Large Language Models' ability to induce specific belief states in other agents through actions rather than conversation, a capability termed Non-Conversational Planning ToM (NCP-ToM). Using the NCP-ExploreToM framework, researchers tested six frontier models and human participants on 600 task instances where agents had to move objects or direct characters to achieve belief goals.