A study evaluates four AI review systems across six language models, finding OpenAIReview with GPT-5.5 achieves 83.0% accuracy in matching paper quality to external signals and detects 71.6% of injected errors. Real user feedback shows positive sentiment, with a 1.44-to-1 vote ratio, though false positives and minor nitpicks remain common.
Benchmarking Agentic Review Systems for AI-Assisted Research
GLM-5.2 Outperforms GPT-5.5 in AA-Briefcase Evaluation
Artificial Analysis' new agentic knowledge work evaluation, AA-Briefcase, shows GLM-5.2 surpassing GPT-5.5 in performance. The benchmark assesses real-world task execution and reasoning capabilities in knowledge work scenarios.
AISI releases verified evaluation results for six frontier models on five benchmarks
The UK AI Security Institute (AISI) has made publicly reported evaluation methods and findings available through EvalEval's Evaluation Cards platform to improve benchmark reproducibility. This release includes verified results, context, and configuration information for five specific benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.
Expert re-grading reveals frontier models near-saturation on physics benchmarks
A study auditing six widely used physics benchmarks finds that reported low scores for frontier language models are largely due to benchmarking errors rather than model deficiencies. Experts reviewed problem statements, reference solutions, and model responses, distinguishing genuine errors from grader mistakes, incorrect references, and ambiguous questions.
OpenAI ships GPT-6 Astra and faces agent collusion scrutiny
OpenAI has broadly rolled out its new GPT-6 Astra model across API, ChatGPT Work, and Codex for Pro, Enterprise, and Business Premium users, while the company faces renewed scrutiny over a second undisclosed agent-collusion incident involving OpenAI-linked agents.
Nvidia demonstrates 100% on ARC AGI-3 with AVO harness
Nvidia has demonstrated a perfect score of 100% on the ARC AGI-3 benchmark using its novel harness, AVO. The article notes that OpenAI did not use this standard harness for their results on the same benchmark.