Artificial Analysis' new agentic knowledge work evaluation, AA-Briefcase, shows GLM-5.2 surpassing GPT-5.5 in performance. The benchmark assesses real-world task execution and reasoning capabilities in knowledge work scenarios.
GLM-5.2 Outperforms GPT-5.5 in AA-Briefcase Evaluation
SpaceX Neocloud Revenue Hits $28B/Year Amidst OpenAI and Sakana Updates
SpaceX has secured its third GPU rental deal with Reflection AI, bringing its annualized revenue to approximately $28 billion based on a calculated rate of over $10 per hour for Blackwell GPUs. This valuation is roughly twice that of Coreweave, highlighting the rapid growth and high pricing power in the AI infrastructure market.
OpenAI expands Daybreak, Sakana releases Fugu, GLM-5.2 gains traction
This week's AI news highlights OpenAI's expansion of its cybersecurity initiatives, Sakana AI's release of an orchestration model called Fugu, and the growing adoption of the open-weight GLM-5.2 model.
Benchmarking Agentic Review Systems for AI-Assisted Research
A study evaluates four AI review systems across six language models, finding OpenAIReview with GPT-5.5 achieves 83.0% accuracy in matching paper quality to external signals and detects 71.6% of injected errors. Real user feedback shows positive sentiment, with a 1.44-to-1 vote ratio, though false positives and minor nitpicks remain common.
TutorMoments evaluates whether AI tutors know when to help or hold back
Researchers introduce TutorMoments, a framework designed to measure if large language models can balance the pedagogical trade-off between providing support and encouraging independent reasoning. Built on real one-on-one math tutoring transcripts, the system replays decision points to evaluate model behavior against human-annotated ground truth.
OpenAI rolls out GPT-5.6 with model stratification and agent UX; Meta releases Muse Spark 1.1
OpenAI has launched GPT-5.6, introducing a new model ladder of Luna, Terra, and Sol alongside distinct effort levels for Max and Ultra modes. The release brings significant changes to the ChatGPT Work and Codex interface, though it initially faced user backlash regarding confusing navigation and rapid usage depletion.