Artificial Analysis' new agentic knowledge work evaluation, AA-Briefcase, shows GLM-5.2 surpassing GPT-5.5 in performance. The benchmark assesses real-world task execution and reasoning capabilities in knowledge work scenarios.
GLM-5.2 Outperforms GPT-5.5 in AA-Briefcase Evaluation
SpaceX Neocloud Revenue Hits $28B/Year Amidst OpenAI and Sakana Updates
SpaceX has secured its third GPU rental deal with Reflection AI, bringing its annualized revenue to approximately $28 billion based on a calculated rate of over $10 per hour for Blackwell GPUs. This valuation is roughly twice that of Coreweave, highlighting the rapid growth and high pricing power in the AI infrastructure market.
OpenAI expands Daybreak, Sakana releases Fugu, GLM-5.2 gains traction
This week's AI news highlights OpenAI's expansion of its cybersecurity initiatives, Sakana AI's release of an orchestration model called Fugu, and the growing adoption of the open-weight GLM-5.2 model.
Benchmarking Agentic Review Systems for AI-Assisted Research
A study evaluates four AI review systems across six language models, finding OpenAIReview with GPT-5.5 achieves 83.0% accuracy in matching paper quality to external signals and detects 71.6% of injected errors. Real user feedback shows positive sentiment, with a 1.44-to-1 vote ratio, though false positives and minor nitpicks remain common.
AISI releases verified evaluation results for six frontier models on five benchmarks
The UK AI Security Institute (AISI) has made publicly reported evaluation methods and findings available through EvalEval's Evaluation Cards platform to improve benchmark reproducibility. This release includes verified results, context, and configuration information for five specific benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.
Frontier models evaluated on log(N)-Questions game over Wikipedia
A study evaluates six frontier language models on a two-agent communication task where a questioner identifies a target from N Wikipedia paragraphs using exactly log2(N) yes/no questions. The experiment ran 408 games across document sets of 4 to 1024 paragraphs, revealing significant performance disparities among the tested models.