An empirical study of 86,156 test-file patches from 33,596 agent-authored PRs reveals that 80.2% of test patches contain weak or no explicit oracle signals. Strong-oracle test files significantly improve merge likelihood (OR = 1.28, p < 0.001) after adjusting for multiple factors, indicating test file presence alone overestimates verification strength.
Oracle Signals in Agent-Authored Test Code
ChronicleBio uses AI to cure POTS; Mind Lab tests Macaron-V1 continual learning
ChronicleBio, founded by former OpenAI exec Fidji Simo, is using AI and 153 terabytes of blood data to cure Postural Orthostatic Tachycardia Syndrome (POTS). Meanwhile, Mind Lab's Macaron-V1 model surpasses GLM-5.2 in benchmarks by dynamically switching between five LoRA expert modules for continual learning.
AMD + Anthropic deal: AMD to supply MI450 chips and invest $5B
AMD and Anthropic have signed a major agreement involving tens of billions of dollars in AI servers, with Anthropic purchasing up to 2 gigawatts of AMD's Instinct MI450 chips starting in the first half of next year. Concurrently, AMD will invest up to $5 billion in Anthropic as specific deployment milestones are met, while both companies collaborate on identifying suitable data centers for the hardware.
Thinking Machines releases Inkling; Anthropic prepares IPO; OpenAI trains GPT-Red
Thinking Machines has released Inkling, a 975B-parameter mixture-of-experts model with 41B active parameters and multimodal reasoning. Simultaneously, Anthropic is preparing for a potential IPO later this year following a $65 billion funding round at a $965 billion valuation. OpenAI has also trained GPT-Red to iteratively generate adversarial prompts, which reduced failures on a prompt-injection benchmark by sixfold for GPT-5.6 Sol.
Claude Code adds in-app browser; Cursor builds general agent
Anthropic has added an in-app browser to Claude Code on desktop, allowing the tool to read, click through, and interact with websites and docs similarly to local dev servers. Concurrently, reports indicate that Cursor is developing a general-purpose AI agent designed to handle emails, texts, spreadsheets, and engineering tasks.
GPT-5 outperforms humans in inducing belief states via planning
A new study evaluates Large Language Models' ability to induce specific belief states in other agents through actions rather than conversation, a capability termed Non-Conversational Planning ToM (NCP-ToM). Using the NCP-ExploreToM framework, researchers tested six frontier models and human participants on 600 task instances where agents had to move objects or direct characters to achieve belief goals.