A new research paper from OpenAI demonstrates how artificial intelligence agents are fundamentally changing the nature of work. The study highlights the capability of these agents to execute longer and more complex tasks than previously possible. This technological advancement is credited with expanding productivity across a wide variety of professional roles. The findings suggest a significant shift in how labor is organized and performed through automation. By handling intricate workflows, AI agents are enabling users to achieve greater efficiency. The paper serves as evidence of the growing impact of autonomous systems on modern employment.
OpenAI Research Shows AI Agents Transforming Work
Open-Ended Optimization lets GPT-5.5 compose its own improvement process
The article introduces Open-Ended Optimization (OEO), a framework that allows a frontier model optimizer to dynamically compose its improvement process online, rather than relying on prescribed optimization pipelines. The authors compare OEO against two staged approaches, SkillOpt and GEPA, across 14 head-to-head comparisons over eight benchmark-target-model settings.
Astra model solves ten open problems in mathematics and theoretical computer science
OpenAI's internal Astra model has generated solutions to ten long-standing open problems across mathematics and theoretical computer science, with the arguments formalized in Lean. These results span high-dimensional geometry, coding theory, group theory, and quantum complexity, addressing questions that had seen no progress for at least a decade.
OpenBioRQ: Benchmark for Agentic Biomedical Research Faithfulness
OpenBioRQ introduces a benchmark of 12,553 unsolved biomedical research questions across 12 domains, designed to test agentic models' faithfulness and abstention. It evaluates models in a tool-using setting without answer keys, using real follow-up evidence rather than parametric knowledge, and reveals significant agentic collapse on the hardest questions where tools are no longer used despite being critical.
MacAgentBench Launches macOS AI Agent Benchmark
MacAgentBench introduces a comprehensive benchmark with 676 tasks across 25 applications, 60% of which involve both GUI and CLI interactions. It uses deterministic rule-based evaluation and fine-grained multi-checkpoint scoring, revealing that Claude Opus 4.6 on OpenClaw achieves 73.7% Pass@1, primarily due to its skill library rather than framework design.
Ohio State University releases open-source Deep Research agent QUEST-35B
Ohio State University's NLP team has released QUEST-35B, an open-source Deep Research agent trained on approximately 32 H100 GPUs using 8,000 synthetic samples. The team open-sourced the training recipe, code, weights, and datasets, with benchmark results showing competitive performance compared to leading closed-source Deep Research systems.