PhoneBuddy combines real and mock app environments to train open models for phone use. It improves task success rates from 36.67% to 45.33% on real phones and from 60.3% to 83.2% on AndroidWorld, showing mock-app training complements but does not replace real-app RL.
PhoneBuddy: Training Open Models for Agentic Phone Use
ThinkingBox measures agent reliability by grading terminal backend state and side effects
Microsoft and Hugging Face released ThinkingBox, a benchmark that evaluates AI agents by grading their impact on isolated database states rather than just their tool calls or final responses. Across 507 stateful business workflows run 20 times each against various LLM models, the study found that while many agents produce valid tool calls, they frequently leave incorrect values, unintended extra effects, or missing required effects in the backend.
ICML 2026 Open Reproductions challenge finds 23% of examined papers have falsified claims
The ICML 2026 Open Reproductions challenge, held from July 15 to August 2, 2026, engaged the community in reproducing accepted papers using AI agents and human oversight. The effort resulted in the largest open, claim-by-claim audit of a machine learning conference to date.
Are We Ready For An Agent-Native Memory System?
A new study decomposes agent memory into four core modules and evaluates 12 systems across five benchmark workloads. It finds no single architecture dominates, with performance dependent on alignment with workload bottlenecks, and reveals that localized maintenance is more cost-efficient than global reorganization.
AgentCIBench Evaluates Privacy Risks in Computer-Use Agents
AgentCIBench introduces a benchmark to assess privacy risks in computer-use agents. It identifies three key failure modes—visual co-location, task-ambiguity overshare, and recipient misalignment—and finds that 11 of 15 evaluated agents leak personal data in over 50% of scenarios, with an average leakage of 67.9%.
LRE: Few-Kilobytes Agent Memory with Zero Neural Cost
LRE is a CPU-only, language-model-free system that learns which interaction history units are load-bearing. It outperforms baselines in accuracy-cost balance, reducing peak context size by up to 52% and improving task completion by 37% in some cases. LRE achieves superior answer quality with 68% fewer tokens and requires no annotations or neural computation for training.