Metis introduces a hierarchical dual-representation memory that combines text and code memory to improve self-evolving agents. It organizes experience into execution plans, facts, and pitfalls, crystallizing reusable plans into validated tools only when justified. Evaluated on AppWorld, Metis achieves up to 20.6% higher task accuracy and 22.8% lower execution cost than ReAct, with better overall balance across accuracy, efficiency, and memory cost.
Metis: Bridging Text and Code Memory for Self-Evolving Agents
NatureBench Evaluates AI Coding Agents' Scientific Discovery Capabilities
NatureBench presents a benchmark of 90 tasks from Nature-family papers to assess AI coding agents' ability to achieve scientific discovery. Under a web-search-disabled protocol, the top model exceeds prior state-of-the-art on only 17.8% of tasks. Agents primarily succeed by translating scientific problems into supervised learning tasks, not through original scientific invention.
CORTIS: Text-Only Adaptation of Spoken Language Models
CORTIS enables task-oriented voice agents to generate structured speech outputs by fine-tuning spoken language models using only text-form task supervision. It outperforms ASR-LLM cascades under acoustic degradation, especially in preserving high-level task semantics, without requiring paired speech-target annotations during training.
Microsoft Releases Open Source FastContext for LLM Coding Agents
Microsoft has open-sourced FastContext-1.0, a lightweight repository-exploration subagent that separates code repository exploration from task solving in LLM coding agents. It uses parallel read-only tool calls to return compact file paths and line ranges, improving end-to-end accuracy and reducing token usage by up to 60.3%, with the 4B-RL model outperforming a 30B-SFT model on SWE-bench Pro.
GLM-5.2 Breakout and Open-Model Progress Highlighted
Zhipu's GLM-5.2 emerged as the top open-weight model, praised for its frontier-adjacent performance in daily use, with improvements in coding tasks and reduced 1M-token inference cost via IndexShare. It outperformed other open models in agentic knowledge work benchmarks, reaching 1266 Elo in Artificial Analysis' AA-Briefcase test, though only 3% of tasks were fully satisfied by top models, indicating persistent challenges in real-world long-horizon agent performance.
Struggling to finish Xiaomi Mimo-v2.5-pro token plan credits before expiry
A user has 24B token credits from a Xiaomi token plan competition, worth $50 but obtained for free. They report heavy token consumption during use, limited tool support, and are now concerned about wasting credits due to expiration in four days. The model is praised for its 90% cache hit rate and 99% price reduction on cache hits, with the user noting it performs well in coding and planning tasks.