A study demonstrates that a language model's memory containing incorrect conclusions is more detrimental than having no memory at all. When models retain stale values while dropping supporting work, they emit confident but wrong answers, whereas empty memories allow for abstention. This phenomenon, termed brittle memory, was observed across seven models where the direction of failure never reversed regardless of task or disposition. The researchers introduced reclaim evaluation to measure correctability by compressing interactions and testing if corrections recover ground truth without using a judge. Results indicate that correctability depends on whether the source information survives compression rather than model capability. A source-first policy, which keeps recomputable sources and drops re-derivable conclusions, restored correctability significantly better than length-matched controls. In chained memory loops, dropped-source errors corrupt downstream steps irreparably, while the proposed fix maintains bounded performance horizons. The findings replicate across three deployed systems and real dialogue data, with a hand-built oracle reaching perfect accuracy.
Reclaim Evaluation Shows Lossy Memory Is Worse Than No Memory
InternScience Releases Agents-A1, a 35B MoE Model with Unbelievable Benchmarks
InternScience has released the Agents-A1 model on Hugging Face, featuring a 35 billion parameter Mixture of Experts (MoE) architecture. The release includes a technical report available on arXiv and is being highlighted for its exceptional benchmark performance.
Tool-Intent Stabilization in Streaming RAG
A study measures tool-intent stabilization in Streaming RAG, defining when speculative tool queries converge to correct answers. On the CRAG benchmark, 73.9% of queries allow substantial latency hiding, with early stabilization observed in questions with verbatim retrievable evidence. Question type significantly predicts early versus late stabilization, informing when speculative triggers are effective.
Technical Taxonomy of LLM Agent Communication Protocols
A new taxonomy classifies LLM agent communication protocols across five dimensions: counterparty, payload, interaction state, discovery mechanism, and schema flexibility. Analysis shows hybrid payloads, session-state persistence, and runtime schema negotiation are common, with decentralized discovery remaining rare. The study predicts short-term convergence toward unified agent-to-agent and agent-to-context protocols, and long-term evolution toward a federated, layered protocol stack.
Open TTS Leaderboard launches scalable multilingual TTS evaluation
The Open TTS Leaderboard has been released to address the fragmentation and lack of standardization in Text-to-Speech (TTS) evaluation by using objective metrics instead of relying solely on slow, arena-based human preference scores. It evaluates models across intelligibility, speed, and speaker similarity using Qwen3 ASR and WavLM embeddings.
Rep2Act aligns VLM representations to improve abstention on new VAD-R benchmark
Researchers introduce Visual Answerability Diagnosis with Rationales (VAD-R), a new benchmark designed to evaluate vision-language models' ability to abstain from unanswerable questions without shortcut cues. The study reveals that while hidden states can distinguish answerability, current models fail to translate this into explicit decisions.