Talos is an open-source tool that automates iterative reanalysis of genomic data to identify rare disease diagnoses. It achieved a 90% recovery rate of in-scope diagnoses with only 1.3 candidate variants per patient, and delivered 241 new diagnoses across 5,000 undiagnosed patients, with most new findings emerging within 32 days of evidence publication.
Talos: Automated Genomic Reanalysis for Rare Disease Diagnosis
Tool-Intent Stabilization in Streaming RAG
A study measures tool-intent stabilization in Streaming RAG, defining when speculative tool queries converge to correct answers. On the CRAG benchmark, 73.9% of queries allow substantial latency hiding, with early stabilization observed in questions with verbatim retrievable evidence. Question type significantly predicts early versus late stabilization, informing when speculative triggers are effective.
M$^3$R-Bench introduces evidence-grounded benchmark for multimodal metaphor understanding
Researchers introduce M$^3$R-Bench, a unified benchmark containing 1,000 image-text instances with human-verified annotations designed to evaluate evidence-grounded multimodal metaphor understanding. The benchmark provides joint annotations for metaphor occurrence, Target-Source mapping, sentiment, and stage-wise explanations based on Conceptual Metaphor Theory.
Microsoft's SkillOpt enables agent skill transfer between Codex and Claude Code
Microsoft, Shanghai Jiao Tong University, Tongji University, and Fudan University researchers developed SkillOpt, a text-space optimizer that trains natural-language skill documents while keeping the target model frozen. The system uses an optimizer model to propose bounded edits based on scored rollouts, exporting the result as a single `best_skill.md` file.
Zing framework improves LLM social intelligence via SoMBench benchmark and Actio grounding
The report introduces Zing, an integrated framework designed to enhance the social intelligence of large language models by measuring, internalizing, and grounding social capabilities. The authors present SoMBench, a psychology-grounded benchmark spanning 17 secondary dimensions and 3,481 expert-verified instances, which reveals that current LLMs have significant room for improvement with no model reaching near-ceiling performance.
SRRM benchmark reveals Qwen2.5-1.5B outperforms 3B in long-context memory
A researcher is seeking an arXiv cs.CL endorsement for a paper introducing Simple Retrieval-Reconstruction Memory (SRRM), a lightweight benchmark designed to evaluate long-context memory in Small Language Models.