Talos is an open-source tool that automates iterative reanalysis of genomic data to identify rare disease diagnoses. It achieved a 90% recovery rate of in-scope diagnoses with only 1.3 candidate variants per patient, and delivered 241 new diagnoses across 5,000 undiagnosed patients, with most new findings emerging within 32 days of evidence publication.
Talos: Automated Genomic Reanalysis for Rare Disease Diagnosis
Tool-Intent Stabilization in Streaming RAG
A study measures tool-intent stabilization in Streaming RAG, defining when speculative tool queries converge to correct answers. On the CRAG benchmark, 73.9% of queries allow substantial latency hiding, with early stabilization observed in questions with verbatim retrievable evidence. Question type significantly predicts early versus late stabilization, informing when speculative triggers are effective.
Rep2Act aligns VLM representations to improve abstention on new VAD-R benchmark
Researchers introduce Visual Answerability Diagnosis with Rationales (VAD-R), a new benchmark designed to evaluate vision-language models' ability to abstain from unanswerable questions without shortcut cues. The study reveals that while hidden states can distinguish answerability, current models fail to translate this into explicit decisions.
Microsoft releases RetroChimera for retrosynthesis prediction
Microsoft Research has published and open-sourced RetroChimera, a new framework for retrosynthesis prediction that combines two complementary models to propose high-quality synthesis routes. The system integrates R-SMILES 2, a Transformer-based de-novo model, with NeuralLoc, a graph neural network (GNN) based model, using a learned ensembling strategy to rank predictions.
Researcher seeks one independent annotator to resolve LLM agreement ambiguity
A researcher is requesting a single independent human annotator to label 100 Turkish narrative scenes in order to determine whether low inter-rater agreement stems from the interpretive nature of the task or from underspecified annotation definitions. The study found that four LLMs and a rule-based detector agreed with each other and the human reference at roughly chance level, with Cohen’s κ scores ranging from 0.000 to 0.185.
Study identifies FragileTokens that fail contextual copying despite passing isolation probes
A study characterizes "FragileTokens," vocabulary entries in open-weight language models that successfully copy when isolated but exhibit errors when embedded in surrounding text. The research highlights that literal identity preservation is not guaranteed by standard isolation tests, as tokens can be deleted, substituted, or truncated within sequences.