BabelJudge introduces an open-source framework to measure four key bias modes in LLM judges across languages and agent trajectories. It reveals a significant reliability drop from Hindi to Swahili—0.714 to 0's 0.550—highlighting cross-lingual degradation invisible to raw accuracy. The framework enables bias-aware evaluation without human labels, using controlled perturbations to create known gold labels, and extends to agentic workflows with new metrics on tool accuracy and hallucination detection.
BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories
Tool-Intent Stabilization in Streaming RAG
A study measures tool-intent stabilization in Streaming RAG, defining when speculative tool queries converge to correct answers. On the CRAG benchmark, 73.9% of queries allow substantial latency hiding, with early stabilization observed in questions with verbatim retrievable evidence. Question type significantly predicts early versus late stabilization, informing when speculative triggers are effective.
Open TTS Leaderboard launches scalable multilingual TTS evaluation
The Open TTS Leaderboard has been released to address the fragmentation and lack of standardization in Text-to-Speech (TTS) evaluation by using objective metrics instead of relying solely on slow, arena-based human preference scores. It evaluates models across intelligibility, speed, and speaker similarity using Qwen3 ASR and WavLM embeddings.
Stale-document poisoning causes outdated retrieval to override correct model answers
Researchers identify "stale-document poisoning," a temporal alignment failure in Retrieval-Augmented Generation (RAG) where outdated external evidence causes models to provide incorrect answers despite knowing the right response without retrieval.
Claude Opus 5 and Tetsu find six hollow CI checks in their own project
On September 27, Claude Opus 5 and the 3B model Tetsu identified six separate continuous integration checks in their public repository that reported green or passed despite failing to verify what they were supposed to measure. The issues ranged from a deployment gate refusing valid restarts due to stale digests to timing benchmarks with hollow thresholds that accepted negligible performance differences.
pop123-ux releases 38 runnable notebooks to learn Hugging Face by removing abstractions
The user pop123-ux has published a learning repository containing 38 runnable notebooks designed to help users understand the Hugging Face stack by progressively removing high-level abstractions. The collection covers a path from core components like transformers and tokenizers into fine-tuning, PEFT, reasoning/RL, and end-to-end project work.