Evaluation & benchmarks — korshunov.ai

Topic · Evaluation & benchmarks

ActiveSAM is a training-free, zero-shot framework that enhances SAM 3 for open-vocabulary semantic segmentation by identifying an image-conditioned active class set. It improves speed-accuracy tradeoff, outperforming SegEarth-OV3 by +1.4 mIoU on average and running up to 5.5x faster on large-vocabulary datasets, with strong robustness to image corruption.

arxiv arXiv cs.LG · 10d ago

ExpRL: Exploratory RL for LLM Mid-Training

ExpRL introduces a novel mid-training approach for LLMs using human-written question-answer data as reward scaffolds. Instead of imitating reference solutions, it constructs problem-specific grading rubrics to reward intermediate reasoning steps, enabling better initialization for sparse-reward RL and outperforming SFT, sparse-reward GRPO, and self-distillation on math reasoning tasks.

arxiv arXiv cs.LG · 10d ago

HABC Improves RL Fine-Tuning of VLAs with Sparse Outcomes

Hierarchical Advantage-Weighted Behavior Cloning (HABC) enhances online RL fine-tuning of vision-language agents by using separate critic heads for viability and efficiency. It combines their outputs via a state-adaptive gate and applies per-transition weights, while intervention-aware credit assignment prevents supervision leakage. In real-robot experiments, HABC boosts success rates to 92%, 88%, and 38% on three bimanual tasks, surpassing SFT baselines of 36%, 44%, and 12%.

media r/LocalLLaMA · 10d ago

HalBench Tests 29 Open Source Models on Sycophancy and Hallucination

HalBench evaluates 29 open-source LLMs on a custom benchmark for sycophancy and hallucination. Qwen 3.6 and Gemma 4 outperform larger models, with Qwen 3.6 achieving 36.6% pushback—higher than GPT-5.4 and Gemini 3.1 Pro. Model size does not correlate with honest responses, indicating that architecture and training data matter more than parameters.

arxiv arXiv cs.LG · 10d ago

Task-Error Residual Learning for Real-Robot Five-Ball Juggling

A residual learning approach using directional task-error supervision achieves stable five-ball juggling on real robots, converging from the second attempt. The system outperforms human practice timelines and relies on both directional feedback and an informative prior, with a fixed-Jacobian Newton update proving most reliable.

arxiv arXiv cs.LG · 10d ago

Post-Hoc Falsification Operators Fail to Improve Accuracy in Small Code Models

A measurement study finds that 26 semantic post-hoc operators do not improve held-out accuracy over Best-of-N in frozen small code models. While some operators reduce compute usage or recover correct programs, none outperform BoN in accuracy, due to systemic limitations like coverage walls and consensus traps. An expression-layer recovery (M1) improves performance on HumanEval+ by 12 tasks, with no harm or leakage, and shows consistent results across model cells.

arxiv arXiv cs.LG · 10d ago

TuneJury: Open Metric for Music Generation Preference Alignment

TuneJury is an open, instance-level pairwise reward model that predicts music preference scores from text prompts and audio clips. It is trained on diverse human-preference data and demonstrates strong generalization, with anchor calibration enabling efficient post-hoc alignment for music generation systems.

arxiv arXiv cs.LG · 10d ago

TokenPilot: Cache-Efficient Context Management for LLM Agents

TokenPilot reduces inference costs by 61% to 87% in both isolated and continuous modes, outperforming prior systems in cost efficiency while maintaining competitive performance. It uses ingestion-aware compaction and lifecycle-aware eviction to stabilize prompt prefixes and manage context segments efficiently.

arxiv arXiv cs.LG · 10d ago

KVEraser: Efficient Localized Context Erasing in LLMs

KVEraser enables efficient localized context erasing in large language models by replacing only the KV cache states of an erased span with learned steering states. It achieves near-full-recomputation performance on in-domain tasks and offers a 24% latency increase versus a 17.6x increase for full recomputation, with up to 3--4x speedup on long-document QA tasks.

arxiv arXiv cs.LG · 10d ago

DP-FL Backdoor Attacks: RING Exploits Privacy for Malicious Signals

A new attack, RING, exploits differential privacy in federated learning to conceal backdoor signals while maximizing impact. It achieves 90.3% attack success against state-of-the-art defenses, up to 26.08x over baseline methods, and reveals a critical security gap in DP-FL due to inherent masking of malicious updates.

arxiv arXiv cs.LG · 10d ago

Phase in Neural Representations: An Internal Oppenheim-Lim Test

Image classifiers like PRISM2D, GFNet, and ViT-B/16 show that phase, not magnitude, drives predictions in hidden layers. ResNet-50 reveals a latent sign code in late blocks, indicating phase/sign identity exists across architectures, though expressed differently due to activation and readout mechanisms.

arxiv arXiv cs.LG · 10d ago

Exact Posterior Score Estimation for Linear Inverse Problems

The paper derives the exact posterior score in closed form for linear Gaussian inverse problems, enabling efficient posterior sampling via denoising. It introduces Exact Posterior Score (EPS), a training objective that preserves pretraining structure and achieves superior performance on fidelity, perceptual, and distributional metrics with fewer denoiser evaluations than gradient-based methods.

media r/LocalLLaMA · 10d ago

vLLM releases new streaming parser for Qwen3+ in nightly

vLLM has introduced a new streaming parser for Qwen3+ available in its nightly build, addressing issues like mid-turn stopping and failed streaming tool calls due to chunk boundaries. The update reportedly resolves these problems in limited testing, improving reliability for agentic workflows.

arxiv arXiv cs.LG · 10d ago

Probabilistic Thinning Decouples Inference from State Updates

A new method decouples ML inference from state persistence in streaming systems using probabilistic thinning. It selectively triggers durable state updates based on event informativeness, reducing persistence path overhead by up to 90% without compromising downstream utility or introducing systemic errors.

arxiv arXiv cs.LG · 10d ago

Dynestyx: Probabilistic Programming for Dynamical Systems

Dynestyx is a probabilistic programming library that provides first-class support for state-space models. It enables users to specify arbitrary priors for discrete- or continuous-time dynamical systems, perform inference on mixed-effect data, and obtain state and parameter estimates with principled uncertainty quantification.

arxiv arXiv cs.LG · 10d ago

Analytic Torsion and Spectral Gap Capture Persistent-Laplacian Performance

A compact spectral representation using Betti numbers, spectral gap, and analytic torsion distills persistent Laplacians into three mathematically grounded invariants. This approach captures essential predictive signals from the full spectrum, outperforms it in some cases, and reduces computational overhead on datasets like MNIST, QM-3D, and SKEMPI WT.

arxiv arXiv cs.LG · 10d ago

Multi-Center Benchmark for Abdominal Disease Diagnosis from Non-Contrast CT

A new multi-center benchmark enables abdominal disease diagnosis and report generation from non-contrast CT by synthesizing contrast-enhanced findings. The dataset includes paired NCCT-CECT studies and reports from two centers, showing NCCT achieves average multi-organ AUCs of 69.1% internally and 63.1% externally. The benchmark and code are publicly released to support research into safer, contrast-free abdominal imaging workflows.

arxiv arXiv cs.LG · 10d ago

PPAD-hardness for min-max optimization of quadratic polynomials

Computing approximate stationary points of min-max optimization over the hypercube is PPAD-hard for quadratic polynomials. This result holds even for multilinear polynomials where each variable appears in at most three monomials, with inverse polynomial approximation factors. As a consequence, two-team zero-sum polymatrix games are proven to be PPAD-hard.

arxiv arXiv cs.LG · 10d ago

Neural EXposure Interaction Search for Interpretable HTE

NEXIS identifies causal heterogeneous treatment effects by discovering Markov-blankets in pre-treatment data. It leverages multi-modal, multi-view measurements and scalable representations with minimal human input, enabling interpretable and actionable policy insights from controlled experiments.

arxiv arXiv cs.LG · 10d ago

Filtered Conformal Ellipsoids for Graph-Native Time Series

A new method called filtered conformal ellipsoids provides prediction sets for multivariate time series by using a frozen state-space filter to generate predictive means and covariances, then applying split-conformal calibration to Mahalanobis scores. The approach achieves coverage under dependence through contraction in an observable predictive-law quotient, with theoretical bounds derived under Gaussian-projection and observability conditions, and shows sharper ellipsoids on graph-native traffic benchmarks compared to static and non-filter baselines.

ActiveSAM: Fast and Accurate Open-Vocabulary Segmentation

ExpRL: Exploratory RL for LLM Mid-Training

HABC Improves RL Fine-Tuning of VLAs with Sparse Outcomes

HalBench Tests 29 Open Source Models on Sycophancy and Hallucination

Task-Error Residual Learning for Real-Robot Five-Ball Juggling

Post-Hoc Falsification Operators Fail to Improve Accuracy in Small Code Models

TuneJury: Open Metric for Music Generation Preference Alignment

TokenPilot: Cache-Efficient Context Management for LLM Agents

KVEraser: Efficient Localized Context Erasing in LLMs

DP-FL Backdoor Attacks: RING Exploits Privacy for Malicious Signals

Phase in Neural Representations: An Internal Oppenheim-Lim Test

Exact Posterior Score Estimation for Linear Inverse Problems

vLLM releases new streaming parser for Qwen3+ in nightly

Probabilistic Thinning Decouples Inference from State Updates

Dynestyx: Probabilistic Programming for Dynamical Systems

Analytic Torsion and Spectral Gap Capture Persistent-Laplacian Performance

Multi-Center Benchmark for Abdominal Disease Diagnosis from Non-Contrast CT

PPAD-hardness for min-max optimization of quadratic polynomials

Neural EXposure Interaction Search for Interpretable HTE

Filtered Conformal Ellipsoids for Graph-Native Time Series