Evaluation & benchmarks — korshunov.ai

Topic · Evaluation & benchmarks

ExpRL introduces a novel mid-training approach for LLMs using human-written question-answer data as reward scaffolds. Instead of imitating reference solutions, it constructs problem-specific grading rubrics to reward intermediate reasoning steps, enabling better initialization for sparse-reward RL and outperforming SFT, sparse-reward GRPO, and self-distillation on math reasoning tasks.

arxiv arXiv cs.LG · 10d ago

HABC Improves RL Fine-Tuning of VLAs with Sparse Outcomes

Hierarchical Advantage-Weighted Behavior Cloning (HABC) enhances online RL fine-tuning of vision-language agents by using separate critic heads for viability and efficiency. It combines their outputs via a state-adaptive gate and applies per-transition weights, while intervention-aware credit assignment prevents supervision leakage. In real-robot experiments, HABC boosts success rates to 92%, 88%, and 38% on three bimanual tasks, surpassing SFT baselines of 36%, 44%, and 12%.

media r/LocalLLaMA · 10d ago

HalBench Tests 29 Open Source Models on Sycophancy and Hallucination

HalBench evaluates 29 open-source LLMs on a custom benchmark for sycophancy and hallucination. Qwen 3.6 and Gemma 4 outperform larger models, with Qwen 3.6 achieving 36.6% pushback—higher than GPT-5.4 and Gemini 3.1 Pro. Model size does not correlate with honest responses, indicating that architecture and training data matter more than parameters.

arxiv arXiv cs.LG · 10d ago

TokenPilot: Cache-Efficient Context Management for LLM Agents

TokenPilot reduces inference costs by 61% to 87% in both isolated and continuous modes, outperforming prior systems in cost efficiency while maintaining competitive performance. It uses ingestion-aware compaction and lifecycle-aware eviction to stabilize prompt prefixes and manage context segments efficiently.

arxiv arXiv cs.LG · 10d ago

KVEraser: Efficient Localized Context Erasing in LLMs

KVEraser enables efficient localized context erasing in large language models by replacing only the KV cache states of an erased span with learned steering states. It achieves near-full-recomputation performance on in-domain tasks and offers a 24% latency increase versus a 17.6x increase for full recomputation, with up to 3--4x speedup on long-document QA tasks.

arxiv arXiv cs.LG · 10d ago

DP-FL Backdoor Attacks: RING Exploits Privacy for Malicious Signals

A new attack, RING, exploits differential privacy in federated learning to conceal backdoor signals while maximizing impact. It achieves 90.3% attack success against state-of-the-art defenses, up to 26.08x over baseline methods, and reveals a critical security gap in DP-FL due to inherent masking of malicious updates.

arxiv arXiv cs.LG · 10d ago

Phase in Neural Representations: An Internal Oppenheim-Lim Test

Image classifiers like PRISM2D, GFNet, and ViT-B/16 show that phase, not magnitude, drives predictions in hidden layers. ResNet-50 reveals a latent sign code in late blocks, indicating phase/sign identity exists across architectures, though expressed differently due to activation and readout mechanisms.

arxiv arXiv cs.LG · 10d ago

Exact Posterior Score Estimation for Linear Inverse Problems

The paper derives the exact posterior score in closed form for linear Gaussian inverse problems, enabling efficient posterior sampling via denoising. It introduces Exact Posterior Score (EPS), a training objective that preserves pretraining structure and achieves superior performance on fidelity, perceptual, and distributional metrics with fewer denoiser evaluations than gradient-based methods.

media r/LocalLLaMA · 10d ago

vLLM releases new streaming parser for Qwen3+ in nightly

vLLM has introduced a new streaming parser for Qwen3+ available in its nightly build, addressing issues like mid-turn stopping and failed streaming tool calls due to chunk boundaries. The update reportedly resolves these problems in limited testing, improving reliability for agentic workflows.

arxiv arXiv cs.LG · 10d ago

Filtered Conformal Ellipsoids for Graph-Native Time Series

A new method called filtered conformal ellipsoids provides prediction sets for multivariate time series by using a frozen state-space filter to generate predictive means and covariances, then applying split-conformal calibration to Mahalanobis scores. The approach achieves coverage under dependence through contraction in an observable predictive-law quotient, with theoretical bounds derived under Gaussian-projection and observability conditions, and shows sharper ellipsoids on graph-native traffic benchmarks compared to static and non-filter baselines.

arxiv arXiv cs.LG · 10d ago

HAMON: Passive Optical Forecasting Core

HAMON uses passive optical diffraction to generate forecasts, outperforming digital baselines on ETTm2 at all horizons and ETTh2 at all but the longest horizon. It achieves up to 14% lower MSE and operates without trainable digital mixing, relying instead on physical optical propagation.

ExpRL: Exploratory RL for LLM Mid-Training

HABC Improves RL Fine-Tuning of VLAs with Sparse Outcomes

HalBench Tests 29 Open Source Models on Sycophancy and Hallucination

TokenPilot: Cache-Efficient Context Management for LLM Agents

KVEraser: Efficient Localized Context Erasing in LLMs

DP-FL Backdoor Attacks: RING Exploits Privacy for Malicious Signals

Phase in Neural Representations: An Internal Oppenheim-Lim Test

Exact Posterior Score Estimation for Linear Inverse Problems

vLLM releases new streaming parser for Qwen3+ in nightly

Filtered Conformal Ellipsoids for Graph-Native Time Series

HAMON: Passive Optical Forecasting Core