Safety & alignment — korshunov.ai

Safety & alignment Page 1 / 11

Thermodynamic Measure of Intelligence

Intelligence is defined as the lawful amplification of rare but valid futures. A framework shows that recursive self-simulation is necessary and nearly sufficient for high thermodynamic intelligence, enabling a universal, measurable scale across systems from matter to humans and AI.

arxiv arXiv cs.AI · 6d ago

MACR: Explicit Conflict Resolution for LLM Inference

MACR introduces a multi-agent reasoning framework to resolve knowledge conflicts in LLM inference by jointly assessing internal and external knowledge. It uses semantic entropy to measure confidence and employs three specialized agents to induce rules, detect conflicts, and resolve inconsistencies across contexts. Empirical results show MACR outperforms state-of-the-art methods and provides interpretable conflict resolutions.

arxiv arXiv cs.AI · 6d ago

Editorial Alignment in LLM-mediated Knowledge Dissemination

A case study with a Nordic public knowledge institution demonstrates how editorial participation can re-align LLM interfaces with editorial standards. The paper introduces editorial alignment as a design practice in Participatory AI, where editorial values are translated into technical alignment objectives. This approach empowers editors with agency in LLM-mediated knowledge dissemination.

arxiv arXiv cs.AI · 6d ago

Confidence-Aware Automated Assessment of Student-Drawn Scientific Models

A vision-based model with parameter-efficient adaptation scores student drawings in science education. It uses confidence-aware scoring to automatically evaluate high-confidence responses while deferring uncertain ones to human review, improving reliability and practicality in large-scale assessments.

arxiv arXiv cs.AI · 6d ago

CRAX: Fast Safe Reinforcement Learning Benchmarking

CRAX introduces a high-fidelity, accelerated safety benchmark for reinforcement learning using MuJoCo XLA. It achieves up to 100x speedups over CPU-based benchmarks via vectorization and hardware acceleration, featuring six environment suites and three agent-specific tasks across three difficulty levels. Evaluation of six safe RL methods shows no single approach dominates, highlighting trade-offs between performance and safety, with curriculum learning and safety transfer improving results.

arxiv arXiv cs.LG · 6d ago

EFIQA: Label-Free Fundus Image Quality Assessment with Explainability

EFIQA proposes a label-free framework for fundus image quality assessment that uses anatomical priors to generate spatial quality maps. It first trains an unsupervised anomaly detector via masked anatomical inpainting to identify missing vasculature, then distills this knowledge into a shallow adapter for quality mapping. Evaluation on external datasets shows EFIQA outperforms supervised methods in both performance and explainability across diverse quality criteria.

arxiv arXiv cs.LG · 6d ago

Federated Conformal Risk Control via Risk-Curve Shrinkage

A new federated conformal risk control method addresses coverage failures in hospital-level predictions. On real brain tumor data from 20 institutions, pooled calibration fails 40% of sites, with one exceeding false-negative targets by 7.8 percentage points. The proposed shrinkage-based protocol uses empirical risk curves and a hyperparameter n0=19 to achieve 2.7/20 coverage violations at 2.0x prediction set stretch, while preserving marginal guarantees and ensuring no patient-level data leaves any site.

arxiv arXiv cs.LG · 6d ago

Effective Dimension Governs Generalization in Quantum Vision Models

Quantum vision models exhibit better generalization with more entanglement or quantum noise, phenomena unified by the effective dimension of the noise-shaped quantum feature kernel. This dimension acts as a regularization mechanism in overfitting regimes, with amplitude damping improving test accuracy by up to 13% along an inverted-U sweet spot.

arxiv arXiv cs.LG · 6d ago

SLiR: Shifting-based Linear Relaxations for Activation Functions

SLiR enables sound, tight linear relaxations of general activation functions using only Lipschitz constants or critical points. It achieves up to 7.8x more verification properties than state-of-the-art methods by efficiently computing upper and lower bounds via a shifting procedure.

arxiv arXiv cs.LG · 6d ago

CRAX: Fast Safe Reinforcement Learning Benchmarking

CRAX introduces a high-fidelity, fast safety benchmark for reinforcement learning using MuJoCo XLA. It achieves up to 100x speedups over CPU-based benchmarks via vectorization and hardware acceleration, featuring six environment suites and three agent-specific tasks across three difficulty levels. Evaluation of six safe RL methods shows no single approach dominates, highlighting trade-offs between performance and safety, with curriculum learning and safety transfer improving results.

lab Claude Code Releases · 6d ago

v2.1.183 Release Notes

v2.1.183 improves auto mode safety by blocking destructive git and destroy commands without explicit user consent. It adds deprecation warnings for models, introduces attribution.sessionUrl to hide session links, and fixes multiple issues including terminal behavior, subagent performance, and input handling in web and tmux environments.

arxiv arXiv cs.CL · 6d ago

Introducing P-CHR AUC and CRR for Semantic Caching

We introduce Precision-Cache Hit Ratio (P-CHR) AUC and Calibration Retention Rate (CRR) to address the calibration gap in semantic caching. These metrics evaluate precision across cache utilization levels and measure how offline ranking quality persists in deployment. Our analysis shows the gap is driven by training objectives, not data scale, and post-hoc calibration only partially resolves it.

arxiv arXiv cs.CL · 6d ago

Sequential DPO Shows Variable Preference Impact Across Settings

A study of sequential Direct Preference Optimization finds that later training does not uniformly degrade earlier learned preferences. The effect varies by objective relationship, signal strength, and training order, ranging from partial degradation to positive transfer. Pair-level analysis reveals heterogeneous changes, with high-confidence preference pairs sometimes improving despite aggregate metric stability.

arxiv arXiv cs.CL · 6d ago

Control-Window Law for Single-Neuron Steering in Language Models

A new framework defines when single-neuron interventions coherently control model behaviors without output collapse. The control window, based on alignment and norm ratios, predicts behavior triggers and collapse ceilings using forward pass data, with high accuracy on held-out neurons. On refusal, control is typed: coherent bypass occurs without actionable content, while genuine actionable reach appears only in specific cases and at later rollout stages.

arxiv arXiv cs.CL · 6d ago

AI-Driven Deliberation: Scaling Inclusivity and Empowering Marginalised Groups

Large Language Models can scale democratic deliberation by scaffolding argumentation and reducing linguistic biases. The chapter uses Systemic-Functional Linguistics to analyze how socio-demographic and communicative variations affect participation, highlighting AI's potential to challenge exclusionary norms while cautioning against over- or under-claiming its capabilities. It calls for ethical safeguards and further research to ensure equitable AI-assisted engagement.

arxiv arXiv cs.CL · 6d ago

REDACT: Multilingual PII Benchmark with Systematic Control

REDACT introduces a systematically controlled multilingual benchmark for personally identifiable information detection, featuring 51 entity types, 4,127 surface-form patterns, and 25 languages. It evaluates five detectors across 1,000 records, revealing that rule-based models fail on high-stakes data while LLMs perform better, especially in high-sensitivity categories. A reference-free LLM assessment confirms sensitivity-tier assignment as the most challenging evaluation axis.

arxiv arXiv cs.CL · 6d ago

Speech Quality Models Fail to Capture Prosodic and F0 Variability

MOS prediction models accurately capture acoustic degradation but fail to detect prosodic errors and speaker-specific characteristics like pitch and speaking rate. Human listeners perceive significant quality drops for these perturbations, while models show strong biases in fundamental frequency and lack sensitivity to speaking rate and F0 variability.

arxiv arXiv cs.CL · 6d ago

Over-Privileged Tool Selection in LLM Agents

LLM agents commonly select higher-privilege tools despite sufficient lower-privilege alternatives. This over-privileged behavior is amplified by transient tool failures and does not reliably improve with general safety alignment. A new privilege-aware post-training defense reduces unnecessary high-privilege tool use while maintaining agent capabilities.

arxiv arXiv cs.CL · 7d ago

No Self-Preference in Model Revision Under Genuine Authorship

A four-model test on IFEval shows no detectable self-preference in large language models when revising their own text. Authors reject verified-good edits at rates comparable to fresh models, with a gap of -5.1 percentage points (95% CI [-12.9, +2.7]). When authors do reject fixes, 97% of reasons are about detecting flaws, not preference.

arxiv arXiv cs.CL · 7d ago

Black-Box Probe Detects Identity Memorization in Text-to-Image Models

A new black-box probe distinguishes whether text-to-image models memorize identities or fabricate them, without needing reference photos or training data. The NAMESAKES dataset includes over one thousand public figures' names and faces, along with less famous perturbed names, to benchmark this capability across state-of-the-art models.