All articles — korshunov.ai

All articles Page 1 / 94

Hyperball Optimization for Faster Language Model Training

Hyperball is a simple optimizer wrapper that sets fixed Frobenius norms for weight matrices and their updates. It improves training speed and learning rate transfer in large models, achieving 20--30% token equivalent speedup over weight decay baselines on up to 1.2B parameter models.

arxiv arXiv cs.LG · 9d ago

Factorized Neural Operators Decompose Dynamic and Persistent Responses

Factorized Neural Operators (FaNO) decompose spectral representations into equivariant dynamic and invariant persistent responses. This factorized structure enables better interpretability, generalization, and consistent predictions across scales, domains, and physical regimes.

arxiv arXiv cs.LG · 9d ago

CEAP Reduces Variance in LLM Circuit Discovery

CEAP, a new circuit discovery method, substantially reduces resampling variance compared to EAP-IG. The paper shows that rephrasing variance arises from prompt templates activating different circuits, suggesting LLMs are inherently hard to steer across diverse inputs. Sample-wise variance is largely benign, as poor unfaithfulness scores result from selective contribution scaling, not circuit defects.

arxiv arXiv cs.LG · 9d ago

Adaptive Functional Gradient Descent with Convergence Guarantees

We propose a new functional gradient descent algorithm that adapts its representation during optimization. The method achieves convergence to a stationary point under smooth losses and to a global minimizer under smoothness and a Polyak-Lojasiewicz condition, despite using finite-dimensional approximations. It outperforms both fixed-approximation FGD and neural network baselines in regression, PDE solving, and computer vision tasks.

arxiv arXiv cs.LG · 9d ago

Unified Causal-Origin Taxonomy of Distributional Shifts in RL

This paper proposes a unified causal-origin taxonomy for distributional shifts in reinforcement learning, linking ID/OOD generalization to non-stationary settings. It decomposes the agent-environment interaction using a POMDP framework, identifying internal, agent-driven, and external, environment-driven shifts, with explicit, implicit, and hybrid types defined by the shifted-time boundary. The work introduces an evaluation framework to measure shift impact through performance degradation and recovery metrics, enabling systematic analysis of RL robustness.

arxiv arXiv cs.LG · 9d ago

Key Properties for Effective Code Interpreter Reasoning

A study identifies extrinsic (crucial tokens) and intrinsic (cognitive behaviors) properties that enhance code interpreter reasoning in large language models. Stronger reasoning models show higher prevalence of verification, backtracking, and backward chaining, with these properties improving performance during inference and training, reducing overthinking and boosting token efficiency.

arxiv arXiv cs.LG · 9d ago

CrossMaps: Confidence-Aware Semantic Mapping for Rover Navigation

CrossMaps is a real-time, confidence-aware semantic mapping pipeline that uses RGB-D data to create language-queryable maps. It integrates multi-scale CLIP embeddings with a dual-memory architecture—Short-Term and Long-Term Memory—to aggregate visual observations and promote coherent, confident cells as persistent semantic landmarks. The system enables natural language queries to guide rover navigation via semantic heatmaps.

arxiv arXiv cs.LG · 9d ago

CircuitLasso: Scalable Circuit Learning for LLM Interpretability

CircuitLasso enables scalable circuit learning in large language models by using sparse linear regression. It recovers circuits with structural accuracy matching state-of-the-art methods at significantly lower computational cost, and demonstrates human-interpretable semantic propagation through model components. The learned circuits achieve comparable performance on a domain-generalization task with reduced cost.

arxiv arXiv cs.LG · 9d ago

A nonparametric two-sample test using PReLU-IPM

The study introduces PReLU-IPM, a new integral probability metric based on a neural network discriminator with a single node. The resulting PReLU-TST test is nonparametric, consistent, and asymptotically equivalent to standard IPM-based tests, showing higher power or competitive performance on simulated and real datasets.

arxiv arXiv cs.LG · 9d ago

Causal Framework for Auditing Synthetic Data Disclosures

A model-agnostic auditing framework detects and distinguishes true and phantom disclosures in synthetic data. It uses only synthetic outputs and a held-out control set to perform statistical testing, offering tighter privacy leakage bounds than prior methods without requiring model access or additional training.

arxiv arXiv cs.LG · 9d ago

Latent space mapping for interpretable molecular coordinates from nanopore signals

A contrastive encoder trained on simulated signals maps DNA barcode signals into an interpretable molecular coordinate system. The method enables molecule identification with a single pass, reducing computational cost by three orders of magnitude and allowing data pooling across devices while remaining invariant to acquisition conditions.

arxiv arXiv cs.LG · 9d ago

Hybrid Convolutional VAE for Crypto Volatility Surfaces

A convolutional variational autoencoder trained on 6,034 Binance Options surfaces for BTC and ETH achieves 0.94-1.56 vol-point RMSE under 10-50% masking. The hybrid predictor reduces error from 7.00 to 0.83 vol points at 50% masking, outperforming parametric re-fit in structured hole patterns and detecting abnormal market events without supervision.

arxiv arXiv cs.LG · 9d ago

Fixed-Size Neural Networks Achieve Arbitrary Sobolev Approximation

A new activation function enables fixed-size neural networks to approximate any function in Sobolev spaces $W^{s,\infty}((a,b)^d)$ with arbitrary accuracy in the $W^{s-1,\infty}$-norm. The results use elementary activations like EUAF and DUAF$_\infty$, with explicit width and depth bounds, and extend to sigmoidal variants $\widetilde{\mathrm{DUAF}}_n$ preserving accuracy for all $1\leq s\leq n$.

arxiv arXiv cs.LG · 9d ago

Task-Error Residual Learning for Real-Robot Five-Ball Juggling

A residual learning approach using directional task-error supervision achieves stable five-ball juggling on real robots, converging from the second attempt. The system outperforms human practice timelines and relies on both directional feedback and an informative prior, with a fixed-Jacobian Newton update proving most reliable.

arxiv arXiv cs.LG · 9d ago

SPaiK: Scalable Pairwise Kernel Learning with Stochastic Vec Trick

SPaiK introduces a scalable kernel learning method for pairwise settings using the stochastic generalized vec trick (sGVT). This innovation reduces computational and memory demands, enabling efficient training on large datasets and making pairwise kernel learning feasible for previously intractable data sizes.

arxiv arXiv cs.LG · 10d ago

Probabilistic Thinning Decouples Inference from State Updates

A new method decouples ML inference from state persistence in streaming systems using probabilistic thinning. It selectively triggers durable state updates based on event informativeness, reducing persistence path overhead by up to 90% without compromising downstream utility or introducing systemic errors.

arxiv arXiv cs.LG · 10d ago

Dynestyx: Probabilistic Programming for Dynamical Systems

Dynestyx is a probabilistic programming library that provides first-class support for state-space models. It enables users to specify arbitrary priors for discrete- or continuous-time dynamical systems, perform inference on mixed-effect data, and obtain state and parameter estimates with principled uncertainty quantification.

arxiv arXiv cs.LG · 10d ago

Fingerprinting agent behavior through procedural trajectories

We introduce a method to identify agents by their procedural behavior fingerprints, achieving 85.7% accuracy in attributing unseen trajectories to correct agents. Using ProcGrep, we analyze coding agent behavior in SWE-Bench, finding that models from similar release periods or distilled from each other exhibit closer behavioral similarity, with a Jensen-Shannon divergence of 0.25.

arxiv arXiv cs.LG · 10d ago

Analytic Torsion and Spectral Gap Capture Persistent-Laplacian Performance

A compact spectral representation using Betti numbers, spectral gap, and analytic torsion distills persistent Laplacians into three mathematically grounded invariants. This approach captures essential predictive signals from the full spectrum, outperforms it in some cases, and reduces computational overhead on datasets like MNIST, QM-3D, and SKEMPI WT.

arxiv arXiv cs.LG · 10d ago

Multi-Center Benchmark for Abdominal Disease Diagnosis from Non-Contrast CT

A new multi-center benchmark enables abdominal disease diagnosis and report generation from non-contrast CT by synthesizing contrast-enhanced findings. The dataset includes paired NCCT-CECT studies and reports from two centers, showing NCCT achieves average multi-organ AUCs of 69.1% internally and 63.1% externally. The benchmark and code are publicly released to support research into safer, contrast-free abdominal imaging workflows.