All articles — korshunov.ai

All articles Page 1 / 129

Local models went from mostly useless to actually useful in one year

Local models transitioned from being primarily privacy-focused toys to practical tools for coding, private document management, and local workflows within a year. While they still fall short of replacing top closed models for complex tasks requiring planning and error correction, the overall improvement in usability and performance is evident.

media r/LocalLLaMA · 14d ago

A Year Building a Fully Local Home Voice Assistant

A developer spent 12 months building a local, open-source voice assistant inspired by Alexa, documenting the challenges and progress. The project aimed to create a privacy-focused alternative using local models, with ongoing improvements and fixes.

media r/LocalLLaMA · 14d ago

GLM-5.2: Built for Long-Horizon Tasks

GLM-5.2 is a language model designed specifically for long-horizon tasks. It aims to better handle complex, multi-step reasoning and long-term planning by improving its ability to maintain context over extended sequences.

github llama.cpp · 14d ago

llama.cpp release b9680: new binaries and Vulkan support

llama.cpp releases version b9680 with updated Vulkan support and new binaries for macOS, Linux, Android, Windows, and openEuler. The release includes CPU and GPU variants for multiple architectures, with support for Vulkan, CUDA, OpenVINO, SYCL, and ROCm.

media r/LocalLLaMA · 14d ago

Rio 3.5 397B likely a failed embezzlement of AI funding

The Rio 3.5 397B AI model was reportedly developed by merging a Nex N2 Pro without additional training, using funds intended for proper model development. The official documentation initially claimed advanced training, but was later updated to admit the shallow merge, while still asserting additional training occurred, and the original model was removed from Hugging Face.

github llama.cpp · 14d ago

llama.cpp releases b9673 with USM system allocations and cross-platform binaries

llama.cpp version b9673 introduces optional USM system allocations for GPU buffers ≥1GB, enabling VRAM overcommit when device support is available. The feature requires GGML_SYCL_USM_SYSTEM environment variable and is disabled by default, falling back to regular allocations if unsupported.

media r/LocalLLaMA · 14d ago

GLM-5.2 Max is currently the third best model

GLM-5.2 Max is ranked as the third best model available, across both open and proprietary models. The assessment is based on performance benchmarks and current evaluations in the field of large language models.

blog Simon Willison · 14d ago

Datasette 1.0a34 Adds Row Editing and Deletion Tools

Datasette 1.0a34 introduces tools to insert, edit, and delete rows within the interface. These features are available on table pages and as action items on row pages, addressing a long-overdue capability in the UI.

media r/LocalLLaMA · 14d ago

Looking for locally hosted tool to create English subtitles from videos

A user is seeking a locally hosted, self-contained app to generate English subtitles (in .srt or .ass format) from video files. They consider Qwen-ASR and Whisper as strong options but report poor subtitle timing in ComfyUI implementations and unreliable performance with older models like those in storytoolkitAI. They ask for recommendations that work well on Windows and can handle multiple languages.

blog Simon Willison · 14d ago

click-to-play — a still that plays

The click-to-play Web Component displays a still image with a click-to-play button that loads a GIF on demand. It supports progressive enhancement, allowing GIFs to be loaded only when users interact with the image.

media Latent Space · 14d ago

GLM-5.2 Claims Top Position in Frontend Coding with Speculative Decoding

GLM-5.2, a 744B parameter model from Z.ai, has been evaluated as the top frontend coding model globally, outperforming all Opus versions including Opus 4.8. This achievement is highlighted in third-party evaluations that validate official offline tests, marking a significant milestone for a model of its size, particularly in the competitive frontend coding domain.

media r/LocalLLaMA · 14d ago

RTX 5060 Ti 16GB vs RX 9060 XT 16GB Benchmark Comparison

A benchmark comparison shows the NVIDIA RTX 5060 Ti 16GB outperforms the AMD RX 9060 XT 16GB across multiple LLM models, with higher response and prompt token speeds. Performance gains are consistent across models like Gemma3, Llama3.2, and Qwen3, with the RTX 5060 Ti showing notably faster prompt processing, especially in larger models.

media r/LocalLLaMA · 14d ago

Elias in the Lighthouse: Diagnosing Low Diversity in LLM Stories

A new study examines the limited diversity in stories generated by large language models, using the recurring character Elias in the lighthouse as a case study. The research highlights how such patterns suggest systemic biases in training data and model outputs.

arxiv arXiv cs.LG · 14d ago

LegalHalluLens: Auditing Hallucinations in Legal AI

LegalHalluLens introduces a framework to audit AI hallucinations in legal contexts by analyzing typed hallucination profiles across four claim categories. It reveals a 38-40 point gap between obligation/numeric and temporal claims, and shows two systems with identical 52% hallucination rates can have opposite risk directions. The framework uses a Risk Direction Index and calibrated debate pipelines to reduce fabricated detections by 45%, offering actionable diagnostics for trustworthy legal AI deployment.

arxiv arXiv cs.LG · 14d ago

Recursive Masked Diffusion Models Introduce New Scaling Axis

Recursive Masked Diffusion Models (R-MDMs) introduce recursive depth as a third scaling axis by reapplying a denoising transformer within each diffusion step. This recursion enables iterative output refinement without increasing parameter count, achieving performance comparable to non-recursive models with up to L times more parameters, where L is the number of iterations. R-MDMs also reduce inference compute by partially replacing denoising steps with recursive refinement.

arxiv arXiv cs.LG · 14d ago

LoopCoder-v2 Achieves Optimal Two-Loop Performance

LoopCoder-v2, a parallel loop Transformer model, achieves superior code generation and reasoning performance with two loops, improving SWE-bench Verified from 43.0 to 64.4 points and Multi-SWE from 14.0 to 31.0 points. Variants with three or more loops perform worse, indicating a non-monotonic loop-count effect due to growing positional mismatch and diminishing returns.

arxiv arXiv cs.LG · 14d ago

Catastrophic Forgetting is Low-Rank: A Function-Space Theory

A function-space theory reveals that catastrophic forgetting in continual adaptation concentrates in a small number of old-task NTK eigenmodes. In frozen-backbone linear-head PEFT-CL, the forgetting vector is exactly predictable up to numerical precision, with a Kronecker scaling rule for the vulnerable rank.

arxiv arXiv cs.LG · 14d ago

INI-VPINN: Physics-Informed Neural Network with Implicit Boundary Handling

INI-VPINN is a variational physics-informed neural network that implicitly enforces Neumann and interface conditions using compact support weighting functions and integration by parts. It achieves higher accuracy and faster convergence than existing PINN methods in solving multi-material problems with geometric singularities and mixed boundary conditions, and is publicly available on GitHub.

arxiv arXiv cs.LG · 14d ago

Baseline Evaluation of Open-Source LLMs for Multi-Label ATT&CK Classification

A ground-truth dataset of 2,076 human-annotated sentences from 83 complex CTI reports was constructed and mapped to 114 ATT&CK techniques with \k{appa} = 0.68 inter-annotator agreement. Seven open-source LLMs ranging from 8B to 236B parameters were evaluated, achieving a maximum micro-averaged F1 score of 0.22. Parameter size showed a statistically significant positive correlation with F1 score, while prompt strategy and temperature did not yield significant improvements, indicating current open-source LLMs are insufficient for production-grade ATT&CK classification.

arxiv arXiv cs.LG · 14d ago

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

We propose a method using velocity-field disagreement to quantify epistemic uncertainty in flow-matching vision-language-action models. This uncertainty estimate enables failure detection during deployment and active fine-tuning via the SAVE framework, which reduces expert demonstrations by at least 22% compared to baselines, with better-calibrated predictions on the LIBERO benchmark.