Research paper — korshunov.ai

Research paper Page 1 / 16

Gold Points Sniper: Self-guided Visual Reasoning for Fine-grained Action Understanding

Gold Points Sniper (GPS) enables lightweight vision-language models to perform self-guided multimodal reasoning for fine-grained human action understanding. By integrating a Gold Points Extractor, Selective Socratic Questioner, and Semantic Entailment Evaluator, GPS achieves performance comparable to GPT-4o while maintaining superior factual accuracy on CAP benchmark-based instruction-tuning data.

lab Hugging Face Blog · 18h ago

NVIDIA NeMo AutoModel Speeds Up Transformer Fine-Tuning

NVIDIA's NeMo AutoModel enables faster fine-tuning of transformer models by automating model selection and optimization. It reduces development time and improves efficiency in training large language models on NVIDIA hardware.

arxiv arXiv cs.AI · 18h ago

LLMs Use Difference-Making Logic to Learn Causal Structure

Large language models learn causal structure through a difference-making logic, akin to the experimental method. This approach identifies which word sequences influence outcomes and which do not, using vast text data during training. Architectural features like token embeddings and self-attention support this inductive process by detecting patterns of variation and indifference in language.

arxiv arXiv cs.AI · 18h ago

DreamUV: End-to-End Flow Matching for Artist-like UV Unwrapping

DreamUV introduces an end-to-end learning framework that treats UV unwrapping as a generative Flow Matching problem. It learns a mesh-conditioned transport process to generate artist-like UV layouts, with boundary-aware training and Model-in-the-Loop fine-tuning to ensure seam geometry and practical validity. Results show straighter seams, tighter axis-aligned islands, and superior alignment with professional artist preferences.

arxiv arXiv cs.AI · 19h ago

A Differentiable Atari VCS for Explainable AI

A fully differentiable emulator of the Atari 2600 VCS is presented, reproducing all 64 ALE games with bit-for-bit accuracy in RAM and screen output. The system enables gradient-based explainable AI by providing a complex, fully known ground truth, with both Julia and JAX implementations validated against a reference emulator and capable of high-throughput differentiable rollouts on GPU.

arxiv arXiv cs.AI · 19h ago

Self-Evolving Cognitive Framework for Embodied Scientific Intelligence

The paper proposes a self-evolving cognitive framework that uses causal world modeling to enable embodied systems to continuously refine their internal models through interaction. It integrates causal modeling, intervention-driven reasoning, and continual refinement, redefining embodied interaction as an epistemic process for causal discovery and knowledge acquisition. The framework supports a shift from predictive to epistemic intelligence, with a new benchmark for evaluating self-evolving embodied scientific intelligence.

arxiv arXiv cs.AI · 19h ago

Character Variety in LLM-Generated Stories

This study compares characters in LLM-generated and human-written stories using narratological dimensions. It finds that while LLMs produce characters with similar basic traits, they lack diversity in complex character features like stylization and wholeness. The research highlights key differences in character depth and variety between human and machine-generated narratives.

arxiv arXiv cs.AI · 19h ago

PRIME: Evaluating Prompt Resolution in Conflicting Instructions

PRIME introduces a framework to analyze how large language models handle conflicting instructions by generating calibrated conflicts in response length, format, and reasoning. The study finds that conflict type has a greater impact on model behavior than model size, revealing diverse failure modes across conflict categories. Results highlight the need for conflict awareness and suggest instruction following cannot be reliably assessed through isolated benchmarks alone.

arxiv arXiv cs.AI · 19h ago

FACTOR Enables Adaptive Verification for Factuality in Long-Form Generation

FACTOR introduces an inference-time model that adapts verification criteria based on claim-level uncertainty. It improves factuality and reduces verification cost by dynamically allocating effort to high-risk claims, demonstrating effective and model-agnostic performance on the FactScore benchmark.

arxiv arXiv cs.AI · 20h ago

VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows

VADAOrchestra introduces a neurosymbolic framework that combines LLM-based workflow orchestration with Datalog+/- symbolic reasoning. It enables adaptive, explainable decision-making by incrementally planning workflows and executing logical inference on demand, offering verifiable traces, auditability, and scalability over large datasets.

media r/LocalLLaMA · 20h ago

My micro-benchmark: how good are LLMs at simulating wetting behaviour?

The author benchmarks LLMs in simulating wetting behaviour using Surface Evolver, a 1992 tool for modeling liquid surfaces. LLMs are evaluated objectively by comparing their generated datafiles against reference implementations, with results showing pass counts and token costs for each model.

arxiv arXiv cs.AI · 20h ago

SCOPE: Self-Adaptive Symbolic Planning for Open-Ended Environments

SCOPE introduces a framework that refines action plans and evolves symbolic world models in open-ended environments. It combines a Symbolic Execution Simulator and a Self-Adaptive Symbolic Memory to improve plan completeness, perturbation resilience, and cross-task adaptability.

arxiv arXiv cs.AI · 20h ago

LLM-Orchestrated Agent for SOI Directional Coupler Design

A large language model orchestrates the design of a silicon-on-insulator 2x2 directional coupler by proposing gap values and assessing convergence. The design is validated through eigenmode and FDTD simulations on a common 2D effective-index model, showing a consistent phase offset of 2.837(11) micrometers that is corrected in a closed-loop process. The final device achieves a 50/50 split with a cross fraction of 0.498, within 0.0017 of the target.

lab Microsoft Research Blog · 21h ago

Talos: Automated Genomic Reanalysis for Rare Disease Diagnosis

Talos is an open-source tool that automates iterative reanalysis of genomic data to identify rare disease diagnoses. It achieved a 90% recovery rate of in-scope diagnoses with only 1.3 candidate variants per patient, and delivered 241 new diagnoses across 5,000 undiagnosed patients, with most new findings emerging within 32 days of evidence publication.

arxiv arXiv cs.AI · 21h ago

Prompt-Side Preprocessing Enhances Edge AI Accuracy

A structured prompt framework improves local LLM accuracy in environmental monitoring by transforming raw sensor data into enriched textual representations. Evaluations on indoor and outdoor datasets show local model accuracy increases from 50.9% to 81.7% indoors and 63.7% to 79.3% outdoors with enriched prompts, while maintaining low latency of nearly 0.22 seconds in no-chain-of-thought mode.

arxiv arXiv cs.AI · 21h ago

Grounded Scaling: Determinism as a Core Limit in Agentic AI

Agentic AI performance degrades exponentially in non-deterministic environments, with k-step success falling as δ^k when per-step determinism δ < 1. The paper introduces a framework linking environment determinism to task success, verifiability, and skill evolution, proposing a Supply Certainty Index and a five-level Determinism Maturity Model. It challenges prevailing views by identifying determinism as a binding constraint across compute, data, embodiment, and alignment.

arxiv arXiv cs.AI · 21h ago

Fed-CausalDiff: Decoupled Synchronization for Federated Do-Simulation

Fed-CausalDiff introduces a federated causal diffusion framework that enables do-simulation in decentralized settings. It decomposes latent state evolution into global and local components, allowing decoupled synchronisation to reduce communication cost while maintaining accurate policy evaluation and ATE estimation.

arxiv arXiv cs.AI · 21h ago

Generative Robust Optimisation Framework

Generative Robust Optimisation (GRO) introduces a deep generative model to define uncertainty sets, capturing nonlinear correlations, asymmetry, and multimodality. A five-point evaluation framework assesses neural network-based uncertainty sets across reconstruction fidelity, distribution matching, latent regularity, robust relevance, and computational tractability, with experiments validating GRO's effectiveness in production planning and facility location.

arxiv arXiv cs.AI · 22h ago

Concept-Constrained Prompt Learning for Few-Shot CLIP Adaptation

CCPL introduces a lightweight framework that anchors class prompts to frozen concept prototypes, improving few-shot CLIP adaptation. It achieves better base-to-new performance on DTD and EuroSAT compared to CoOp, with consistent gains from text-space concept regularization, while maintaining neutrality on OxfordPets. The method uses concept dropout and controllable ensemble fusion at inference, with results sensitive to dataset semantics and protocol.

arxiv arXiv cs.AI · 22h ago

Context-Aware Distillation and Ablation for Text2DSL

A new Text2DSL system uses context-aware distillation with a structured context of BNF grammar, API specification, and closed identifier vocabulary. Ablation studies show that the vocabulary has the largest impact on semantic quality, while API and BNF significantly improve structural validity, confirming structured context as a critical, load-bearing component.