vLLM has released version 0.27.1, a patch update built on top of the previous v0.27.0 release.
- The update introduces support for quantized DSpark Markov heads.
vLLM has released version 0.27.1, a patch update built on top of the previous v0.27.0 release.
The llama.cpp b10361 release addresses a critical bug where Sliding Window Attention (SWA) was not enabled for the LGAI-EXAONE 4.5 model due to incorrect parameter ordering during loading.
ALTK-Evolve, a system for agentic memory that converts agent trajectories into reusable guidelines without weight updates, matches or exceeds the accuracy of ACE while using significantly fewer tokens. Unlike ACE, which injects a comprehensive playbook on every step, ALTK-Evolve uses selective retrieval to deliver only relevant guidelines based on model capacity.
Researchers identify Relevant Visual Information Shift (RVIS) during decoding as the primary cause for the failure of existing visual token pruning methods in Multimodal Large Language Models (MLLMs). To address this, they propose Decoding-stage Shift-aware Token Pruning (DSTP), a training-free framework that aligns visual tokens with shifting reasoning requirements.
webAI has released TwIL-LM, a two-model family of formal-logic reasoners with 1.7B and 3B parameters designed for autoformalization on local hardware. The models translate English into first-order logic and check conclusion validity, with the 3B variant achieving a macro gate score of 0.4218 on in-domain formal logic tasks.
Meta released Muse Glimmer, a 30B dense multimodal model optimized for local agent deployment under Apache 2.0, while also promising future weights for Muse Spark 1.2.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy