The vLLM project released version 0.29.0, which includes a core bugfix regarding the dense prefix cache.
- The update applies the dense prefix cache default setting specifically to hybrid models.
The vLLM project released version 0.29.0, which includes a core bugfix regarding the dense prefix cache.
The llama.cpp project released version b10891, which addresses a critical stability issue on PowerVR GPUs. The update modifies the Vulkan backend to fall back to shared-memory reduction for dequantized matrix-vector multiplication (dmmv) operations.
DeepSeek AI has released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model featuring a 1-million-token context window and an FP4 KV cache that reduces global cache footprint to 890 bytes per token. The model utilizes a causal encoder-decoder architecture and Compressed Sparse Attention 2 (CSA2) to significantly lower memory requirements while maintaining performance.
The llama.cpp project has released build b10889, which includes a memory optimization to avoid allocating the V cache for the indexer since it is not used.
The llama.cpp project has released build b10888, which includes a new feature to add command-buffer debug labels for GPU profilers in the Vulkan backend.
The llama.cpp project released build b10886, which introduces Q1_0 vector intrinsic support for the s390x architecture. This update includes the implementation of `ggml_vec_dot_q1_0_q8_0` and updates documentation to reflect the new capability.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy