Prism-ML has released the Ternary-Bonsai-27B model, available as GGUF checkpoints on Hugging Face.
The release provides a 27-billion parameter variant of the Qwen 3.6 architecture optimized for local inference through the GGUF format.
Prism-ML has released the Ternary-Bonsai-27B model, available as GGUF checkpoints on Hugging Face.
The release provides a 27-billion parameter variant of the Qwen 3.6 architecture optimized for local inference through the GGUF format.
The llama.cpp b10677 release addresses a critical bug in the Vulkan backend where missing view-alias dependencies caused the graph optimizer to incorrectly reorder nodes.
The llama.cpp release b10676 corrects a bug in `conv_transpose_2d` where only the first batch was computed, leaving subsequent batches as zero. This fix applies to both the core ggml library and the Metal backend implementation.
The llama.cpp project released build b10673, which includes specific performance optimizations for Apple M4 hardware. This update introduces fa_vec_tuned_table records into the ggml-metal-tuning module to improve Metal backend efficiency.
The llama.cpp project released version b10670, which includes a specific optimization for Intel's Xe2 architecture. The update routes the quantized key-value (KV) decode operation to use the TILE backend exclusively for BMG hardware, while keeping the VEC backend on other architectures until they are validated.
The llama.cpp project released build b10672, which updates the OpenVINO backend to version 2026.3.1 and adds support for Whisper.cpp models on the NPU.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy