The CUDA implementation has been updated to prioritize whole-tile scheduling strategies for FlashAttention. This change aims to optimize the execution flow of attention mechanisms within the framework.
CUDA prefers whole-tile FlashAttention scheduling
Eqx-zoo provides verified Equinox ports of Hugging Face models
The author has built eqx-zoo, a library that loads Hugging Face Hub checkpoints directly as plain Equinox modules and verifies each model against the transformers library. The project currently supports Llama, Qwen2, Qwen3, Qwen3-MoE for generation, and BERT, RoBERTa, XLM-RoBERTa for embeddings.
llama.cpp b11401 release fixes logging and enables ANSI colors on Windows
The llama.cpp b11401 release addresses logging issues in the server router mode and adds support for ANSI colors on the Windows console. This update ensures that child process logs are correctly separated from state commands and renders colored output properly on Windows terminals.
Aleph Alpha releases Kolibri, a 78.1B English-German MoE model with 3.46B active parameters
Aleph Alpha has released Kolibri-1, an open-weight Mixture-of-Experts language model designed for bilingual English and German tasks. The model features 78.1 billion total parameters but activates only 3.46 billion per token, allowing it to run on a single B200, B300, or H200 GPU.
llama.cpp b11390 fixes CUDA MMQ memory fault
The llama.cpp project has released version b11390, which includes a fix for a CUDA MMQ memory fault that occurred when the number of experts was significantly larger than the micro-batch size.
llama.cpp b11382 adds f16 support to WebGPU fill/set_rows
The llama.cpp project has released build b11382, which includes a key update to its WebGPU backend. This release adds 16-bit floating-point (f16) support specifically for the `fill` and `set_rows` operations within the WebGPU implementation.