The llama.cpp project released build b11454, which introduces support for the K2 Horizon model architecture, including both dense and Mixture-of-Value Attention (MoVA) variants.
- Added K2 Horizon GGUF conversion code, tensor loading, compute graph implementation, and tokenizer registration.
- Fixed Unicode regex splitting to correctly handle K2-Horizon's specific requirements on Windows and other platforms.
- Implemented Jinja template support for sequence indices in selectattr and rejectattr filters.
- Enabled chat support for K2 Horizon reasoning and tool calls, including response schema enforcement and YaRN beta metadata loading.
This update allows users to run K2 Horizon models locally using llama.cpp across various backends.