The ggml-openvino backend for llama.cpp has been updated to version 2026.4.1, introducing performance optimizations, expanded operator support, and improved device enumeration.
- Fusion of Mixture-of-Experts (MoE) routing and GDN QK normalization, with GPU MoE fusion enabled by default.
- Support for fused gate-up weights in models like gemma-4, achieving a prefill throughput increase from 66.16 to 1608.73 t/s on Arc B390.
- Fix for rank-3 axis handling in stateful execution, resolving graph build aborts for MoE models.
- Implementation of PRD-compliant device enumeration and memory reporting, ensuring accurate GPU/IGPU distinction.
- Addition of the k-requant option q4_asym64 and cache_only mode for importing compiled models directly from disk.
These changes improve inference performance for MoE architectures and provide more reliable hardware detection and configuration for OpenVINO users.