The llama.cpp project released build b11284, which includes a fix for the OpenVINO backend to correctly handle GET_ROWS operations on quantized weight views.
- The update resolves view_src when collecting weight Constants, preventing views over quantized weights from becoming dynamic typed Parameters.
- The row offset of the view is now folded into the gather indices instead of slicing the dequantization subgraph.
- This change lifts the rejection of quantized src0 views with nonzero offsets, allowing them to work with OpenVINO.
The release provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU (CUDA, Vulkan, ROCm, SYCL), and NPU backends.