The llama.cpp project released build b11266, which includes a Vulkan optimization that loads F32 A matrices two elements at a time when they are 2-aligned. This change addresses performance issues on Intel hardware where loading floats one by one is inefficient, and it also corrects a missing validation for `a_offset` alignment in the original patch.
- The modification leverages existing 2-aligned load logic in `mul_mat_vec` to improve throughput on Intel BMG architectures.
- Benchmark results on an Intel B60 show significant speedups for specific matrix multiplication shapes, with performance increasing from 153.08 GFLOPS to 221.66 GFLOPS for one test case and up to 1.63 TFLOPS for another.
- The release provides binaries for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL, Snapdragon), Android, Windows, and openEuler.
This update improves inference performance on Intel GPUs using the Vulkan backend by optimizing memory access patterns for floating-point matrix operations.