The llama.cpp b11224 release addresses incorrect computation results in the Vulkan backend when matrix multiplication operations read slices of larger, strided tensors such as KV caches.
- The batch stride for in-place src0 and src1 tensors is now correctly derived from nb[2] instead of ne00*ne01, ensuring proper row access across all heads.
- In-place A and B ranges are sized by their strided extent using ggml_nbytes or ggml_vk_subbuffer to prevent reading zeroed-out data or hanging on NVIDIA hardware without coopmat2.
- mul_mat_id and mul_mm paths are updated to handle strided expert views and partial tiles correctly, avoiding hangs on NVFP4 operations.
This fix ensures accurate inference results for decoder self-attention mechanisms that rely on strided memory layouts in the Vulkan backend.