The llama.cpp project released version b11310, which includes a fix for the CUDA backend's IQ4_NL dequantization logic. The update addresses a memory safety issue where threads could read and write past the end of a row when processing blocks shorter than QK_K.

  • Guards the iq4_nl dequantize row kernel against short rows by skipping sub-blocks that start at or past the row length.
  • Provides pre-built binaries for macOS (Apple Silicon and Intel), Linux, Windows, Android, and openEuler across CPU, CUDA, Vulkan, ROCm, OpenVINO, SYCL, and Snapdragon backends.

This fix prevents potential out-of-bounds memory access during inference with IQ4_NL quantized models on CUDA devices.