The llama.cpp project released build b10693, which introduces support for Qualcomm Hexagon NPUs. This update enables runtime discovery of available NPU cores and allows for the creation of inference sessions on demand.

  • Added lazy session allocation and cleanup interfaces for Hexagon devices.
  • Implemented runtime discovery of available NPU cores.
  • Configured early rejection of non-existing devices during initialization.
  • Provided binaries for macOS, Linux, Windows, Android, and openEuler across CPU, Vulkan, ROCm, OpenVINO, SYCL, and CUDA backends.