The llama.cpp project has released build b10150, which includes adjustments to the ggml logic for offloading operations to the weight's backend. This update also addresses fixes for the llama dsv4 graph.

  • ggml: adjust logic for offloading ops to weight's backend
  • llama: dsv4 graph fixes

This release provides updated binaries for macOS, Linux, Windows, Android, and iOS across various hardware backends including CPU, CUDA, Vulkan, ROCm, and OpenVINO.