The llama.cpp project released build b11351, introducing new methods to the ggml backend buffer type interface. The update adds `alloc_buffer_n` and `get_alloc_size_n` to allow for more granular control over buffer allocation and size querying.

  • Added `ggml_backend_buft_alloc_buffer_n` public API and corresponding callback to `ggml_backend_buffer_type_i`.
  • Implemented default handling for multi-buffer splitting and tensor allocation via `ggml_tallocr`.
  • Added optional `get_alloc_size_n` callback to share tensor-to-buffer planning logic.
  • Replaced unchecked realloc with std::vector in the default implementation.
  • Removed temporary `ggml_backend_meta_alloc_ctx_tensors_from_buft` function.

This release provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU (CUDA, Vulkan, ROCm, OpenCL), NPU (Hexagon), and other accelerators.