The llama.cpp b11228 release introduces support for left and circular padding in the Metal backend's GGML_OP_PAD operation, aligning it with CPU, CUDA, and Vulkan implementations. This change fixes right padding issues for permuted sources and ensures consistent behavior across different hardware backends.

  • The implementation shifts source coordinates by left paddings and wraps them around for circular padding modes.
  • A test case has been added to cover the new functionality.
  • The f32_4 kernel is dropped because it is slower and fails two padding cases when enabled.
  • The code uses a function constant for the circular pad variant to compile the kernel once and specialize it per pipeline.

The release provides binaries for macOS, iOS, Linux, Windows, Android, and openEuler across various architectures and backend configurations including CUDA, Vulkan, ROCm, and OpenVINO.