The llama.cpp project released build b11238, introducing a unified left-padding implementation via the new `ggml_pad_ext` function. This change consolidates previous methods used by audio encoders like Parakeet, LFM2-Audio, Granite Speech, and Gemma 4, ensuring bit-identical embeddings for Gemma 4 while simplifying backend support.

  • The `ggml_pad_ext` node handles left padding in a single operation now that all backends support it, replacing complex roll or zero-fill concatenation logic.
  • The update includes an optimization to skip DFlash2 taps that only read padding data, as terms at or past block_size are zero.
  • Pre-built binaries are available for macOS (Apple Silicon and Intel), iOS, Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL, Snapdragon), Android, Windows, and openEuler.

This release improves consistency across different model architectures by standardizing padding operations and provides updated binaries for a wide range of hardware platforms.