The llama.cpp project released build b11238, introducing a unified left-padding implementation via the new `ggml_pad_ext` function. This change consolidates previous methods used by audio encoders like Parakeet, LFM2-Audio, Granite Speech, and Gemma 4, ensuring bit-identical embeddings for Gemma 4 while simplifying backend support.
- The `ggml_pad_ext` node handles left padding in a single operation now that all backends support it, replacing complex roll or zero-fill concatenation logic.
- The update includes an optimization to skip DFlash2 taps that only read padding data, as terms at or past block_size are zero.
- Pre-built binaries are available for macOS (Apple Silicon and Intel), iOS, Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL, Snapdragon), Android, Windows, and openEuler.
This release improves consistency across different model architectures by standardizing padding operations and provides updated binaries for a wide range of hardware platforms.