The llama.cpp project released version b11190, which includes a fix for the mel preprocessor used in the LFM2 audio model. This change addresses discrepancies that caused different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances.
- Uses log(x + 2^-24) instead of clamping to the log floor.
- Implements a symmetric Hann window, equivalent to torch.hann_window(periodic=False).
- Adds normalization epsilon to the standard deviation rather than inside the square root.
- Reduces mel relative L2 error versus liquid-audio from ~3-4% to ~2e-6 median for both English and Japanese.
The update ensures that the lfm2a preprocessor matches the reference implementation's output, improving transcription accuracy across multiple languages and precision formats.