The llama.cpp project has released build b10649, which introduces a new feature for synthetic speculative acceptance. This addition is specifically designed for benchmarking purposes and is available in both the llama-server and llama-cli components.
- Adds benchmark-only synthetic speculative acceptance to llama-server and llama-cli.
- Includes macOS, iOS, Linux, Windows, Android, and openEuler binaries across CPU, GPU, and various accelerator backends.
- Disables KleidiAI support on macOS Apple Silicon for this build.
This release provides updated binaries for a wide range of platforms while adding specific tooling for evaluating speculative decoding performance.