The llama.cpp project has released build b10649, which introduces a new feature for synthetic speculative acceptance. This addition is specifically designed for benchmarking purposes and is available in both the llama-server and llama-cli components.

  • Adds benchmark-only synthetic speculative acceptance to llama-server and llama-cli.
  • Includes macOS, iOS, Linux, Windows, Android, and openEuler binaries across CPU, GPU, and various accelerator backends.
  • Disables KleidiAI support on macOS Apple Silicon for this build.

This release provides updated binaries for a wide range of platforms while adding specific tooling for evaluating speculative decoding performance.