The llama.cpp project released build b11368, which introduces probabilistic sampling for simple draft and MTP (Multi-Token Prediction) workflows. This update modifies the drafting mechanism to use probabilistic selection while verifying targets via rejection sampling.

  • The drafter now uses probabilistic sampling, with target verification handled by rejection sampling.
  • Stale spec_draft_q buffers are dropped before the drafting process begins.
  • Grammar-constrained requests now support rejection sampling and fallback to argmax sampling.
  • A new flag enables probabilistic draft sampling, defaulting to greedy behavior.
  • The draft sampler no longer shares the target's RNG stream, and distributions are renormalized after masking.

This release provides binaries for macOS, Linux, Windows, Android, and openEuler across various hardware backends including CUDA, ROCm, Vulkan, OpenVINO, and SYCL.