AdaSpark is a scheduler for block drafters like DSpark that learns verify widths and acceptance probabilities during serving without prior profiling or calibration. It fits verify time as a function of context and candidate acceptance to the target's outcomes, allowing drafted and text-derived candidates to compete in a single best-first order.
- On six public datasets with three dense targets and one mixture-of-experts target, AdaSpark decodes 1.5-3.1x faster than llama.cpp's DSpark using the same drafters.
- The imparo engine with AdaSpark is 1.17-1.52x faster than running with a default three-token chain, a gain attributed solely to the scheduler.
- Without width sweeps, performance remains within 0.3% of the best pinned tree width on dense targets and ties the best pinned width on mixture-of-experts models.
This approach eliminates the need for offline calibration or sweeps while providing significant speedups over standard scheduling methods.