Liquid AI has released an experimental DSpark draft model for its vision-language model LFM2.5-VL-3B, enabling faster decoding on edge devices and GPUs without changing output quality.
The drafter adds approximately 8.9% to the deployed model's parameter count and achieves decoding throughput improvements of up to 2.66x on H100 GPUs and 3.13x on Apple Silicon edge devices.
End-to-end throughput gains reach 2.27x on GPUs and 2.62x on edge hardware, with the system maintaining a throughput advantage across varying concurrency levels in SGLang.
The authors note that speculative decoding only accelerates the decode phase, meaning vision encoding and prefill costs remain unchanged, which limits overall latency improvements on resource-constrained edge devices.