Liquid AI has released LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The drafter adds approximately 280M parameters and accelerates token generation without altering the model's output distribution.

  • The 279.5M parameter drafter uses a simplified attention-only architecture with 4 layers, increasing the total deployed parameter count by 8.9%.
  • Decoding speeds reach up to 3.13x on Apple M5 Max silicon and 2.66x on NVIDIA H100 GPUs, though end-to-end gains are lower due to fixed image encoding times.
  • Weights are available in Safetensors and GGUF formats with day-one support for SGLang, MLX-VLM, and llama.cpp under the LFM Open License v1.0.
  • Output remains identical to the base model under greedy decoding, while higher temperatures reduce acceptance rates and throughput.

The release allows users to significantly speed up inference on compatible hardware, with Liquid AI noting that prefill and vision encoding phases remain unaccelerated.