Liquid AI has released LFM2.5-VL-3B, a 3-billion parameter vision-language model designed to deliver improved and faster performance on edge devices. The model pairs a SigLIP2 400M NaFlex vision encoder with the backbone of its text-only counterpart, trained on approximately 34 trillion tokens including four times more vision data than previous releases.

  • Key capabilities include enhanced screen and UI understanding, improved grounding for object detection, multi-image reasoning, and stronger function calling in both text-only and vision-text contexts.
  • The model fits in about 3 GB of memory, decoding at 228 tokens/s on an M5 Max and 116 tokens/s on a Ryzen AI Max+ 395, while reaching 20 tokens/s on a Galaxy S26 Ultra.
  • On GPU inference, it achieves roughly 11K tokens per second at high concurrency, which is approximately twice the throughput of larger 4B-class models and ahead of smaller 2B-class models.

The release aims to provide strong, general-purpose vision-language intelligence for high-volume on-device workloads, with day-one support for inference frameworks like llama.cpp, MLX, vLLM, SGLang, and ONNX.