LFM2.5-VL-3B is a new vision-language model that delivers competitive performance against larger models while maintaining low latency for real-time and on-device applications. It builds on the previous LFM2-VL-3B release with significant enhancements in screen understanding, grounding, function calling, and multi-image input capabilities.
- Screen/UI understanding averages 80.7 on ScreenSpot-v2, outperforming Gemma-4-E4B (51.2) and Qwen 3.5 4B (78.5).
- Function calling scores more than double from 26.4 to 59.5 on ToolSandbox and climb to 32.5 on BFCL v4.
- Grounding precision@1 on RefCOCO rises from 57.1 to 87.9 through scaled synthetic data.
- Multi-image reasoning improves BLINK scores from 50.2 to 61.5 and MUIRBench from 34.9 to 58.3.
- The model decodes 228 tokens/s on Apple M5 Max and 20 tokens/s on Galaxy S26 Ultra, using approximately 3 GB of memory.
The non-reasoning architecture ensures fast first-token responses, making it suitable for private, on-device inference and high-throughput GPU serving.