Liquid AI Releases LFM2.5-VL-3B: On-Device Multimodal Model with 80.7 ScreenSpot Score


Liquid AI has officially launched LFM2.5-VL-3B, a compact 3.1-billion-parameter vision-language model (VLM) engineered to run entirely on-device across consumer laptops, mobile silicon, and embedded edge systems. Available on Hugging Face under the LFM Open License v1.0 in standard PyTorch and quantized GGUF formats, the model is built on Liquid AI's proprietary hybrid architecture combining adaptive linear operators with grouped-query attention. The entire model operates within a 3.3-gigabyte memory footprint, enabling local deployment without dedicated data center accelerators. According to benchmark results published by Liquid AI and analyzed by MarkTechPost and Unite.AI, LFM2.5-VL-3B achieved an 80.7 average accuracy on ScreenSpot-v2—outperforming larger 8B dense vision models such as Gemma-4-E4B by 29.5 points on UI element identification. It also recorded 87.9 precision@1 on RefCOCO visual object grounding and doubled ToolSandbox performance to 59.5 for agentic function calling. On local hardware benchmarks, the non-reasoning direct-answer model demonstrated inference speeds of 228 tokens per second on an Apple M5 Max and 116 tokens per second on AMD Ryzen AI Max+ 395 silicon, positioning it as an ultra-fast perception engine for local desktop automation agents.


