Liquid AI LFM2.5-VL-3B Deep Dive: Sub-4GB Edge Vision, ScreenSpot Mastery, and 228 tok/s Execution


While frontier AI labs race to scale multi-trillion-parameter models across multi-megawatt data centers, a critical operational bottleneck remains unsolved: autonomous desktop agents need real-time visual perception with sub-second latency and zero per-frame API costs. On August 12, 2026, Liquid AI addressed this bottleneck with the release of LFM2.5-VL-3B, a 3.1-billion-parameter open-weights vision-language model engineered to run directly on consumer laptops. By bypassing external cloud round-trips and fitting comfortably into 3.3 gigabytes of unified system RAM, the model demonstrates that edge-native architectures can match or exceed cloud giants on focused desktop automation benchmarks.
Vitals & Architecture Specifications
| Specification | Details | |---|---| | Provider | Liquid AI, Inc. | | Model Identifier | LiquidAI/LFM2.5-VL-3B (PyTorch) / LFM2.5-VL-3B-GGUF | | Architecture Tier | Hybrid Liquid Neural Network + Grouped-Query Attention (GQA) | | Parameter Count | 3.1 billion total parameters | | Memory Footprint | ~3.3 GB (Quantized 4-bit / 8-bit GGUF formats) | | Context Window | 131,072 tokens (128K native) | | Image Input Resolution | Dynamic patch encoding up to 1344×1344 pixels | | Inference Throughput | 228 tokens/sec (Apple M5 Max) / 116 tokens/sec (AMD Ryzen AI Max+ 395) | | Open Licensing | LFM Open License v1.0 (Commercial self-hosting supported) | | Modalities | Text, High-Resolution Images, UI Screenshots, Visual Grounding | | Serving Ecosystem | llama.cpp, Ollama, vLLM, Hugging Face Transformers |
Benchmark Breakdown: Screen Reading and UI Grounding
Technical evaluations published in Liquid AI's official release paper and verified by Developers Digest and MarkTechPost reveal specialized advantages in digital screen interaction:
- ScreenSpot-v2 (Self-Reported by Liquid AI): LFM2.5-VL-3B achieved an 80.7 average accuracy across web, desktop, and mobile interface parsing. This score surpasses Google's larger 8B Gemma-4-E4B (which scored 51.2) by 29.5 percentage points, highlighting the effectiveness of Liquid AI's dense UI-token pre-training.
- RefCOCO Precision@1 (Visual Object Grounding): The model recorded 87.9%, reliably outputting exact normalized bounding box coordinates ([ymin, xmin, ymax, xmax]) for referenced UI controls and interactive buttons.
- ToolSandbox (Function Calling): Scored 59.5, more than doubling the baseline set by previous compact models and confirming its ability to translate visual observations into structured JSON tool calls.
- Where It Trails (Broad Abstract Reasoning): On general academic knowledge evaluations like MMLU, LFM2.5-VL-3B scored moderately at 62.4%, trailing dense cloud models. Because it functions as a direct-answer model rather than a chain-of-thought reasoning engine, it excels at fast perception rather than multi-step mathematical theorem proving.
The Advantage of Hybrid Linear-Attention on Edge Hardware
Traditional vision transformers incur quadratic compute costs as image resolutions and context lengths increase. Liquid AI's hybrid architecture replaces standard self-attention in intermediate layers with adaptive linear operators:
- Constant Memory Footprint During Generation: The hybrid state-space operators maintain linear memory scaling, enabling smooth multi-screenshot tracking during long desktop automation sessions.
- Instant Local Time-to-First-Token (TTFT): With pre-fill and decode speeds exceeding 200 tokens per second on Apple silicon, local agents can capture a screenshot, identify a button, and issue a mouse-click command in under 80 milliseconds.
- Complete Offline Privacy: Highly regulated enterprise environments (healthcare, legal, defense) can deploy the GGUF checkpoint without streaming proprietary client desktop screens to third-party cloud APIs.
Production Implementation Blueprints
Developers integrating LFM2.5-VL-3B into local desktop automation loops can utilize the following structured blueprints:
Blueprint 1: Local llama.cpp Screenshot Grounding
from llama_cpp import Llama
from llama_cpp.llama_chat_format import Llava15ChatHandler
chat_handler = Llava15ChatHandler(clip_model_path="lfm2.5-vl-mmproj.bin")
llm = Llama(
model_path="lfm2.5-vl-3b-q4_k_m.gguf",
chat_handler=chat_handler,
n_ctx=4096,
n_gpu_layers=-1
)
response = llm.create_chat_completion(
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Locate the 'Export to CSV' button and return its bounding box."},
{"type": "image_url", "image_url": "file:///tmp/screen_capture.png"}
]
}
]
)
print(response["choices"][0]["message"]["content"])
Blueprint 2: Structured Tool-Call Schema for Local OS Navigation
<system>
You are an on-device OS navigation engine. Output JSON actions only: {"action": "click"|"type", "coordinates": [x, y], "text": "optional"}.
</system>
<user>
Screenshot attached. Click the blue 'Deploy Staging' button located in the navigation header.
</user>
Strategic Implications
LFM2.5-VL-3B illustrates how open-weight edge models are carving out indispensable niches alongside massive frontier systems. While cloud flagships like GPT-6 Astra and Claude Opus 5 handle high-level project orchestration, lightweight 3B vision engines like LFM2.5-VL-3B provide the fast, low-cost local eyes necessary to make autonomous desktop agents practical.


