Aakib Ansari.
Back to articles
Deep Dive

Liquid AI LFM2.5-VL-3B Deep Dive: Sub-4GB Edge Vision, ScreenSpot Mastery, and 228 tok/s Execution

Md Aakib Ansari
Md Aakib AnsariWeb Developer & AI Tools Reviewer
•6 min read•Model: Liquid AI LFM2.5-VL-3B•Company: Liquid AI
Liquid AI LFM2.5-VL-3B Deep Dive: Sub-4GB Edge Vision, ScreenSpot Mastery, and 228 tok/s Execution

While frontier AI labs race to scale multi-trillion-parameter models across multi-megawatt data centers, a critical operational bottleneck remains unsolved: autonomous desktop agents need real-time visual perception with sub-second latency and zero per-frame API costs. On August 12, 2026, Liquid AI addressed this bottleneck with the release of LFM2.5-VL-3B, a 3.1-billion-parameter open-weights vision-language model engineered to run directly on consumer laptops. By bypassing external cloud round-trips and fitting comfortably into 3.3 gigabytes of unified system RAM, the model demonstrates that edge-native architectures can match or exceed cloud giants on focused desktop automation benchmarks.

Vitals & Architecture Specifications

| Specification | Details | |---|---| | Provider | Liquid AI, Inc. | | Model Identifier | LiquidAI/LFM2.5-VL-3B (PyTorch) / LFM2.5-VL-3B-GGUF | | Architecture Tier | Hybrid Liquid Neural Network + Grouped-Query Attention (GQA) | | Parameter Count | 3.1 billion total parameters | | Memory Footprint | ~3.3 GB (Quantized 4-bit / 8-bit GGUF formats) | | Context Window | 131,072 tokens (128K native) | | Image Input Resolution | Dynamic patch encoding up to 1344×1344 pixels | | Inference Throughput | 228 tokens/sec (Apple M5 Max) / 116 tokens/sec (AMD Ryzen AI Max+ 395) | | Open Licensing | LFM Open License v1.0 (Commercial self-hosting supported) | | Modalities | Text, High-Resolution Images, UI Screenshots, Visual Grounding | | Serving Ecosystem | llama.cpp, Ollama, vLLM, Hugging Face Transformers |

Benchmark Breakdown: Screen Reading and UI Grounding

Technical evaluations published in Liquid AI's official release paper and verified by Developers Digest and MarkTechPost reveal specialized advantages in digital screen interaction:

  • ScreenSpot-v2 (Self-Reported by Liquid AI): LFM2.5-VL-3B achieved an 80.7 average accuracy across web, desktop, and mobile interface parsing. This score surpasses Google's larger 8B Gemma-4-E4B (which scored 51.2) by 29.5 percentage points, highlighting the effectiveness of Liquid AI's dense UI-token pre-training.
  • RefCOCO Precision@1 (Visual Object Grounding): The model recorded 87.9%, reliably outputting exact normalized bounding box coordinates ([ymin, xmin, ymax, xmax]) for referenced UI controls and interactive buttons.
  • ToolSandbox (Function Calling): Scored 59.5, more than doubling the baseline set by previous compact models and confirming its ability to translate visual observations into structured JSON tool calls.
  • Where It Trails (Broad Abstract Reasoning): On general academic knowledge evaluations like MMLU, LFM2.5-VL-3B scored moderately at 62.4%, trailing dense cloud models. Because it functions as a direct-answer model rather than a chain-of-thought reasoning engine, it excels at fast perception rather than multi-step mathematical theorem proving.

The Advantage of Hybrid Linear-Attention on Edge Hardware

Traditional vision transformers incur quadratic compute costs as image resolutions and context lengths increase. Liquid AI's hybrid architecture replaces standard self-attention in intermediate layers with adaptive linear operators:

  1. Constant Memory Footprint During Generation: The hybrid state-space operators maintain linear memory scaling, enabling smooth multi-screenshot tracking during long desktop automation sessions.
  2. Instant Local Time-to-First-Token (TTFT): With pre-fill and decode speeds exceeding 200 tokens per second on Apple silicon, local agents can capture a screenshot, identify a button, and issue a mouse-click command in under 80 milliseconds.
  3. Complete Offline Privacy: Highly regulated enterprise environments (healthcare, legal, defense) can deploy the GGUF checkpoint without streaming proprietary client desktop screens to third-party cloud APIs.

Production Implementation Blueprints

Developers integrating LFM2.5-VL-3B into local desktop automation loops can utilize the following structured blueprints:

Blueprint 1: Local llama.cpp Screenshot Grounding

from llama_cpp import Llama
from llama_cpp.llama_chat_format import Llava15ChatHandler
chat_handler = Llava15ChatHandler(clip_model_path="lfm2.5-vl-mmproj.bin")
llm = Llama(
    model_path="lfm2.5-vl-3b-q4_k_m.gguf",
    chat_handler=chat_handler,
    n_ctx=4096,
    n_gpu_layers=-1
)
response = llm.create_chat_completion(
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Locate the 'Export to CSV' button and return its bounding box."},
                {"type": "image_url", "image_url": "file:///tmp/screen_capture.png"}
            ]
        }
    ]
)
print(response["choices"][0]["message"]["content"])

Blueprint 2: Structured Tool-Call Schema for Local OS Navigation

<system>
You are an on-device OS navigation engine. Output JSON actions only: {"action": "click"|"type", "coordinates": [x, y], "text": "optional"}.
</system>
<user>
Screenshot attached. Click the blue 'Deploy Staging' button located in the navigation header.
</user>

Strategic Implications

LFM2.5-VL-3B illustrates how open-weight edge models are carving out indispensable niches alongside massive frontier systems. While cloud flagships like GPT-6 Astra and Claude Opus 5 handle high-level project orchestration, lightweight 3B vision engines like LFM2.5-VL-3B provide the fast, low-cost local eyes necessary to make autonomous desktop agents practical.

Frequently Asked Questions

Can Liquid AI LFM2.5-VL-3B run on standard consumer laptops?
Yes. In quantized GGUF format, the model requires approximately 3.3 GB of RAM, allowing it to run smoothly alongside everyday applications on Apple M-series Macs and modern PC laptops.
How does LFM2.5-VL-3B handle digital screens compared to cloud models?
With an 80.7 score on ScreenSpot-v2, LFM2.5-VL-3B outperforms larger 8B open models and approaches cloud-tier accuracy for UI control detection, while offering sub-80ms local response times with zero API fees.
Is LFM2.5-VL-3B permitted for commercial enterprise use?
Yes, Liquid AI released the model under the LFM Open License v1.0, permitting commercial self-hosting, on-premise deployment, and commercial product integration.

Related Articles