LFM2.5-VL-3B: a lightweight VLM for screens, documents and tool calls
Melvin Vivas · X post · 2026-08-12 · Open on X
Topics: LLM Fundamentals, AI Agents, Tool Use & MCP, Industry Trends & Job Market · Level: intermediate
Summary
Liquid AI released LFM2.5-VL-3B, a lightweight vision-language model. It reads mobile, web and desktop screens, documents and the physical world, and it can call tools from text or image input. The creator plans to pair it with the LFM2.5-2.6B text model.
Key points
- LFM2.5-VL-3B understands digital screens across mobile, web and desktop.
- It can ground objects to coordinates, which is useful for computer-use or GUI agents.
- It reads text and charts in documents.
- It can call tools from either text or image input.
- It can be combined with the LFM2.5-2.6B language model for a small local setup.
Resources mentioned
- LFM2.5-VL-3B · tool · huggingface.co · free
Liquid AI's lightweight vision-language model for screen and document understanding, grounding and tool calling.
Also in: OCR Testing a Small Vision Model with LLM-Made Ground Truth (Melvin Vivas on X · notes), AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers (Melvin Vivas on X · notes), AIBackends Adds Support for LFM2.5-VL-3B (Melvin Vivas on X · notes), Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide) (Melvin Vivas on X · notes) and 1 more - LFM2.5-2.6B · tool · huggingface.co · free
Liquid AI's small language model, which can be paired with the LFM2.5-VL-3B vision model.
Also in: Coworker: Open-Source Work Agent That Runs on Small Local Models (Melvin Vivas on X · notes), Coworker with Liquid AI LFM2.5-2.6B via LM Studio on a Mac (Melvin Vivas on X · notes), Zero-Cost Coworker Setup: OpenRouter Free Models + Local LFM2.5 (Melvin Vivas on X · notes), Coworker: A Subscription-Free AI Agent on Local LFM2.5-2.6B (Melvin Vivas on X · notes) and 9 more - Liquid AI (@liquidai) on X · website · x.com · free
An AI company that builds efficient foundation models. The quoted post shows its PII handling working on Japanese text.
Also in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), Liquid AI LFM 2.5 Encoder: a CPU-friendly encoder model (Melvin Vivas on X · notes), Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes), Liquid AI's LFM2-Longevity models for aging-data analysis (Melvin Vivas on X · notes) and 19 more
Try this
- Combine LFM2.5-VL-3B (vision) with LFM2.5-2.6B (text) into a small local multimodal assistant.
More in LLM Fundamentals
- Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide)
- Cohere's North Micro Vision: a small open-source vision model for documents
- Quick test of Liquid AI's LFM2.5-VL-3B vision model
- NVIDIA Nemotron 3.5 Lightning: Fast Open MoE Model for Agents
- Claude Sonnet 5 Pricing Made Permanent ($2/$10 per M Tokens)
- Running Liquid AI Models On-Device with the Apollo iPhone App