AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers
Melvin Vivas · X post · 2026-08-16 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
Release notes for version 0.4.0 of AIBackends, the creator's open-source Python package. It adds support for Liquid AI's LFM2.5-2.6B (on llama.cpp and Transformers) and LFM2.5-VL-3B (vision, on llama.cpp), plus CPU benchmark reports for both models.
Key points
- Install with: pip install aibackends
- LFM2.5-2.6B runs on both the llama.cpp and Transformers backends.
- LFM2.5-VL-3B adds vision support through llama.cpp.
- The release includes CPU benchmark reports for both models, useful for choosing a model for CPU-only deployment.
Resources mentioned
- AIBackends (PyPI 0.4.0) · tool · pypi.org · free
The creator's Python package for running local AI model backends. - AIBackends · repo · aibackends.com · free
Open-source API server runtime for common AI use cases that supports many models and providers (Ollama, LM Studio, OpenRouter, OpenAI, Anthropic).
Also in: Building a production website with Opus 5.5 and TanStack (Melvin Vivas on X · notes), CamelFlow: open-source visual viewer for Apache Camel routes (Melvin Vivas on X · notes), AI Backends: A Production AI Workflow Engineering Site (Link Share) (Melvin Vivas on X · notes), Demo: Claude Opus 5.5 Generating a Motion-Graphics Video for AIBackends (Melvin Vivas on X · notes) and 16 more - Liquid AI (@liquidai) on X · website · x.com · free
An AI company that builds efficient foundation models. The quoted post shows its PII handling working on Japanese text.
Also in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), Liquid AI LFM 2.5 Encoder: a CPU-friendly encoder model (Melvin Vivas on X · notes), Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes), Liquid AI's LFM2-Longevity models for aging-data analysis (Melvin Vivas on X · notes) and 19 more - LFM2.5-2.6B · tool · huggingface.co · free
Liquid AI's small language model, which can be paired with the LFM2.5-VL-3B vision model.
Also in: Coworker: Open-Source Work Agent That Runs on Small Local Models (Melvin Vivas on X · notes), Coworker with Liquid AI LFM2.5-2.6B via LM Studio on a Mac (Melvin Vivas on X · notes), Zero-Cost Coworker Setup: OpenRouter Free Models + Local LFM2.5 (Melvin Vivas on X · notes), Coworker: A Subscription-Free AI Agent on Local LFM2.5-2.6B (Melvin Vivas on X · notes) and 9 more - LFM2.5-VL-3B · tool · huggingface.co · free
Liquid AI's lightweight vision-language model for screen and document understanding, grounding and tool calling.
Also in: OCR Testing a Small Vision Model with LLM-Made Ground Truth (Melvin Vivas on X · notes), AIBackends Adds Support for LFM2.5-VL-3B (Melvin Vivas on X · notes), Using LFM2.5-VL-3B's Vision Capabilities (Liquid AI Guide) (Melvin Vivas on X · notes), Quick test of Liquid AI's LFM2.5-VL-3B vision model (Melvin Vivas on X · notes) and 1 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - Hugging Face Transformers · tool · github.com · free
Open-source library for loading and running pretrained models, which now supports GGUF directly.
Also in: Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), Run GGUF models directly in Hugging Face Transformers (Melvin Vivas on X · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes)
Try this
- Install with pip install aibackends and try the LFM2.5 models.
More in LLMOps, Deployment & Monitoring
- Hosted Qwen3.8 27B Costs More Than GPT 5.6 Luna, So Run It Locally
- Running Qwen3.8 27B Locally with Hermes
- AIBackends: Python Library for Local-Model AI Workflows
- AIBackends Adds Support for LFM2.5-VL-3B
- Running LFM2.5-2.6B Q4_K_M Locally with llama.cpp and Pi
- Run Qwen3.8-27B locally with llama.cpp (llama-server)