Running LFM2.5-2.6B Q4_K_M Locally with llama.cpp and Pi
Melvin Vivas · X post · 2026-08-15 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Agents, Tool Use & MCP · Level: intermediate
Summary
The creator runs Liquid AI's LFM2.5-2.6B at Q4_K_M quantization through llama.cpp's llama-server and connects it to the Pi agent. On a 16GB M1 MacBook it uses about 2.4GB of RAM, and it can read a website and turn the content into markdown. He includes the llama-server command to start router mode before configuring Pi.
Key points
- Model: LFM2.5-2.6B quantized to Q4_K_M (GGUF) and served with llama.cpp.
- Hardware: MacBook M1 with 16GB RAM; memory use is about 2.4GB.
- Working task: read a website's content and convert it into markdown.
- Start llama-server in router mode before configuring Pi: llama-server --host 127.0.0.1 --port 8080 --models-max 1
- --models-max 1 keeps only one model loaded at a time in router mode.
Resources mentioned
- LFM2.5-2.6B · tool · huggingface.co · free
Liquid AI's small language model, which can be paired with the LFM2.5-VL-3B vision model.
Also in: Coworker: Open-Source Work Agent That Runs on Small Local Models (Melvin Vivas on X · notes), Coworker with Liquid AI LFM2.5-2.6B via LM Studio on a Mac (Melvin Vivas on X · notes), Zero-Cost Coworker Setup: OpenRouter Free Models + Local LFM2.5 (Melvin Vivas on X · notes), Coworker: A Subscription-Free AI Agent on Local LFM2.5-2.6B (Melvin Vivas on X · notes) and 9 more - Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more
Try this
- Start llama-server with --host 127.0.0.1 --port 8080 --models-max 1, then configure Pi to use it.
- Try a Q4_K_M quant of LFM2.5-2.6B on a laptop.
- Build a local web-to-markdown agent using a small quantized model with llama.cpp and Pi.
More in LLMOps, Deployment & Monitoring
- AIBackends: Python Library for Local-Model AI Workflows
- AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers
- AIBackends Adds Support for LFM2.5-VL-3B
- Run Qwen3.8-27B locally with llama.cpp (llama-server)
- Unsloth releases Qwen3.8-27B GGUF quantizations
- Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp