Run LFM2.5-2.6B locally with llama.cpp
Melvin Vivas · X post · 2026-08-09 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
The creator quotes his own post that runs Liquid AI's LFM2.5-2.6B GGUF with llama-server. It uses about 6.9 GB of VRAM with a Q8_0 quant and a 128k context.
Key points
- Load the Q8_0 GGUF straight from Hugging Face with
llama-server -hf LiquidAI/LFM2.5-2.6B-GGUF:Q8_0. - Key flags: --jinja, --reasoning-format auto, -ngl all, -fa on, -c 128000, --port 8000.
- VRAM usage is about 6.9 GB.
Resources mentioned
- LiquidAI/LFM2.5-2.6B-GGUF · tool · huggingface.co · free
The Hugging Face repo with GGUF quantized weights of Liquid AI's LFM2.5-2.6B model.
Also in: Serving LFM2.5-2.6B with llama-server: full command (Melvin Vivas on X · notes) - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more
Try this
- Serve LFM2.5-2.6B locally with llama-server using the quoted command.
More in LLMOps, Deployment & Monitoring
- Running Liquid AI LFM2.5-2.6B Locally with llama-server
- Adding LFM2.5-2.6B support to AIBackends with Cursor
- Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder
- Serving LFM2.5-2.6B with llama-server: full command
- Personal AI Computer to Run DeepSeek V4-Flash Locally
- Running DeepSeek V4 Flash Locally: RAM Needs for 4-bit and 3-bit Quants