Serving LFM2.5-2.6B with llama-server: full command
Melvin Vivas · X post · 2026-08-09 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
A full llama-server command for running Liquid AI's LFM2.5-2.6B GGUF (Q8_0) on a GPU. It uses about 6.9 GB of VRAM, a 128k context, and the recommended sampling settings. It needs a CUDA-enabled build of llama.cpp.
Key points
- Command: llama-server -hf LiquidAI/LFM2.5-2.6B-GGUF:Q8_0 --alias LiquidAI/LFM2.5-2.6B --port 8000 --jinja --reasoning-format auto -ngl all -fa on -c 128000.
- Sampling: --temp 0.1 --top-k 50 --repeat-penalty 1.1.
- -ngl all moves every layer to the GPU; -fa on turns on flash attention.
- VRAM usage is about 6.9 GB at Q8_0 with a 128k context.
- Requires a CUDA-enabled build of llama-server.
Resources mentioned
- LiquidAI/LFM2.5-2.6B-GGUF · tool · huggingface.co · free
The Hugging Face repo with GGUF quantized weights of Liquid AI's LFM2.5-2.6B model.
Also in: Run LFM2.5-2.6B locally with llama.cpp (Melvin Vivas on X · notes) - Liquid AI (@liquidai) on X · website · x.com · free
An AI company that builds efficient foundation models. The quoted post shows its PII handling working on Japanese text.
Also in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), Liquid AI LFM 2.5 Encoder: a CPU-friendly encoder model (Melvin Vivas on X · notes), Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes), Liquid AI's LFM2-Longevity models for aging-data analysis (Melvin Vivas on X · notes) and 19 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - @aivandroid · person · x.com · free
The developer who maintains the geocine/llama-swap fork.
Also in: Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code (Melvin Vivas on X · notes)
Try this
- Build llama.cpp with CUDA support.
- Run the llama-server command to serve LFM2.5-2.6B on port 8000.
More in LLMOps, Deployment & Monitoring
- Adding LFM2.5-2.6B support to AIBackends with Cursor
- Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder
- Run LFM2.5-2.6B locally with llama.cpp
- Personal AI Computer to Run DeepSeek V4-Flash Locally
- Running DeepSeek V4 Flash Locally: RAM Needs for 4-bit and 3-bit Quants
- aibackends 0.3.0: Model Caching Speeds Up PII and OCR Inference