AI Engineer Study Library

Running Liquid AI LFM2.5-2.6B Locally with llama-server

Melvin Vivas · X post · 2026-08-10 · Open on X

Topics: LLMOps, Deployment & Monitoring · Level: intermediate

Summary

This post shares a llama.cpp llama-server command for serving Liquid AI's LFM2.5-2.6B from its Q8_0 GGUF. The server exposes the model on port 8000 with a 128k context. It uses about 6.9 GB of VRAM, so it fits on modest GPUs.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring