Local Speech-to-Text with LFM2.5-Audio-1.5B and llama.cpp
Melvin Vivas · X video post · 2026-09-08 · 0:16 · 600 views · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
This is a 16-second demo with no speech. It shows a speech-to-text (audio-to-text) CLI that runs entirely on a laptop. The tool serves Liquid AI's LFM2.5-Audio-1.5B model with llama.cpp. The quoted post links to the open-source code in the Liquid4All cookbook, which was recently updated to use the LFM2.5 model family.
Key points
- LFM2.5-Audio-1.5B is a small (1.5B-parameter) audio model from Liquid AI's LFM2.5 family that can do speech-to-text.
- llama.cpp can serve the model on a laptop, so transcription runs 100% locally with no cloud API.
- The demo is a command-line tool that transcribes any audio input you give it.
- The example code is in the Liquid4All cookbook repo under examples/audio-transcription-cli.
- The demo was recently switched to the LFM2.5 family, replacing an earlier model version.
- Small audio models plus llama.cpp are a practical way to get private, offline, low-cost transcription.
Resources mentioned
- Liquid4All Cookbook – audio-transcription-cli example · repo · github.com · free
Open-source example code for a fully local audio-to-text CLI that serves LFM2.5-Audio-1.5B with llama.cpp. - LFM2.5-Audio-1.5B · tool · huggingface.co · free
Liquid AI's 1.5B-parameter audio model from the LFM2.5 family, used here for speech-to-text. - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - Liquid AI (@liquidai) on X · website · x.com · free
An AI company that builds efficient foundation models. The quoted post shows its PII handling working on Japanese text.
Also in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), Liquid AI LFM 2.5 Encoder: a CPU-friendly encoder model (Melvin Vivas on X · notes), Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes), Liquid AI's LFM2-Longevity models for aging-data analysis (Melvin Vivas on X · notes) and 19 more
Try this
- Clone the Liquid4All cookbook and run the audio-transcription-cli example.
- Serve LFM2.5-Audio-1.5B with llama.cpp on your laptop and transcribe your own audio files.
- Build a fully local, private audio-to-text transcription CLI using LFM2.5-Audio-1.5B and llama.cpp.
More in LLMOps, Deployment & Monitoring
- How inference engines work: the full life of an LLM request
- SGLang v0.5.19 release: new models and beam search
- llama.cpp's built-in web UI for testing local models
- Daytona Offers GPU Sandboxes (H100, RTX 4090/5090)
- Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark
- Docker Template Bundling Coding Agents for GPU Cloud Hosting