Run LFM2.5-2.6B locally with llama-cpp-python in Colab
Melvin Vivas · X post · 2026-08-16 · Open on X
Topics: LLM Fundamentals, AI Agents, Tool Use & MCP, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
The creator shares a Colab notebook that runs the small LFM2.5-2.6B model with llama-cpp-python. It includes a sample inference and a tool-calling example. He also posts quick benchmarks from a free T4 GPU using Q4_K_M quantization.
Key points
- The notebook runs LFM2.5-2.6B through llama-cpp-python (GGUF / llama.cpp backend).
- It includes a basic inference example and a tool-calling example.
- Benchmarks were run on a T4 GPU in Google Colab.
- Quantization was Q4_K_M.
- Average latency was 1.40 s.
- Average end-to-end throughput was 91.55 tokens/s.
- The repo collects notebooks for fine-tuning and running local models in Colab.
Resources mentioned
- donvito/notebooks · repo · github.com · free
The creator's notebooks for fine-tuning and running local models, which you can run in Google Colab, including a GLiNER2.5-Decide intent classification example.
Also in: Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes), Free Colab Notebooks for Local Model Fine-Tuning and Inference (Melvin Vivas on X · notes), Intent classification for support using GLiNER2.5-Decide notebook (Melvin Vivas on X · notes), Getting started with GLiNER2.5-Decide in a Colab notebook (Melvin Vivas on X · notes) and 3 more - LFM2.5-2.6B · tool · huggingface.co · free
Liquid AI's small language model, which can be paired with the LFM2.5-VL-3B vision model.
Also in: Coworker: Open-Source Work Agent That Runs on Small Local Models (Melvin Vivas on X · notes), Coworker with Liquid AI LFM2.5-2.6B via LM Studio on a Mac (Melvin Vivas on X · notes), Zero-Cost Coworker Setup: OpenRouter Free Models + Local LFM2.5 (Melvin Vivas on X · notes), Coworker: A Subscription-Free AI Agent on Local LFM2.5-2.6B (Melvin Vivas on X · notes) and 9 more - llama-cpp-python · tool · github.com · free
Python bindings for llama.cpp for running quantized GGUF models locally. - Google Colab · tool · colab.research.google.com · free · recommended by both Bashiri Smith & Melvin Vivas
Free GPU notebooks.
Also in: Run Notebooks on a Free GPU with Google Colab (T4, 15GB VRAM) (Melvin Vivas on X · notes), Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes), Free Colab Notebooks for Local Model Fine-Tuning and Inference (Melvin Vivas on X · notes), Intent classification for support using GLiNER2.5-Decide notebook (Melvin Vivas on X · notes) and 12 more
Try this
- Open the donvito/notebooks repo and run the LFM2.5-2.6B notebook on a Colab T4 GPU.
- Try the tool-calling example with a small local model.
- Benchmark different quantizations (e.g., Q4_K_M vs Q8_0) of a small model on a Colab T4 and compare latency and throughput.
- Build a local tool-calling agent on top of a small quantized model with llama-cpp-python.
More in LLM Fundamentals
- Find Discounted Models with OpenRouter's New Filter
- Cost-Aware Model Routing in Codex: Default to Luna, Escalate to Sol
- Qwen 3.8 27B Matches GPT 5.6 Luna (max) on Artificial Analysis
- GLM-5.3 free weekend trial announcement
- Swapping AI SDK for pi-ai as the LLM provider layer
- Open Models Aren't Always Local Models