Serve LFM2.5-2.6B with vLLM and connect it to Hermes
Melvin Vivas · X post · 2026-08-09 · Open on X
Topics: AI Agents, Tool Use & MCP, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
A step-by-step guide to serving Liquid AI's LFM2.5-2.6B with vLLM, with tool calling turned on, and using it as a custom provider in the Hermes agent. It also covers adding Firecrawl for web search and checking that tool calls work.
Key points
- vllm serve "LiquidAI/LFM2.5-2.6B" --enable-auto-tool-choice --tool-call-parser lfm2 --reasoning-parser qwen3.
- If vLLM throws errors, ask Codex to set it up for you.
- In Hermes, run
hermes model, choose "custom provider", and enter your vLLM URL; confirm the model in the dashboard. - For web search, run
hermes tools, select web search and scraping, and add a Firecrawl API key. - Send a test message; seeing tools get called means the setup works.
Resources mentioned
- LFM2.5-2.6B · tool · huggingface.co · free
Liquid AI's small language model, which can be paired with the LFM2.5-VL-3B vision model.
Also in: Coworker: Open-Source Work Agent That Runs on Small Local Models (Melvin Vivas on X · notes), Coworker with Liquid AI LFM2.5-2.6B via LM Studio on a Mac (Melvin Vivas on X · notes), Zero-Cost Coworker Setup: OpenRouter Free Models + Local LFM2.5 (Melvin Vivas on X · notes), Coworker: A Subscription-Free AI Agent on Local LFM2.5-2.6B (Melvin Vivas on X · notes) and 9 more - Liquid AI (@liquidai) on X · website · x.com · free
An AI company that builds efficient foundation models. The quoted post shows its PII handling working on Japanese text.
Also in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), Liquid AI LFM 2.5 Encoder: a CPU-friendly encoder model (Melvin Vivas on X · notes), Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes), Liquid AI's LFM2-Longevity models for aging-data analysis (Melvin Vivas on X · notes) and 19 more - vLLM (@vllm_project) · tool · x.com · free
An open-source, high-throughput engine for serving LLMs, with tool-call and reasoning parsers. - Hermes Agent · tool · github.com · free
Nous Research's open-source AI agent with CLI, TUI and desktop interfaces. It now supports hands-free activation with a wake word.
Also in: Inspecting Coding-Agent Traces Live with JSONL Viewer (Codex, Claude Code) (Melvin Vivas on X · notes), Hermes Agent: each bot is its own profile (Melvin Vivas on X · notes), An X research bot built with Hermes (Melvin Vivas on X · notes), Run Your Hermes Agent on Free LFM2.5-2.6B via OpenRouter (Melvin Vivas on X · notes) and 55 more - Firecrawl · tool · x.com · free
An open-source web data API with Search, Scrape and Interact features that turn web pages into LLM-ready markdown or structured data for AI agents.
Also in: Firecrawl Keyless: Free Web Search & Scraping for AI Agents, No API Key (Melvin Vivas on X · notes), Coworker: Open-Source Work Agent That Runs on Small Local Models (Melvin Vivas on X · notes), Running Coworker on Free OpenRouter Models and Its Agent Stack (Melvin Vivas on X · notes), Coworker: Open-Source Grok Bot Clone Built on Pi and CopilotKit (Melvin Vivas on X · notes) - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more - @helloiamleonie · person · x.com · free
The Liquid AI team member who helped with the Hermes setup.
Try this
- Serve LFM2.5-2.6B with vLLM using the lfm2 tool-call parser.
- Add the vLLM endpoint to Hermes as a custom provider.
- Add a Firecrawl API key through `hermes tools` for web search.
- Send a test message and check that tools get called.
- Run a self-hosted agent on a small local model with tool calling and web search.
More in AI Agents, Tool Use & MCP
- Agents Aren't Cloud Native Because They Rely on Filesystems
- Pi Coding Agent: Four Built-In Tools (read, bash, edit, write)
- Using a Hermes agent to watch the terminal and drive Codex
- LFM2.5-2.6B Matches DeepSeek-V4-Flash on Tool Calling; LEAP Fine-Tuning
- Claude Code Agents Can Now Talk to Each Other Natively
- Prime Agent: A Self-Improving RLM Coding Harness with Code-Based Tool Calls