AIBackends v0.7.0: Prompt Routing with Liquid AI's LFM2.5 Encoder
Melvin Vivas · X post · 2026-08-27 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI System Design & Architecture, LLM Fundamentals · Level: intermediate
Summary
Melvin Vivas released version 0.7.0 of his AIBackends Python library, which adds prompt routing with Liquid AI's LFM2.5-Encoder-350M. The router classifies each prompt zero-shot (without task-specific training), so complex tasks go to expensive, smarter models and simple ones go to small or free models.
Key points
- Install with: pip install "aibackends[routing]".
- Routing uses Liquid AI's LFM2.5-Encoder-350M, a small encoder that classifies prompts zero-shot.
- Why it matters: as models get more expensive, save the most capable models for complex tasks.
- Send simple prompts to smaller, cheaper or free models to cut costs.
Resources mentioned
- aibackends 0.7.0 (PyPI) · tool · pypi.org · free
PyPI page for the AIBackends Python package, version 0.7.0, which adds prompt routing. - Release aibackends v0.7.0 — LFM2.5 Prompt Routing (donvito/aibackends) · repo · github.com · free
GitHub release notes for AIBackends v0.7.0 with LFM2.5 prompt routing. - LFM2.5-Encoder-350M · tool · huggingface.co · free
A 350M-parameter encoder model from Liquid AI that does zero-shot prompt classification and routing against categories you supply at runtime.
Also in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), Model Routing Fine-Tuned on Your Agent Harness Traces (Melvin Vivas on X · notes), Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder (Melvin Vivas on X · notes) - Liquid AI (@liquidai) on X · website · x.com · free
An AI company that builds efficient foundation models. The quoted post shows its PII handling working on Japanese text.
Also in: Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M (Melvin Vivas on X · notes), Liquid AI LFM 2.5 Encoder: a CPU-friendly encoder model (Melvin Vivas on X · notes), Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes), Liquid AI's LFM2-Longevity models for aging-data analysis (Melvin Vivas on X · notes) and 19 more
Try this
- Install aibackends[routing] and try routing prompts to different models.
- Build a cost-aware router that sends simple prompts to a small or free model and hard prompts to a frontier model.
More in LLMOps, Deployment & Monitoring
- Low-Cost Agent Run: DeepSeek V4 Flash via OpenRouter in ohmypi
- Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference
- Running GLM 5.3 Flash on Baseten with the Pi Coding Agent
- GLM 5.3 Flash Now Available on Baseten
- Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M
- Novita AI Spot GPU Instances: Cheap GPU Compute