aibackends 0.3.0: Model Caching Speeds Up PII and OCR Inference
Melvin Vivas · X post · 2026-07-18 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Safety, Security & Guardrails · Level: intermediate
Summary
The creator's Python library aibackends now caches models. GLiNER-PII inference is about 20x faster, and OCR with Qwen 4B VL saves about 9 seconds per call once the model has loaded the first time. The main lesson is that loading a model once and reusing it removes most of the per-call delay.
Key points
- Install: pip install aibackends==0.3.0
- Model caching makes GLiNER-PII inference about 20x faster
- Qwen 4B VL OCR saves about 9 seconds per call after a one-time model load
- Lesson: load the model once and reuse it instead of reloading it on every request
Resources mentioned
- AIBackends · repo · aibackends.com · free
Open-source API server runtime for common AI use cases that supports many models and providers (Ollama, LM Studio, OpenRouter, OpenAI, Anthropic).
Also in: Building a production website with Opus 5.5 and TanStack (Melvin Vivas on X · notes), CamelFlow: open-source visual viewer for Apache Camel routes (Melvin Vivas on X · notes), AI Backends: A Production AI Workflow Engineering Site (Link Share) (Melvin Vivas on X · notes), Demo: Claude Opus 5.5 Generating a Motion-Graphics Video for AIBackends (Melvin Vivas on X · notes) and 16 more - GLiNER-PII · tool · huggingface.co · free
NVIDIA's GLiNER-based model that finds personal data (PII) in text, so it can be redacted locally.
Also in: Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models (Melvin Vivas on X · notes), Detect PII and PHI with NVIDIA's GLiNER-PII Model (Melvin Vivas on X · notes) - Qwen 4B VL · tool · huggingface.co · free
Qwen's small 4B vision-language model, used here for OCR.
Try this
- Install aibackends with pip install aibackends==0.3.0
- Cache loaded models in your own inference code to cut latency
More in LLMOps, Deployment & Monitoring
- Serving LFM2.5-2.6B with llama-server: full command
- Personal AI Computer to Run DeepSeek V4-Flash Locally
- Running DeepSeek V4 Flash Locally: RAM Needs for 4-bit and 3-bit Quants
- Monitor GPU Usage With nvtop Instead of nvidia-smi
- Cost-Saving Model Fallback Chain: Grok → Codex → OpenRouter DeepSeek
- Running a Personal AI Agent for $0.69/Day with DeepSeek V4 Flash