Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models
Melvin Vivas · X video post · 2026-04-26 · 0:07 · 485 views · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Safety, Security & Guardrails, AI Dev Tools & Productivity · Level: intermediate
Summary
Melvin Vivas shows a short demo of AIBackends, his project for running AI tasks locally. He rebuilt it as a Python library over a weekend, using Cursor with GPT 5.4/5.5 as coding models. It runs open models such as Google DeepMind's and Alibaba's Qwen on your own machine. One of its features is redacting personal data (PII) with NVIDIA's gliner-PII model. The caption is cut off after "Image", so the rest of the feature list is unknown.
Key points
- Open models from Google DeepMind (likely Gemma) and Alibaba Qwen are now good enough to run AI locally, without cloud APIs.
- AIBackends was rebuilt as a Python library over a single weekend using an AI coding assistant (Cursor) with GPT 5.4/5.5.
- PII redaction is handled locally by gliner-PII, NVIDIA's model for finding personal data, so sensitive text never leaves your machine.
- Wrapping several local AI tasks (PII redaction, image tasks and others) in one Python library makes them reusable inside your own apps.
- The caption is cut off after 'Image', so the rest of the feature list is unknown. The 7-second video has no speech.
Resources mentioned
- AIBackends · repo · aibackends.com · free
Open-source API server runtime for common AI use cases that supports many models and providers (Ollama, LM Studio, OpenRouter, OpenAI, Anthropic).
Also in: Building a production website with Opus 5.5 and TanStack (Melvin Vivas on X · notes), CamelFlow: open-source visual viewer for Apache Camel routes (Melvin Vivas on X · notes), AI Backends: A Production AI Workflow Engineering Site (Link Share) (Melvin Vivas on X · notes), Demo: Claude Opus 5.5 Generating a Motion-Graphics Video for AIBackends (Melvin Vivas on X · notes) and 16 more - GLiNER-PII · tool · huggingface.co · free
NVIDIA's GLiNER-based model that finds personal data (PII) in text, so it can be redacted locally.
Also in: aibackends 0.3.0: Model Caching Speeds Up PII and OCR Inference (Melvin Vivas on X · notes), Detect PII and PHI with NVIDIA's GLiNER-PII Model (Melvin Vivas on X · notes) - Gemma 4 · tool · ai.google.dev · free
Google's family of open-weight models in several sizes, built to run on devices and offline, with multimodal and agentic abilities, and open to fine-tuning.
Also in: Gemma 4 Runs Locally On-Device in the Antigravity SDK (Melvin Vivas on X · notes), On-Device AI: Running Gemma 4 E2B Offline on an iPhone with LiteRT (Melvin Vivas on X · notes), Running Gemma 4 Models Offline on an iPhone (Melvin Vivas on X · notes), Fine-tuning Gemma4-E2B on your own tweet style with Unsloth (Melvin Vivas on X · notes) and 25 more - Qwen models (Alibaba) · tool · github.com · free
Alibaba's family of open models (LLM and multimodal) that can run locally. - Cursor · tool · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
AI-native code editor (VS Code with AI built in).
Also in: GLM 5.3 and GLM 5.3 Flash now in Cursor (Melvin Vivas on X · notes), Grok Bot Can Now Hand Off Coding Tasks to Cursor (Melvin Vivas on X · notes), Why Plan Mode Still Matters in AI Coding Assistants (Codex, Claude, Cursor) (Melvin Vivas on X · notes), Why Plan Mode Still Matters in AI Coding Agents (Codex, Claude, Cursor) (Melvin Vivas on X · notes) and 121 more - GPT 5.5 · tool · openai.com · paid
OpenAI models the creator used as the coding model inside Cursor.
Also in: Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Running GPT-5.5 via Codex as Hermes's Main Model (Melvin Vivas on X · notes), Conductor Walkthrough: Running Parallel Coding Agents in Isolated Git Worktrees (Melvin Vivas on X · notes), Composer 2.5 as the Default Coding Model in Cursor (Melvin Vivas on X · notes) and 2 more - Qwen (@Alibaba_Qwen) on X · person · x.com · free
Official X account of Alibaba's Qwen team, which posts model releases and demos.
Also in: Qwen-Audio-3.1: Alibaba's Five-Model Audio Stack (Melvin Vivas on X · notes), Demo: Qwen3.8-27B Running Locally with Pi and llama.cpp (Melvin Vivas on X · notes), Audio-Visual Vibe Coding with Qwen 3.5 Omni: Spoken Specs to Web App (Melvin Vivas on X · notes) - Google DeepMind (@GoogleDeepMind) on X · person · x.com · free
Google DeepMind's official X account, which posts research and model releases such as Gemma.
Also in: Gemma 4 and Google DeepMind's Open-Source Push (Melvin Vivas on X · notes), Gemma 4 E4B: A Local Agentic Edge Model in 6GB RAM (Melvin Vivas on X · notes) - NVIDIA AI (@NVIDIAAI) on X · person · x.com · free
NVIDIA AI's official X account, which posts model releases and endpoint availability.
Also in: Using Free OpenRouter Models (Nemotron 3.5) With Coworker (Melvin Vivas on X · notes), AI News Roundup: Claude Fable 5, Scientist AI, ZCode, NVIDIA RL (Melvin Vivas on X · notes), Jensen Huang on NVIDIA's Long-Term Commitment to Open Nemotron Models (Melvin Vivas on X · notes), Qwen3.5 Multimodal Model Now on NVIDIA AI Endpoints (Melvin Vivas on X · notes)
Try this
- Open the original X post to read the full AIBackends feature list, since the caption is cut off.
- Try running an open model (Gemma or Qwen) locally.
- Try gliner-PII to redact personal data locally before sending text to any cloud LLM.
- Build a Python library that wraps local AI tasks such as PII redaction and image processing behind one simple API, using open models like Gemma, Qwen and gliner-PII.
- Build a privacy filter that redacts PII locally with gliner-PII before forwarding prompts to a hosted LLM.
More in LLMOps, Deployment & Monitoring
- Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code
- Runpod Flash Reaches GA: Deploy AI Workloads from Python
- Z.ai's Lessons from Serving GLM-5 for Coding Agents at Scale
- Run Gemma 4 Locally with llama.cpp in Two Commands
- LM Studio's LM Link: Use a Local Model on Another Machine