Self-Host a Coding Model on QuickPod with llama-swap and Use It in Claude Code
Melvin Vivas · X video post · 2026-05-02 · 0:24 · 457 views · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity, LLM Fundamentals · Level: advanced
Summary
This is a short demo with no narration. Melvin Vivas runs his own coding model on a rented GPU from QuickPod. He serves a Qwen 27B model distilled from Claude Opus through Aivan Monceller's (@aivandroid) fork of llama-swap, which adds some extra features. He then points Claude Code at that self-hosted model instead of Anthropic's hosted models.
Key points
- llama-swap is a proxy that swaps models in and out for any local server that speaks the OpenAI or Anthropic API, such as llama.cpp or vLLM.
- The geocine/llama-swap fork by @aivandroid adds extra features. Because it is Anthropic-compatible, Claude Code can talk to it directly.
- The model in the demo is a Qwen 27B model distilled from Claude Opus, an open-weight model tuned to behave more like Opus.
- The model runs on a GPU pod rented from QuickPod, so you don't need a strong local GPU.
- Claude Code connects to the model hosted on QuickPod. This lets you keep the Claude Code agent workflow while running your own model.
- The video has no speech, so you need the linked repo and the QuickPod console to reproduce the setup.
Resources mentioned
- geocine/llama-swap (GitHub) · repo · github.com · free
A fork of llama-swap that reliably swaps models for any local server compatible with the OpenAI or Anthropic API, such as llama.cpp or vLLM. - QuickPod · tool · console.quickpod.io · paid
A cloud service for renting GPU pods to host and run your own models.
Also in: Reusable GPU devbox: PyTorch/CUDA plus six coding agents (Melvin Vivas on X · notes), Docker Template Bundling Coding Agents for GPU Cloud Hosting (Melvin Vivas on X · notes) - @aivandroid · person · x.com · free
The developer who maintains the geocine/llama-swap fork.
Also in: Serving LFM2.5-2.6B with llama-server: full command (Melvin Vivas on X · notes) - llama-swap (original by mostlygeek) · repo · github.com · free
The original open-source llama-swap proxy for swapping models on demand on local LLM servers. - Qwen 27B Opus-distilled model · tool · huggingface.co · free
An open-weight Qwen model of about 27B parameters, distilled from Claude Opus outputs for coding. - Claude Code · tool · code.claude.com · paid · recommended by both Bashiri Smith & Melvin Vivas
Build agents and pipelines from the terminal; the guide's main agentic coding tool.
Also in: Create Claude Code Plugins with /plugin-authoring (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes), SkillsBento: Free Plugin Marketplace for Codex and Claude Code (Melvin Vivas on X · notes) and 101 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more
Try this
- Rent a GPU pod on QuickPod.
- Install the geocine/llama-swap fork on the pod and configure it to serve a coding model, such as the Qwen 27B Opus-distilled model.
- Point Claude Code at the Anthropic-compatible endpoint that llama-swap exposes.
- Self-host an open-weight coding model on a rented GPU and use it as the backend for Claude Code. Compare its quality, speed and cost with hosted models.
More in LLMOps, Deployment & Monitoring
- LM Studio MLX v1.8.1: Vision Model Batching and Better Caching
- OpenRouter Pareto Code: cost-optimized coding router
- Speeding Up Gemma 4 Inference: MTP (3x) vs. DFlash Speculative Decoding (6x)
- Runpod Flash Reaches GA: Deploy AI Workloads from Python
- Z.ai's Lessons from Serving GLM-5 for Coding Agents at Scale
- Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models