Plan With a Frontier Model, Code With a Local Quantized Qwen
Melvin Vivas · X post · 2026-08-30 · Open on X
Topics: AI Dev Tools & Productivity, LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
The creator thinks about a hybrid coding setup: use Codex with a frontier model (Sol) for planning, and hand the coding to Qwen3.8-27B quantized to Q4_K_M running locally. A local model has no usage limits, so it can run in a loop for free even if it is slower.
Key points
- Split the work: a strong hosted model does the planning, a local model does the coding.
- Qwen3.8-27B at Q4_K_M quantization is the local model being considered.
- Local inference has no rate or usage limits, so you can run agent loops without paying per token.
- The trade-off is slower coding in exchange for free, unlimited runs.
- The creator asked ChatGPT whether the local model compares to the hosted 'Luna' model, which is not a real benchmark.
Resources mentioned
- Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more
Try this
- Build a planner/executor coding loop: a hosted frontier model writes the plan and a locally hosted quantized Qwen carries it out.
More in AI Dev Tools & Productivity
- GPT-5.6 Luna at Highest Reasoning Handled a Release in Codex
- Running Local Qwen3.8-27B in Codex on a Long Goal
- Codex + Local Qwen3.8-27B (Q4_K_M) with /goal
- Synthetic (synthetic.new): A Clear Way to Show LLM Subscription Usage Limits
- herdr: Managing Codex and Claude Agents Across Projects
- Debug Animated UI Bugs by Having Codex or Claude Code Watch a Screen Recording