Hosted Qwen3.8 27B Costs More Than GPT 5.6 Luna, So Run It Locally
Melvin Vivas · X post · 2026-08-19 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
The creator found that Qwen3.8 27B costs more on OpenRouter than OpenAI's GPT 5.6 Luna, and decided to run the open-weight model on his own machine instead. The lesson: compare per-token prices across hosted providers before assuming an open model is the cheap option, and consider local inference when you have the hardware.
Key points
- On OpenRouter, a hosted open-weight model (Qwen3.8 27B) can cost more than a small proprietary model (GPT 5.6 Luna).
- Open weights don't mean cheap: hosted price depends on the provider, not on the license.
- A 27B open model can run locally, so your only cost is your own hardware and electricity.
Resources mentioned
- Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more - OpenRouter · tool · openrouter.ai · free
A single OpenAI-compatible API that routes requests to hundreds of models from many providers, with one bill and model fallbacks.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), The Jev model is now on OpenRouter (Melvin Vivas on X · notes), Kev-4B model, an alternative to Jev, now available on OpenRouter (Melvin Vivas on X · notes), Space Bunny Alpha: stealth 1M-context flash model on OpenRouter (Melvin Vivas on X · notes) and 42 more - GPT 5.6 Luna · tool · openai.com · paid
The OpenAI model used inside Codex for the demo. The transcript gives the variant name as 'Soul', which is unclear.
Also in: Set Codex subagent model and reasoning to save usage limits (Melvin Vivas on X · notes), Use GPT-5.6 Luna in Codex for Terminal Tasks (Melvin Vivas on X · notes), Match Reasoning Effort to Task Length in Codex (Astra/Sol) (Melvin Vivas on X · notes), Run Coworker desktop agents cheaply with GPT-5.6 Luna on OpenRouter (Melvin Vivas on X · notes) and 24 more
Try this
- Check per-token prices on OpenRouter before picking a hosted model.
- Consider running mid-size open models locally to cut API costs.
More in LLMOps, Deployment & Monitoring
- Red Hat's Quantized Qwen3.8-2.4T-A95B Variant (NVFP4/FP8)
- One-Click Deploy Link for Qwen3.8 27B on Baseten
- Self-Hosting Qwen3.8 27B FP8 on an H100 for $6.50/hr
- Running Qwen3.8 27B Locally with Hermes
- AIBackends: Python Library for Local-Model AI Workflows
- AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers