Running Qwen3.8 27B Locally with Hermes
Melvin Vivas · X post · 2026-08-18 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Agents, Tool Use & MCP · Level: intermediate
Summary
The creator runs Qwen3.8 27B locally with Hermes and calls the model really good. Its downside is slow thinking (reasoning), but it costs nothing to use.
Key points
- Qwen3.8 27B runs on local hardware and gives strong results.
- The trade-off is speed: the reasoning phase is slow on a local setup.
- Running locally means zero API cost.
Resources mentioned
- Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more - Hermes · tool · hermes-agent.nousresearch.com · free
The AI agent the creator uses to automate making explainer videos. It is probably Nous Research's Hermes Agent, but the post does not say so.
Also in: OpenAI DevDay recap: Dots, GPT-6.1 Sol, Codex Cloud, Agents API (Melvin Vivas on X · notes), Coworker: free open-source desktop AI coworker app (Melvin Vivas on X · notes), Agent Monitor: see traces, tokens and costs of your coding agents (Melvin Vivas on X · notes), Coworker: An Open-Source Desktop Agent App for Local and Cloud Models (Melvin Vivas on X · notes) and 13 more
Try this
- Try running Qwen3.8 27B locally if you want a free model and can accept slower reasoning.
More in LLMOps, Deployment & Monitoring
- One-Click Deploy Link for Qwen3.8 27B on Baseten
- Self-Hosting Qwen3.8 27B FP8 on an H100 for $6.50/hr
- Hosted Qwen3.8 27B Costs More Than GPT 5.6 Luna, So Run It Locally
- AIBackends: Python Library for Local-Model AI Workflows
- AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers
- AIBackends Adds Support for LFM2.5-VL-3B