Running a Hermes agent locally on Qwen 3.8 27B (Q4_K_M) on an RTX 3090
Melvin Vivas · X post · 2026-08-17 · Open on X
Topics: AI Agents, Tool Use & MCP, LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
Short post showing that a 27B Qwen 3.8 model, quantized to Q4_K_M, can run an agent (Hermes Agent) on a single 24 GB RTX 3090. It's a data point on what consumer hardware can handle for local agents.
Key points
- Model used: Qwen 3.8 27B with Q4_K_M quantization (a 4-bit GGUF quant).
- Hardware: one RTX 3090 (24 GB VRAM).
- 4-bit quantization lets a ~27B model fit on one consumer GPU for agent testing.
- Agent framework being tested: Hermes Agent.
Resources mentioned
- Hermes Agent · tool · github.com · free
Nous Research's open-source AI agent with CLI, TUI and desktop interfaces. It now supports hands-free activation with a wake word.
Also in: Inspecting Coding-Agent Traces Live with JSONL Viewer (Codex, Claude Code) (Melvin Vivas on X · notes), Hermes Agent: each bot is its own profile (Melvin Vivas on X · notes), An X research bot built with Hermes (Melvin Vivas on X · notes), Run Your Hermes Agent on Free LFM2.5-2.6B via OpenRouter (Melvin Vivas on X · notes) and 55 more - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more
Try this
- Run a local agent with a quantized ~27B model (Q4_K_M) on a single 24 GB GPU
More in AI Agents, Tool Use & MCP
- AutoDesign: Improving the Harness Around the Model with a Meta-Harness Loop
- Running Qwen 3.8 27B locally in Hermes agent on an RTX 3090
- GAIA: free Docker installer for multiple Hermes agents
- DonvitoCodes Colab notebooks for running local models (LFM2.5-2.6B)
- Multi-agent delegation to lower-cost models to save Codex limits
- Adding Agent Capabilities to ai-backends with Pi