Running Qwen 3.8 27B locally in Hermes agent on an RTX 3090
Melvin Vivas · X post · 2026-08-17 · Open on X
Topics: AI Agents, Tool Use & MCP, LLM Fundamentals · Level: intermediate
Summary
The creator is impressed with Qwen 3.8 27B running in the Hermes agent on a single RTX 3090. He had the agent search the web for recently released open models in the same size class that fit his GPU. It shows that a mid-size open model can do useful agentic research locally.
Key points
- Qwen 3.8 27B, an open-weight 27B model, runs on a single RTX 3090 (24GB VRAM).
- Hermes agent can use a local model for tool-using tasks such as web search.
- Example task: ask the agent to find currently released open models in its class that fit your GPU.
Resources mentioned
- Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more - Hermes Agent · tool · github.com · free
Nous Research's open-source AI agent with CLI, TUI and desktop interfaces. It now supports hands-free activation with a wake word.
Also in: Inspecting Coding-Agent Traces Live with JSONL Viewer (Codex, Claude Code) (Melvin Vivas on X · notes), Hermes Agent: each bot is its own profile (Melvin Vivas on X · notes), An X research bot built with Hermes (Melvin Vivas on X · notes), Run Your Hermes Agent on Free LFM2.5-2.6B via OpenRouter (Melvin Vivas on X · notes) and 55 more
Try this
- Try running a ~27B open model locally in an agent and give it a research task.
- Use a local agent to research and compare open models that fit your GPU's VRAM.
More in AI Agents, Tool Use & MCP
- Agent Loops in ML Engineering: Liquid AI's Vibe-Coded Tokenizer Trainer
- Comfy MCP Goes Local and Open-Source: Agents That See Your Install
- AutoDesign: Improving the Harness Around the Model with a Meta-Harness Loop
- GAIA: free Docker installer for multiple Hermes agents
- Running a Hermes agent locally on Qwen 3.8 27B (Q4_K_M) on an RTX 3090
- DonvitoCodes Colab notebooks for running local models (LFM2.5-2.6B)