Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent
Melvin Vivas · X post · 2026-08-11 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Agents, Tool Use & MCP · Level: intermediate
Summary
Melvin Vivas tests Meta's Muse Glimmer 30B open model with the Hermes agent. He runs it locally through llama.cpp using Unsloth's 4-bit GGUF quant. The quoted guide lists his hardware and setup, so you can judge whether your own machine can run it.
Key points
- Model: unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL (Unsloth dynamic 4-bit quant).
- Served with llama.cpp and used as the backend for the Hermes agent.
- Setup: WSL Ubuntu 24.04.1 on Windows 11 Pro with an RTX 3090.
- VRAM use is about 23 GB at full context, so it fits on a single 24 GB card.
- He says it should also work with OpenClaw and Pi, but he hasn't tested those.
Resources mentioned
- unsloth/Muse-Glimmer-30B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantizations of Meta's Muse Glimmer 30B open-weights model, used here with the UD-Q4_K_XL quant.
Also in: Three Local GGUF Models That Fit on an RTX 3090 (24GB) (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes) - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - Hermes Agent · tool · github.com · free
Nous Research's open-source AI agent with CLI, TUI and desktop interfaces. It now supports hands-free activation with a wake word.
Also in: Inspecting Coding-Agent Traces Live with JSONL Viewer (Codex, Claude Code) (Melvin Vivas on X · notes), Hermes Agent: each bot is its own profile (Melvin Vivas on X · notes), An X research bot built with Hermes (Melvin Vivas on X · notes), Run Your Hermes Agent on Free LFM2.5-2.6B via OpenRouter (Melvin Vivas on X · notes) and 55 more - OpenClaw · tool · github.com · free
An open-source, self-hostable personal AI agent that can run on local or hosted LLMs.
Also in: Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes), Running Local Models for Agents: Tool Use, Context and Quantization (Melvin Vivas on X · notes), Ollama Is Now an Official Provider for OpenClaw (Melvin Vivas on X · notes) and 2 more - Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more
Try this
- Follow the guide to serve Muse Glimmer 30B (UD-Q4_K_XL) with llama.cpp and connect it to an agent such as Hermes.
- Build a fully local agent stack: llama.cpp serving Muse Glimmer 30B on a 24 GB GPU, driven by Hermes, OpenClaw or Pi.
More in LLMOps, Deployment & Monitoring
- Run Qwen3.8-27B locally with llama.cpp (llama-server)
- Unsloth releases Qwen3.8-27B GGUF quantizations
- Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp
- Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent
- Running Liquid AI LFM2.5-2.6B Locally with llama-server
- Adding LFM2.5-2.6B support to AIBackends with Cursor