Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp
Melvin Vivas · X post · 2026-08-13 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Agents, Tool Use & MCP · Level: intermediate
Summary
Shares a guide to running Muse Glimmer 30B on one RTX 3090 with llama.cpp under WSL Ubuntu on Windows 11. It uses Unsloth's Q4_K_XL GGUF quantization and needs about 23GB of VRAM at full context. The guide's author tested it with the Hermes agent and expects it to work with OpenClaw and Pi.
Key points
- Setup: Windows 11 Pro, WSL Ubuntu 24.04.1, RTX 3090 (24GB).
- Model: unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL.
- VRAM usage is about 23GB at full context, so it just fits on a 24GB card.
- Serve the model with llama.cpp and connect it to an agent harness (tested with Hermes agent; OpenClaw and Pi untested).
Resources mentioned
- llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - unsloth/Muse-Glimmer-30B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantizations of Meta's Muse Glimmer 30B open-weights model, used here with the UD-Q4_K_XL quant.
Also in: Three Local GGUF Models That Fit on an RTX 3090 (24GB) (Melvin Vivas on X · notes), Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes) - Hermes Agent · tool · github.com · free
Nous Research's open-source AI agent with CLI, TUI and desktop interfaces. It now supports hands-free activation with a wake word.
Also in: Inspecting Coding-Agent Traces Live with JSONL Viewer (Codex, Claude Code) (Melvin Vivas on X · notes), Hermes Agent: each bot is its own profile (Melvin Vivas on X · notes), An X research bot built with Hermes (Melvin Vivas on X · notes), Run Your Hermes Agent on Free LFM2.5-2.6B via OpenRouter (Melvin Vivas on X · notes) and 55 more - OpenClaw · tool · github.com · free
An open-source, self-hostable personal AI agent that can run on local or hosted LLMs.
Also in: Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes), Running Local Models for Agents: Tool Use, Context and Quantization (Melvin Vivas on X · notes), Ollama Is Now an Official Provider for OpenClaw (Melvin Vivas on X · notes) and 2 more - Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more
Try this
- Install llama.cpp in WSL Ubuntu and load the UD-Q4_K_XL GGUF on a 24GB GPU.
- Point an agent harness (e.g., Hermes agent) at the local llama.cpp server.
- Run a fully local coding agent on Muse Glimmer 30B with a single consumer GPU.
More in LLMOps, Deployment & Monitoring
- Running LFM2.5-2.6B Q4_K_M Locally with llama.cpp and Pi
- Run Qwen3.8-27B locally with llama.cpp (llama-server)
- Unsloth releases Qwen3.8-27B GGUF quantizations
- Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent
- Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent
- Running Liquid AI LFM2.5-2.6B Locally with llama-server