Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent
Melvin Vivas · X post · 2026-08-11 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Agents, Tool Use & MCP, LLM Fundamentals · Level: intermediate
Summary
This post is a step-by-step guide to serving Unsloth's quantized Muse Glimmer 30B GGUF model on one RTX 3090 (24 GB) using llama.cpp built with CUDA under WSL Ubuntu on Windows 11. It covers building llama-server, the launch flags and sampling settings, a health check, and pointing an agent (Hermes Agent) at the local OpenAI-compatible endpoint. The author says OpenClaw and Pi should also work but hasn't tested them.
Key points
- Setup: Windows 11 Pro + WSL Ubuntu 24.04.1, RTX 3090. Uses about 23 GB VRAM at full 65,536-token context with the unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL quantization.
- Build llama.cpp with CUDA: git clone the ggml-org/llama.cpp repo, then run
cmake -S . -B build-cuda -DGGML_CUDA=ON -DLLAMA_OPENSSL=ONandcmake --build build-cuda --config Release --target llama-server -j 8. - Launch with
./build-cuda/bin/llama-server -hf unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL. The -hf flag downloads the model straight from Hugging Face. - Server flags: --ctx-size 65536, --parallel 1, --alias unsloth-Muse-Glimmer-30B-GGUF, --host 127.0.0.1, --port 8001. Raise ctx-size if you have spare VRAM; any port works.
- Sampling settings: --temp 1.0, --top-p 0.95, --top-k 64.
- Check the server with
curl -fsS http://localhost:8001/health. The OpenAI-compatible base URL is http://localhost:8001/v1. - Hermes Agent setup: choose Custom endpoint, Base URL http://localhost:8001/v1, leave the API key empty, Model = the alias (unsloth-Muse-Glimmer-30B-GGUF), Context length 65536, API mode chat_completions.
- Any agent that can use an OpenAI-compatible endpoint (the author names OpenClaw and Pi, untested) should work with the same settings.
Resources mentioned
- llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - unsloth/Muse-Glimmer-30B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantizations of Meta's Muse Glimmer 30B open-weights model, used here with the UD-Q4_K_XL quant.
Also in: Three Local GGUF Models That Fit on an RTX 3090 (24GB) (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent (Melvin Vivas on X · notes) - Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more - Hermes Agent · tool · github.com · free
Nous Research's open-source AI agent with CLI, TUI and desktop interfaces. It now supports hands-free activation with a wake word.
Also in: Inspecting Coding-Agent Traces Live with JSONL Viewer (Codex, Claude Code) (Melvin Vivas on X · notes), Hermes Agent: each bot is its own profile (Melvin Vivas on X · notes), An X research bot built with Hermes (Melvin Vivas on X · notes), Run Your Hermes Agent on Free LFM2.5-2.6B via OpenRouter (Melvin Vivas on X · notes) and 55 more - OpenClaw · tool · github.com · free
An open-source, self-hostable personal AI agent that can run on local or hosted LLMs.
Also in: Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent (Melvin Vivas on X · notes), Running Local Models for Agents: Tool Use, Context and Quantization (Melvin Vivas on X · notes), Ollama Is Now an Official Provider for OpenClaw (Melvin Vivas on X · notes) and 2 more - Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more - WSL (Windows Subsystem for Linux) · tool · learn.microsoft.com · free
Runs a Linux environment such as Ubuntu on Windows, here used to build llama.cpp with CUDA.
Also in: Any OS works for AI dev: Mac, Windows + WSL (Melvin Vivas on X · notes), Running Gemma 4 E2B Video Understanding Locally in WSL (Melvin Vivas on X · notes), Installing Unsloth Studio on Windows via WSL (Melvin Vivas on X · notes) - Melvin Vivas · person · x.com · free
A creator who shares experiments with AI agents and developer tools on X.
Also in: How LLMs Work: A Motion-Graphics Explainer Made in One Shot with Claude Opus 5.5 (Melvin Vivas on X · notes), Concept Demo: An Agent Monitor Built on OpenAI Codex Traces (Melvin Vivas on X · notes), Running MiniMax H3 Locally on a Single RTX 3090 (Demo) (Melvin Vivas on X · notes)
Try this
- Build llama.cpp's llama-server with CUDA (-DGGML_CUDA=ON) under WSL Ubuntu.
- Start llama-server with unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL, 65536 context, temp 1.0, top-p 0.95, top-k 64 on port 8001.
- Raise --ctx-size if your GPU has spare VRAM.
- Check the server with curl http://localhost:8001/health.
- In Hermes Agent, add a Custom endpoint: base URL http://localhost:8001/v1, empty API key, model alias, context 65536, API mode chat_completions.
- Follow @melvindvivas for more guides.
- Run a fully local coding/agent setup: llama.cpp serving a 30B GGUF model on a 24 GB GPU, used as the backend for Hermes Agent, OpenClaw or Pi.
More in LLMOps, Deployment & Monitoring
- Unsloth releases Qwen3.8-27B GGUF quantizations
- Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp
- Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent
- Running Liquid AI LFM2.5-2.6B Locally with llama-server
- Adding LFM2.5-2.6B support to AIBackends with Cursor
- Zero-Shot Prompt Routing by Task Complexity with LFM2.5-Encoder