AI Engineer Study Library

Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent

Melvin Vivas · X post · 2026-08-11 · Open on X

Topics: LLMOps, Deployment & Monitoring, AI Agents, Tool Use & MCP, LLM Fundamentals · Level: intermediate

Summary

This post is a step-by-step guide to serving Unsloth's quantized Muse Glimmer 30B GGUF model on one RTX 3090 (24 GB) using llama.cpp built with CUDA under WSL Ubuntu on Windows 11. It covers building llama-server, the launch flags and sampling settings, a health check, and pointing an agent (Hermes Agent) at the local OpenAI-compatible endpoint. The author says OpenClaw and Pi should also work but hasn't tested them.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring