AI Engineer Study Library

Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp

Melvin Vivas · X post · 2026-08-13 · Open on X

Topics: LLMOps, Deployment & Monitoring, AI Agents, Tool Use & MCP · Level: intermediate

Summary

Shares a guide to running Muse Glimmer 30B on one RTX 3090 with llama.cpp under WSL Ubuntu on Windows 11. It uses Unsloth's Q4_K_XL GGUF quantization and needs about 23GB of VRAM at full context. The guide's author tested it with the Hermes agent and expects it to work with OpenClaw and Pi.

Key points

Resources mentioned

Try this

More in LLMOps, Deployment & Monitoring

All of LLMOps, Deployment & Monitoring