Muse Glimmer 30B: Unsloth GGUF Release and Run Guide
Melvin Vivas · X post · 2026-08-10 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring, Industry Trends & Job Market · Level: intermediate
Summary
Meta released Muse Glimmer, a 30B open model under the Apache 2.0 license. It supports vision, runs in about 18 GB of RAM, and is described as the strongest agentic model for its size. Unsloth published GGUF quants and a guide for running and training it.
Key points
- Muse Glimmer is a 30B open model from Meta under the Apache 2.0 license.
- It supports vision input and is described as the strongest agentic model for its size.
- It runs in about 18 GB of RAM using Unsloth's GGUF quants.
- Unsloth provides both GGUF files (huggingface.co/unsloth/Muse-Glimmer-30B-GGUF) and a docs guide for running and training it.
Resources mentioned
- Muse Glimmer 30B GGUF on Hugging Face (Unsloth) · tool · huggingface.co · free · open in a browser to verify
Unsloth's GGUF quantizations of Muse Glimmer 30B for running locally. - Unsloth Documentation: Muse Glimmer guide · docs · unsloth.ai · free
Unsloth's guide and free notebooks for fine-tuning and GRPO-training Muse Glimmer 30B.
Also in: Fine-Tune Muse Glimmer 30B for Free with Unsloth (incl. GRPO) (Melvin Vivas on X · notes) - Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more
Try this
- Download the Unsloth GGUF and try running Muse Glimmer 30B locally (about 18 GB RAM).
More in LLM Fundamentals
- NVIDIA Nemotron 3.5 Lightning: Fast Open MoE Model for Agents
- Claude Sonnet 5 Pricing Made Permanent ($2/$10 per M Tokens)
- Running Liquid AI Models On-Device with the Apollo iPhone App
- Transferable KV Cache: Reusing One Model's Cache in Another
- DeepSeek V4 Flash at 90% Off on Nous Portal
- Qwen3.8-Max on the Frontend Code Arena cost-performance frontier