Three Local GGUF Models That Fit on an RTX 3090 (24GB)
Melvin Vivas · X post · 2026-09-09 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
The creator lists three quantized GGUF models he has run on one RTX 3090 with 24GB of VRAM. The list shows which model sizes and quantization levels (Q4_K_M, Q4_K_XL) fit on a consumer GPU, including a 35B mixture-of-experts model with about 3B active parameters.
Key points
- Hardware: one RTX 3090 with 24GB of VRAM.
- ggml-org/Qwen3.8-27B-GGUF at Q4_K_M quantization.
- unsloth/Muse-Glimmer-30B-GGUF at Q4_K_XL quantization.
- ornith-ai/Ornith-1.5-35B-A3B-GGUF: a 35B mixture-of-experts model with about 3B active parameters (A3B).
- At around 4-bit quantization, models of roughly 27-35B parameters can fit in 24GB of VRAM.
Resources mentioned
- ggml-org/Qwen3.8-27B-GGUF · tool · huggingface.co · free
A GGUF-quantized build of Qwen3.8 27B for llama.cpp-style local inference. - unsloth/Muse-Glimmer-30B-GGUF · tool · huggingface.co · free
Unsloth's GGUF quantizations of Meta's Muse Glimmer 30B open-weights model, used here with the UD-Q4_K_XL quant.
Also in: Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes) - ornith-ai/Ornith-1.5-35B-A3B-GGUF · tool · huggingface.co · free
A GGUF build of Ornith 1.5, a 35B mixture-of-experts model with about 3B active parameters.
More in LLM Fundamentals
- When a higher reasoning level is worth the extra cost
- Gemma 4 as a Strong Small Model for Local and On-Device Use
- Running Qwen3.8 27B Locally with llama.cpp for Writing
- Comparing GPT Models in Codex by Intelligence and Cost per Task
- Open-weight Nemotron models for finance and healthcare
- Qwen3.8 27B Quantization Benchmark: 4-Bit Is Enough for Agentic Coding