AI Engineer Study Library

Three Local GGUF Models That Fit on an RTX 3090 (24GB)

Melvin Vivas · X post · 2026-09-09 · Open on X

Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: intermediate

Summary

The creator lists three quantized GGUF models he has run on one RTX 3090 with 24GB of VRAM. The list shows which model sizes and quantization levels (Q4_K_M, Q4_K_XL) fit on a consumer GPU, including a 35B mixture-of-experts model with about 3B active parameters.

Key points

Resources mentioned

More in LLM Fundamentals

All of LLM Fundamentals