Running DeepSeek V4 Flash Locally: RAM Needs for 4-bit and 3-bit Quants
Melvin Vivas · X post · 2026-08-01 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: advanced
Summary
Quoting Unsloth, the creator jokes about the hardware needed to run DeepSeek V4 Flash 0731 locally. The lossless 4-bit quant needs about 168GB RAM and the 3-bit quant about 110GB. You can run it with Unsloth or llama.cpp from GGUF files, and smaller quants were promised.
Key points
- DeepSeek V4 Flash 0731 can run locally as GGUF quants
- Lossless 4-bit quant needs ~168GB RAM; 3-bit needs ~110GB RAM
- Run it with Unsloth or llama.cpp
- Unsloth claims V4 Flash 0731 outperforms V4 Pro
- Smaller quants (less RAM) were announced as coming soon
- Quantization trades bits per weight for memory: fewer bits means less RAM, with some possible quality loss
Resources mentioned
- Unsloth Documentation – DeepSeek V4 guide · docs · unsloth.ai · free
Unsloth's guide to running DeepSeek V4 models locally with quantized GGUFs. - unsloth/DeepSeek-V4-Flash-0731-GGUF (Hugging Face) · repo · huggingface.co · free · open in a browser to verify
Hugging Face repository with the GGUF quantized weights of DeepSeek V4 Flash 0731. - Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - DeepSeek V4 Flash 0731 · tool · huggingface.co · free
DeepSeek's open-weight V4 Flash model, released as quantized GGUFs for local use.
Also in: DeepSeek V4 Flash 0731 Available on OpenRouter (Melvin Vivas on X · notes), DeepSeek V4 Flash 0731 Agentic Benchmarks and Use with Hermes Agent (Melvin Vivas on X · notes)
Try this
- Check your RAM against the quant sizes (168GB at 4-bit, 110GB at 3-bit) before trying to run it locally
- Follow the Unsloth guide to run it with Unsloth or llama.cpp
- Watch for the smaller quants if your machine has less memory
More in LLMOps, Deployment & Monitoring
- Run LFM2.5-2.6B locally with llama.cpp
- Serving LFM2.5-2.6B with llama-server: full command
- Personal AI Computer to Run DeepSeek V4-Flash Locally
- aibackends 0.3.0: Model Caching Speeds Up PII and OCR Inference
- Monitor GPU Usage With nvtop Instead of nvidia-smi
- Cost-Saving Model Fallback Chain: Grok → Codex → OpenRouter DeepSeek