Run Qwen3.8-Flash-Next (125B MoE) Locally with Unsloth GGUFs
Melvin Vivas · X post · 2026-08-27 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
Shares news that Qwen3.8-Flash-Next, a 125B mixture-of-experts model, can now run locally using Unsloth's GGUF quantizations. The post says it needs about 75GB of RAM and that the model is built so CPU RAM or unified-memory setups get close to VRAM speeds. It links Unsloth's guide and the GGUF files.
Key points
- Qwen3.8-Flash-Next is a 125B-parameter Mixture-of-Experts (MoE) model.
- The post says it outperforms Claude-Opus-4.6 (Max). This is a claim from the quoted post, not independently verified.
- Unsloth GGUF quantizations let it run in about 75GB of RAM.
- The model is designed so CPU RAM or unified-memory machines (e.g., Macs) reach near-VRAM inference speeds.
- Unsloth publishes step-by-step docs for running new models locally.
Resources mentioned
- Unsloth LLM Tutorials (Unsloth Documentation) · docs · unsloth.ai · free
Unsloth's guides for running and fine-tuning open LLMs locally, including the Qwen3.8-Next guide.
Also in: Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs (Melvin Vivas on X · notes) - Unsloth Qwen3.8-Flash-Next-GGUF (Hugging Face) · tool · huggingface.co · free
Quantized GGUF weights of Qwen3.8-Flash-Next published by Unsloth on Hugging Face. - Qwen3.8-Flash-Next · tool · huggingface.co · free
Open 125B MoE model from Qwen that is tuned to run fast on CPU RAM or unified memory. - Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more
Try this
- Follow the Unsloth guide to run Qwen3.8-Flash-Next locally if you have about 75GB of RAM or unified memory.
More in LLM Fundamentals
- Colab Notebook: Entity Extraction with GLiNER2.5 in aibackends
- Cohere Parse Beats Frontier LLMs at Receipt Parsing
- You don't need frontier LLMs for everything: use SLMs
- GPT 5.6 Sol (medium) in ChatGPT for planning tasks
- GLM-5.2 Vision on Baseten: Turning Images into Code
- Try the Free Stealth Model Ox Alpha on OpenRouter