Unsloth MTP GGUFs make Qwen3.6 run 1.4x faster locally
Melvin Vivas · X post · 2026-05-15 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
Unsloth released experimental GGUFs of Qwen3.6 that use Multi-Token Prediction (MTP). They report a speed-up of more than 1.4x over the original GGUFs with no change in accuracy. The creator says he got about 200 tokens/s, which shows how speculative or multi-token decoding speeds up local inference.
Key points
- MTP (Multi-Token Prediction) GGUFs let a model predict several tokens per step, so generation is faster.
- Qwen3.6 27B MTP: about 140 tokens/s.
- Qwen3.6 35B-A3B (MoE, about 3B active parameters) MTP: about 220 tokens/s on a single GPU.
- Reported speed-up of more than 1.4x over the original GGUFs, with the same accuracy.
- The creator got about 200 tok/s in his own test.
- The release is experimental. Unsloth published a guide along with the GGUFs.
Resources mentioned
- Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more - Qwen3.6-27B · tool · huggingface.co · free
Dense open-weight Qwen model. The MTP GGUF version runs at about 140 tokens/s.
Also in: Meta Muse Glimmer-30B: Open Weights Model That Beats Qwen3.6 37B on Agentic Tasks (Melvin Vivas on X · notes), Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama (Melvin Vivas on X · notes), Qwen 3.6 27B: An Open Local Model with Benchmarks Near Claude Opus 4.5 (Melvin Vivas on X · notes), Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs (Melvin Vivas on X · notes) and 1 more - Qwen 3.6 35B (MTP) · tool · huggingface.co · free
Qwen 3.6 mixture-of-experts model (35B total, about 3B active parameters) with multi-token prediction for local inference.
Also in: Use a Local Model for Confidential Data with Your Agent (Melvin Vivas on X · notes), Running Qwen 3.6 35B locally on an RTX 3090 for agent tool calling (Melvin Vivas on X · notes), Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Run Hermes Agent locally with Qwen 3.6 35B MTP in LM Studio (Melvin Vivas on X · notes) and 4 more
Try this
- Try Unsloth's MTP GGUFs and follow their guide to speed up local inference.
More in LLMOps, Deployment & Monitoring
- Tracking LLM Spend, Tokens & Guardrails with OpenRouter's Activity Explorer
- Bonsai Image Model Running In-Browser with WebGPU
- Running Qwen3.6-27B Fully in the Browser With WebGPU and wllama
- Cheap TTS serving: Qwen3-TTS on vLLM-Omni at $3 per 1M characters
- LM Studio MLX v1.8.1: Vision Model Batching and Better Caching
- OpenRouter Pareto Code: cost-optimized coding router