Gemma 4 Gets Up to 3x Faster with MTP Drafters
Melvin Vivas · X post · 2026-05-06 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
Gemma 4 now has Multi-Token Prediction (MTP) drafters for speculative decoding. They give up to 3x more tokens per second with the same reasoning output. Support was available from day one in Transformers, MLX and vLLM, under an Apache 2.0 license.
Key points
- MTP drafters use speculative decoding to speed up generation.
- Up to 3x more tokens/sec than normal Gemma 4.
- Same reasoning output, just faster.
- Day-0 support in Hugging Face Transformers, MLX and vLLM.
- Apache 2.0 license.
Resources mentioned
- Gemma 4 · tool · ai.google.dev · free
Google's family of open-weight models in several sizes, built to run on devices and offline, with multimodal and agentic abilities, and open to fine-tuning.
Also in: Gemma 4 Runs Locally On-Device in the Antigravity SDK (Melvin Vivas on X · notes), On-Device AI: Running Gemma 4 E2B Offline on an iPhone with LiteRT (Melvin Vivas on X · notes), Running Gemma 4 Models Offline on an iPhone (Melvin Vivas on X · notes), Fine-tuning Gemma4-E2B on your own tweet style with Unsloth (Melvin Vivas on X · notes) and 25 more - Hugging Face Transformers · tool · github.com · free
Open-source library for loading and running pretrained models, which now supports GGUF directly.
Also in: Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), Run GGUF models directly in Hugging Face Transformers (Melvin Vivas on X · notes), AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers (Melvin Vivas on X · notes) - MLX · repo · github.com · free
Apple's open-source array and ML framework for efficient inference on Apple Silicon.
Also in: LM Studio MLX v1.8.1: Vision Model Batching and Better Caching (Melvin Vivas on X · notes), 1-bit Bonsai 8B Runs On-Device on iPhone at 40+ tok/s (Melvin Vivas on X · notes), Qwen 3.5 2B Runs On-Device on iPhone with MLX (Melvin Vivas on X · notes) - vLLM · repo · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Open-source, high-throughput LLM inference and serving engine with continuous batching and efficient KV-cache management (PagedAttention).
Also in: Ollama vs vLLM: From Local AI Demo to Production Inference Serving (Bashiri Smith on Facebook · notes), Serve GLM-5.2 NVFP4 with vLLM on NVIDIA Blackwell (Melvin Vivas on X · notes)
Try this
- Compare Gemma 4's tokens/sec with and without MTP speculative decoding in Transformers, MLX or vLLM.
- Benchmark Gemma 4 speed with and without MTP drafters on your own hardware.
More in LLM Fundamentals
- Qwen3.7-Max: Qwen's flagship model for agents
- MiniCPM-V 4.6 1.3B: Small Open-Source Vision/OCR Model for Edge Devices
- Enable Gemma 4 MTP Speculative Decoding on iPhone
- Hy-MT1.5-1.8B-1.25bit: A 440MB Offline Phone Translation Model
- Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs
- Qwen3.6-27B: Dense Open Model for Local Agentic Coding