A Visual Guide to Gemma 4 Architecture (Maarten Grootendorst)
Melvin Vivas · X post · 2026-04-04 · Open on X
Topics: LLM Fundamentals · Level: intermediate
Summary
The creator recommends Maarten Grootendorst's illustrated guide to how the Gemma 4 models are built. The article uses nearly 40 custom visuals to explain Mixture of Experts, the vision encoder, per-layer embeddings and the audio encoder in Google DeepMind's new open models.
Key points
- The guide has almost 40 custom visuals explaining Gemma 4's internals.
- It covers Mixture of Experts (MoE) layers.
- It covers the vision encoder used for image inputs.
- It covers Per-Layer Embeddings (PLE), used in the small Gemma 4 models.
- It covers the audio encoder for speech and audio inputs.
Resources mentioned
- A Visual Guide to Gemma 4 · article · newsletter.maartengrootendorst.com · free
Illustrated walkthrough of the Gemma 4 architecture: MoE, vision encoder, per-layer embeddings and audio encoder.
Also in: Gemma 4 Architecture: MoE, Encoders and Per-Layer Embeddings (Melvin Vivas on X · notes) - Maarten Grootendorst · person · x.com · free
Author of visual guides to LLMs and co-author of Hands-On Large Language Models; worth following for clear explanations of model architectures. - Exploring Language Models (Maarten Grootendorst's Substack) · newsletter · newsletter.maartengrootendorst.com · free
Maarten Grootendorst's newsletter of visual guides to LLM concepts and architectures.
Try this
- Read A Visual Guide to Gemma 4 to understand MoE, vision/audio encoders and per-layer embeddings.
More in LLM Fundamentals
- Using Gemma 4 via the Gemini API and Google AI Studio
- Run Gemma 4 on Your Phone with Google AI Edge Gallery
- Gemma 4 E4B: Local Image Understanding at 131K Context in 6GB VRAM
- Gemma 4 Architecture: MoE, Encoders and Per-Layer Embeddings
- Local Audio Transcription with Cohere Transcribe on WebGPU
- Gemma 4 26B for OCR in LM Studio