Running Gemma 4 E2B Video Understanding Locally in WSL
Melvin Vivas · X post · 2026-04-16 · Open on X
Topics: LLM Fundamentals, AI Dev Tools & Productivity · Level: intermediate
Summary
The creator reports getting video understanding working locally with Google's small Gemma 4 E2B model on Windows Subsystem for Linux (WSL). It shows that small open models can do multimodal video tasks on a local machine.
Key points
- Gemma 4 E2B is a small Gemma 4 model that can do video understanding.
- The creator got it running locally inside WSL on Windows.
- Small open models make it possible to run multimodal (video) tasks locally.
Resources mentioned
- Gemma 4 · tool · ai.google.dev · free
Google's family of open-weight models in several sizes, built to run on devices and offline, with multimodal and agentic abilities, and open to fine-tuning.
Also in: Gemma 4 Runs Locally On-Device in the Antigravity SDK (Melvin Vivas on X · notes), On-Device AI: Running Gemma 4 E2B Offline on an iPhone with LiteRT (Melvin Vivas on X · notes), Running Gemma 4 Models Offline on an iPhone (Melvin Vivas on X · notes), Fine-tuning Gemma4-E2B on your own tweet style with Unsloth (Melvin Vivas on X · notes) and 25 more - WSL (Windows Subsystem for Linux) · tool · learn.microsoft.com · free
Runs a Linux environment such as Ubuntu on Windows, here used to build llama.cpp with CUDA.
Also in: Any OS works for AI dev: Mac, Windows + WSL (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes), Installing Unsloth Studio on Windows via WSL (Melvin Vivas on X · notes)
Try this
- Run Gemma 4 E2B locally (for example in WSL) and test it on your own video clips.
More in LLM Fundamentals
- Gemma 4 Now Available Through the Gemini API (Dev Use Only)
- Run Gemma 4 Offline on an iPhone with the Locally AI App
- OpenRouter Adds Video Generation: One API for Veo, Seedance, Wan and Sora
- Using Gemma 4 via the Gemini API and Google AI Studio
- Run Gemma 4 on Your Phone with Google AI Edge Gallery
- Gemma 4 E4B: Local Image Understanding at 131K Context in 6GB VRAM