Gemma 4 as a Strong Small Model for Local and On-Device Use
Melvin Vivas · X post · 2026-09-11 · Open on X
Topics: LLM Fundamentals · Level: beginner
Summary
A short post saying that Google's Gemma 4 is still one of the strongest small models you can run locally and on-device.
Key points
- Gemma 4 is recommended as a strong small model for running locally or on-device.
- Consider small open models when you need local inference.
Resources mentioned
- Gemma 4 · tool · ai.google.dev · free
Google's family of open-weight models in several sizes, built to run on devices and offline, with multimodal and agentic abilities, and open to fine-tuning.
Also in: Gemma 4 Runs Locally On-Device in the Antigravity SDK (Melvin Vivas on X · notes), On-Device AI: Running Gemma 4 E2B Offline on an iPhone with LiteRT (Melvin Vivas on X · notes), Running Gemma 4 Models Offline on an iPhone (Melvin Vivas on X · notes), Fine-tuning Gemma4-E2B on your own tweet style with Unsloth (Melvin Vivas on X · notes) and 25 more
More in LLM Fundamentals
- DeepSeek 4.1 Flash speed on the official API: about 325 tokens/s
- Read the GPT-6 Astra launch article to learn what the model can do
- When a higher reasoning level is worth the extra cost
- Running Qwen3.8 27B Locally with llama.cpp for Writing
- Three Local GGUF Models That Fit on an RTX 3090 (24GB)
- Comparing GPT Models in Codex by Intelligence and Cost per Task