Run Gemma 4 Offline on an iPhone with the Locally AI App
Melvin Vivas · X video post · 2026-04-17 · 0:43 · 564 views · Open on X
Topics: LLM Fundamentals, AI Dev Tools & Productivity · Level: beginner
Summary
This short demo shows how to run Google's open Gemma 4 model on an iPhone for free with the Locally AI app. You download the app, get Gemma 4 from Manage Models, and select it. After that you can chat with the model fully offline, even in airplane mode, with no API calls, data plan or subscription.
Key points
- Install the Locally AI app on your iPhone.
- Open the app, go to Manage Models, scroll to Gemma 4, download it, then select it.
- Once downloaded, the model runs on the device, so it works in airplane mode with no internet.
- Running locally means no API calls, no data plan and no monthly fees, and your prompts stay on the phone.
- The quoted post says Gemma 4 handles long context even on the phone.
- Demo prompt: 'Could you explain how quantum computing works in simple terms?' The model answered offline.
Resources mentioned
- Gemma 4 · tool · ai.google.dev · free
Google's family of open-weight models in several sizes, built to run on devices and offline, with multimodal and agentic abilities, and open to fine-tuning.
Also in: Gemma 4 Runs Locally On-Device in the Antigravity SDK (Melvin Vivas on X · notes), On-Device AI: Running Gemma 4 E2B Offline on an iPhone with LiteRT (Melvin Vivas on X · notes), Running Gemma 4 Models Offline on an iPhone (Melvin Vivas on X · notes), Fine-tuning Gemma4-E2B on your own tweet style with Unsloth (Melvin Vivas on X · notes) and 25 more - Locally AI · tool · locallyai.app · free
An iPhone app for downloading open models and running them on the device, offline.
Also in: 1-bit Bonsai 8B Runs On-Device on iPhone at 40+ tok/s (Melvin Vivas on X · notes)
Try this
- Download the Locally AI app on your iPhone.
- In Manage Models, download Gemma 4 and select it.
- Turn on airplane mode and send a prompt to confirm the model runs offline.
More in LLM Fundamentals
- Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs
- Qwen3.6-27B: Dense Open Model for Local Agentic Coding
- Gemma 4 Now Available Through the Gemini API (Dev Use Only)
- OpenRouter Adds Video Generation: One API for Veo, Seedance, Wan and Sora
- Running Gemma 4 E2B Video Understanding Locally in WSL
- Using Gemma 4 via the Gemini API and Google AI Studio