Gemma 4 Now Available Through the Gemini API (Dev Use Only)
Melvin Vivas · X post · 2026-04-19 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: beginner
Summary
Google's open Gemma 4 models can now be called through the Gemini API. The creator says it suits development and prototyping only, because rate limits make it unsuitable for production for now. The post links to the official docs.
Key points
- Gemma 4 can be called through the Gemini API, so you can try it without hosting the model yourself.
- Rate limits currently make it a development and prototyping option, not a production one.
- For production, consider self-hosting Gemma or using a production-grade Gemini model.
- The official guide is 'Run Gemma with the Gemini API' on Google AI for Developers.
Resources mentioned
- Run Gemma with the Gemini API | Google AI for Developers · docs · ai.google.dev · free
Official guide to calling Gemma models through the Gemini API. - Gemma 4 · tool · ai.google.dev · free
Google's family of open-weight models in several sizes, built to run on devices and offline, with multimodal and agentic abilities, and open to fine-tuning.
Also in: Gemma 4 Runs Locally On-Device in the Antigravity SDK (Melvin Vivas on X · notes), On-Device AI: Running Gemma 4 E2B Offline on an iPhone with LiteRT (Melvin Vivas on X · notes), Running Gemma 4 Models Offline on an iPhone (Melvin Vivas on X · notes), Fine-tuning Gemma4-E2B on your own tweet style with Unsloth (Melvin Vivas on X · notes) and 25 more - Gemini API · docs · ai.google.dev · free
Google's developer API and docs for calling Gemini and Gemma models.
Also in: Gemini 3.5 Live Translate: Real-Time Speech Translation via the Live API (Melvin Vivas on X · notes), Using Gemma 4 via the Gemini API and Google AI Studio (Melvin Vivas on X · notes)
Try this
- Read the 'Run Gemma with the Gemini API' docs.
- Use Gemma 4 through the Gemini API for development and prototyping only, not production.
More in LLM Fundamentals
- Hy-MT1.5-1.8B-1.25bit: A 440MB Offline Phone Translation Model
- Run Qwen3.6-27B Locally in 18GB RAM with Unsloth GGUFs
- Qwen3.6-27B: Dense Open Model for Local Agentic Coding
- Run Gemma 4 Offline on an iPhone with the Locally AI App
- OpenRouter Adds Video Generation: One API for Veo, Seedance, Wan and Sora
- Running Gemma 4 E2B Video Understanding Locally in WSL