Gemini Omni: Google's Any-to-Video Multimodal Model (Google I/O)
Melvin Vivas · X video post · 2026-05-19 · 0:07 · 59 views · Open on X
Topics: Industry Trends & Job Market, LLM Fundamentals · Level: beginner
Summary
This is a short reaction post. Melvin Vivas quotes Google's Google I/O announcement of Gemini Omni, a model meant to create anything from any input, starting with video. You can mix images, audio, video and text as input to generate high-quality videos, or use drawings to guide what gets generated. The post is a trend signal about multimodal generation and has no tutorial.
Key points
- Gemini Omni was announced at Google I/O (#GoogleIO) and is described as able to 'create anything from any input'.
- Video is the first output it supports.
- Inputs can be images, audio, video and text, combined in one request, to generate high-quality videos.
- Drawings or sketches can be used to steer what gets generated so it matches your vision.
- The trend: generative models are moving from text-only to any-input, any-output multimodal systems.
- The 7-second clip has no speech, so the quoted announcement text is the only content.
Resources mentioned
- Gemini Omni · tool · deepmind.google · check price
Google's multimodal generative model that takes combined image, audio, video, text and drawing inputs and generates high-quality video.
Also in: Opinion: Seedance 2.0 Still Beats Gemini Omni for Video (Melvin Vivas on X · notes) - Google I/O · other · io.google · free
Google's yearly developer conference, where Gemini Omni was announced.
More in Industry Trends & Job Market
- Gemini 3.5 Flash Released
- Qwen 3.5 Live Translate: Real-Time Speech Translation Demo
- Qwen3.5-LiveTranslate: Real-Time Audio-Visual Interpretation Model
- Cursor Launches Composer 2.5 with Doubled Usage for a Week
- Qwen 3.7 Plus Preview Ranks #16 in the Vision Arena
- Figure F.03 Livestream: Humanoid Robots Working Nonstop