Gemini Embedding 2: One Vector Space for Five Modalities
Melvin Vivas · X post · 2026-04-23 · Open on X
Topics: Embeddings & Vector Databases, Retrieval-Augmented Generation (RAG) · Level: intermediate
Summary
Gemini Embedding 2 is now generally available. It is a single embedding model that maps text, images, video, audio and PDFs into one shared vector space, which allows search across modalities.
Key points
- Gemini Embedding 2 is generally available (GA).
- It embeds 5 modalities (text, images, video, audio, PDFs) into one unified space.
- It supports up to 8,192 input tokens and 100+ languages.
- It embeds audio natively, with no transcription step.
- The output size is flexible.
Resources mentioned
- Gemini Embedding 2 · tool · blog.google · check price · open in a browser to verify
Google's multimodal embedding model, with one vector space for text, image, video, audio and PDF.
Try this
- Build a multimodal search/RAG system that searches text, images, audio and PDFs in one Gemini Embedding 2 index.
More in Embeddings & Vector Databases
- How RAG Finds the Right Document Fast: Graph-Based Vector Search
- Sentence Transformers v6: ColBERT-Style Multi-Vector Models Fully Supported
- Building Interactive AI Concept Demos with Claude Design (Embeddings Example)
- Fine-Tuning MiniLM Embeddings with Synthetic Data, Running on CPU
- Fine-tuning all-MiniLM-L6-v2 to Beat OpenAI Embeddings