AI Engineer Study Library

Multimodal AI Pointer: Combining Voice, Mouse Pointing and Vision with Gemini

Melvin Vivas · X video post · 2026-05-13 · 0:57 · 86 views · Open on X

Topics: AI Agents, Tool Use & MCP, Industry Trends & Job Market, Prompt & Context Engineering · Level: beginner

Summary

Melvin Vivas shares a Google DeepMind demo of a prototype that puts Gemini behind the mouse pointer. You point at things on screen and speak deictic words like "this", "that", "here" or "there". The model combines your voice, where the pointer is, and what it sees on screen to work out what you mean, then acts across apps. It shows a new kind of multimodal, agent-style interface where the pointer becomes a way to give the model context.

Key points

Resources mentioned

Try this

More in AI Agents, Tool Use & MCP

All of AI Agents, Tool Use & MCP