Demo: A Real-Time Vision AI Assistant That Coaches Your Diet
Melvin Vivas · X video post · 2026-05-12 · 1:00 · 94 views · Open on X
Topics: Industry Trends & Job Market · Level: beginner
Summary
A one-minute demo of a voice-and-camera AI assistant that watches what the user picks up and gives diet advice in real time. The user sets a goal at the start (no caffeine, no sweets). The assistant then recognizes objects on camera (a coffee cup, green cans, sugary sticks, an apple) and reacts to each one. It shows how a live multimodal assistant can keep a goal in mind across a conversation and update its view when the user corrects it. The post doesn't name the model or product behind the demo.
Key points
- The user states a goal at the start ('keep me away from caffeinated drinks or sweet treats'), and the assistant applies it to everything it sees afterward.
- The assistant identifies objects from the live camera feed (a black cup that looks like coffee, a green can, sugary sticks, a green apple) and judges each one against the goal.
- It suggests healthier swaps on its own, such as herbal tea instead of coffee.
- When the user explains the cans are unsweetened green tea, the assistant accepts it and changes its advice, while still flagging the caffeine late in the day.
- It explains why each choice is good or bad (for example, the apple has 'no added sugar, no caffeine'), which makes the advice easier to trust.
- The pattern to notice: goal from the user, then live vision, then spoken feedback that keeps a running memory of the conversation. This is what the newest multimodal assistants can do.
Try this
- Build a camera-based diet coach: the user sets food or drink rules, and a vision-capable multimodal model checks webcam frames and gives spoken or written advice that follows those rules.
More in Industry Trends & Job Market
- Figure F.03 Livestream: Humanoid Robots Working Nonstop
- Humanoid Robots Run a Full 8-Hour Shift with Helix-02
- Claude Plans Get Monthly Credit for Agent SDK and claude -p
- Self-Flow: Training Multimodal Generative Models Without an External Encoder
- "Veo Omni?": A 10-Second Teaser Clip of a Possible New Google Video Model
- OpenAI launches the OpenAI Deployment Company