Google Launches Guided Vision in Gemini Live to Describe the World in Real Time
Google has rolled out Guided Vision on Android, allowing Gemini Live to provide real-time spoken descriptions and answer questions about objects captured on camera.

Google Brings Real-Time Visual Assistance to Gemini Live
Google has launched Guided Vision, an accessibility feature integrated into Gemini Live for compatible Android smartphones. Designed to assist users who are blind, have low vision, or need situational visual help, the tool uses artificial intelligence to generate spoken, real-time descriptions of whatever is captured through the device's camera.
Interactive Descriptions, Audio Cues, and System Integration
Guided Vision operates when users share their camera feed directly within Gemini Live. The AI can read fine print, identify and locate nearby objects, detail physical surroundings, and describe specific attributes of items in view.
The tool supports interactive, conversational follow-ups. For example, after Gemini identifies a product on a shelf, a user can ask it to locate and read out the expiration date. To address framing challenges, the software delivers directional audio cues that guide users to reposition their phone if the subject they are asking about is outside the camera shot.
Guided Vision is accessible directly within the Gemini app, via Google’s TalkBack screen reader, or through a configurable accessibility shortcut in the Android Settings menu on devices running Android 9 and newer. The capability follows an initial announcement in recent Pixel software updates. Google explicitly warns that Guided Vision should not be used as a replacement for a cane or mobility aid, cautioning users against relying on the software for navigation, safe travel, or obstacle detection.
The Push for Multimodal Accessibility
The release brings Google's AI assistant in direct alignment with competitors leveraging computer vision for accessibility. Apple offers a similar capability with its VoiceOver Live Recognition tool across iPhones and the Vision Pro headset, reflecting a broader industry push to deploy low-latency, multimodal AI models as assistive aids for real-world visual interpretation.



