Photo and depth
The camera captures a frame while the LiDAR sensor measures the distance to everything in view. When the phone has no LiDAR, a machine-learning model estimates the depth instead.
- LiDAR
- Depth map
- ML model (fallback)
When you ask “what's ahead?”, the phone in your hand fires up the whole machine, from the photo and depth measurement, through AI models, to voice and vibration. Scroll to see every step.
The camera captures a frame while the LiDAR sensor measures the distance to everything in view. When the phone has no LiDAR, a machine-learning model estimates the depth instead.
Before anything leaves the phone, two models analyse the same frame: one recognises whether you're outdoors or inside, the other, from feature points, whether the scene changed since the last frame. If it didn't, there's no need to send anything.
The frame, depth map and conversation context go to OpenAI's GPT Realtime model in the cloud. App Attest confirms it's the real app connecting, and everything is encrypted. This work is measured in tokens, and that's where the price comes from.
Post, on the right, very close.
The answer comes back as speech: short and predictable, what, which side, how far. Near an obstacle the phone also vibrates, and that feature works even without the internet.
All of this machinery sits under a single screen with four buttons. The guide walks you through each of them.