Technology

What happens in the second between question and answer

When you ask “what's ahead?”, the phone in your hand fires up the whole machine, from the photo and depth measurement, through AI models, to voice and vibration. Scroll to see every step.

Step one

Photo and depth

The camera captures a frame while the LiDAR sensor measures the distance to everything in view. When the phone has no LiDAR, a machine-learning model estimates the depth instead.

  • LiDAR
  • Depth map
  • ML model (fallback)
Step two

On-device analysis

Before anything leaves the phone, two models analyse the same frame: one recognises whether you're outdoors or inside, the other, from feature points, whether the scene changed since the last frame. If it didn't, there's no need to send anything.

  • ML models
  • Feature points
  • On-device
Step three

Secure send to the AI

The frame, depth map and conversation context go to OpenAI's GPT Realtime model in the cloud. App Attest confirms it's the real app connecting, and everything is encrypted. This work is measured in tokens, and that's where the price comes from.

  • GPT Realtime
  • App Attest
  • Encryption
  • Tokens
Step four

Voice and vibration

Post, on the right, very close.

The answer comes back as speech: short and predictable, what, which side, how far. Near an obstacle the phone also vibrates, and that feature works even without the internet.

  • Speech synthesis
  • Haptics
  • VoiceOver

See what it looks like on screen

All of this machinery sits under a single screen with four buttons. The guide walks you through each of them.