Technology
Attention is a pattern, not a point.
No single signal says what a driver is doing. AutoSight combines the eyes, head, hands, body, and phone motion, weighs each by its quality in the moment, and reads them across time.
Signals
Eight signals, one estimate.
Built on Google's open-source MediaPipe face, hand, and pose models, which track 478 points on the face alone, including the irises.[5] Everything runs on the phone.
Eye gaze
Where the eyes point, tracked separately from the head.
Head pose
Yaw, pitch, and roll of the head, from facial geometry.
Eyelids
Blink length and eyelid openness, for drowsiness.
Hand position
Whether hands stay in the wheel region or move toward the lap and console.
Upper-body posture
Shoulder and arm position. A lowered arm and dropped shoulder point to the lap.
Screen light
At night, light from a phone below the camera's view shows up on the face.
Phone motion
If the mounted phone is picked up or handled mid-drive, its motion sensors know.
Time
Duration, repetition, and the order in which attention moves.
Architecture
Five layers, one question: is attention still on the road?
- LAYER 1
Perception
Tracks face landmarks, irises, eyelids, head pose, hands, and upper-body pose from the front camera, plus the phone's own motion sensors.
gazeheadeyelidshandsposture - LAYER 2
Temporal model
Reads behavior across seconds: how long attention stays somewhere, how often it returns, and the sequence it follows.
durationtransitionsrepetition - LAYER 3
Signal-quality fusion
Scores every signal for quality, frame by frame. Glare, sunglasses, and low light shift weight toward the signals that remain clear.
qualityocclusionfusion - LAYER 4
Attention state
Estimates whether the driver is attentive, transitioning, distracted, or drowsy, and checks that the signals agree over time.
attentiondrowsinessconsistency - LAYER 5
Response
Matches the response to the moment: keep monitoring, verify, warn, or escalate. A quick glance and a sustained lapse never get the same response.
monitorverifywarnescalate
Temporal analysis
Three seconds tell a story one frame can't.
Good drivers look away from the windshield constantly: mirrors, blind spots, intersections, instruments. AutoSight treats that as driving. What it watches is where attention went, how long it stayed, and what came before and after.
In NHTSA's 100-Car study, off-road glances adding up to more than two seconds at least doubled crash and near-crash risk.[2] Duration sits at the center of the response logic.
- 0.0 sRoadDriving.
- 0.4 sMirrorMirror check. No response.
- 0.8 sRoadBack on the road.
- 1.4 sDownGlance down. Duration tracking starts.
- 2.1 sDownStill down. State moves to transitioning.
- 2.8 sDownSustained off-road attention. Alert.
Hard cases
Built for the situations that matter most.
Sunglasses and eyewear
Eyes hidden. Still covered.
When lenses or glare hide the pupils, AutoSight shifts weight to head pose, hand position, posture, screen light, and timing. A hand leaving the wheel while the head tips toward the lap reads as distraction with or without the eyes.
A phone below the camera
Read the driver, not the device.
The question isn't whether the camera can see a phone. It's whether the driver's behavior shows attention has left the road. Gaze into the lap, a lowered arm, and repeated downward glances answer that without the phone ever appearing in frame.
Natural scanning
Mirror checks aren't distraction.
Duration and sequence logic separate the scanning every good driver does from the lapses that matter, so alerts stay rare and meaningful.
Gaming the system
Looking compliant isn't enough.
AutoSight checks that signals agree over time. A face held toward the road while the eyes keep dropping, or a hand that keeps leaving the wheel, doesn't pass as attention.
Where it runs
One model, built to move between platforms.
| Platform | Sensors | Use |
|---|---|---|
| Smartphone app | Front camera, motion sensors | New drivers and families. The first product. |
| Dedicated driver camera | Infrared camera | Stronger night and sunglasses performance. |
| Fleet | Phone or dedicated camera | Attention coaching and analytics for commercial drivers. |
| SDK and vehicle integration | Existing in-cabin cameras | The attention model as a layer inside other systems. |
Work on the hard parts with us.
If you work in computer vision, human factors, or driver safety, we want your critique.