Waymark · Spatial · Model 04 / 04

Waymark Gaze

Eye tracking from a webcam: gaze direction and blinks, in the browser or over the API

It looks at your eyes through a camera and says which way they point and, after a short calibration, where on the screen you are looking.

For web pages and apps that follow a person’s gaze with their consent, and for anyone curious to try eye tracking.

Released In the browser and on the API since 2026-10-01; measured on synthetic eyes only Spatial v1.0.0

Waymark Gaze is a 420K-parameter CNN that reads two eye patches, in the browser or over the API, and after a nine-dot calibration places a gaze on the screen.

Overview #

Waymark Gaze follows your eyes through a laptop webcam and estimates the point on the screen you are looking at. You open the demo, allow the camera, look at nine dots so the page can calibrate to you, and a dot then follows your gaze and pauses while you blink. In the demo everything runs inside the browser tab: no video, picture or measurement is sent anywhere. The same model also answers over the Falcon API, for hosts that cut the eye patches themselves or send a photo of a face.

Waymark Gaze is a 420,221-parameter convolutional network that reads one gray eye patch of 36 by 60 pixels and returns the direction that eye is pointing as a pitch and a yaw angle, how open the eye is, and where the pupil sits in the patch. It is one stage of a pipeline that runs per camera frame: a face-landmark model finds the eye corners, the page cuts a normalised patch around each eye, Waymark Gaze reads both patches, and a mapping fitted on the nine calibration dots turns the two gaze directions into a point on the screen.

It is the fourth model of the Waymark series and the only one that reads pixels. It runs in two places on the same weights: in the visitor’s browser, where the demo needs no key and sends nothing, and behind the Falcon API, on three keyed routes that take eye patches, a photo, or calibration samples.

Waymark Gaze was trained and measured on computer-drawn eyes only. Its accuracy on real faces has not been measured; the figures on this page say how well it reads the eyes its own renderer draws.

Intended use #

  • Trying webcam eye tracking in the demo: seeing an estimate of where you look, on your own screen, after a short calibration.
  • Web pages and apps that follow the gaze of the person using them, with that person’s knowledge and consent, as a convenience and never as the only way to operate something.
  • Reading the gaze direction in a photo of a face over the API, when the person in the photo has agreed to it.
  • Detecting blinks of the person at the screen from the openness output.
  • Research and teaching on appearance-based gaze estimation, with the synthetic-only training stated wherever a figure is quoted.

Out of scope #

  • Medical, diagnostic or clinical use, and anything where a wrong gaze point has safety consequences.
  • Accessibility-critical control, where a person would depend on the estimate to communicate or to operate a device.
  • Watching, recording or scoring people without their knowledge: attention monitoring, proctoring, covert observation. Waymark Gaze is for the person at the screen, on their own computer.
  • Identifying anyone. It reads the direction of an eye and identifies nobody; it is not a face-recognition or iris-recognition model.
  • Phones, external cameras far from the screen, more than one face, or use without the nine-dot calibration.

Choose Waymark Gaze when #

  • The question is where the person at this screen is looking, and the answer may be approximate.
  • Everything must stay in the browser, with no upload and no server to run: use the demo’s approach and the weights file.
  • The host is not a browser, or already has the eye corners: send the two patches to /v1/gaze, or a photo to /v1/gaze/photo.
  • The question is about a photographed scene rather than an eye — choose Waymark for a layout answer, or Waymark Extra for a walk through the scene.
  • The question is where a person is walking, from an observed track — choose LIM3D-XL.

Specification #

Parameters420,221
Input36 × 60 gray eye patch, values 0–255, standardised per patch inside the model
BackboneThree blocks of two 3 × 3 convolutions with batch normalisation and ReLU, each followed by 2 × 2 max pooling: 24, 48 and 96 channels
HeadFlatten → linear to 96 → ReLU → linear to 5
Outputspitch, yaw (radians); openness (sigmoid, 0–1); pupil x / 10, pupil y / 10 (patch pixels)
Blink thresholdopenness below 0.3
Weightsgaze.onnx, 1,686,037 bytes (ONNX opset 17, float32, dynamic batch)
Runs onThe Falcon API process (a TypeScript forward pass over the same weights), or ONNX Runtime Web 1.30.0 on WebAssembly in the visitor’s browser
Needs beside itA face-landmark model for the eye corners (the demo and the photo route use MediaPipe Face Landmarker) and a nine-dot calibration

Try it #

You send
Open the demo at falconlab.app/gaze on a laptop, allow the camera and look at nine dots, one after another.
You get back
A dot on the screen that follows where you look and pauses while you blink. The camera picture stays in your browser.

The same exchange as the API sees it:

json
{
  "patches": {
    "type": "float32",
    "dims": [2, 1, 36, 60],
    "data": "2 × 36 × 60 gray values from 0 to 255: the right eye patch row by row, then the mirrored left eye patch"
  }
}
json
{
  "gaze": {
    "type": "float32",
    "dims": [2, 5],
    "data": [0.062, -0.118, 0.94, 2.87, 1.74, 0.058, 0.131, 0.91, 3.05, 1.79]
  }
}

Limits & safety #

Out of scope: clinical measurement, driver or operator monitoring, attention scoring, and any decision about a person taken from where they looked. A host that embeds the model must ask for the camera plainly, keep the picture in the browser unless the person has agreed to its being sent, and offer another way to do whatever the gaze controls. A host that calls the API must have the consent of the person whose eyes or face it sends, and must not send pictures taken without that person’s knowledge.

  • Accuracy on real faces has not been measured. Every metric on this page is synthetic, and real eyes are expected to be harder by an unknown amount.
  • It needs calibration. Without the nine dots there is no screen point, and the calibration holds only while the head stays roughly where it was. The API keeps no calibration: a host that loses the mapping must calibrate again.
  • It needs a laptop-style webcam near the screen and decent light. A dark, backlit or low-resolution picture gives worse landmarks and worse patches; the demo warns about low light and a missing face.
  • Glasses, half-closed eyes and large gaze angles raise the error even on synthetic eyes (3.99°, 4.53° and 4.17° against 3.49° overall on gaze-valid patches).
  • It tracks where the person using the page is looking, with their consent, and identifies nobody. It is not for monitoring, proctoring or any covert use.
  • It is not a medical device, not a safety system and not an assistive technology anyone should depend on.
  • The model can have learned artefacts of its renderer: skin without pores or wrinkles, flat glasses, one lid model, mirrored left eyes.
  • On the API, the patch route trusts the patch it is given. A patch cut with another corner distance, centre, mirroring or gray conversion is answered without an error, with angles that mean something else.
  • The photo route reads one face, the largest, from one still picture of at most 6,000,000 bytes. How well its landmark step places the eye corners on real photographs has not been measured, and neither has the route’s end-to-end error.
  • Every frame is one call and one spatial unit. A calibration is about ninety calls, and following a gaze frame by frame uses the allowance quickly; the browser demo, which calls nothing, is the path for continuous tracking.
  • A mapping fitted on one route’s features, one person, one screen and one seat does not carry over to another.
Spatial Released
Trained on Indoor + outdoor

Give it a written sketch of a photo and ask where something is; it answers in one short sentence.

Sketch-to-layout language model for viewer-relative scene answers

v1.1.0
Spatial Released
Trained on Outdoor

Give it a sketch of a scene and a goal like “walk to the door”, and it describes the walk step by step.

Movement traces from a scene sketch and a goal, with a plain-words summary

v1.0.0
Intention Released

The more careful version of the movement reader: it looks at every object and every step before it answers.

Attention over every object and every step of an observed motion

v1.0.0

Latest versions #

VersionDateStatusNote
1.0.0ReleasedFirst released version: the epoch-14 checkpoint of the 2026-09-30 run, served with the demo at /gaze and, from the same day, on the API. Evaluated on synthetic eyes only.

Read the full documentation

Nine chapters: architecture, inputs and outputs, training, evaluation, API, runtime, limits and versions.

Full documentation