Waymark Gaze
Eye tracking from a webcam: gaze direction and blinks, in the browser or over the API
Four to six sentences from Cognitio, written on request. Generated text: the page is the reference.
On this page
Overview #
Waymark Gaze follows your eyes through a laptop webcam and estimates the point on the screen you are looking at. You open the demo, allow the camera, look at nine dots so the page can calibrate to you, and a dot then follows your gaze and pauses while you blink. In the demo everything runs inside the browser tab: no video, picture or measurement is sent anywhere. The same model also answers over the Falcon API, for hosts that cut the eye patches themselves or send a photo of a face.
Waymark Gaze is a 420,221-parameter convolutional network that reads one gray eye patch of 36 by 60 pixels and returns the direction that eye is pointing as a pitch and a yaw angle, how open the eye is, and where the pupil sits in the patch. It is one stage of a pipeline that runs per camera frame: a face-landmark model finds the eye corners, the page cuts a normalised patch around each eye, Waymark Gaze reads both patches, and a mapping fitted on the nine calibration dots turns the two gaze directions into a point on the screen.
It is the fourth model of the Waymark series and the only one that reads pixels. It runs in two places on the same weights: in the visitor’s browser, where the demo needs no key and sends nothing, and behind the Falcon API, on three keyed routes that take eye patches, a photo, or calibration samples.
Waymark Gaze was trained and measured on computer-drawn eyes only. Its accuracy on real faces has not been measured; the figures on this page say how well it reads the eyes its own renderer draws.
Intended use #
- Trying webcam eye tracking in the demo: seeing an estimate of where you look, on your own screen, after a short calibration.
- Web pages and apps that follow the gaze of the person using them, with that person’s knowledge and consent, as a convenience and never as the only way to operate something.
- Reading the gaze direction in a photo of a face over the API, when the person in the photo has agreed to it.
- Detecting blinks of the person at the screen from the openness output.
- Research and teaching on appearance-based gaze estimation, with the synthetic-only training stated wherever a figure is quoted.
Out of scope #
- Medical, diagnostic or clinical use, and anything where a wrong gaze point has safety consequences.
- Accessibility-critical control, where a person would depend on the estimate to communicate or to operate a device.
- Watching, recording or scoring people without their knowledge: attention monitoring, proctoring, covert observation. Waymark Gaze is for the person at the screen, on their own computer.
- Identifying anyone. It reads the direction of an eye and identifies nobody; it is not a face-recognition or iris-recognition model.
- Phones, external cameras far from the screen, more than one face, or use without the nine-dot calibration.
Choose Waymark Gaze when #
- The question is where the person at this screen is looking, and the answer may be approximate.
- Everything must stay in the browser, with no upload and no server to run: use the demo’s approach and the weights file.
- The host is not a browser, or already has the eye corners: send the two patches to
/v1/gaze, or a photo to/v1/gaze/photo. - The question is about a photographed scene rather than an eye — choose Waymark for a layout answer, or Waymark Extra for a walk through the scene.
- The question is where a person is walking, from an observed track — choose LIM3D-XL.
Architecture #
| Parameters | 420,221 |
|---|---|
| Input | 36 × 60 gray eye patch, values 0–255, standardised per patch inside the model |
| Backbone | Three blocks of two 3 × 3 convolutions with batch normalisation and ReLU, each followed by 2 × 2 max pooling: 24, 48 and 96 channels |
| Head | Flatten → linear to 96 → ReLU → linear to 5 |
| Outputs | pitch, yaw (radians); openness (sigmoid, 0–1); pupil x / 10, pupil y / 10 (patch pixels) |
| Blink threshold | openness below 0.3 |
| Weights | gaze.onnx, 1,686,037 bytes (ONNX opset 17, float32, dynamic batch) |
| Runs on | The Falcon API process (a TypeScript forward pass over the same weights), or ONNX Runtime Web 1.30.0 on WebAssembly in the visitor’s browser |
| Needs beside it | A face-landmark model for the eye corners (the demo and the photo route use MediaPipe Face Landmarker) and a nine-dot calibration |
Waymark Gaze is a small convolutional network. The first operation standardises the patch: it subtracts the patch’s own mean and divides by its standard deviation, so the caller passes raw gray values and exposure differences between cameras matter less. Three blocks follow, each made of two 3 × 3 convolutions with batch normalisation and ReLU and a 2 × 2 max pooling, at 24, 48 and 96 channels; the 36 × 60 patch is 4 × 7 after the third. A linear layer of 96 units and a final linear layer of five outputs complete it: pitch and yaw as plain regressions, openness through a sigmoid, and the pupil centre in tenths of a patch pixel.
The gaze angles describe the optical axis of the eyeball in the patch’s own frame: x to the right, y down, z away from the camera. The direction is g = (cos p · sin y, −sin p, −cos p · cos y) for pitch p and yaw y, so both at zero mean the eye looks straight at the camera. Left eyes are cropped mirrored, so the model always sees the nose side at +x; the caller negates the yaw of a left eye. The offset between the optical axis and the line of sight differs from person to person and is left to the calibration.
What is deliberately absent: no face detector and no landmark model (the caller supplies the eye corners), no head-pose input, no memory of earlier frames, and no screen geometry. Turning two eye directions into a screen point is the calibration’s job, not the network’s.
Inputs & outputs #
Input #
| Field | Type | Required | Description | Limit |
|---|---|---|---|---|
patches | float32 [N, 1, 36, 60] | Yes | Gray eye patches, raw values 0 to 255 (0.299 R + 0.587 G + 0.114 B); the model standardises each patch itself. The demo sends two per frame: row 0 the right eye, row 1 the left eye mirrored, so the nose side is at +x in both. | Any batch size; 36 rows by 60 columns, eye corners 36 px apart on the middle row |
A patch is cut from the unmirrored camera frame by a similarity transform: the midpoint of the two eye corners goes to the patch centre (30, 18), the line from the outer to the inner corner becomes the +x axis, and the corner distance becomes 36 patch pixels. The demo takes the corners from MediaPipe Face Landmarker (landmarks 33 and 133 for the right eye, 263 and 362 for the left), converts to gray and samples bilinearly. The tensor for one frame, with the 4,320 values abbreviated:
{
"patches": {
"type": "float32",
"dims": [2, 1, 36, 60],
"data": "2 × 36 × 60 gray values from 0 to 255: the right eye patch row by row, then the mirrored left eye patch"
}
}Limits on the input:
- The patch geometry is part of the contract. A patch cut with a different corner distance, centre or mirroring gives angles that mean something else.
- The source should give the eye at least half a camera pixel per patch pixel; in the synthetic evaluation the error grows from 2.7° to 4.6° as the eye gets smaller in the frame.
- Values are raw gray levels, not normalised to 0–1 and not standardised by the caller.
Output #
| Field | Type | Description |
|---|---|---|
gaze | float32 [N, 5] | Per patch: pitch and yaw in radians (pitch above zero is up, yaw above zero is toward the nose side of the patch), openness from 0 to 1, and the pupil centre as x / 10 and y / 10 in patch pixels. |
{
"gaze": {
"type": "float32",
"dims": [2, 5],
"data": [0.062, -0.118, 0.94, 2.87, 1.74, 0.058, 0.131, 0.91, 3.05, 1.79]
}
}How to read the result:
- Each row is one patch: pitch, yaw, openness, pupil x / 10, pupil y / 10. The first row above is the right eye, the second the mirrored left eye.
- Pitch and yaw are radians. Negate the yaw of the left eye before combining the two, because its patch was mirrored.
- Openness is the largest lid gap divided by the corner distance, scaled so that an open eye reads about 1; below 0.3 counts as a blink. Narrow eyes read below 1 when fully open.
- The pupil centre is in patch pixels after multiplying by 10; for a left eye, x is counted from the mirrored side.
- The values above are illustrative of the shape; exact numbers depend on the frame.
Training #
Waymark Gaze was trained from scratch on 2026-09-30 on an Apple M5 Pro laptop, with PyTorch on the MPS backend: 14 epochs in 721 seconds, about 12 minutes. No cloud service and no network access were used at any stage.
The training data is entirely synthetic. A procedural renderer written in-house ray-casts a three-dimensional eye — eyeball, refracting cornea, iris, pupil, lids, lashes, skin, brows, optional glasses, lights and webcam degradations — from formulas and seeded random numbers. The run used 250,000 training patches from 31,250 generated identities, 10,000 validation patches and 10,000 test patches from 1,250 identities each. No photograph, scan or camera capture of a person was used, for training or for testing, and no public gaze dataset.
Recipe: AdamW, batch 256, peak learning rate 3e-3, weight decay 1e-4, light photometric augmentation. The run was planned for 30 epochs and shortened to 14 after epoch 4 to fit a local time budget: a one-cycle schedule for the first four epochs, then a cosine decay to 1 % of the peak rate. The released weights are the epoch-14 checkpoint, the best on validation (3.50° mean on gaze-valid patches); validation error was still falling slowly when the run ended. The loss combines a smooth L1 term on the two angles, an L1 term on openness and an L1 term on the pupil centre.
What it was not trained on: real eyes of any kind; real webcam frames; left eyes as such (a left eye is a mirrored right eye); contact lenses, makeup, skin pores, hair or wrinkles; lens refraction through glasses, which are drawn as a flat rim with glare.
Evaluation #
| Metric | Value | Source |
|---|---|---|
| Accuracy on real faces | Not measured | No evaluation on real people, real webcam frames or a public gaze dataset has been done; every figure below is synthetic |
| Angular error, all test patches (synthetic) | 4.75° mean / 3.30° median | eval.json of the 2026-09-30 run — 10,000 rendered test patches from 1,250 new identities, blinks and mostly hidden irises included |
| Angular error, open-eye patches (synthetic) | 3.69° mean / 3.00° median | eval.json of the 2026-09-30 run — 8,597 patches whose openness label is at least 0.3; the closest to the frames the demo uses |
| Angular error, gaze-valid patches (synthetic) | 3.49° mean / 2.91° median | eval.json of the 2026-09-30 run — 8,077 patches with at least 30 % of the iris visible; the most favourable set, and the one the checkpoint was chosen on |
| Constant-prediction baseline, all test patches (synthetic) | 17.69° mean | eval.json of the 2026-09-30 run — the error of always answering the mean gaze |
| Blink accuracy (synthetic) | 0.979 | eval.json of the 2026-09-30 run — openness below 0.3 counted as a blink; recall 0.91, precision 0.94; 14 % of test patches are blinks |
| Openness error (synthetic) | 0.031 MAE | eval.json of the 2026-09-30 run — all 10,000 test patches |
| Pupil-centre error (synthetic) | 0.44 px median | eval.json of the 2026-09-30 run — gaze-valid patches, in patch pixels |
| On-screen error per frame, head still (simulated users) | 1.16 cm | eval.json of the 2026-09-30 run — 16 simulated users at 60 cm after a nine-dot calibration; 1.25 cm with correlated landmark noise |
| On-screen error per frame, head moved (simulated users) | 1.55 cm | eval.json of the 2026-09-30 run — the same users after the head moved; 1.67 cm with correlated landmark noise |
| ONNX export against PyTorch | 1.8e-06 max difference | export.json of the 2026-09-30 run — tolerance 1e-4, checked with onnxruntime 1.19.2 on CPU |
| API forward pass against PyTorch | 8.6e-07 max difference | the API’s own check on 2026-10-01 — the eight procedural test patches of the trainer’s parity fixture, tolerance 1e-4 |
Every figure is synthetic. The test split is new identities from the same renderer and the same distribution as the training data: it measures how well the model interpolates within what the renderer can draw, not how it behaves on anything else. Accuracy on real faces has not been measured, and neither has accuracy on eyes drawn outside the training distribution. Real webcam eyes are expected to be harder; by how much is not known.
The three angular figures are the same model on three sets of test patches. All patches is every test patch, blinks included. Open-eye patches leaves out the blinks and is the closest to what the demo uses, because the page drops a frame only when its blink detector calls an eye closed. Gaze-valid patches keeps only eyes with at least 30 % of the iris visible; it is the most favourable set and the one the checkpoint was selected on. On patches it was trained on the model is 0.5° to 0.6° better, so part of the test error is a gap even inside the same distribution.
| Slice, gaze-valid patches only | Mean angular error |
|---|---|
| Eye smallest in the frame (0.4–0.6 camera px per patch px) | 4.64° |
| Eye largest in the frame (1.3–1.8 camera px per patch px) | 2.73° |
| With glasses | 3.99° |
| Without glasses | 3.32° |
| Half-closed eyes (openness label below 0.5) | 4.53° |
| Gaze more than 30° off the camera | 4.17° |
The on-screen figures come from simulated users, not people: 16 generated identities sitting 60 cm from a 30.2 × 19.6 cm screen, calibrated on nine targets and tested on 25. The simulation renders eye patches directly and shares only the ray and calibration arithmetic with the demo; it does not run the face-landmark model or the page’s image crop, and its landmark noise is assumed rather than measured. Its figures are an optimistic bound. Averaged over the ten frames of a fixation the error is 0.61 cm with the head still and 1.15 cm after it moved, which holds only if the noise between frames is independent; with correlated noise those become 0.80 cm and 1.33 cm.
Known gaps. No real person, real camera frame or public dataset was used. Nothing outside the training distribution was measured. The frame rate and the per-frame time on real hardware have not been recorded. Only Chromium has been exercised by the automated tests.
API #
| Endpoint | Auth | Body limit | Description |
|---|---|---|---|
POST /v1/gaze | Bearer preview key (fln_) or session token (fls_) | 128 KB; two patches of 2,160 values | One or two 36 × 60 eye patches (+ optional head pose and mapping) → pitch, yaw, openness and blink per eye, the seven features, and a screen point when a mapping is sent. Runs in the API process. |
POST /v1/gaze/photo | Bearer preview key (fln_) or session token (fls_) | 24 MB; decoded image ≤ 6,000,000 bytes | A JPEG, PNG or WebP of one face (+ optional mapping) → the same per-eye answers for the largest face, its head pose, the seven features, and a screen point when a mapping is sent. The image is processed in memory and not kept. |
POST /v1/gaze/calibrate | Bearer preview key (fln_) or session token (fls_) | 512 KB; 5–2,000 samples | Calibration samples (features + the screen point looked at) and the screen size → a fitted mapping and its leave-one-target-out error. Nothing is stored: the caller keeps the mapping and sends it with later calls. |
Waymark Gaze is public on the Falcon API at POST /v1/gaze, POST /v1/gaze/photo and POST /v1/gaze/calibrate. All three count against the spatial preview quota bucket, shared with /v1/spatial, /v1/move and the airspace routes; the patch route’s body is limited to 128 KB. The base URL, envelope, authentication and retry guidance are documented in API conventions. The model also runs without the API, in the browser: that path is described under In the browser below.
Nothing is stored between calls. No patch, image, sample or mapping is written to disk, logged or kept in memory after the response is sent, and none is used for training. A calibration exists only in the response that returns it and in the requests that send it back.
Eye patches: /v1/gaze #
POST /v1/gaze HTTP/1.1
Authorization: Bearer fln_…
Content-Type: application/json{
"left": "m5ueoaaqrrK0uLq7vL2/wMHCwsLDw8PDxMLBwcLBwL+/v76+vr27ube1s7Ct…",
"right": "raytrqyurbK0tba4ubq6urq6u7q5ubi4ubi2tra2trS1tLSzsrKxsbGwsLGw…",
"head": {"yaw": 0.02, "pitch": -0.05, "roll": 0.01}
}{
"ok": true,
"engine": "waymark-gaze",
"model": "waymark-gaze",
"eyes": {
"left": {"pitch": -0.182708, "yaw": -0.287324, "pitch_deg": -10.468397, "yaw_deg": -16.462453, "openness": 0.934952, "blink": false},
"right": {"pitch": 0.0664, "yaw": 0.122169, "pitch_deg": 3.80444, "yaw_deg": 6.999768, "openness": 0.999328, "blink": false}
},
"eyes_used": 2,
"features": [-0.182708, -0.287324, 0.0664, 0.122169, 0.02, -0.05, 0.01],
"feature_names": ["left_pitch", "left_yaw", "right_pitch", "right_yaw", "head_yaw", "head_pitch", "head_roll"]
}The request takes left and right, the patches of the subject’s own left and right eye, each either a base64 string of exactly 2,160 raw bytes or an array of 2,160 numbers from 0 to 255 (fractions are kept), and at least one of them: send null, or leave the field out, for an eye that is closed, hidden or out of the frame. head is optional: the head’s yaw, pitch and roll in radians from the host’s own landmark model, used only as calibration features. mapping is optional: a mapping returned by /v1/gaze/calibrate.
The response gives, per eye, pitch and yaw in radians and in degrees, openness from 0 to 1, and blink, true when openness is below 0.3. Pitch above zero is up. Yaw above zero is toward the +x side of the camera image, the subject’s left, for both eyes: the route has already negated the left eye’s yaw, which the network reports mirrored, so the two eyes can be compared or averaged directly. Both at zero mean the eye looks straight at the camera. features is the vector the calibration works on, in the order of feature_names: the head terms are 0 when no head was sent, and when only one eye was sent its two angles fill both eyes’ places and eyes_used is 1. The values above come from two unrelated test patches and are illustrative of the shape.
How to produce a patch. The geometry is part of the contract; a patch cut any other way gives angles that mean something else.
- Start from the camera frame as the camera delivers it, not a mirrored selfie preview, with pixel coordinates counted from the top-left corner: pixel k covers k to k + 1, and its value sits at its centre, k + 0.5.
- Find the two corners of the eye in that frame: the outer corner (toward the ear) and the inner corner (toward the nose). With MediaPipe Face Landmarker these are landmarks 33 and 133 for the subject’s right eye and 263 and 362 for the left. Right and left are the subject’s own; in an unmirrored frame the subject’s right eye is on the image’s left.
- Let m be the midpoint of the two corners, d their distance in camera pixels, and u the unit vector from the outer corner to the inner one. Let v be u turned a quarter turn so that it points down the image: (−u.y, u.x) for the right eye, (u.y, −u.x) for the left.
- The patch is 36 rows by 60 columns. The value at row i, column j is the image sampled at m + (d / 36) · ((j + 0.5 − 30) · u + (i + 0.5 − 18) · v), bilinearly, with edge pixels repeated outside the frame. So the midpoint lands on the patch centre (30, 18), the outer corner on (12, 18), the inner corner on (48, 18), and the nose side is at +x for both eyes: the left eye’s patch is a mirror image, the same as cropping it like a right eye and flipping the result left to right.
- Each value is the gray level 0.299 R + 0.587 G + 0.114 B of the 8-bit frame, from 0 to 255. Do not scale to 0–1, standardise or equalise: the network standardises each patch itself.
- Write the rows top to bottom, each left to right: 2,160 values. Round them to bytes and base64-encode them (2,880 characters), or send the numbers as an array.
The eye should span at least half a camera pixel per patch pixel, a corner distance of about 18 camera pixels or more; above about 65, shrink the eye region before sampling so that fine detail is averaged rather than skipped, as the demo does.
A photo: /v1/gaze/photo #
{ "image": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQ…" }{
"ok": true,
"engine": "waymark-gaze",
"model": "waymark-gaze",
"route": "/v1/gaze/photo",
"face": true,
"image": {"width": 1280, "height": 720},
"eyes": {
"left": {"pitch": -0.05, "yaw": 0.11, "pitch_deg": -2.864789, "yaw_deg": 6.302536, "openness": 0.93, "blink": false, "center": {"x": 742.4, "y": 301.5}},
"right": {"pitch": -0.04, "yaw": 0.09, "pitch_deg": -2.291831, "yaw_deg": 5.15662, "openness": 0.91, "blink": false, "center": {"x": 537.8, "y": 303.1}}
},
"head": {"yaw_deg": 1.1, "pitch_deg": -1.7, "roll_deg": 0.6, "distance_cm": 58.2},
"features": [-0.05, 0.11, -0.04, 0.09, 0.0192, -0.0297, 0.0105],
"feature_names": ["left_pitch", "left_yaw", "right_pitch", "right_yaw", "head_yaw", "head_pitch", "head_roll"]
}The body takes image (alias image_base64): base64, or a data: URL, of a JPEG, PNG or WebP of at most 6,000,000 bytes once decoded, and an optional mapping. A separate service finds the face, locates the eye corners with a face-landmark model, cuts the two patches by the recipe above and runs the same network. When the picture holds several faces, only the largest is read. eyes carries each eye’s answer with its center in image pixels, head the estimated head pose and distance, and features the same seven numbers the patch route returns. When no face is found the route answers 200 with face: false and eyes, head and features set to null; that call still counts one unit, because the detector ran. The values above are illustrative of the shape.
What happens to the image: it is decoded and read in memory for that one request, and it is not stored, not logged and not used for training; the same holds for the patches of the patch route. The patch route receives two crops of 36 by 60 gray pixels and never a face; the photo route receives a picture of a face. A host that calls either one must have the consent of the person in the picture and must not use the routes to watch anyone without their knowledge.
Calibration: /v1/gaze/calibrate #
The model’s angles are not a screen position. A mapping must be fitted per person and per sitting, from samples the host collects while the person looks at known points:
- Show nine targets on a 3 × 3 grid, at 6 %, 50 % and 94 % of the screen width and 8 %, 50 % and 92 % of its height, one at a time, the centre first.
- While the person looks at each target, call
/v1/gaze(or/v1/gaze/photo) for ten frames or more and keep each response’sfeatureswith the target’s position in pixels. Leave out frames in which either eye reportsblink. - Send the samples, with the size of the screen in the same pixels, to
/v1/gaze/calibrate. It returns amappingand afit. - Send that
mappingwith later/v1/gazeor/v1/gaze/photocalls. Each response then carriesscreen. - Calibrate again when the person moves away, the window or the screen changes, or the session ends. The API keeps no mapping: the host holds it, and it is gone when the host discards it.
{
"samples": [
{"features": [0.1563, -0.1687, 0.1282, -0.2607, 0.0026, -0.0215, 0.0006], "screen": {"x": 86, "y": 72}},
{"features": [0.1527, -0.1696, 0.1218, -0.2666, 0.0086, -0.0156, 0.0018], "screen": {"x": 86, "y": 72}},
{"features": [0.1542, 0.0533, 0.122, -0.0329, 0.0078, -0.0299, -0.0007], "screen": {"x": 720, "y": 72}}
],
"screen": {"width": 1440, "height": 900}
}{
"ok": true,
"engine": "waymark-gaze",
"mapping": {
"kind": "ridge",
"version": 1,
"feature_names": ["left_pitch", "left_yaw", "right_pitch", "right_yaw", "head_yaw", "head_pitch", "head_roll"],
"mean": [0.0197, 0.0501, -0.0102, -0.0402, 0.0099, -0.0197, 0.0001],
"scale": [0.1102, 0.1798, 0.1103, 0.1798, 0.1, 0.1, 0.1],
"weights": [[-0.8, 263.6, -0.6, 253.3, -6.8, 5.9, -12.9, 720], [-156, -0.2, -152.2, -0.3, -2.3, 6.1, -0.8, 450]],
"lambda": 0.003,
"screen": {"width": 1440, "height": 900}
},
"fit": {"samples": 108, "targets": 9, "loo_error_px": 10.244042, "loo_error_fraction_of_diagonal": 0.006033}
}The request above is abbreviated to three of its 108 samples, and the numbers in the mapping are rounded for the page; both come from a simulated sitter, not a person. samples takes 5 to 2,000 entries covering at least three and at most 100 distinct screen points; nine targets with ten or more frames each is the recommended set. Samples that share a screen point are treated as one target, so every sample of a target must carry exactly the same point. The fit is a ridge regression on the standardised features with an unpenalised intercept, the same arithmetic as the demo: the strength lambda is the one of nine candidates with the lowest error when each target in turn is left out, and a feature that barely moved during calibration (the head, held still) is given a minimum scale so that it cannot throw the estimate far when it moves later. fit.loo_error_px is that leave-one-target-out error, the mean distance per sample in pixels, and loo_error_fraction_of_diagonal the same divided by the screen diagonal; both are null when too few samples remain for any fold. They describe that sitting and nothing more.
In a mapping, mean and scale standardise the seven features, and each row of weights holds seven weights and then the intercept: the first row gives x, the second y, in the caller’s pixels. A response to a call that carried a mapping adds:
{ "screen": {"x": 741.295124, "y": 429.232706, "on_screen": true} }The point is never clamped. A gaze that maps outside the screen is returned as computed, with on_screen false, so the host can tell a glance away from a point on the edge. On the photo route screen is present only when a face was found. Fit and apply a mapping with features from the same route: a mapping fitted on patch-route features belongs with patch-route calls.
Metering and errors #
Each call to any of the three routes counts one unit in the spatial bucket once it has reached the model, or, for calibration, once the fit has run: calibration is arithmetic only and is counted the same way to keep one rule. A photo in which no face is found counts. Validation errors, 401, 402, 429 and every 503 do not count. A nine-target calibration at ten frames each is therefore ninety calls and one more for the fit.
| Status | Code | Meaning |
|---|---|---|
| 400 | bad_json | The body is not valid JSON, is not an object, or is over the route’s body limit. |
| 400 | patches_required | /v1/gaze was called with neither left nor right. |
| 400 | bad_patch | A patch is not base64 or an array, does not hold exactly 2,160 values, or holds a value outside 0–255; eye names the eye and detail says why. |
| 400 | bad_head | head is not an object with finite yaw, pitch and roll. |
| 400 | bad_mapping | mapping is not a mapping this version returns: another kind or version, other feature names, a wrong number of values, or a value that is not a finite number; detail names the field. |
| 400 | image_required | /v1/gaze/photo was called without an image. |
| 400 | bad_image | The image is not base64 of a JPEG, PNG or WebP, or could not be decoded. |
| 400 | image_too_large | The decoded image is over 6,000,000 bytes; maxBytes is included. |
| 400 | samples_required | /v1/gaze/calibrate was called with fewer than five samples. |
| 400 | bad_samples | A sample is malformed (detail names it), there are more than 2,000, or they cover fewer than three or more than 100 distinct screen points. |
| 400 | bad_screen | screen is not {width, height} with both above zero. |
| 401 | invalid_credentials | Missing, malformed or revoked bearer credential. |
| 402 | payment_required | The credential’s owner is not in good standing with the preview. |
| 429 | quota_exceeded | The spatial bucket for the current UTC calendar month is exhausted; kind is spatial. |
| 502 | gaze_failed | The route failed after a valid request; safe to retry once. |
| 503 | gaze_warming | /v1/gaze/photo: the photo service is starting after a quiet period; retry after the Retry-After interval (20 s, or 10 s while no instance is free). |
| 503 | gaze_unavailable | The weights are not loaded (/v1/gaze) or no photo service is configured (/v1/gaze/photo); retry once after 2.5 s. |
There is no keyless route for this model. The public playground does not offer it, because pictures of faces and crops of eyes are not accepted from anonymous visitors; the portal’s sandbox, for signed-in organisation users, runs the three routes.
In the browser #
The weights are also a static file, /gaze/model/gaze.onnx (1,686,037 bytes, ONNX opset 17), which the demo loads and runs in the visitor’s browser with no key, no quota and no request to any server. It has one input, patches, float32 of shape [N, 1, 36, 60], and one output, gaze, float32 of shape [N, 5], as described under Inputs & outputs. The demo loads it with ONNX Runtime Web on the WebAssembly backend:
import * as ort from "/gaze/vendor/onnxruntime-web-1.30.0/ort.wasm.min.mjs";
ort.env.wasm.wasmPaths = "/gaze/vendor/onnxruntime-web-1.30.0/";
ort.env.wasm.numThreads = 1;
const model = await fetch("/gaze/model/gaze.onnx").then((r) => r.arrayBuffer());
const session = await ort.InferenceSession.create(new Uint8Array(model), { executionProviders: ["wasm"] });
// gray: Float32Array of 2 × 36 × 60 values, right eye patch first, then the mirrored left eye patch
const { gaze } = await session.run({ patches: new ort.Tensor("float32", gray, [2, 1, 36, 60]) });/gaze/model/gaze.json is the trainer’s own record of the export: the patch geometry, the tensor names, the landmark indices and the synthetic metrics. Its version and status fields match this page; name is the trainer’s internal name for the network. The raw tensor reports a left eye’s yaw mirrored; a host that runs the file itself negates it, as the API does.
The demo calibrates the same way as the routes above, on its own: nine targets, about twenty frames each, several mappings fitted and the one with the lowest leave-one-dot-out error kept, a One Euro filter on the mapped point, and a hold while either eye is closed. It keeps the calibration in the tab’s memory only and turns the camera off when the tab is hidden. The page reports its own leave-one-dot-out error after calibration and can measure the error on nine new dots; those numbers describe that sitting on that computer and nothing more.
Runtime & deployment #
| Kind | In the browser and in-process |
|---|---|
| Resident | In the Falcon API process, a 1.7 MB bundle of the weights loaded on first use; in the browser, the ONNX file held by ONNX Runtime Web until the tab closes |
| Serving | The Falcon API routes /v1/gaze (in-process), /v1/gaze/calibrate (arithmetic only) and /v1/gaze/photo (a private CPU service); static files under /gaze/ for the browser demo |
| Cold start | None for the patch route beyond the API’s own start; the photo service sleeps when idle and the call that wakes it may answer 503 gaze_warming with Retry-After 20 |
| Concurrency | One patch at a time in the API process; in the browser, one camera frame at a time with both eyes in one batch |
| Timeout | 20 s on the call to the photo service |
Waymark Gaze runs in two places on the same weights.
On the Falcon API. The network runs in-process: the API holds a 1.7 MB bundle of the weights and computes the forward pass itself, with no model service behind /v1/gaze. Against the trainer’s reference outputs the in-process pass differs by at most 8.6e-07 on eight test patches, and one eye takes about 7 ms on a development laptop; the time on the API’s own hardware has not been measured. /v1/gaze/calibrate is arithmetic on the request and loads nothing. /v1/gaze/photo is the exception: the face and its landmarks are found by a separate CPU service, private to the Falcon API, that sleeps when idle. The call that wakes it waits while it starts and may answer 503 gaze_warming with a Retry-After of 20 s; how long a start takes has not been measured yet.
- Metering. One spatial unit per call that reaches the model, on all three routes; a photo with no face counts, a
503does not. - Retention. None. Patches, images, samples and mappings are held for the one request and are not stored, logged or used for training.
- Retries. Retry
503 gaze_warmingafterRetry-After; retry503 gaze_unavailableand502 gaze_failedonce after 2.5 s. - Playground. Not offered without a key. The portal’s sandbox runs the three routes for a signed-in user, with a photo limited to its 256 KB envelope; the sandbox page has no ready-made example for them.
In the browser. The demo at /gaze is a static page with its own scripts, the gaze model, the face-landmark model and two vendored libraries, all served from this site; it loads nothing from any other address, and its content security policy forbids it. After the files have loaded the page makes no network request, and it does not call the routes above.
- Camera. The browser asks for permission; the picture is read inside the tab and never uploaded, recorded or stored. The camera is switched off when the tab is hidden or the page is left.
- Calibration. Nine dots; the result is kept in the tab’s memory and gone on reload.
- Blinks. The demo takes blinks from the landmark model’s eye-blink scores by default and holds the dot while an eye is closed.
- Fallback. If the gaze model cannot be loaded, the page says so and falls back to mappings that use the landmarks alone.
What the visitor’s computer must provide for the demo:
- A laptop or desktop with a webcam above the screen, ideally 1280 × 720; a lower resolution works with a warning.
- Light on the face, not behind it, and a seat about an arm’s length from the screen.
- A current browser with WebAssembly and JavaScript modules, on a secure address. Chromium-based browsers have been tested; Safari and Firefox have not.
| Setting | Value |
|---|---|
| Patch route | /v1/gaze, in the API process; body limit 128 KB |
| Photo route | /v1/gaze/photo, a private CPU service that scales to zero; image ≤ 6,000,000 bytes; 503 gaze_warming, Retry-After 20 s while it starts |
| Calibration route | /v1/gaze/calibrate, arithmetic only; 5–2,000 samples |
| Quota bucket | spatial |
| Stored between calls | nothing |
| Browser demo | /gaze on this site; /gaze/model/gaze.onnx, 1.7 MB |
| Browser runtime | ONNX Runtime Web 1.30.0, WebAssembly, one thread |
| Landmarks | MediaPipe Face Landmarker, Tasks Vision 0.10.35 |
| First download of the demo | about 31 MB before compression |
Limits & safety #
- Accuracy on real faces has not been measured. Every metric on this page is synthetic, and real eyes are expected to be harder by an unknown amount.
- It needs calibration. Without the nine dots there is no screen point, and the calibration holds only while the head stays roughly where it was. The API keeps no calibration: a host that loses the mapping must calibrate again.
- It needs a laptop-style webcam near the screen and decent light. A dark, backlit or low-resolution picture gives worse landmarks and worse patches; the demo warns about low light and a missing face.
- Glasses, half-closed eyes and large gaze angles raise the error even on synthetic eyes (3.99°, 4.53° and 4.17° against 3.49° overall on gaze-valid patches).
- It tracks where the person using the page is looking, with their consent, and identifies nobody. It is not for monitoring, proctoring or any covert use.
- It is not a medical device, not a safety system and not an assistive technology anyone should depend on.
- The model can have learned artefacts of its renderer: skin without pores or wrinkles, flat glasses, one lid model, mirrored left eyes.
- On the API, the patch route trusts the patch it is given. A patch cut with another corner distance, centre, mirroring or gray conversion is answered without an error, with angles that mean something else.
- The photo route reads one face, the largest, from one still picture of at most 6,000,000 bytes. How well its landmark step places the eye corners on real photographs has not been measured, and neither has the route’s end-to-end error.
- Every frame is one call and one spatial unit. A calibration is about ninety calls, and following a gaze frame by frame uses the allowance quickly; the browser demo, which calls nothing, is the path for continuous tracking.
- A mapping fitted on one route’s features, one person, one screen and one seat does not carry over to another.
Out of scope: clinical measurement, driver or operator monitoring, attention scoring, and any decision about a person taken from where they looked. A host that embeds the model must ask for the camera plainly, keep the picture in the browser unless the person has agreed to its being sent, and offer another way to do whatever the gaze controls. A host that calls the API must have the consent of the person whose eyes or face it sends, and must not send pictures taken without that person’s knowledge.
Fixed weights per version; the model does not learn from requests. Nothing the camera sees changes them, and nothing is collected from which they could be retrained: the demo sends nothing, and the API routes keep nothing.
Versions #
| Version | Date | Status | Note |
|---|---|---|---|
| 1.0.0 | Released | First released version: the epoch-14 checkpoint of the 2026-09-30 run, served with the demo at /gaze and, from the same day, on the API. Evaluated on synthetic eyes only. |
Compatibility. A major version bump means the contract changes — the patch geometry, the tensor names or shapes, the meaning or units of an output — and a host must be updated to use it. A minor bump is a retrain with the same contract: angles for the same patch may differ, and a host should calibrate again. A patch bump changes only metadata or the page around the model.
Current weights: the epoch-14 checkpoint of the 2026-09-30 run, exported as gaze.onnx, 1,686,037 bytes, opset 17, float32, 420,221 parameters, sha256 beginning 3e77b9fc. The file is served with the demo at /gaze/model/gaze.onnx; the API runs the same weights with batch normalisation folded into the convolutions, which changes no output beyond rounding. A change to the shape of a mapping is announced by its version field: a route refuses a mapping of another version with 400 bad_mapping, and the host calibrates again.
The weights and the page’s code belong to Ducky Software and are served so that the demo can run; other use needs permission. The demo also loads MediaPipe Tasks Vision and its Face Landmarker model (Apache-2.0) and ONNX Runtime Web (MIT).