Screen perception for reasoning models

Screen recordings in.
Structured state out.

Näky turns a screen recording into a chronological text representation: what text was observed, where it was, how it persisted, and when the screen changed.

Loading recording, one shared playhead

See the pixels and the state together

What the model receives

Loading the public recording gallery…

A text model receives the exact state stream on the right. The recording remains visible here only as a reference.

01

What the screen shows

Complete Näky overlay
Retained element Recent change Observed activity
Loading authentic output…
0:00 / 0:13
02

What the model receives

EventsRaw state
@ timeS screen= + ~ > − set · add · change · move · remove! observed activity

What survives the conversion

A screen becomes a small, evolving document.

Identity + geometry

Text stays anchored

Stable IDs associate OCR observations across time. Boxes keep each observation tied to its screen position.

Time-ordered deltas

Change stays explicit

Additions, edits, moves, removals, and observed visual activity update the state instead of repeating every frame.

Honest boundaries

Pixels still know more

Appearance, icons, color, fine visual state, and anything OCR misses may remain available only in pixels.

Read the overlay literally

Observations, not inferred actions.

OCR-derived text and boxes are observations and can be wrong. Geometry and stable IDs show what Näky associated over time.

The overlay draws every element retained at the playhead, every recent structural change, and every displayed activity region. Anything extraction did not retain remains visible only in the pixels.

Visual activity means pixels changed in a region during an interval. It is not a click, scroll, focus, selection, action, intent, or cause label.

The released comparison found higher quality from screenshots. Näky scored 82.61% versus 88.21% for screenshots across 3,468 tasks, while using 50.91% fewer logical three-trial input tokens and an estimated 46.50% lower cache-aware reader cost. Read the full matched comparison.

Demo source

Select a recording to see its source. Third-party names, marks, and interface content belong to their owners. Full attribution and transformation notes.