Absolute task quality, paired differences, exact reader input tokens, estimated API cost, latency, and predefined subgroup results from the authenticated three-trial benchmark.
Matched reader policy: GPT-6 Luna, low reasoning, 16,384 maximum output tokens, high image detail, no storage, and high service tier. The screenshot arm uses 20 frames at a 768-pixel long edge sampled nearest each equal-presentation-duration-bin midpoint.
Näky accuracy82.61%2,865 / 3,468 tasks
20-screenshot accuracy88.21%3,059 / 3,468 tasks
Paired difference-5.59 pp95% CI -6.82 pp to -4.38 pp
Input-token saving50.91%exact unique body counts
Quality
Arm
Correct
Task accuracy
Scenario macro
Pooled-call accuracy
Invalid tasks
Näky structured output
2,865 / 3,468
82.61%
81.81%
82.36%
38
20 screenshots
3,059 / 3,468
88.21%
87.77%
87.75%
10
Paired outcomes: 131 Näky wins, 325 losses, 2,994 ties, 18 unresolved. Confidence intervals use the score’s predefined scenario-stratified media-cluster bootstrap.
Tokens, cost, and latency
Arm
Unique input tokens
3-trial planned tokens
Observed input tokens
Estimated cost
p50 ms
p95 ms
Näky structured output
12,413,476
37,240,428
37,235,285
$5.2647
954
5,326
20 screenshots
25,289,408
75,868,224
75,868,224
$9.8410
1,938
4,716
Cost is the score’s cache-aware estimate under its frozen price table, not a provider invoice. Latency covers authenticated responses in the complete evaluable cohort.
Product processing and evidence size
Metric
Value
Evaluable recordings
1,819
Normalized video
27,124.27 s
ScreenEvents
1,378,711
Record wall-time sum
31,011.66 s
Record wall-time mean
17.05 s
Record wall-time p50
11.01 s
Record wall-time p95
52.75 s
Video seconds per record wall second
0.875
Peak record RSS
679,280 KiB (663.36 MiB)
Näky state text
13,639,728 (13.01 MiB)
Screenshot PNG evidence
5,941,847,499 (5.53 GiB)
Record wall-time values are per-record compute work, not parallel batch makespan. Evidence sizes cover the complete 1,819-record evaluable corpus and are separate from reasoning-model input tokens.
Complete-case sensitivity
exclude the task pair affected by ambiguous transport. Excluded missing tasks: 1.
Arm
Accuracy
Correct
Tasks
Näky structured output
82.61%
2864
3467
20 screenshots
88.20%
3058
3467
Complete-case paired delta: -5.60 pp.
Interpretation
terminal-tolerant sensitivity with one authenticated ambiguous transport; no noninferiority or equivalence claim
Reproducibility scope
The 1,819-record equivalence run executed the authenticated unstripped binary. The release executable is its twice-reproduced GNU strip --strip-all derivative; the stripped file was not rerun across the full corpus.
Strict evaluated equivalence uses an authenticated mixed-host composition: 1,815 records from the base run and four AVX2/FMA/no-AVX-512 replacements. The base-to-replacement diagnostic differed in 4 of 1,819 event streams and 2 of 1,819 screen bodies. The receipt does not bind the base host ISA, so cross-CPU byte invariance is not claimed.