Accuracy changelog
Every release of the recogniser publishes two numbers here, measured on the curated set of real stream frames that the model was never trained on, at the default confidence threshold:
- Top-1 accuracy: how often the card shown is the right card.
- False-positive rate: how often something is shown when it should not be (an empty frame, a card back, or the wrong card presented as certain).
Below-threshold results are never presented as certain: they are suppressed or shown as “tap to confirm”. The gate for public beta is at least 90% top-1 and at most 2% false positives.
2026-09-29 — Pre-release measurement (internal build, not released)
Test set: 61 real stream frames from two clips the model was never trained on (a demo recording
and one live-auction recording), English Pokémon cards, at the default thresholds
(v3-calibrated-2026-09-29).
- Right card name at the top of the list: 73.9% of frames (exact printing 52.9%).
- Shown as certain (“full” band): 16.4% of frames, with no wrong name among them.
- False positives: 0 frames where an empty frame or a card back was shown as a card.
- On a clip whose localisation the model trained on, the right name is at the top on 98.7% of frames and 82.4% are shown as certain, again with no wrong name.
What changed: the artwork model was fine-tuned on stream-like crops, cards named by text join the candidate list, and the confidence bands were calibrated on 195 named real cards. Look-alike printings score alike, so “certain” means the name is certain; the app offers the tied printings.
2026-09-29 — Format of this page
No public release yet. From the first release onward, each entry lists the release, the test set (frame count and clips), top-1 accuracy, false-positive rate, and what changed in the model or its thresholds.