I build Pooka, a Pokémon card collection app for iOS, Android and web. Until version 2.0, scanning a card meant uploading a photo: the server turned it into a CLIP embedding and ran a vector search over every card image. It worked, but every scan waited on the network and cost server time.
For 2.0 the whole pipeline runs on the phone, in Flutter. This is what that took, including the parts that went wrong.
The pipeline
- A corner detector finds every card in the camera frame and its four corners, so a binder page with nine cards becomes nine upright crops.
- An embedding model turns each crop into a 256-dimensional vector.
- That vector is matched against an index of about 47,000 English and Japanese cards that ships inside the app.
- An accept rule decides whether to add the card straight away or show a short list to pick from.
No OCR anywhere. Card text is small, often Japanese, and frequently behind a glossy sleeve; the artwork carries far more signal.
Training data that doesn't exist
The hardest part wasn't the model, it was data. There is no labelled set of real phone photos covering 47,000 cards, and there never will be.
So I wrote a simulator that turns clean catalog scans into fake phone photos: perspective, tables and playmats as backgrounds, lighting changes, shadows, glare and blur, single cards, overlapping piles and binder pages. Every training image comes out of it, with the exact card and its corners as free labels.
To check the model against something closer to reality, I rendered cards in Blender: in sleeves, in toploaders, in binder pockets, under different lights. That gave 605 evaluation images with 992 cards whose exact identity is known. A small set of hand-labelled real photos sits on top of that.
Honest caveat: the evaluation is mostly synthetic. Only 39 of the test cards are real photos, so one real photo moves that sub-score by 2.6 points.
Picking a backbone
I trained three candidates under the same 60-minute budget: DINOv2-small, a Perception Encoder model and MobileNetV4. DINOv2-small and the PE model tied on accuracy; MobileNetV4 fell far behind on the real photos. DINOv2-small won because it was 1.6x faster and smaller after export.
Fine-tuned, it gets the artwork right 97% of the time and the exact print 91% of the time across the evaluation set.
Getting it onto the phone
The model is exported from PyTorch to LiteRT (the runtime formerly known as TensorFlow Lite) and runs through flutter_litert.
- fp32: 87.3 MB, reproduces PyTorch exactly.
- int8 weights: 23.2 MB, no measurable accuracy loss. This is what ships.
Two things cost me time here:
Dynamic int8 produced wrong embeddings. In Python, the dynamic-range quantized model scored the same as the others. In the app's runtime it ran without any error and returned vectors that matched nothing: 0 out of 12 test cards on the iOS simulator. The weight-only int8 model matched Python to five decimals and got 12 out of 12. Lesson: test the quantized model in the runtime you actually ship, not in the one you exported with.
XNNPACK was not on by default for this model. Without explicitly adding the XNNPACK delegate, one inference took about 1,060 ms on an M-series Mac. With XNNPackDelegate(numThreads: 4) it dropped to about 150 ms.
The corner detector is a separate small model I trained for this. It replaced an AGPL-licensed YOLO model the app used before.
A 15 MB index for 47,000 cards
The card index stores one int8 vector per card plus a per-row scale, followed by the card's id, name, set and number. For 46,952 cards that is 14.9 MB, and the int8 index costs 0.2 points of top-1 accuracy compared to the float one.
New sets keep coming out, so the app doesn't wait for an app update: it downloads a base file plus small delta files with only the new cards, and swaps them into the running scanner without a restart.
The reprint problem
This is the part I underestimated. About 13,700 groups of cards share the same artwork: reprints, promos, the same illustration in another set. A model can be completely sure about the picture and still pick the wrong card.
So the app doesn't trust the top score. It only adds a card on its own when the best match beats the best candidate with different artwork by a clear margin. Otherwise it shows the few likely prints and lets you pick.
That margin matters more than the absolute score: real photos score lower similarities than renders, but the gap to the nearest different artwork stays readable. On the evaluation set, that rule accepts about 95% of scans automatically, and 99.6% of those have the right artwork.
What changed for users
- Recognition works without a connection, and no photo leaves the phone.
- A whole binder page, nine cards, can be scanned in one photo.
- Japanese cards are recognised too, not just English ones.
The app is at pooka.app. I'm happy to go deeper into the simulator, the index format or the accept rule in the comments.