Guandan Planner
A mobile app that reads a full 27-card Guandan hand through the phone camera and suggests how to arrange it into playable groups. Two steps: scan, then results — cards lock in one by one as evidence accumulates across frames, a misread is fixed by tapping the card, and everything runs on the device.
Motivation
Guandan is a Chinese climbing card game played with two full decks; each player holds 27 cards, and every card can legally appear twice. Arranging such a hand well is a real skill — and recognizing one is a genuinely hard vision problem: heavy occlusion in a fan, duplicate faces that defeat the usual "already seen it" sanity check, and no public dataset anywhere. That combination — a small, concrete product wrapped around several unsolved-for-this-domain technical problems — is why the project exists.
Architecture
Three subsystems, each built from scratch:
- On-device recognition — a 109-class YOLO (54 card faces × two corner roles, plus the hand silhouette), followed by geometric corner pairing, multi-frame log-odds evidence tracking, and a ledger that admits a duplicate card only when both copies are visible in the same frame.
- Arrangement engine — a pure-Dart exact-cover search with fail-first pivoting and a playability scorer; three strategies come from a single search pass, in a background isolate.
- Data and training infrastructure — a local studio that turns photos of a physical deck into rectified assets, renders labeled scenes through one shared pinhole camera, manages datasets with content-addressed fingerprints, dispatches training to rented GPUs, and evaluates against ground-truthed real photos.
Key decisions
| Decision | Choice | Rationale |
|---|---|---|
| Where inference runs | On-device (TFLite / Core ML) | The scan loop needs 8+ fps sustained; photos stay on the phone |
| What to detect | Corner indices, not whole cards | Fanned cards overlap at IoU ≈ 0.8 — whole-card boxes collapse under NMS |
| Label space | 109 classes (54 faces × 2 corner roles + hand) | Corner roles make card counting decidable with duplicates legal |
| Training data | 100% synthetic, rendered | No dataset exists; rendered labels are exact and regenerable overnight |
| Acceptance metric | Whole-hand exact rate on real photos | mAP once endorsed a scheme that read zero real hands correctly |
| Arrangement objective | Playability over fewest hands | Minimizing hand count buries aces inside straights |
Evaluation
Models are judged on ground-truthed real photographs — 82 images, 2,111 cards, annotated as card multisets — with hash-checked isolation from training data. The primary metric is the fraction of hands read exactly right: with 27 cards a scan, per-card accuracy flatters, and the exponent is merciless.
| Measure | Value |
|---|---|
| Per-card accuracy | 98.7% |
| Whole-hand exact rate | 69.5% |
| Model size (Android / iOS) | 3.0 MB / 2.4 MB |
| Training data | 130k rendered images, zero hand labels |
| Training span | 21 runs over 19 days |
Iteration highlights
- A "cleaner" 54-class labeling scheme posted the best synthetic mAP of the project and read zero real hands correctly — rolled back within a day, and the reason the acceptance metric is what it is.
- The single largest accuracy jump came from an audit of 84 real photos, not from the model: the renderer drew landscape scenes while reality was 79/84 portrait. Rotating the canvas (plus matching real framing and image quality) moved whole-hand accuracy by ~20 points.
- An inference-time preprocessing stack (grayscale / binarize / retry) was measured against annotated clips: zero changed answers. Deleted; the rule "robustness lives in training data" is now written into the repo's engineering notes.
- A late, purely geometric post-processing rule — discard detections that are both far from the corner cluster and duplicates — added 14 points of whole-hand accuracy with no retraining.
Status
Feature-complete, unreleased — by decision, not by accident. Testing at a real table exposed the scenario's flaw: one hand holds the fan, the other holds the phone, and every correction makes you the slow player. The engineering survives the product: the data pipeline, the evaluation harness, and the write-ups.
Takeaways
- Build the whole loop before polishing any stage — day one ended at 7 cards out of 27, and that honest number set the agenda.
- Measure with the product's own metric; every proxy that disagreed with it turned out to be the one lying.
- Synthetic data works if its assumptions are audited against reality — mine failed silently until measured.
- Knowing when not to ship is part of the craft.
The decisions, failures, and numbers are documented in detail: Guandan Planner write-ups →