Guandan Planner

2026 · Flutter · iOS & Android · built solo

A mobile app that reads a full 27-card Guandan hand through the phone camera and suggests how to arrange it into playable groups. Two steps: scan, then results — cards lock in one by one as evidence accumulates across frames, a misread is fixed by tapping the card, and everything runs on the device.

The result stage of an early build: the scanned photo with per-card markers, and an arrangement plan laid out below
An early build: the scanned photo, the recognized cards, and an arrangement plan — one whose flaws later rewrote the scoring function.

Motivation

Guandan is a Chinese climbing card game played with two full decks; each player holds 27 cards, and every card can legally appear twice. Arranging such a hand well is a real skill — and recognizing one is a genuinely hard vision problem: heavy occlusion in a fan, duplicate faces that defeat the usual "already seen it" sanity check, and no public dataset anywhere. That combination — a small, concrete product wrapped around several unsolved-for-this-domain technical problems — is why the project exists.

Architecture

Three subsystems, each built from scratch:

  • On-device recognition — a 109-class YOLO (54 card faces × two corner roles, plus the hand silhouette), followed by geometric corner pairing, multi-frame log-odds evidence tracking, and a ledger that admits a duplicate card only when both copies are visible in the same frame.
  • Arrangement engine — a pure-Dart exact-cover search with fail-first pivoting and a playability scorer; three strategies come from a single search pass, in a background isolate.
  • Data and training infrastructure — a local studio that turns photos of a physical deck into rectified assets, renders labeled scenes through one shared pinhole camera, manages datasets with content-addressed fingerprints, dispatches training to rented GPUs, and evaluates against ground-truthed real photos.

Key decisions

DecisionChoiceRationale
Where inference runs On-device (TFLite / Core ML) The scan loop needs 8+ fps sustained; photos stay on the phone
What to detect Corner indices, not whole cards Fanned cards overlap at IoU ≈ 0.8 — whole-card boxes collapse under NMS
Label space 109 classes (54 faces × 2 corner roles + hand) Corner roles make card counting decidable with duplicates legal
Training data 100% synthetic, rendered No dataset exists; rendered labels are exact and regenerable overnight
Acceptance metric Whole-hand exact rate on real photos mAP once endorsed a scheme that read zero real hands correctly
Arrangement objective Playability over fewest hands Minimizing hand count buries aces inside straights

Evaluation

Models are judged on ground-truthed real photographs — 82 images, 2,111 cards, annotated as card multisets — with hash-checked isolation from training data. The primary metric is the fraction of hands read exactly right: with 27 cards a scan, per-card accuracy flatters, and the exponent is merciless.

MeasureValue
Per-card accuracy98.7%
Whole-hand exact rate69.5%
Model size (Android / iOS)3.0 MB / 2.4 MB
Training data130k rendered images, zero hand labels
Training span21 runs over 19 days

Iteration highlights

  • A "cleaner" 54-class labeling scheme posted the best synthetic mAP of the project and read zero real hands correctly — rolled back within a day, and the reason the acceptance metric is what it is.
  • The single largest accuracy jump came from an audit of 84 real photos, not from the model: the renderer drew landscape scenes while reality was 79/84 portrait. Rotating the canvas (plus matching real framing and image quality) moved whole-hand accuracy by ~20 points.
  • An inference-time preprocessing stack (grayscale / binarize / retry) was measured against annotated clips: zero changed answers. Deleted; the rule "robustness lives in training data" is now written into the repo's engineering notes.
  • A late, purely geometric post-processing rule — discard detections that are both far from the corner cluster and duplicates — added 14 points of whole-hand accuracy with no retraining.

Status

Feature-complete, unreleased — by decision, not by accident. Testing at a real table exposed the scenario's flaw: one hand holds the fan, the other holds the phone, and every correction makes you the slow player. The engineering survives the product: the data pipeline, the evaluation harness, and the write-ups.

Takeaways

  • Build the whole loop before polishing any stage — day one ended at 7 cards out of 27, and that honest number set the agenda.
  • Measure with the product's own metric; every proxy that disagreed with it turned out to be the one lying.
  • Synthetic data works if its assumptions are audited against reality — mine failed silently until measured.
  • Knowing when not to ship is part of the craft.

The decisions, failures, and numbers are documented in detail: Guandan Planner write-ups →