Cardvision: collect & ingest real card images (data flywheel from live scans) #496
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
Cardvision's recognition index is built entirely from TCGplayer stock scans
(
build_indexincardvision.pydownloads each catalog row'simage_url, CLIP-encodesit, and extracts ORB descriptors — one reference view per print). Those stock images are
pristine, evenly lit, dead-on, no sleeve, no glare. Real phone photos are none of those.
The synthetic augmentation we do (180° twins, warps) papers over geometry, not the
appearance gap: holo glare, lighting colour cast, wear, sleeves/toploaders, off-angle,
and phone-camera characteristics.
That gap is exactly why match margins are tight on look-alikes and EN/JP twins — the
stock reference doesn't look like what the camera sees. We should start matching against
real images of real cards, collected from actual use.
The core idea: a data flywheel
Every scan is already a real, about-to-be-labelled photo that we currently throw away.
/identifyreceives the real photo, and moments later the user resolves the true card —by accepting the match, picking from the variant/language picker, or getting a cloud
(Gemini) answer. Today
/identifykeeps only a 160px thumbnail in a ring buffer anddiscards the full image; the confirmed identity never comes back.
Capture that pair — (real photo) → (confirmed card id) — and we get a continuously
growing, correctly-labelled reference set for free, weighted toward exactly the hard
cases (the scans that needed the picker or the cloud are the ones where the stock
reference was insufficient).
Why this is feasible without rewriting the matcher
embedding + ORB set each);
Index.dupes_of(i)groups rows that are the same card, andthe margin test already excludes a card's own group from the runner-up (
cardvision.py~L580). So a real photo added as another row with the same card id, grouped via
dupes_of, becomes a matchable reference view and does not count against the card's
own confidence margin. No change to
identify().calibrate(index_dir, photos_dir)(L1082)already runs a folder of
<card_id>.jpgreal photos through the pipeline and reportstop-1 accuracy + false-accept rates per threshold. That's our validation gate, ready to
reuse.
Proposed pipeline
1. Capture — persist the real upload + provisional label.
POST /identify(envCARDVISION_CAPTURE=1, off by default):write the raw image to
/mnt/user/appdata/cardvision/captures/<capture_id>.jpgplus asidecar JSON
{capture_id, ts, status, top_candidate_id, inliers, margin, ratio, source}.capture_idin the/identifyresponse so the client can attach a label later.2. Label — attach ground truth.
POST /capture/label {capture_id, confirmed_id, source}called by the app/backendonce the user lands on a card.
source ∈ {auto_highconf, picker, cloud, manual}.manual/picker > auto_highconf > cloud. Store it; awrong label poisons the index, so unlabeled captures are never ingested.
3. Curate — quality + human review.
grouped by card — thumbnail, provisional vs confirmed id, scores, source — with
approve / reject / relabel.
cap N diverse views per card (e.g. 5–10; we want variety, not 200 of one card).
4. Ingest — add approved captures as reference rows.
{id: <confirmed card id>, source: real, ...}, CLIP-encode + ORB-extract the real photo, and put it in the card's dupes_of group.build_indexis ~8h (seereproject_allnote); appending rows + re-runningcompute_art_groupsfor touchedcards is the target. Fall back to periodic full rebuild from an augmented manifest (stock
URL + local
real/<id>/*.jpg) if append proves fiddly.5. Deploy — roll the new index to the fleet.
/mnt/user/appdata/cardvision/data(mounted
rointo all 4 workers), then restart the workers (they load the index atstartup; the warmer re-warms within a cycle). Zero-downtime-ish; keep the prior index dir
for instant rollback.
6. Measure — close the loop / don't regress.
calibratebefore vs after ingest to confirmtop-1 accuracy up and false-accepts not up. This is the ship gate.
fallback (the dashboard already has the raw signal).
The crux: ground-truth quality
The whole thing lives or dies on label correctness. A mislabelled capture teaches the
index that a photo is the wrong card — worse than no data. Hence: never ingest an
unlabeled/low-trust capture without human review, keep the
sourcetrust ranking, andgate every rollout behind the
calibrateregression check.Suggested phasing (one epic, phased commits per AGENTS.md)
/identifyopt-in save +capture_id+/capture/label;app wires the confirmed id back. Just accumulate labelled data; no ingest yet. Low risk,
starts the flywheel immediately.
calibrategate.calibrate, swap if green; humans only review the marginal ones.
Open questions / decisions
vendor/staff, opt-in flag, documented. OK as-is, or need explicit per-user consent?
full rebuilds to start?
Gambit backend (which already resolves text-match/cloud and knows the final card)?
Backend-mediated is probably cleaner and centralises trust.
Out of scope (for now)
training the encoder.
Two clarifications: Railway topology + capture quality
Where this runs (Railway vs Unraid). Cardvision runs on the Unraid box (RTX 3090),
not Railway. For a cardvision scan the phone uploads the photo directly to cardvision
(
cardvision.ufind.app→ tunnel → box); the raw image never reaches Railway — theRailway backend only receives the resolved text identity for pricing. So capture,
storage, and index rebuild all stay on Unraid (that's where the GPU + index live; Railway
has no persistent volume for a growing corpus and no GPU). What round-trips back to
cardvision is the label, not the photo, and the app is the right source of it
(it knows the final card after the picker/cloud). Edge case: on cloud-fallback scans the
photo does go to Railway→Gemini, but cardvision already holds that same photo (it saw it
first) — we just need Gemini's answer routed back as the label. Net: Railway stays a thin
identity/pricing resolver; the flywheel doesn't load it.
What quality we'd actually be learning from. The app does not send a pristine photo.
Before upload (
downscaleForCardvision) it is: cropped to the card, EXIF-baked upright,capped at 1100px longest edge, re-JPEG'd at q0.85. And the engine then squashes a
cropped upload to a fixed 488×680 canvas before extracting a single feature — so the
matcher's real operating resolution for our scans is 488×680.
Implications:
inliers vs 3024px→290). A captured 1100px crop is plenty to build a 488×680 reference,
and it is domain-matched — references would look exactly like real queries, which is
better for matching than pristine hi-res stock.
#1 job of the curation gate), (2) double-JPEG artifacts (minor), (3) it's lossy /
un-redoable — 1100px q0.85 caps us if we ever move to a higher-res matcher.
Recommendation: Phase 1 captures the
/identifybytes as-is (free, domain-matched,sufficient for today's engine). If we want future-proof archival references, have the app
send one higher-res original only for the subset approved for ingest, not every scan —
keeps bandwidth/storage sane without locking us to 1100px forever.
Chosen capture approach: dual-send (fast small match + quiet full-res training upload)
Refines the Phase 1 capture design. Instead of persisting the
/identifybytes (which arealready downscaled to 1100px and lossy), send two versions of the same crop:
POST /identify→ near-instant match.full-res crop + the confirmed label to a new
POST /captureon cardvision.Why this is nearly free in the app:
cropListerPhotoUrialready produces a full-rescrop;
downscaleForCardvisiononly makes the 1100px copy for/identify. Both versionsalready exist — foreground sends the small one, background sends the full one. No extra
image processing.
Why it's better than capturing
/identifybytes:higher-res matcher or want to re-derive references, the pixels exist.
arrive together in one request. No orphan captures, no separate
/capture/labelround-trip to reconcile. (Abandoned scans upload nothing, which is correct — no label.)
Decisions this forces:
and flush on wifi. Best-effort, never blocks the scan or UI.
dedup + retention window so we don't hoard the 51st view of a common card.
stays.
Railway: unchanged — the background upload goes phone → cardvision (Unraid) directly, never
through Railway.
This supersedes "capture the /identify bytes" in Phase 1 and folds in the old
/capture/labelstep (label now travels with the full-res upload).Phase 1 built (capture + label) — mobile committed, Tower deploy staged
Mobile (branch
feature/496-cardvision-training-capture, commit 9b8fc70): at commit(Buy / inventory batch / direct inventory), each scanned card's full-res crop +
user-confirmed tcgplayer id is background-uploaded to cardvision
POST /train-capture.Commit is the ground-truth moment — variant picks are resolved by then; slabs, sealed,
non-scan entries, and unresolved picks never train, so discarded wrong matches can't
poison the future index. Fire-and-forget: awaited cost is one local file copy. Gate logic
pinned by
lib/scan/trainingCapture.test.js; mobile tsc green.Server (agg.py updated in ~/ai/cardvision, NOT yet on Tower):
/train-captureon theaggregator (8220) — token tripwire (
tokenform field must match TRAIN_TOKEN), per-cardcap 10, stores
captures/labeled/<card_id>/<ts>_<id>.jpg+labeled.jsonl;/capture-statsnow also reports labeled counts.Tower deploy steps (agent is classifier-blocked from writing to the box):
scp ~/ai/cardvision/agg.py root@192.168.1.95:/mnt/user/appdata/cardvision/agg.pyserver {}block (no basic auth — token-gated in the app):location = /train-capture { proxy_pass http://127.0.0.1:8220; proxy_read_timeout 30s; }docker exec cardvision-lb nginx -t && docker restart cardvision-dash cardvision-lb-F image=@warm.jpg -F card_id=TEST123 -F token=<TRAIN_TOKEN>→{"stored": true}; thenrm -rf captures/labeled+ restart dash to clear the test row.Wifi-only gating skipped (no netinfo dep; ~1–3MB per committed scan, best-effort). Ingest
(enroll into the index) stays Phase 3 behind the calibrate gate.
Server side is LIVE on Tower: /train-capture deployed (agg.py + all 3 lb confs), smoke-tested — bad token 403 via public URL, stores + per-card counts survive restart, test rows cleaned. Scan path unaffected (healthz 200 via LB and public). Mobile branch feature/496-cardvision-training-capture awaits PR.
Phase 1 SHIPPED + a Phase 3 requirement
Phase 1 is live end to end: PR #337 merged (main a08b249), OTA on runtime 1.3.6, /train-capture live on Tower. Committed scans now accumulate in captures/labeled/.
Phase 3 requirement (from review discussion): ingest-time auto-verify. The picker is a mis-tap surface — a user can confirm the wrong card (worst case a wrong-language twin, the exact poison #489 fought). Same-art print-variant mistaps are near-harmless visually, but cross-card/cross-language mislabels are not. So before any capture is ingested as a reference row: run it back through cardvision itself; if the claimed card_id is not among the visual top-k candidates, do NOT auto-ingest — route it to the human review queue with the mismatch shown. Cheap (one /identify pass per capture at ingest time) and it automatically catches "tapped a completely different card".
LIVE ENROLL SHIPPED (Phase 3, gated) — the flywheel is closed
Cardvision now learns in near-real-time from confirmed scans. Deployed to the Tower fleet tonight:
How it works: /train-capture (labeled capture at commit/accept) now forwards each photo to a worker's new token-gated
/enroll. The worker runs a 1-vs-1 verify gate — ORB/RANSAC inliers of the photo against the CLAIMED card's own stock art, both orientations, floor 20 inliers (CARDVISION_ENROLL_INLIER_MIN). Pass → the photo becomes a live reference row: one self-contained .npz in a shared rw mount (enrolled/), embedding + index-side ORB descriptors, grouped with the card's stock rows so real views never hurt the card's own margin. All 4 workers poll the dir (5s) and atomically swap in an OverlayIndex — no rebuild, no restart. Fail → the capture STAYS in captures/labeled/ flaggedenrolled:falseas the human-review queue (hard cases like foil-textured cards whose stock scan doesn't match reality land here — the cards that need real references most).Why the gate is sound even though recognition missed: recognition is 1-vs-87k with a margin rule — misses are usually the right card scoring well but a twin too close (Ursaluna: 88 vs 58 inliers → cloud). Verification asks only "does this photo match the named card's art" — no competition. A mislabeled photo is different artwork and scores ~0-15.
E2E proof on prod: wrong label (Tangrowth photo claimed as Alakazam) → verify_failed at 15 inliers, held for review. True label → enrolled at 78 inliers, all 4 workers picked it up within one poll. Re-identify then matched via the REAL reference at 281 inliers (vs 119 stock). Within minutes of deploy, 5 organic commits flowed through: 4 enrolled live, 1 blurry duplicate correctly held back.
Limits/controls: cap 5 enrolled views/card, token tripwire, employees-only traffic, rollback = delete the .npz (workers drop it next poll) or
enrolled/wholesale. Periodiccalibrateregression audit still recommended (Phase 4). Cosmetic known issue: enrolled rows leak their meta (real/ts/source/group_rows) into /identify candidate JSON — harmless to the app.Loose end: 9 pre-deploy captures in captures/labeled/ never attempted enroll (rows without an
enrolledfield) — backfill by POSTing them to /enroll, or leave for the review surface.Live-enroll incident + fix (same evening): cross-photo noise, resolved
Kris's first real session exposed a defect: same-session sleeved photos cross-match at 15-25 spurious ORB inliers (sleeve texture/desk, a noise fit — ratio ~0.1-0.2), so one card's enrolled reference could tank ANOTHER card's margin (more cloud fallbacks) or even win outright (an Elder Dragon scan matched Dragon's Rage's enrolled row). Meanwhile a TRUE photo-vs-photo match of the same card lands 90-280 inliers — the bands don't overlap.
Fixes (deployed + verified on the failing scans):
Re-tested the exact failing scans: Elder Dragon → Elder Dragon (42/62 inl via its real reference), Premonition → itself (228), no cross-card wins.
Misaligned-scan rescue deployed (no app change needed)
Kris's field report: ignoring the align guide (card far from the reticle) made scans miss. Root cause: the app sends its reticle crop with already_cropped=1, so the engine trusted the flag, squashed a mostly-desk crop to card shape, and matched noise — and the CLIP shortlist itself was poisoned by the desk pixels, so every fallback stage verified only wrong candidates (the deep sweep that used to cover this is deliberately disabled for latency).
Fix, server-side only: already_cropped is now a HINT with a graceful fall-through — (1) stage B2 gives a failed cropped query the aspect-preserving hi-res frame path + match-guided refinement; (2) the contour-quad detector now also runs on cropped queries (free when aligned — a border-touching card yields no quads); (3) NEW quad_pass: when nothing matched, the best quad crop is re-embedded for a fresh shortlist and verified — geometry-guided rescue that works when the shortlist was poisoned.
Verified: synthetic far-away crop (card at 40% of frame on a noisy desk) went low_confidence/wrong-16-inliers → match, 104 inliers, ratio 0.92, margin 6.5, quad_pass=true, 469ms. Aligned fast path unchanged (121ms). Note: a rescued scan lands near the app's 500ms bail budget — if far scans still occasionally fall to cloud on real phones (network adds ~50-100ms), bump CARDVISION_BUDGET_MS in the app (~900ms) via OTA.
Native on-device auto-crop (tap → Vision rectangle detect → send tight crop) remains a possible future UX upgrade (needs a native build); the server rescue removes the accuracy cliff without it.
Reference hygiene: enrolled references are now card-localized (tight crops)
Decision (Kris + agent): index references must be the CARD, not the scene — background robustness is the localizer's job; the on-card variation (glare/lighting/wear) is what real references exist to capture. The ORIGINAL uncropped capture stays in captures/labeled/, so a future neural-detector approach that wants messy-background data keeps its raw material; only the derived index reference is cropped.
/enroll now: (1) the verify GATE tries the whole capture AND quad-detected card warps, both orientations — so far/misaligned captures (rescued by quad_pass at scan time) can enroll instead of dying at the door; (2) the stored reference is chosen among border-trim + quad warps + nested quads (sleeve-edge case) by whichever verifies best against the claimed card's own stock art; (3) each enrollment saves its exact reference view as a .jpg beside the .npz for audit/review.
Also from tonight's re-runs: the identity-conflict cross-check retro-caught a second wrong-label commit (90452 vs 611747) which on inspection is the EN/JP Wailmer art-twin pair — with the stronger quad-aware gate the photo proves 83 inliers against the user's chosen EN print, so it enrolls under the user's label (the designed twin behavior; the app's language picker remains the EN/JP defense). Store rebuilt end-to-end: 30 tight references live across the fleet. Verified visually: a synthetic far-away desk photo enrolls as a clean card-localized color reference at 88 inliers.