Cardvision: collect & ingest real card images (data flywheel from live scans) #496

Open
opened 2026-08-28 04:20:17 +00:00 by gambit-admin · 9 comments
Owner

Problem

Cardvision's recognition index is built entirely from TCGplayer stock scans
(build_index in cardvision.py downloads each catalog row's image_url, CLIP-encodes
it, and extracts ORB descriptors — one reference view per print). Those stock images are
pristine, evenly lit, dead-on, no sleeve, no glare. Real phone photos are none of those.
The synthetic augmentation we do (180° twins, warps) papers over geometry, not the
appearance gap: holo glare, lighting colour cast, wear, sleeves/toploaders, off-angle,
and phone-camera characteristics.

That gap is exactly why match margins are tight on look-alikes and EN/JP twins — the
stock reference doesn't look like what the camera sees. We should start matching against
real images of real cards, collected from actual use.

The core idea: a data flywheel

Every scan is already a real, about-to-be-labelled photo that we currently throw away.
/identify receives the real photo, and moments later the user resolves the true card —
by accepting the match, picking from the variant/language picker, or getting a cloud
(Gemini) answer. Today /identify keeps only a 160px thumbnail in a ring buffer and
discards the full image; the confirmed identity never comes back.

Capture that pair — (real photo) → (confirmed card id) — and we get a continuously
growing, correctly-labelled reference set for free, weighted toward exactly the hard
cases (the scans that needed the picker or the cloud are the ones where the stock
reference was insufficient).

Why this is feasible without rewriting the matcher

  • Multi-reference-per-card already fits. The index is a flat list of rows (one
    embedding + ORB set each); Index.dupes_of(i) groups rows that are the same card, and
    the margin test already excludes a card's own group from the runner-up (cardvision.py
    ~L580). So a real photo added as another row with the same card id, grouped via
    dupes_of
    , becomes a matchable reference view and does not count against the card's
    own confidence margin. No change to identify().
  • A real-photo harness already exists. calibrate(index_dir, photos_dir) (L1082)
    already runs a folder of <card_id>.jpg real photos through the pipeline and reports
    top-1 accuracy + false-accept rates per threshold. That's our validation gate, ready to
    reuse.

Proposed pipeline

1. Capture — persist the real upload + provisional label.

  • Add opt-in capture to POST /identify (env CARDVISION_CAPTURE=1, off by default):
    write the raw image to /mnt/user/appdata/cardvision/captures/<capture_id>.jpg plus a
    sidecar JSON {capture_id, ts, status, top_candidate_id, inliers, margin, ratio, source}.
  • Return capture_id in the /identify response so the client can attach a label later.
  • Provisional label = cardvision's top candidate (may be wrong/empty — not trusted yet).

2. Label — attach ground truth.

  • New POST /capture/label {capture_id, confirmed_id, source} called by the app/backend
    once the user lands on a card. source ∈ {auto_highconf, picker, cloud, manual}.
  • Trust ranking for later ingest: manual/picker > auto_highconf > cloud. Store it; a
    wrong label poisons the index, so unlabeled captures are never ingested.

3. Curate — quality + human review.

  • Extend the (now password-protected) fleet dashboard with a review surface: captures
    grouped by card — thumbnail, provisional vs confirmed id, scores, source — with
    approve / reject / relabel.
  • Auto-filters before review: drop unlabeled, blur/quality gate, near-duplicate dedupe,
    cap N diverse views per card (e.g. 5–10; we want variety, not 200 of one card).

4. Ingest — add approved captures as reference rows.

  • For each approved capture: append an index row {id: <confirmed card id>, source: real, ...}, CLIP-encode + ORB-extract the real photo, and put it in the card's dupes_of group.
  • Prefer incremental append over full rebuild — a full build_index is ~8h (see
    reproject_all note); appending rows + re-running compute_art_groups for touched
    cards is the target. Fall back to periodic full rebuild from an augmented manifest (stock
    URL + local real/<id>/*.jpg) if append proves fiddly.
  • Tag real rows so we can weight/measure them and roll back.

5. Deploy — roll the new index to the fleet.

  • Build into a staged index dir, atomic-swap into /mnt/user/appdata/cardvision/data
    (mounted ro into all 4 workers), then restart the workers (they load the index at
    startup; the warmer re-warms within a cycle). Zero-downtime-ish; keep the prior index dir
    for instant rollback.

6. Measure — close the loop / don't regress.

  • Hold out a labelled real-photo set; run calibrate before vs after ingest to confirm
    top-1 accuracy up and false-accepts not up. This is the ship gate.
  • Track whether cards with added real views subsequently match at higher rate / lower cloud
    fallback (the dashboard already has the raw signal).

The crux: ground-truth quality

The whole thing lives or dies on label correctness. A mislabelled capture teaches the
index that a photo is the wrong card — worse than no data. Hence: never ingest an
unlabeled/low-trust capture without human review, keep the source trust ranking, and
gate every rollout behind the calibrate regression check.

Suggested phasing (one epic, phased commits per AGENTS.md)

  • Phase 1 — Capture only. /identify opt-in save + capture_id + /capture/label;
    app wires the confirmed id back. Just accumulate labelled data; no ingest yet. Low risk,
    starts the flywheel immediately.
  • Phase 2 — Curate. Dashboard review surface + auto-filters.
  • Phase 3 — Ingest + deploy. Append-to-index + staged swap + calibrate gate.
  • Phase 4 — Automate. Nightly: auto-approve high-trust captures, incremental ingest,
    calibrate, swap if green; humans only review the marginal ones.

Open questions / decisions

  • Consent/scope. Captures are users' own inventory photos. Scope to Gambit
    vendor/staff, opt-in flag, documented. OK as-is, or need explicit per-user consent?
  • Append vs rebuild for ingest (Phase 3) — invest in incremental, or accept periodic
    full rebuilds to start?
  • Where the label round-trips — from the mobile app directly to cardvision, or via the
    Gambit backend (which already resolves text-match/cloud and knows the final card)?
    Backend-mediated is probably cleaner and centralises trust.
  • Per-card view cap and the blur/dedupe thresholds — tune on real volume.

Out of scope (for now)

  • Retraining/fine-tuning CLIP itself. This is about adding real reference views, not
    training the encoder.
  • Changing the matcher/thresholds (tracked separately; thresholds are live-tuned via env).
## Problem Cardvision's recognition index is built entirely from **TCGplayer stock scans** (`build_index` in `cardvision.py` downloads each catalog row's `image_url`, CLIP-encodes it, and extracts ORB descriptors — one reference view per print). Those stock images are pristine, evenly lit, dead-on, no sleeve, no glare. Real phone photos are none of those. The synthetic augmentation we do (180° twins, warps) papers over *geometry*, not the *appearance gap*: holo glare, lighting colour cast, wear, sleeves/toploaders, off-angle, and phone-camera characteristics. That gap is exactly why match margins are tight on look-alikes and EN/JP twins — the stock reference doesn't look like what the camera sees. We should start matching against **real images of real cards**, collected from actual use. ## The core idea: a data flywheel **Every scan is already a real, about-to-be-labelled photo that we currently throw away.** `/identify` receives the real photo, and moments later the user *resolves the true card* — by accepting the match, picking from the variant/language picker, or getting a cloud (Gemini) answer. Today `/identify` keeps only a 160px thumbnail in a ring buffer and discards the full image; the confirmed identity never comes back. Capture that pair — **(real photo) → (confirmed card id)** — and we get a continuously growing, correctly-labelled reference set *for free*, weighted toward exactly the hard cases (the scans that needed the picker or the cloud are the ones where the stock reference was insufficient). ## Why this is feasible without rewriting the matcher - **Multi-reference-per-card already fits.** The index is a flat list of rows (one embedding + ORB set each); `Index.dupes_of(i)` groups rows that are the same card, and the margin test already *excludes* a card's own group from the runner-up (`cardvision.py` ~L580). So a real photo added as **another row with the same card id, grouped via dupes_of**, becomes a matchable reference view and does **not** count against the card's own confidence margin. No change to `identify()`. - **A real-photo harness already exists.** `calibrate(index_dir, photos_dir)` (L1082) already runs a folder of `<card_id>.jpg` real photos through the pipeline and reports top-1 accuracy + false-accept rates per threshold. That's our validation gate, ready to reuse. ## Proposed pipeline **1. Capture** — persist the real upload + provisional label. - Add opt-in capture to `POST /identify` (env `CARDVISION_CAPTURE=1`, off by default): write the raw image to `/mnt/user/appdata/cardvision/captures/<capture_id>.jpg` plus a sidecar JSON `{capture_id, ts, status, top_candidate_id, inliers, margin, ratio, source}`. - Return `capture_id` in the `/identify` response so the client can attach a label later. - Provisional label = cardvision's top candidate (may be wrong/empty — not trusted yet). **2. Label** — attach ground truth. - New `POST /capture/label {capture_id, confirmed_id, source}` called by the app/backend once the user lands on a card. `source ∈ {auto_highconf, picker, cloud, manual}`. - Trust ranking for later ingest: `manual/picker > auto_highconf > cloud`. Store it; a wrong label poisons the index, so unlabeled captures are never ingested. **3. Curate** — quality + human review. - Extend the (now password-protected) fleet dashboard with a review surface: captures grouped by card — thumbnail, provisional vs confirmed id, scores, source — with approve / reject / relabel. - Auto-filters before review: drop unlabeled, blur/quality gate, near-duplicate dedupe, cap **N diverse views per card** (e.g. 5–10; we want variety, not 200 of one card). **4. Ingest** — add approved captures as reference rows. - For each approved capture: append an index row `{id: <confirmed card id>, source: real, ...}`, CLIP-encode + ORB-extract the real photo, and put it in the card's dupes_of group. - Prefer **incremental append** over full rebuild — a full `build_index` is ~8h (see `reproject_all` note); appending rows + re-running `compute_art_groups` for touched cards is the target. Fall back to periodic full rebuild from an augmented manifest (stock URL + local `real/<id>/*.jpg`) if append proves fiddly. - Tag real rows so we can weight/measure them and roll back. **5. Deploy** — roll the new index to the fleet. - Build into a **staged** index dir, atomic-swap into `/mnt/user/appdata/cardvision/data` (mounted `ro` into all 4 workers), then restart the workers (they load the index at startup; the warmer re-warms within a cycle). Zero-downtime-ish; keep the prior index dir for instant rollback. **6. Measure** — close the loop / don't regress. - Hold out a labelled real-photo set; run `calibrate` before vs after ingest to confirm top-1 accuracy up and false-accepts not up. This is the ship gate. - Track whether cards with added real views subsequently match at higher rate / lower cloud fallback (the dashboard already has the raw signal). ## The crux: ground-truth quality The whole thing lives or dies on label correctness. A mislabelled capture teaches the index that a photo *is* the wrong card — worse than no data. Hence: never ingest an unlabeled/low-trust capture without human review, keep the `source` trust ranking, and gate every rollout behind the `calibrate` regression check. ## Suggested phasing (one epic, phased commits per AGENTS.md) - **Phase 1 — Capture only.** `/identify` opt-in save + `capture_id` + `/capture/label`; app wires the confirmed id back. Just accumulate labelled data; no ingest yet. Low risk, starts the flywheel immediately. - **Phase 2 — Curate.** Dashboard review surface + auto-filters. - **Phase 3 — Ingest + deploy.** Append-to-index + staged swap + `calibrate` gate. - **Phase 4 — Automate.** Nightly: auto-approve high-trust captures, incremental ingest, calibrate, swap if green; humans only review the marginal ones. ## Open questions / decisions - **Consent/scope.** Captures are users' own inventory photos. Scope to Gambit vendor/staff, opt-in flag, documented. OK as-is, or need explicit per-user consent? - **Append vs rebuild** for ingest (Phase 3) — invest in incremental, or accept periodic full rebuilds to start? - **Where the label round-trips** — from the mobile app directly to cardvision, or via the Gambit backend (which already resolves text-match/cloud and knows the final card)? Backend-mediated is probably cleaner and centralises trust. - Per-card view cap and the blur/dedupe thresholds — tune on real volume. ## Out of scope (for now) - Retraining/fine-tuning CLIP itself. This is about adding real *reference views*, not training the encoder. - Changing the matcher/thresholds (tracked separately; thresholds are live-tuned via env).
gambit-admin added the featureapi labels 2026-08-28 04:20:17 +00:00
Author
Owner

Two clarifications: Railway topology + capture quality

Where this runs (Railway vs Unraid). Cardvision runs on the Unraid box (RTX 3090),
not Railway. For a cardvision scan the phone uploads the photo directly to cardvision
(cardvision.ufind.app → tunnel → box); the raw image never reaches Railway — the
Railway backend only receives the resolved text identity for pricing. So capture,
storage, and index rebuild all stay on Unraid
(that's where the GPU + index live; Railway
has no persistent volume for a growing corpus and no GPU). What round-trips back to
cardvision is the label, not the photo, and the app is the right source of it
(it knows the final card after the picker/cloud). Edge case: on cloud-fallback scans the
photo does go to Railway→Gemini, but cardvision already holds that same photo (it saw it
first) — we just need Gemini's answer routed back as the label. Net: Railway stays a thin
identity/pricing resolver; the flywheel doesn't load it.

What quality we'd actually be learning from. The app does not send a pristine photo.
Before upload (downscaleForCardvision) it is: cropped to the card, EXIF-baked upright,
capped at 1100px longest edge, re-JPEG'd at q0.85. And the engine then squashes a
cropped upload to a fixed 488×680 canvas before extracting a single feature
— so the
matcher's real operating resolution for our scans is 488×680.

Implications:

  • Resolution is a non-issue for the current matcher (measured in-code: 800px→286
    inliers vs 3024px→290). A captured 1100px crop is plenty to build a 488×680 reference,
    and it is domain-matched — references would look exactly like real queries, which is
    better for matching than pristine hi-res stock.
  • The real quality risks are (1) the auto-crop — a bad crop poisons the reference (the
    #1 job of the curation gate), (2) double-JPEG artifacts (minor), (3) it's lossy /
    un-redoable
    — 1100px q0.85 caps us if we ever move to a higher-res matcher.

Recommendation: Phase 1 captures the /identify bytes as-is (free, domain-matched,
sufficient for today's engine). If we want future-proof archival references, have the app
send one higher-res original only for the subset approved for ingest, not every scan —
keeps bandwidth/storage sane without locking us to 1100px forever.

### Two clarifications: Railway topology + capture quality **Where this runs (Railway vs Unraid).** Cardvision runs on the Unraid box (RTX 3090), not Railway. For a cardvision scan the phone uploads the photo *directly* to cardvision (`cardvision.ufind.app` → tunnel → box); **the raw image never reaches Railway** — the Railway backend only receives the resolved text identity for pricing. So **capture, storage, and index rebuild all stay on Unraid** (that's where the GPU + index live; Railway has no persistent volume for a growing corpus and no GPU). What round-trips back to cardvision is the **label, not the photo**, and the **app** is the right source of it (it knows the final card after the picker/cloud). Edge case: on cloud-fallback scans the photo does go to Railway→Gemini, but cardvision already holds that same photo (it saw it first) — we just need Gemini's answer routed back as the label. Net: Railway stays a thin identity/pricing resolver; the flywheel doesn't load it. **What quality we'd actually be learning from.** The app does not send a pristine photo. Before upload (`downscaleForCardvision`) it is: cropped to the card, EXIF-baked upright, **capped at 1100px** longest edge, **re-JPEG'd at q0.85**. And the engine then **squashes a cropped upload to a fixed 488×680 canvas before extracting a single feature** — so the matcher's real operating resolution for our scans is 488×680. Implications: - **Resolution is a non-issue for the current matcher** (measured in-code: 800px→286 inliers vs 3024px→290). A captured 1100px crop is plenty to build a 488×680 reference, and it is **domain-matched** — references would look exactly like real queries, which is better for matching than pristine hi-res stock. - The real quality risks are **(1) the auto-crop** — a bad crop poisons the reference (the #1 job of the curation gate), **(2) double-JPEG artifacts** (minor), **(3) it's lossy / un-redoable** — 1100px q0.85 caps us if we ever move to a higher-res matcher. **Recommendation:** Phase 1 captures the `/identify` bytes as-is (free, domain-matched, sufficient for today's engine). If we want future-proof archival references, have the app send one higher-res original **only for the subset approved for ingest**, not every scan — keeps bandwidth/storage sane without locking us to 1100px forever.
Author
Owner

Chosen capture approach: dual-send (fast small match + quiet full-res training upload)

Refines the Phase 1 capture design. Instead of persisting the /identify bytes (which are
already downscaled to 1100px and lossy), send two versions of the same crop:

  • Foreground (unchanged): the 1100px crop → POST /identify → near-instant match.
  • Background: once the card is confirmed (picker/accept/cloud), quietly upload the
    full-res crop + the confirmed label to a new POST /capture on cardvision.

Why this is nearly free in the app: cropListerPhotoUri already produces a full-res
crop
; downscaleForCardvision only makes the 1100px copy for /identify. Both versions
already exist — foreground sends the small one, background sends the full one. No extra
image processing.

Why it's better than capturing /identify bytes:

  • Archival fidelity — solves the "lossy / un-redoable" risk; if we ever move to a
    higher-res matcher or want to re-derive references, the pixels exist.
  • Simpler labeling — upload happens after the card is confirmed, so image + label
    arrive together in one request. No orphan captures, no separate /capture/label
    round-trip to reconcile. (Abandoned scans upload nothing, which is correct — no label.)

Decisions this forces:

  1. Bandwidth — full-res crops are ~1–5MB vs ~150KB. Prefer wifi; queue on cellular
    and flush on wifi. Best-effort, never blocks the scan or UI.
  2. Storage growth — unbounded ≈ 20GB/day at 10k scans. Server-side per-card cap +
    dedup + retention window
    so we don't hoard the 51st view of a common card.
  3. Crop quality still matters — full-res doesn't fix a bad auto-crop; the curation gate
    stays.

Railway: unchanged — the background upload goes phone → cardvision (Unraid) directly, never
through Railway.

This supersedes "capture the /identify bytes" in Phase 1 and folds in the old
/capture/label step (label now travels with the full-res upload).

### Chosen capture approach: dual-send (fast small match + quiet full-res training upload) Refines the Phase 1 capture design. Instead of persisting the `/identify` bytes (which are already downscaled to 1100px and lossy), **send two versions of the same crop**: - **Foreground (unchanged):** the 1100px crop → `POST /identify` → near-instant match. - **Background:** once the card is *confirmed* (picker/accept/cloud), quietly upload the **full-res crop + the confirmed label** to a new `POST /capture` on cardvision. Why this is nearly free in the app: `cropListerPhotoUri` already produces a **full-res crop**; `downscaleForCardvision` only makes the 1100px copy for `/identify`. Both versions already exist — foreground sends the small one, background sends the full one. No extra image processing. Why it's better than capturing `/identify` bytes: - **Archival fidelity** — solves the "lossy / un-redoable" risk; if we ever move to a higher-res matcher or want to re-derive references, the pixels exist. - **Simpler labeling** — upload happens *after* the card is confirmed, so image + label arrive together in one request. No orphan captures, no separate `/capture/label` round-trip to reconcile. (Abandoned scans upload nothing, which is correct — no label.) Decisions this forces: 1. **Bandwidth** — full-res crops are ~1–5MB vs ~150KB. Prefer **wifi**; queue on cellular and flush on wifi. Best-effort, never blocks the scan or UI. 2. **Storage growth** — unbounded ≈ 20GB/day at 10k scans. Server-side **per-card cap + dedup + retention window** so we don't hoard the 51st view of a common card. 3. **Crop quality still matters** — full-res doesn't fix a bad auto-crop; the curation gate stays. Railway: unchanged — the background upload goes phone → cardvision (Unraid) directly, never through Railway. This supersedes "capture the /identify bytes" in Phase 1 and folds in the old `/capture/label` step (label now travels with the full-res upload).
gambit-admin added the claimed:kris label 2026-08-28 20:02:41 +00:00
Author
Owner

Phase 1 built (capture + label) — mobile committed, Tower deploy staged

Mobile (branch feature/496-cardvision-training-capture, commit 9b8fc70): at commit
(Buy / inventory batch / direct inventory), each scanned card's full-res crop +
user-confirmed tcgplayer id is background-uploaded to cardvision POST /train-capture.
Commit is the ground-truth moment — variant picks are resolved by then; slabs, sealed,
non-scan entries, and unresolved picks never train, so discarded wrong matches can't
poison the future index. Fire-and-forget: awaited cost is one local file copy. Gate logic
pinned by lib/scan/trainingCapture.test.js; mobile tsc green.

Server (agg.py updated in ~/ai/cardvision, NOT yet on Tower): /train-capture on the
aggregator (8220) — token tripwire (token form field must match TRAIN_TOKEN), per-card
cap 10, stores captures/labeled/<card_id>/<ts>_<id>.jpg + labeled.jsonl;
/capture-stats now also reports labeled counts.

Tower deploy steps (agent is classifier-blocked from writing to the box):

  1. scp ~/ai/cardvision/agg.py root@192.168.1.95:/mnt/user/appdata/cardvision/agg.py
  2. Add to lb.conf + lb-scale.conf + lb-demo.conf inside the server {} block (no basic auth — token-gated in the app):
    location = /train-capture { proxy_pass http://127.0.0.1:8220; proxy_read_timeout 30s; }
  3. docker exec cardvision-lb nginx -t && docker restart cardvision-dash cardvision-lb
  4. Smoke: bad token → 403; -F image=@warm.jpg -F card_id=TEST123 -F token=<TRAIN_TOKEN> → {"stored": true}; then rm -rf captures/labeled + restart dash to clear the test row.

Wifi-only gating skipped (no netinfo dep; ~1–3MB per committed scan, best-effort). Ingest
(enroll into the index) stays Phase 3 behind the calibrate gate.

### Phase 1 built (capture + label) — mobile committed, Tower deploy staged **Mobile (branch `feature/496-cardvision-training-capture`, commit 9b8fc70):** at commit (Buy / inventory batch / direct inventory), each scanned card's full-res crop + user-confirmed tcgplayer id is background-uploaded to cardvision `POST /train-capture`. Commit is the ground-truth moment — variant picks are resolved by then; slabs, sealed, non-scan entries, and unresolved picks never train, so discarded wrong matches can't poison the future index. Fire-and-forget: awaited cost is one local file copy. Gate logic pinned by `lib/scan/trainingCapture.test.js`; mobile tsc green. **Server (agg.py updated in ~/ai/cardvision, NOT yet on Tower):** `/train-capture` on the aggregator (8220) — token tripwire (`token` form field must match TRAIN_TOKEN), per-card cap 10, stores `captures/labeled/<card_id>/<ts>_<id>.jpg` + `labeled.jsonl`; `/capture-stats` now also reports labeled counts. **Tower deploy steps (agent is classifier-blocked from writing to the box):** 1. `scp ~/ai/cardvision/agg.py root@192.168.1.95:/mnt/user/appdata/cardvision/agg.py` 2. Add to lb.conf + lb-scale.conf + lb-demo.conf inside the `server {}` block (no basic auth — token-gated in the app): `location = /train-capture { proxy_pass http://127.0.0.1:8220; proxy_read_timeout 30s; }` 3. `docker exec cardvision-lb nginx -t && docker restart cardvision-dash cardvision-lb` 4. Smoke: bad token → 403; `-F image=@warm.jpg -F card_id=TEST123 -F token=<TRAIN_TOKEN>` → `{"stored": true}`; then `rm -rf captures/labeled` + restart dash to clear the test row. Wifi-only gating skipped (no netinfo dep; ~1–3MB per committed scan, best-effort). Ingest (enroll into the index) stays Phase 3 behind the calibrate gate.
Author
Owner

Server side is LIVE on Tower: /train-capture deployed (agg.py + all 3 lb confs), smoke-tested — bad token 403 via public URL, stores + per-card counts survive restart, test rows cleaned. Scan path unaffected (healthz 200 via LB and public). Mobile branch feature/496-cardvision-training-capture awaits PR.

Server side is LIVE on Tower: /train-capture deployed (agg.py + all 3 lb confs), smoke-tested — bad token 403 via public URL, stores + per-card counts survive restart, test rows cleaned. Scan path unaffected (healthz 200 via LB and public). Mobile branch feature/496-cardvision-training-capture awaits PR.
Author
Owner

Phase 1 SHIPPED + a Phase 3 requirement

Phase 1 is live end to end: PR #337 merged (main a08b249), OTA on runtime 1.3.6, /train-capture live on Tower. Committed scans now accumulate in captures/labeled/.

Phase 3 requirement (from review discussion): ingest-time auto-verify. The picker is a mis-tap surface — a user can confirm the wrong card (worst case a wrong-language twin, the exact poison #489 fought). Same-art print-variant mistaps are near-harmless visually, but cross-card/cross-language mislabels are not. So before any capture is ingested as a reference row: run it back through cardvision itself; if the claimed card_id is not among the visual top-k candidates, do NOT auto-ingest — route it to the human review queue with the mismatch shown. Cheap (one /identify pass per capture at ingest time) and it automatically catches "tapped a completely different card".

### Phase 1 SHIPPED + a Phase 3 requirement Phase 1 is live end to end: PR #337 merged (main a08b249), OTA on runtime 1.3.6, /train-capture live on Tower. Committed scans now accumulate in captures/labeled/. **Phase 3 requirement (from review discussion): ingest-time auto-verify.** The picker is a mis-tap surface — a user can confirm the wrong card (worst case a wrong-language twin, the exact poison #489 fought). Same-art print-variant mistaps are near-harmless visually, but cross-card/cross-language mislabels are not. So before any capture is ingested as a reference row: run it back through cardvision itself; if the claimed card_id is not among the visual top-k candidates, do NOT auto-ingest — route it to the human review queue with the mismatch shown. Cheap (one /identify pass per capture at ingest time) and it automatically catches "tapped a completely different card".
Author
Owner

LIVE ENROLL SHIPPED (Phase 3, gated) — the flywheel is closed

Cardvision now learns in near-real-time from confirmed scans. Deployed to the Tower fleet tonight:

How it works: /train-capture (labeled capture at commit/accept) now forwards each photo to a worker's new token-gated /enroll. The worker runs a 1-vs-1 verify gate — ORB/RANSAC inliers of the photo against the CLAIMED card's own stock art, both orientations, floor 20 inliers (CARDVISION_ENROLL_INLIER_MIN). Pass → the photo becomes a live reference row: one self-contained .npz in a shared rw mount (enrolled/), embedding + index-side ORB descriptors, grouped with the card's stock rows so real views never hurt the card's own margin. All 4 workers poll the dir (5s) and atomically swap in an OverlayIndex — no rebuild, no restart. Fail → the capture STAYS in captures/labeled/ flagged enrolled:false as the human-review queue (hard cases like foil-textured cards whose stock scan doesn't match reality land here — the cards that need real references most).

Why the gate is sound even though recognition missed: recognition is 1-vs-87k with a margin rule — misses are usually the right card scoring well but a twin too close (Ursaluna: 88 vs 58 inliers → cloud). Verification asks only "does this photo match the named card's art" — no competition. A mislabeled photo is different artwork and scores ~0-15.

E2E proof on prod: wrong label (Tangrowth photo claimed as Alakazam) → verify_failed at 15 inliers, held for review. True label → enrolled at 78 inliers, all 4 workers picked it up within one poll. Re-identify then matched via the REAL reference at 281 inliers (vs 119 stock). Within minutes of deploy, 5 organic commits flowed through: 4 enrolled live, 1 blurry duplicate correctly held back.

Limits/controls: cap 5 enrolled views/card, token tripwire, employees-only traffic, rollback = delete the .npz (workers drop it next poll) or enrolled/ wholesale. Periodic calibrate regression audit still recommended (Phase 4). Cosmetic known issue: enrolled rows leak their meta (real/ts/source/group_rows) into /identify candidate JSON — harmless to the app.

Loose end: 9 pre-deploy captures in captures/labeled/ never attempted enroll (rows without an enrolled field) — backfill by POSTing them to /enroll, or leave for the review surface.

### LIVE ENROLL SHIPPED (Phase 3, gated) — the flywheel is closed Cardvision now learns in near-real-time from confirmed scans. Deployed to the Tower fleet tonight: **How it works:** /train-capture (labeled capture at commit/accept) now forwards each photo to a worker's new token-gated `/enroll`. The worker runs a **1-vs-1 verify gate** — ORB/RANSAC inliers of the photo against the CLAIMED card's own stock art, both orientations, floor 20 inliers (`CARDVISION_ENROLL_INLIER_MIN`). Pass → the photo becomes a live reference row: one self-contained .npz in a shared rw mount (`enrolled/`), embedding + index-side ORB descriptors, grouped with the card's stock rows so real views never hurt the card's own margin. All 4 workers poll the dir (5s) and atomically swap in an OverlayIndex — no rebuild, no restart. Fail → the capture STAYS in captures/labeled/ flagged `enrolled:false` as the human-review queue (hard cases like foil-textured cards whose stock scan doesn't match reality land here — the cards that need real references most). **Why the gate is sound even though recognition missed:** recognition is 1-vs-87k with a margin rule — misses are usually the right card scoring well but a twin too close (Ursaluna: 88 vs 58 inliers → cloud). Verification asks only "does this photo match the named card's art" — no competition. A mislabeled photo is different artwork and scores ~0-15. **E2E proof on prod:** wrong label (Tangrowth photo claimed as Alakazam) → verify_failed at 15 inliers, held for review. True label → enrolled at 78 inliers, all 4 workers picked it up within one poll. Re-identify then matched via the REAL reference at 281 inliers (vs 119 stock). Within minutes of deploy, 5 organic commits flowed through: 4 enrolled live, 1 blurry duplicate correctly held back. **Limits/controls:** cap 5 enrolled views/card, token tripwire, employees-only traffic, rollback = delete the .npz (workers drop it next poll) or `enrolled/` wholesale. Periodic `calibrate` regression audit still recommended (Phase 4). Cosmetic known issue: enrolled rows leak their meta (real/ts/source/group_rows) into /identify candidate JSON — harmless to the app. **Loose end:** 9 pre-deploy captures in captures/labeled/ never attempted enroll (rows without an `enrolled` field) — backfill by POSTing them to /enroll, or leave for the review surface.
Author
Owner

Live-enroll incident + fix (same evening): cross-photo noise, resolved

Kris's first real session exposed a defect: same-session sleeved photos cross-match at 15-25 spurious ORB inliers (sleeve texture/desk, a noise fit — ratio ~0.1-0.2), so one card's enrolled reference could tank ANOTHER card's margin (more cloud fallbacks) or even win outright (an Elder Dragon scan matched Dragon's Rage's enrolled row). Meanwhile a TRUE photo-vs-photo match of the same card lands 90-280 inliers — the bands don't overlap.

Fixes (deployed + verified on the failing scans):

  • CARDVISION_REAL_ROW_MIN=40: an enrolled row's verification below 40 inliers is discarded outright — below the floor a real row can't win, can't be a margin runner-up, can't crowd candidates. Stock-only behavior is untouched.
  • Reference border-trim (CARDVISION_ENROLL_TRIM=0.07): enrolled references store only the card interior.
  • Stock co-verify: when CLIP picks a real row as a group's verify representative, the group's best stock row is verified too, so a weak photo can't eat its own card's only slot.
  • Real rows excluded from the print-ambiguity list (they're extra views, not competing prints), and a real winner's ratio gate uses max(own, best stock member's).
  • All enrollments wiped + re-enrolled from stored captures through the fixed pipeline: 13 references / 11 cards live, 4 verify-fails in the review queue — including one suspected wrong-print commit (a capture labeled 672815 whose photo verifies as its 672817 Manga sibling at 40 inl but scores 6 against its own claimed print): the gate refusing that enroll is the poison-protection working.

Re-tested the exact failing scans: Elder Dragon → Elder Dragon (42/62 inl via its real reference), Premonition → itself (228), no cross-card wins.

### Live-enroll incident + fix (same evening): cross-photo noise, resolved Kris's first real session exposed a defect: same-session sleeved photos cross-match at 15-25 spurious ORB inliers (sleeve texture/desk, a noise fit — ratio ~0.1-0.2), so one card's enrolled reference could tank ANOTHER card's margin (more cloud fallbacks) or even win outright (an Elder Dragon scan matched Dragon's Rage's enrolled row). Meanwhile a TRUE photo-vs-photo match of the same card lands 90-280 inliers — the bands don't overlap. Fixes (deployed + verified on the failing scans): - **CARDVISION_REAL_ROW_MIN=40**: an enrolled row's verification below 40 inliers is discarded outright — below the floor a real row can't win, can't be a margin runner-up, can't crowd candidates. Stock-only behavior is untouched. - **Reference border-trim** (CARDVISION_ENROLL_TRIM=0.07): enrolled references store only the card interior. - **Stock co-verify**: when CLIP picks a real row as a group's verify representative, the group's best stock row is verified too, so a weak photo can't eat its own card's only slot. - **Real rows excluded from the print-ambiguity list** (they're extra views, not competing prints), and a real winner's ratio gate uses max(own, best stock member's). - All enrollments wiped + re-enrolled from stored captures through the fixed pipeline: 13 references / 11 cards live, 4 verify-fails in the review queue — including one suspected wrong-print commit (a capture labeled 672815 whose photo verifies as its 672817 Manga sibling at 40 inl but scores 6 against its own claimed print): the gate refusing that enroll is the poison-protection working. Re-tested the exact failing scans: Elder Dragon → Elder Dragon (42/62 inl via its real reference), Premonition → itself (228), no cross-card wins.
Author
Owner

Misaligned-scan rescue deployed (no app change needed)

Kris's field report: ignoring the align guide (card far from the reticle) made scans miss. Root cause: the app sends its reticle crop with already_cropped=1, so the engine trusted the flag, squashed a mostly-desk crop to card shape, and matched noise — and the CLIP shortlist itself was poisoned by the desk pixels, so every fallback stage verified only wrong candidates (the deep sweep that used to cover this is deliberately disabled for latency).

Fix, server-side only: already_cropped is now a HINT with a graceful fall-through — (1) stage B2 gives a failed cropped query the aspect-preserving hi-res frame path + match-guided refinement; (2) the contour-quad detector now also runs on cropped queries (free when aligned — a border-touching card yields no quads); (3) NEW quad_pass: when nothing matched, the best quad crop is re-embedded for a fresh shortlist and verified — geometry-guided rescue that works when the shortlist was poisoned.

Verified: synthetic far-away crop (card at 40% of frame on a noisy desk) went low_confidence/wrong-16-inliers → match, 104 inliers, ratio 0.92, margin 6.5, quad_pass=true, 469ms. Aligned fast path unchanged (121ms). Note: a rescued scan lands near the app's 500ms bail budget — if far scans still occasionally fall to cloud on real phones (network adds ~50-100ms), bump CARDVISION_BUDGET_MS in the app (~900ms) via OTA.

Native on-device auto-crop (tap → Vision rectangle detect → send tight crop) remains a possible future UX upgrade (needs a native build); the server rescue removes the accuracy cliff without it.

### Misaligned-scan rescue deployed (no app change needed) Kris's field report: ignoring the align guide (card far from the reticle) made scans miss. Root cause: the app sends its reticle crop with already_cropped=1, so the engine trusted the flag, squashed a mostly-desk crop to card shape, and matched noise — and the CLIP shortlist itself was poisoned by the desk pixels, so every fallback stage verified only wrong candidates (the deep sweep that used to cover this is deliberately disabled for latency). Fix, server-side only: already_cropped is now a HINT with a graceful fall-through — (1) stage B2 gives a failed cropped query the aspect-preserving hi-res frame path + match-guided refinement; (2) the contour-quad detector now also runs on cropped queries (free when aligned — a border-touching card yields no quads); (3) NEW quad_pass: when nothing matched, the best quad crop is re-embedded for a fresh shortlist and verified — geometry-guided rescue that works when the shortlist was poisoned. Verified: synthetic far-away crop (card at 40% of frame on a noisy desk) went low_confidence/wrong-16-inliers → **match, 104 inliers, ratio 0.92, margin 6.5, quad_pass=true, 469ms**. Aligned fast path unchanged (121ms). Note: a rescued scan lands near the app's 500ms bail budget — if far scans still occasionally fall to cloud on real phones (network adds ~50-100ms), bump CARDVISION_BUDGET_MS in the app (~900ms) via OTA. Native on-device auto-crop (tap → Vision rectangle detect → send tight crop) remains a possible future UX upgrade (needs a native build); the server rescue removes the accuracy cliff without it.
Author
Owner

Reference hygiene: enrolled references are now card-localized (tight crops)

Decision (Kris + agent): index references must be the CARD, not the scene — background robustness is the localizer's job; the on-card variation (glare/lighting/wear) is what real references exist to capture. The ORIGINAL uncropped capture stays in captures/labeled/, so a future neural-detector approach that wants messy-background data keeps its raw material; only the derived index reference is cropped.

/enroll now: (1) the verify GATE tries the whole capture AND quad-detected card warps, both orientations — so far/misaligned captures (rescued by quad_pass at scan time) can enroll instead of dying at the door; (2) the stored reference is chosen among border-trim + quad warps + nested quads (sleeve-edge case) by whichever verifies best against the claimed card's own stock art; (3) each enrollment saves its exact reference view as a .jpg beside the .npz for audit/review.

Also from tonight's re-runs: the identity-conflict cross-check retro-caught a second wrong-label commit (90452 vs 611747) which on inspection is the EN/JP Wailmer art-twin pair — with the stronger quad-aware gate the photo proves 83 inliers against the user's chosen EN print, so it enrolls under the user's label (the designed twin behavior; the app's language picker remains the EN/JP defense). Store rebuilt end-to-end: 30 tight references live across the fleet. Verified visually: a synthetic far-away desk photo enrolls as a clean card-localized color reference at 88 inliers.

### Reference hygiene: enrolled references are now card-localized (tight crops) Decision (Kris + agent): index references must be the CARD, not the scene — background robustness is the localizer's job; the on-card variation (glare/lighting/wear) is what real references exist to capture. The ORIGINAL uncropped capture stays in captures/labeled/, so a future neural-detector approach that wants messy-background data keeps its raw material; only the derived index reference is cropped. /enroll now: (1) the verify GATE tries the whole capture AND quad-detected card warps, both orientations — so far/misaligned captures (rescued by quad_pass at scan time) can enroll instead of dying at the door; (2) the stored reference is chosen among border-trim + quad warps + nested quads (sleeve-edge case) by whichever verifies best against the claimed card's own stock art; (3) each enrollment saves its exact reference view as a .jpg beside the .npz for audit/review. Also from tonight's re-runs: the identity-conflict cross-check retro-caught a second wrong-label commit (90452 vs 611747) which on inspection is the EN/JP Wailmer art-twin pair — with the stronger quad-aware gate the photo proves 83 inliers against the user's chosen EN print, so it enrolls under the user's label (the designed twin behavior; the app's language picker remains the EN/JP defense). Store rebuilt end-to-end: 30 tight references live across the fleet. Verified visually: a synthetic far-away desk photo enrolls as a clean card-localized color reference at 88 inliers.
gambit-admin added the 01 · scan & recognition label 2026-09-04 13:36:52 +00:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: gambit/gambit#496