Binder page scan: 9-pocket live-view crop → parallel CardVision identify #692

Open
opened 2026-09-12 15:44:31 +00:00 by gambit-admin · 0 comments
Owner

Kris's spec (2026-09-12): the user holds the phone over a 9-pocket binder page, a fixed 3×3 guide is drawn on the live camera, they line the page up physically and tap once. No detection, no post-capture crop adjustment. The phone crops the 9 pockets by the guide alone, fires 9 /identify calls in parallel, and fills a 3×3 result grid as results land. Target: whole page done in < 2 s.

Bench evidence (SlabTest /slabtest/binder, 2 real Crown Zenith pages, 18 cards): guide-only crop framed every card upright even with the guide a few % off; 9 parallel identifies complete in 1.0–1.6 s wall; cells at 640px/q75 (~50 KB each, ~450 KB/page) match the accuracy of 1100px cells; tight crop (0 bleed) slightly beats the single-scan 5% bleed; CardVision's own in-cell detection gains nothing and costs +50%. Auto-accept 6/9 per page today, 8/9 once the two CardVision gate issues (blurry false-reject; 40–60-inlier margin) land.

Scope (mobile, mobile/rork-recon-card-scanner/expo/):

  • Binder mode entry from the scan screen; fixed 3×3 overlay on the live preview (portrait page).
  • Tap → capture → 9 guide-rect crops (reuse cropListerPhotoUri / photoCrop helpers; 0 bleed; ~640–800 px max edge) → 9 concurrent /identify (already_cropped=1) → /scan/cardvision-resolve (raise its 4-id cap or batch 3 calls).
  • 3×3 result grid: match / low / blurry / error per cell, tap a red cell to retake that pocket via the normal single-card flow; accept-all hands 9 results into the existing commit path.
  • Skip the per-scan language OCR probe + training capture for binder cells (or throttle) so a page doesn't fire 9× side effects.

Related: SlabTest bench page + run log under ~/.slabtest/runs/binder-*.


Filed from a Claude Code session

Kris's spec (2026-09-12): the user holds the phone over a 9-pocket binder page, a fixed 3×3 guide is drawn on the live camera, they line the page up physically and tap once. No detection, no post-capture crop adjustment. The phone crops the 9 pockets **by the guide alone**, fires 9 `/identify` calls in parallel, and fills a 3×3 result grid as results land. Target: whole page done in **< 2 s**. **Bench evidence** (SlabTest `/slabtest/binder`, 2 real Crown Zenith pages, 18 cards): guide-only crop framed every card upright even with the guide a few % off; 9 parallel identifies complete in 1.0–1.6 s wall; cells at 640px/q75 (~50 KB each, ~450 KB/page) match the accuracy of 1100px cells; tight crop (0 bleed) slightly beats the single-scan 5% bleed; CardVision's own in-cell detection gains nothing and costs +50%. Auto-accept 6/9 per page today, 8/9 once the two CardVision gate issues (blurry false-reject; 40–60-inlier margin) land. **Scope (mobile, `mobile/rork-recon-card-scanner/expo/`):** - Binder mode entry from the scan screen; fixed 3×3 overlay on the live preview (portrait page). - Tap → capture → 9 guide-rect crops (reuse `cropListerPhotoUri` / photoCrop helpers; 0 bleed; ~640–800 px max edge) → 9 concurrent `/identify` (already_cropped=1) → `/scan/cardvision-resolve` (raise its 4-id cap or batch 3 calls). - 3×3 result grid: match / low / blurry / error per cell, tap a red cell to retake that pocket via the normal single-card flow; accept-all hands 9 results into the existing commit path. - Skip the per-scan language OCR probe + training capture for binder cells (or throttle) so a page doesn't fire 9× side effects. Related: SlabTest bench page + run log under `~/.slabtest/runs/binder-*`. --- Filed from a Claude Code session
gambit-admin added the featureuiclaimed:kris labels 2026-09-12 15:45:09 +00:00
Sign in to join this conversation.