name_search corrupts Japanese names: diacritic stripper drops dakuten/handakuten (イーブイ → イーフイ) on 14.5k rows #579

Open
opened 2026-09-02 23:09:40 +00:00 by gambit-admin · 1 comment
Owner

What

card_search._strip_diacritics (backend/app/services/card_search.py:45-53) normalizes to NFD and drops every char whose Unicode category is Mn. In NFD, voiced katakana/hiragana decompose into base kana + a combining dakuten (U+3099) or handakuten (U+309A) — both Mn — so they are deleted:

  • イーブイ → イーフイ (the #578 row)
  • ポケモン → ホケモン, リザードン → リサートン, ピカチュウ → ヒカチュウ

Every writer routes through _name_search: scrydex_canonical_ingest.py:202,214, tcgcsv_canonical_ingest.py:274,285, canonical_write.py:279, justtcg_history_rescue.parse_slug. So canonical_card.name_search AND card_cache.name_search are wrong, and the trigram GIN index (main.py:421) is built over the corrupted text.

Prod impact (measured 2026-09-02)

  • 14,535 of 22,035 JP pokemon card_cache rows have name_search <> lower(name) (every name containing a voiced kana). Plus 10 stray rows in lorcana/gundam/one-piece.
  • Kana search still appears to work because the reader (routes/scan.py:72-76 _scan_name_search) applies the same strip to the query — but it uses NFKD while the writer uses NFD, so reader/writer already diverge on compatibility chars, and ブ/フ, ポ/ホ, バ/ハ etc. now collide.
  • Anything that compares names exactly against name_search fails for JP: the graded ladder's name alignment (external_scan.py tier a/a2), _attach_existing_canonical (canonical_write.py:191).

Fix shape (do not fix yet — scoping only)

  1. Make the stripper skip U+3099/U+309A (or only strip Mn for Latin base chars), and have reader + writer share ONE normalizer.
  2. Backfill: UPDATE canonical_card SET name_search = f(name) then reproject_all (or a direct card_cache UPDATE + REINDEX of the trgm index).
  3. Regression test with the four examples above.

Found during the 2026-09-02 card-data review. Related: #578 (one of its three compounding causes).


Filed from a Claude Code session

## What `card_search._strip_diacritics` (`backend/app/services/card_search.py:45-53`) normalizes to **NFD** and drops every char whose Unicode category is `Mn`. In NFD, voiced katakana/hiragana decompose into base kana + a combining dakuten (U+3099) or handakuten (U+309A) — both `Mn` — so they are deleted: - イーブイ → イーフイ (the #578 row) - ポケモン → ホケモン, リザードン → リサートン, ピカチュウ → ヒカチュウ Every writer routes through `_name_search`: `scrydex_canonical_ingest.py:202,214`, `tcgcsv_canonical_ingest.py:274,285`, `canonical_write.py:279`, `justtcg_history_rescue.parse_slug`. So `canonical_card.name_search` AND `card_cache.name_search` are wrong, and the trigram GIN index (`main.py:421`) is built over the corrupted text. ## Prod impact (measured 2026-09-02) - **14,535 of 22,035** JP pokemon `card_cache` rows have `name_search <> lower(name)` (every name containing a voiced kana). Plus 10 stray rows in lorcana/gundam/one-piece. - Kana search still *appears* to work because the reader (`routes/scan.py:72-76 _scan_name_search`) applies the same strip to the query — but it uses **NFKD** while the writer uses NFD, so reader/writer already diverge on compatibility chars, and ブ/フ, ポ/ホ, バ/ハ etc. now collide. - Anything that compares names exactly against `name_search` fails for JP: the graded ladder's name alignment (`external_scan.py` tier a/a2), `_attach_existing_canonical` (`canonical_write.py:191`). ## Fix shape (do not fix yet — scoping only) 1. Make the stripper skip U+3099/U+309A (or only strip `Mn` for Latin base chars), and have reader + writer share ONE normalizer. 2. Backfill: `UPDATE canonical_card SET name_search = f(name)` then `reproject_all` (or a direct `card_cache` UPDATE + REINDEX of the trgm index). 3. Regression test with the four examples above. Found during the 2026-09-02 card-data review. Related: #578 (one of its three compounding causes). --- Filed from a Claude Code session
gambit-admin added the bugapi01 · scan & recognition labels 2026-09-02 23:09:40 +00:00
gambit-admin added the claimed:lucas label 2026-09-03 20:02:16 +00:00
Author
Owner

STATUS — PR Gambit-Inc/gambit#398 open, verify green, changes requested.

Blocker found in review: routes/public_catalog.py:34 and routes/api_explorer.py:24 still
hold byte-identical copies of the pre-fix stripper, applied to the user's query and then
ILIKE'd against card_cache.name_search. Today both sides are consistently corrupted so JP
search works; fixing only the write side turns q=ポケモン into a guaranteed zero-row result
on the public catalog. Two one-line imports of the shared _name_search.

Also established in review, and it changes this issue's urgency: the 14,535 rows self-heal.
scrydex_canonical_ingest.py:201 reassigns name_search unconditionally on every resolve,
and the daily pull_all re-walk covers pokemon (EN+JP). Exposure is one sweep, ≤24h — the
backfill script is an accelerator, not a gate on the promote. Lucas pushed review fixes at
20:07 UTC.

STATUS — PR Gambit-Inc/gambit#398 open, `verify` green, **changes requested**. Blocker found in review: `routes/public_catalog.py:34` and `routes/api_explorer.py:24` still hold byte-identical copies of the pre-fix stripper, applied to the **user's query** and then ILIKE'd against `card_cache.name_search`. Today both sides are consistently corrupted so JP search works; fixing only the write side turns `q=ポケモン` into a guaranteed zero-row result on the public catalog. Two one-line imports of the shared `_name_search`. Also established in review, and it changes this issue's urgency: **the 14,535 rows self-heal.** `scrydex_canonical_ingest.py:201` reassigns `name_search` unconditionally on every resolve, and the daily `pull_all` re-walk covers `pokemon` (EN+JP). Exposure is one sweep, ≤24h — the backfill script is an accelerator, not a gate on the promote. Lucas pushed review fixes at 20:07 UTC.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: gambit/gambit#579