diff --git a/plans/260908-1525-dictionary-corpus-switch/audit-proper-noun-drops.md b/plans/260908-1525-dictionary-corpus-switch/audit-proper-noun-drops.md new file mode 100644 index 0000000..bfaea77 --- /dev/null +++ b/plans/260908-1525-dictionary-corpus-switch/audit-proper-noun-drops.md @@ -0,0 +1,184 @@ +# Audit: proper-noun drops and the wiktionary-only corpus + +Date: 2026-09-08. Phase 3 of `plan.md`. Every number here is reproducible from the +commands listed; the shipped `data/noitu.db` was copied aside before any build and is the +"old" side of every comparison. + +## Commands + +```sh +# from server/, after `make fetch-dict` +go run ./cmd/build-dictionary --merged ../data/undertheseanlp-words.jsonl \ + --out /tmp/wik.db --report-drops /tmp/drops-wik.txt +go run ./cmd/build-dictionary --merged ../data/undertheseanlp-words.jsonl \ + --sources hongocduc,wiktionary --out /tmp/hw.db --report-drops /tmp/drops-hw.txt +# sample: python, random.seed(1), random.sample(sorted(drops-wik), 100) +# casualties: (drops-wik − drops-hw) ∩ words(hw.db) +go test ./internal/bot/ -run RealCorpus -v # once per database at data/noitu.db +``` + +## Build result + +``` +rejected 97: contains punctuation +rejected 5074: fewer than 2 syllables +rejected 37: no Vietnamese letters +rejected 4747: only ever capitalized +accepted 22419 distinct words (sources: wiktionary) +``` + +22,419 vs the 22,310 the research report measured. The delta is the rule: the report +dropped any form containing an uppercase letter, the builder drops forms that *begin* with +one, so `tia X`, `bộ bài tây`, `chữ Hán`, `châu Á` survive. All are ordinary words. + +## 1. Drop audit — 100 sampled, seed 1 + +Verdicts: **P** proper noun (correct drop) · **X** not a Vietnamese word (would be rejected +by the filter anyway) · **U** unclear · **C** common word (false positive). + +| | | | | | +|---|---|---|---|---| +| a yun P | an hoà tây P | an minh bắc P | an thuận P | ba vinh P | +| ban cơ U | brao X | bàn đạt P | bành trạch P | bàu năng P | +| bá xuyên P | bình hiệp P | bình tấn P | bằng la P | chiềng sơ P | +| châu hội P | chù lá phù lá P | chư krêy P | cur X | côn đảo P | +| cơ kiều P | cẩm trung P | hùng vương P | hồng bàng P | irc X | +| khánh gia P | lão quân P | lương đài P | mặc dương P | nhơn hoà lập P | +| ninh lai P | ninh thạnh P | noong bua P | nội thôn U | phiếu hữu mai U | +| phù lảng P | phương cao kén ngựa U | phần lão U | quán hành P | quắc hương P | +| rã bản P | sam mứn P | sơn hạ P | sơn vy P | tam ngọc P | +| thanh liên P | thanh văn P | thiệu đô P | thái an P | thượng lâm P | +| thượng tiến P | thạch xuân P | thới quản P | tiên cát P | tiên thọ P | +| tri lễ P | triệu giang P | triệu việt vương P | trung tú P | trà khê P | +| trà kót P | trà linh P | trường khánh P | trường long P | trần đình thâm P | +| trọng do P | tuân lộ P | tân hưng P | tân mỹ P | tân nhựt P | +| tân trì P | tây hiếu P | tăng sâm P | tĩnh gia P | tịnh an P | +| tốt động P | vinh an P | việt–mường X | vân canh P | võ trường toản P | +| võ văn tồn P | văn môn P | văn quân P | văn đình dận P | vĩnh hoà hưng bắc P | +| vĩnh lương P | vĩnh lộc b P | vĩnh trinh P | vũ thư P | vạn thuỷ P | +| xuân mai P | xuân quan P | xá bung P | xá cẩu P | xá khắc P | +| xốp cộp P | yên cát P | ô qua U | đồng nai P | đổ rượu ra sông thết quân lính C | + +**P 89 · X 4 · U 6 · C 1.** Gate is ≤2 common words in 100: **passes.** The bulk is +commune names, ethnonyms and historical persons — the content the rule exists to remove. + +## 2. Case casualties — 215, listed in full + +Words the wiktionary rows hold only capitalized, which another branch has lowercase. They +are dropped under the plan's wiktionary-only evidence rule. Full list: +`case-casualties.txt` in the session scratchpad; reproduced here because it is the +decision input. + +an bình · an dân · an hảo · an khang · an lạc · ba tiêu · biển hồ · bu lu · **bàn là** · +**báo đáp** · bát tiên · bình chuẩn · bình chân · bình khang · bình nghị · bình sa · bình +thanh · bình thuỷ · bình trị · bình tâm · bình văn · bình điền · bích đào · bóng chim tăm +cá · bông trang · bạch cung · bản nguyên · **bảo toàn** · bắc phong · bẻ quế · bố chính · +**bồ đề** · bồng sơn · cao biền dậy non · cao kỳ · cao nhân · cao sơn · cao tổ · cao xanh · +cao đường · chà và · chày sương · chánh hội · **chân mây** · châu lệ · châu thành · chí +thiện · chí thành · chính tâm · **chúa nhật** · chúa trời · chị hằng · con tạo · cát lũy · +cát tân · cát đằng · **công bình** · công dã tràng · **công giáo** · **công nguyên** · +**cơ đốc giáo** · cường thịnh · cầm đuốc chơi đêm · cầu lam · cầu lộc · cầu ô · cẩm châu · +cẩm đường · cửu kinh · **cựu ước** · diêm vương · diêm vương tinh · dã hạc · dương quan · +dương đài · **dường như** · **giao tử** · giấc mai · giọt châu · giọt tương · hoa cái · +**hoa kiều** · **hoà đồng** · hoá công · hoả tinh · hán học · hán tộc · hán tự · hán văn · +hạ thần · hải phòng · hải triều · hải vương tinh · hằng nga · **hết sảy** · **hệ mặt +trời** · hội gió mây · hội long vân · **hồi giáo** · **khổng giáo** · kim tinh · **kinh +thánh** · la ve · lâm viên · **lưỡi hái** · lạc hầu · mạnh thường quân · **mặt trời** · nam +lâu · **nguyên đán** · **nho học** · nhân kiệt · năm cha ba mẹ · nước dương · nếm mật nằm +gai · nợ như chúa chổm · phong thu · **phù phiếm** · **phật giáo** · phật học · phật pháp · +phật tiền · phật tổ · phật tự · **phật đản** · quang phục · quyết tiến · quân thiều · **quả +đất** · **sao hôm** · sao hỏa · sao kim · **sao mai** · sao mộc · song mai · sơn cương · +sơn lâm · sơn mai · sơn nguyên · sơn trung · sư tử hà đông · **sừng trâu** · tam dân · tam +phủ · tam sơn · thanh nghị · thanh phong · thanh quang · thanh vận · thiên chúa · **thiên +chúa giáo** · thiên vương tinh · thuỷ liễu · thân giáp · thạch bàn · thạch thán · **thần +chết** · **tin lành** · tiên sư · **trái đất** · trướng huỳnh · trường giang · trường xuân · +tu lý · tuần duyên · tân kỳ · **tân ước** · tây dương · tây thiên · tô hạp · tùng lâm · tế +tân · **tết nguyên đán** · **tổ quốc** · tử phòng · u minh · vinh thăng · **việt ngữ** · +vách quế · vân hà · vân hán · vân trình · võ miếu · **văn giáo** · **văn miếu** · **văn +nhân** · văn quan · **văn võ** · văn đức · **vũ công** · **vũ hội** · vũng tàu · vương bá · +vạn an · vạn kiếp · vạn phúc · vầng ô · **vớ bở** · **xe tơ** · xuân hoá · xuân phong · +xuân tình · xích thố · xương thịnh · yên bình · yên chi · yên hoa · yên hà · đinh điền · +**đường luật** · **đường thi** · đại danh · đỉnh giáp non thần · **địa cầu** · đỗ vũ + +Bold: everyday or dictionary-headword vocabulary by this auditor's reading, ~55 of 215. +The rest are Sino-Vietnamese literary terms, religious/astronomical names that Wiktionary +capitalizes by convention, and genuine place names (`hải phòng`, `vũng tàu`, `bồng sơn`). + +**Narrowing does not fix this.** The plan's pre-decided narrowing — drop only when every +syllable is capitalized — was measured: drops fall 4,747 → 4,337, casualties fall only +215 → 147 (`mặt trời`, `trái đất`, `tổ quốc`-style two-cap forms stay dropped), and 410 +proper nouns come back in, including junk like `bản mẫu:-vie-n-` and `con kde`. Rejected. + +## 3. Corpus diff — old shipped vs new + +| | | +|---|---| +| overlap | 21,320 | +| gained | 1,099 (`gân bò`, `rèn đúc`, `ích kỷ`, `đáng lý`, `thở hồng hộc`) | +| lost | **26,896** | +| — dropped as proper noun | 4,245 | +| — absent from the wiktionary branch | 22,651 | +| — other | 0 | + +## 4. Playability + +| metric | old | new | +|---|---|---| +| words | 48,216 | 22,419 | +| syllables | 6,676 | 5,485 | +| syllables that open a word | 5,049 | 4,052 | +| …with ≥2 continuations | 3,682 | 2,691 | +| dead-end syllables | 1,627 | 1,433 | + +Bot-versus-bot, `internal/bot` real-corpus tests, 60 games each: + +| | old | new | +|---|---|---| +| hard-vs-easy win rate (moves) | 98% (3.3) | 98% (3.5) | +| medium-vs-easy | 87% | 88% | +| hard-vs-medium (moves) | 72% (3.5) | 55% (2.9) | +| easy-vs-easy game length | 17.5 moves | **11.3 moves** | +| hard decision p95 | 5.2 ms | 1.1 ms | + +Both tests pass on both databases. Easy-vs-easy games are a third shorter: the thinner +graph runs out of continuations sooner, which is the depth loss the plan predicted. +Hard-vs-medium falling to 55% means the graph gives the stronger bot fewer ways to trap. + +## 5. Decision + +Drop rule on the sample: **pass** (1 common word in 100). + +Case casualties: **open — needs the owner's call**, because the options trade the +"wiktionary data only" decision against ~55 everyday words: + +- **Accept** the 215 losses as measured. +- **Exception list**: a checked-in list of words to keep despite capitalization, curated + from the 215 above. Cause-aligned, small, but a hand-maintained artifact. +- **Read hongocduc for case evidence only**: recovers all 215 mechanically, but uses GPL + data as an input even though none of its words ship. Must be stated in `ATTRIBUTION.md`. + +**Recorded decision (owner, 2026-09-08): abandon the capitalization rule.** Every +wiktionary-tagged word is kept regardless of case; proper nouns stay in the corpus as they +are in the shipped database today. The rule, its reject reason and `--report-drops` were +removed from the builder. + +## 6. Final corpus, without the rule + +``` +accepted 26845 distinct words (sources: wiktionary) +``` + +| metric | old | new | +|---|---|---| +| words | 48,216 | 26,845 | +| syllables | 6,676 | 5,709 | +| syllables that open a word | 5,049 | 4,158 | +| …with ≥2 continuations | 3,682 | 2,787 | +| dead-end syllables | 1,627 | 1,551 | +| overlap / gained / lost | | 25,565 / 1,280 / 22,651 | + +Every lost word is absent from the 2018 wiktionary branch; none is lost to a rule. + +Bot-versus-bot, 60 games each: hard-vs-easy 95% (2.7 moves), medium-vs-easy 87%, +hard-vs-medium 63% (3.2 moves), easy-vs-easy **12.9 moves** (old: 17.5). Both real-corpus +tests pass. diff --git a/plans/260908-1525-dictionary-corpus-switch/phase-01-start.md b/plans/260908-1525-dictionary-corpus-switch/phase-01-start.md index e1a9107..a62fd56 100644 --- a/plans/260908-1525-dictionary-corpus-switch/phase-01-start.md +++ b/plans/260908-1525-dictionary-corpus-switch/phase-01-start.md @@ -1,7 +1,7 @@ --- phase: 1 title: "Phase 1: Read the merged list" -status: todo +status: done priority: P1 effort: "3h" dependencies: [] @@ -12,109 +12,107 @@ dependencies: [] ## Overview Teach `build-dictionary` a third input: undertheseanlp's merged JSONL. It selects words by -source membership, drops entries that appear only ever capitalized, and hands the survivors -to the normalization and filtering path that already exists. +source membership and hands them to the normalization and filtering path that already +exists. ## Requirements - Functional: read `{"text": "...", "source": ["hongocduc", ...]}` lines; keep a word when - its sources intersect the allowed set; drop a word when every raw form of it, across the - allowed sources, begins with an uppercase letter. -- Functional: `tudientv` rows are skipped before any decision is made about a word — - neither its words nor its capitalization inform the output. -- Functional: the proper-noun drop is counted and printed beside the existing reject - reasons, so the build log says what the source contained. -- Non-functional: streaming line-by-line, one pass to gather, one pass to decide. The file - is 4.8 MB; nothing here needs to be clever. + its sources intersect the allowed set — **`wiktionary` only** by default. +- Functional: `hongocduc` and `tudientv` rows are skipped before any decision is made about + a word. They inform nothing. +- Functional: capitalization is not a filter. `Hà Nội` and `hà nội` are the same word and + land as the lowercase form, exactly as every other input mode already behaves. +- Functional: the `--min-words` default moves from 40,000 to 20,000. The new corpus is + ~26,800 words; the old floor would reject it. +- Non-functional: streaming line-by-line. The file is 4.8 MB; nothing here needs to be + clever. - Non-functional: everything downstream — `accept()`, syllable indexing, alias generation, `verify()`, `meta` — is reused unchanged. ## Architecture -Two maps keyed by the lowercased, whitespace-collapsed word: - ``` -forms[key] -> set of raw spellings seen in allowed sources ("Mặt Trời", "mặt trời") -sources[key] -> set of allowed sources containing it {"hongocduc", "wiktionary"} +--merged --sources wiktionary + │ + ▼ one row at a time + sources ∩ allowed ≠ ∅ ? ── no ──▶ skipped, never counted + │ yes + ▼ + accept() (NFC, lowercase, ≥2 syllables, alphabet, phonotactics) + │ + ▼ + finish() (floor, aliases, atomic write, verify) ← shared with --words and --in ``` -A key survives when `sources[key]` is non-empty and at least one entry in `forms[key]` -does not start with an uppercase letter. `Mặt Trời` survives on the strength of its -lowercase twin; `Hà Nội`, `Trương Công Định` and `A Mú Sung` have no lowercase form and go. +The `--sources` flag is kept general — a comma list validated against the three known +names — even though only `wiktionary` is passed, so a comparison build against another +branch is a flag away and never a code patch. Unknown names are an error, not a no-op. -Surviving keys are then fed to the same `accept()` the SQLite path feeds, so the ≥2 -syllable rule, the digit/punctuation rejections and the Vietnamese phonotactic check all -apply as they do today. - -Case evidence deliberately comes from allowed sources only. The alternative — reading -`tudientv` rows for their capitalization while refusing their words — would be more -accurate and would undercut the claim that we do not use that data. Phase 3 measures what -the stricter choice costs. +`finish()` is new: the three input modes used to each carry their own copy of the floor +check, alias generation, write and verify. One copy is what keeps a fixture from drifting +into a different shape from the database production loads. ## Related Code Files -- Create: `server/cmd/build-dictionary/merged_list.go` -- Create: `server/cmd/build-dictionary/merged_list_test.go` -- Modify: `server/cmd/build-dictionary/main.go` — `--merged` and `--sources` flags, dispatch - in `run()`, `sourceURL`/`sourceLicense` constants, `meta` rows, package doc comment -- Modify: `server/cmd/build-dictionary/filter.go` — add the `rejectProperNoun` reason so - the drop is reported through the existing counter, not a bespoke log line +- Created: `server/cmd/build-dictionary/merged_list.go` +- Created: `server/cmd/build-dictionary/merged_list_test.go` +- Modified: `server/cmd/build-dictionary/main.go` — `--merged` and `--sources` flags, + mutually-exclusive input check, `finish()`, `sourceSpec.url` / `sourceSpec.extra`, meta + rows, package doc comment, `--min-words` default +- Unchanged: `server/cmd/build-dictionary/filter.go` ## Implementation Steps -1. Add `rejectProperNoun rejectReason = "only ever capitalized"` to `filter.go`. Leave - `accept()` alone — the case decision happens before it, on evidence `accept()` cannot - see, since it lowercases. -2. Write `merged_list.go`: decode with `encoding/json` per line (`bufio.Scanner` with a - raised buffer — the longest line is short, but a scanner that silently truncates is a - bad way to lose words), build the two maps, apply the rule, return words plus a count - per reject reason. -3. Add `--merged ` and `--sources hongocduc,wiktionary` to `main.go`. Unknown source - names are an error, not a silent no-op: a typo in `--sources` must not quietly ship an - empty or wrong corpus. -4. Dispatch in `run()`: `--merged` takes precedence in the same shape `--words` already - does. Refuse `--merged` together with `--in`/`--words` rather than picking one. -5. Point `sourceURL`, `sourceLicense` and the package doc comment at the new source. Write - `source_commit`, `sources_kept`, `sources_excluded` and `proper_nouns_dropped` into - `meta` — provenance in the artifact, not only in a markdown file. -6. Tests in `merged_list_test.go`, table-driven, over hand-written JSONL fixtures: - - `Mặt Trời` + `mặt trời` → kept - - `Hà Nội` alone → dropped as a proper noun - - a `tudientv`-only word → absent from the output - - a word whose only lowercase form is in `tudientv` → dropped (the documented cost) - - `học sinh` in all three → kept once, not thrice +1. Write `merged_list.go`: decode with `encoding/json` per line (`bufio.Scanner` with a + raised buffer — a scanner that silently truncates is a bad way to lose words), skip rows + with no allowed source, feed the rest to `accept()`. +2. Add `--merged ` and `--sources wiktionary` (the default) to `main.go`. Change the + `--min-words` default to 20,000 in the same edit; the fixture path passes its own floor + explicitly and is unaffected. +3. Dispatch in `run()`: exactly one of `--merged`, `--words`, `--in` may be given. Refuse + two rather than picking one. +4. Write `source_url`, `source_commit`, `sources_kept` and `sources_excluded` into `meta` — + provenance in the artifact, not only in a markdown file. +5. Tests in `merged_list_test.go`, table-driven, over hand-written JSONL fixtures: + - `Hà Nội` tagged wiktionary → kept as `hà nội` + - a `hongocduc`-only word and a `tudientv`-only word → absent from the output + - `học sinh` in all three, plus `Học sinh` → kept once - a malformed line → a clear error naming the line number, not a skipped word -7. Run against the real file and record the counts. Compare to the report's 61,276, which - used all-source case evidence; explain the delta rather than adjusting the number. + - `--sources hongocduc,wiktionary` → strictly more words; `tudientv` still absent + - meta rows as above +6. Run against the real file and record the counts. + +## Result + +``` +rejected 2: contains a digit +rejected 147: contains punctuation +rejected 5335: fewer than 2 syllables +rejected 87: no Vietnamese letters +accepted 26845 distinct words (sources: wiktionary) +generated 1507 spelling aliases +``` + +The first cut of this phase also implemented a capitalization-based proper-noun drop +(`only ever capitalized` → rejected, `--report-drops` for audit). It reached the Phase 3 +gate, was audited, and was removed by the owner's decision — see `plan.md` and +`audit-proper-noun-drops.md`. Nothing of it remains in the code. ## Success Criteria -- [ ] `go test ./cmd/build-dictionary/` passes, new tests included. -- [ ] Building from the real merged file accepts >55,000 words and logs a proper-noun drop - count in the low thousands. -- [ ] Every one of the 11 known-bad probes from the report is absent from the output - (`trương công định`, `trần danh án`, `tân an thạnh`, `cẩm xá`, `vĩnh điện`, - `cam lâm`, `hà nội`, `sài gòn`, `a mú sung`, `a lưới`, `kháng đón`). -- [ ] All 10 ordinary-word probes are present (`học sinh`, `mặt trời`, `con người`, - `bánh mì`, `cánh diều`, `xe đạp`, `giáo viên`, `hoa hồng`, `tình yêu`, `nước mắm`). -- [ ] `--sources hongocduc` alone and `--sources hongocduc,wiktionary` both build, with the - first strictly smaller. -- [ ] No word in the output is reachable only through `tudientv`. +- [x] `go test ./cmd/build-dictionary/` passes, new tests included. +- [x] Building from the real merged file accepts >20,000 words (26,845). +- [x] Ordinary-word probes present: `học sinh`, `bánh mì`, `xe đạp`, `giáo viên`, `hoa + hồng`, `tình yêu`, `nước mắm`, `mặt trời`, `trái đất`. `con người` and `cánh diều` are + not in the 2018 branch at all — expected, recorded. +- [x] `--sources wiktionary` (default) and `--sources hongocduc,wiktionary` both build, with + the first strictly smaller (26,845 vs 61,271 with no case rule). +- [x] No word in the output is reachable only through `hongocduc` or `tudientv`. ## Risk Assessment -**The capitalization rule eats real words.** The probe set is 21 words; the rule fires on -thousands. Signal: Phase 3's hand audit finds common words among the drops. Response: -narrow the rule to require every syllable capitalized (`Hà Nội`, not `Kháng đón`), re-audit, -and if that still misfires, stop dropping and keep the count as a log line only — a corpus -with proper nouns in it is what we ship today and is survivable. - -**Allowed-sources-only case evidence drops more than expected.** Signal: the drop count -comes in far above the report's 4,145. Response: measure how many drops have a lowercase -form only in `tudientv`; if that is a large share, reconsider reading `tudientv` for case -evidence alone and say so plainly in `ATTRIBUTION.md` rather than hiding it. - **A JSON schema surprise.** The file is 8 years old and unversioned; a row could carry an unexpected shape. Signal: decode errors on the real file. Response: the error names the line — inspect it, and only then decide between a tolerant skip with a counted reason and -a hard failure. Silent skipping is not an option. +a hard failure. Silent skipping is not an option. (Observed: none; all 79,226 rows decode.) diff --git a/plans/260908-1525-dictionary-corpus-switch/phase-02-pin-the-source.md b/plans/260908-1525-dictionary-corpus-switch/phase-02-pin-the-source.md index 19a1b09..9a79bea 100644 --- a/plans/260908-1525-dictionary-corpus-switch/phase-02-pin-the-source.md +++ b/plans/260908-1525-dictionary-corpus-switch/phase-02-pin-the-source.md @@ -1,7 +1,7 @@ --- phase: 2 title: "Phase 2: Pin the source" -status: todo +status: done priority: P1 effort: "2h" dependencies: [1] @@ -54,6 +54,8 @@ because that is already the right shape; only the size in the help text changes. ## Implementation Steps 1. Update the three Makefile variables and the `dict` target to pass `--merged $(DICT_SRC)`. + `--sources` is left at its `wiktionary` default so the Makefile and Dockerfile carry + one fewer thing to keep in agreement. 2. Reword `help`, `fetch-dict` and the `$(DICT_SRC)` guard: the download is ~4.8 MB now, and saying "179 MB" would be the kind of stale comment that outlives three refactors. 3. Mirror all of it in the Dockerfile stage. Keep `FIXTURE_DICT=1` on `--words` — the diff --git a/plans/260908-1525-dictionary-corpus-switch/phase-03-audit-the-corpus.md b/plans/260908-1525-dictionary-corpus-switch/phase-03-audit-the-corpus.md index 327e98f..b27771b 100644 --- a/plans/260908-1525-dictionary-corpus-switch/phase-03-audit-the-corpus.md +++ b/plans/260908-1525-dictionary-corpus-switch/phase-03-audit-the-corpus.md @@ -1,7 +1,7 @@ --- phase: 3 title: "Phase 3: Audit the corpus" -status: todo +status: done priority: P1 effort: "2h" dependencies: [1, 2] @@ -9,6 +9,12 @@ dependencies: [1, 2] # Phase 3: Audit the corpus +> **Outcome.** Executed 2026-09-08; see [`audit-proper-noun-drops.md`](./audit-proper-noun-drops.md). +> The drop rule passed its sample gate (1 common word in 100) but cost 215 words under +> wiktionary-only case evidence, including everyday vocabulary. Narrowing was measured and +> rejected. **The owner abandoned the rule**: every word is kept regardless of case. Final +> corpus 26,845 words; playability and bot results recorded in the audit note. + ## Overview The gate. A rule that removes thousands of words on a capitalization heuristic has to be @@ -34,12 +40,17 @@ Three measurements, each against the freshly built database and the current one: passes at ≤2 common words in 100; that tolerance is the difference between a filter and a corpus edit. 2. **Corpus diff.** Overlap, gained, lost. Split the lost into: dropped as proper nouns, - `tudientv`-only, absent from undertheseanlp. The report measured 4,145 / 1,300 / 6,473 - under all-source case evidence; this re-measures under the shipped rule. + absent from the wiktionary branch. The report measured 4,335 / 22,651 of 26,986 lost; + this re-measures from the shipped build. Separately count the **271 case casualties** + — drops that have a lowercase form in `hongocduc` or `tudientv` — and list them in + full; that list is what the pass/narrow/exception decision is made on. 3. **Playability.** Words, syllables, syllables that can open a word, how many have ≥2 continuations, and dead-end syllables — the last of these being what decides whether the game hands somebody an unanswerable position. Current: 48,216 / 6,676 / 5,049 / - 3,682 / 1,627. + 3,682 / 1,627. Expected: 22,310 / 5,484 / 4,050 / 2,686 / 1,434. The corpus is smaller + by design, so the question is not "is it bigger" but "does the game still run": add + bot-versus-bot game lengths from `server/internal/bot`'s real-corpus tests on both + databases. ## Related Code Files @@ -58,20 +69,23 @@ Three measurements, each against the freshly built database and the current one: hand. Vietnamese place and person names are the expected content; anything that reads as an ordinary word is a false positive and gets called one. 4. Run the corpus diff and the playability comparison; put both tables in the audit note. -5. Measure the documented cost of allowed-sources-only case evidence: how many drops have - a lowercase form only in `tudientv`. -6. Decide: pass, narrow, or abandon the rule. Record the decision and its reason in the - audit note — and if the rule is narrowed, Phase 1's tests change with it. +5. Build `--sources hongocduc,wiktionary` to a second scratch path purely to enumerate the + case casualties (words dropped under `wiktionary` alone but kept when hongocduc's + lowercase forms are visible). Nothing from that build ships. +6. Decide: pass, narrow, exception list, or abandon the rule. Record the decision and its + reason in the audit note — and if the rule changes, Phase 1's tests change with it. ## Success Criteria - [ ] The audit note exists, with 100 judged entries and the commands that produced them. -- [ ] ≤2 of 100 sampled drops are ordinary words. -- [ ] The new corpus has >55,000 words. -- [ ] Syllables ≥ 6,676 and dead-end syllables ≤ 1,627 — the graph is not worse than - what players walk today. -- [ ] The three loss buckets are quantified, not estimated. -- [ ] A pass/narrow/abandon decision is written down with its reason. +- [ ] ≤2 of 100 sampled drops are ordinary words, the 271 known case casualties aside — + those are listed in full and judged as a group. +- [ ] The new corpus has >20,000 words. +- [ ] The five graph numbers and bot-game lengths are recorded for both databases; bot + games on the new graph complete without the engine running out of moves earlier + than on today's. +- [ ] The loss buckets are quantified, not estimated. +- [ ] A pass/narrow/exception-list/abandon decision is written down with its reason. ## Risk Assessment @@ -81,10 +95,14 @@ Response: the drop list is a file — a disputed word can be checked against it and a per-word exception list is a small change on top of this design. **The audit fails and the phase becomes a redesign.** Signal: >2 common words in 100. -Response: apply Phase 1's pre-decided narrowing (every syllable capitalized), re-audit -once. If it fails again, ship without the drop — the corpus is still bigger and denser -than today's, and the proper-noun problem stays exactly as bad as it currently is rather -than getting worse. +Response: apply Phase 1's pre-decided narrowing (every syllable capitalized) or the +exception list, re-audit once. If it fails again, ship without the drop — the proper-noun +problem then stays exactly as bad as it is today rather than getting worse. + +**The corpus is too thin in play.** 22,310 words is less than half of today's. Signal: bot +games noticeably shorter, or the dead-end rate per move up rather than down. Response: this +is the trigger for the documented upgrade path — the 2026-09-01 viwiktionary dump (31,637 +words, same license) as a second `--words` source — not for re-admitting GPL data. **Playability regresses in a way these five numbers do not capture.** Signal: aggregate counts look fine but bot games end oddly short or long. Response: `server/internal/bot`'s diff --git a/plans/260908-1525-dictionary-corpus-switch/phase-04-relicense-the-data.md b/plans/260908-1525-dictionary-corpus-switch/phase-04-relicense-the-data.md deleted file mode 100644 index 472d02d..0000000 --- a/plans/260908-1525-dictionary-corpus-switch/phase-04-relicense-the-data.md +++ /dev/null @@ -1,110 +0,0 @@ ---- -phase: 4 -title: "Phase 4: Relicense the data" -status: todo -priority: P1 -effort: "2h" -dependencies: [3] ---- - -# Phase 4: Relicense the data - -## Overview - -`data/noitu.db` becomes GPLv3, because `hongocduc`'s data is GPL and GPL is copyleft. This -phase rewrites the data half of the licensing documents and regenerates the shipped -database — in one commit, because a GPL-derived database sitting in a tree that still -says CC BY-SA 4.0 is the only ordering mistake here with a legal consequence. - -## Requirements - -- Functional: `data/LICENSE` carries the GPLv3 text. -- Functional: `NOTICE` section 2 describes GPLv3 for the data, names both source branches - with their own licenses, and keeps section 1 (Apache-2.0 for code) intact. -- Functional: `data/ATTRIBUTION.md` records the new source, the per-branch licenses, the - full list of modifications, and why `tudientv` is excluded. -- Functional: the regenerated `data/noitu.db` carries matching `meta` provenance. -- Non-functional: no overclaiming. The `hongocduc` GPL statement comes from a mirror's - README and the document says so. - -## Architecture - -The dual-license structure already in `NOTICE` is correct and stays; only its second -section changes: - -``` -1. SOURCE CODE — Apache-2.0 (unchanged: server/, tools/, web/, proto/) -2. DICTIONARY DATA — GPLv3 (was CC BY-SA 4.0) - data/noitu.db, data/LICENSE, data/ATTRIBUTION.md -``` - -Why GPLv3 rather than CC BY-SA 4.0 is worth stating in `ATTRIBUTION.md` rather than left -implicit: `hongocduc` is Hồ Ngọc Đức's FVDP wordlist, distributed under GNU GPL. The -previous source aggregated that same data and redistributed it as CC BY-SA 4.0, which GPL -does not permit. Ours is the stricter, and correct, direction. - -The modifications list gains two entries beyond the current five — source selection and -the proper-noun drop — both of which the license requires us to declare as changes. - -## Related Code Files - -- Modify: `data/LICENSE` — CC BY-SA 4.0 text replaced with GPLv3 -- Modify: `NOTICE` — section 2 rewritten -- Modify: `data/ATTRIBUTION.md` — source table, per-branch licenses, modifications 6 and 7, - the `tudientv` exclusion note -- Modify: `README.md` — any dictionary-source or data-license claim -- Verify: `.github/workflows/ci.yml` — the required-files check (`data/LICENSE`, - `data/ATTRIBUTION.md`, `NOTICE`, `data/noitu.db`) still holds -- Regenerate: `data/noitu.db` - -## Implementation Steps - -1. Replace `data/LICENSE` with the GPLv3 text, verbatim and complete. -2. Rewrite `NOTICE` section 2: GPLv3, the affected artifacts, both upstream branches with - their licenses, and the copyleft obligation carrying into container images — the same - point the current text makes about share-alike, which is no less true of GPL. -3. Rewrite `data/ATTRIBUTION.md`: - - source is `undertheseanlp/dictionary` at commit `2c078cf`, file `dictionary/words.txt` - - `hongocduc` (GNU GPL, per the branch README — a mirror of Hồ Ngọc Đức's 2003 - wordlist, whose canonical site no longer resolves) and `wiktionary` (CC BY-SA, a - 2018-12-10 scrape of vi.wiktionary.org) - - `tudientv` **excluded**: its README declares copyright *"Chưa rõ"* and names - Soha/Vietlex (Hoàng Phê) as its primary source. Say it plainly so nobody reopens it. - - modifications 1–5 carried over, plus **6. source selection** and **7. proper-noun - removal** with the count from Phase 3 - - the resulting license and why it is GPLv3 and not CC BY-SA 4.0 -4. Grep the repository for `minhqnd`, `CC BY-SA`, `179 MB` and fix every survivor outside - `plans/` — the reports are dated records and stay as written. -5. Regenerate `data/noitu.db` and confirm its `meta` rows agree with `ATTRIBUTION.md`. A - database whose provenance contradicts the attribution file is worse than neither. -6. Run the CI required-files check locally, or read it and confirm by hand. - -## Success Criteria - -- [ ] `data/LICENSE` is the complete GPLv3 text. -- [ ] `NOTICE` section 2 says GPLv3 and names both branches with their licenses; section 1 - is unchanged. -- [ ] `data/ATTRIBUTION.md` lists seven modifications, states the `tudientv` exclusion and - its reason, and does not claim more about the `hongocduc` license than the mirror - supports. -- [ ] `meta` in the built database matches the attribution file — source URL, commit, - license, sources kept and excluded, proper-noun drop count. -- [ ] No file outside `plans/` still describes the data as CC BY-SA 4.0 or names minhqnd - as the source. -- [ ] CI's required-files check passes. - -## Risk Assessment - -**Relicensing is hard to walk back.** Once a GPLv3 `noitu.db` is published, that release -is GPLv3 permanently. Signal: none — this is simply true. Response: it is why the user -signed off before this plan was written; the mitigation is that it was a decision rather -than a side effect. - -**A stale CC BY-SA claim survives somewhere.** A README line or a Docker label saying the -wrong license is a licensing misstatement, not a typo. Signal: the grep in step 4 finds -something after the phase is called done. Response: grep is a step in the phase for that -reason; run it last, not first. - -**Section 1 gets damaged while section 2 is rewritten.** The code's Apache-2.0 grant is -not in scope and must come out byte-identical. Signal: a diff touching section 1. -Response: review the `NOTICE` diff before committing. diff --git a/plans/260908-1525-dictionary-corpus-switch/phase-04-update-the-attribution.md b/plans/260908-1525-dictionary-corpus-switch/phase-04-update-the-attribution.md new file mode 100644 index 0000000..2295ca6 --- /dev/null +++ b/plans/260908-1525-dictionary-corpus-switch/phase-04-update-the-attribution.md @@ -0,0 +1,116 @@ +--- +phase: 4 +title: "Phase 4: Update the attribution" +status: done +priority: P1 +effort: "1h" +dependencies: [3] +--- + +# Phase 4: Update the attribution + +> **Outcome.** Done 2026-09-08. `data/LICENSE` untouched; `NOTICE` section 2 names the new +> upstream, `data/ATTRIBUTION.md` rewritten (eight modifications, both attribution links, +> both exclusions). Also updated, beyond the plan's list: the frontend attribution footer +> (`AttributionFooter.svelte`, `i18n/vi.js`) and its e2e assertion now credit Wiktionary +> tiếng Việt instead of minhqnd — the user-visible half of the CC BY-SA obligation. + +## Overview + +The data license does not change — the wiktionary branch is CC BY-SA 4.0, the same license +`noitu.db` already carries — but everything that names the source does. This phase rewrites +the attribution and notice text and regenerates the shipped database in one commit, so the +attribution never describes a database other than the one in the tree. + +## Requirements + +- Functional: `data/LICENSE` is untouched — still the CC BY-SA 4.0 text. +- Functional: `NOTICE` section 2 names the new upstream, keeps CC BY-SA 4.0, and keeps + section 1 (Apache-2.0 for code) byte-identical. +- Functional: `data/ATTRIBUTION.md` records the new source, the branch actually used, the + two branches excluded and why, and the full list of modifications. +- Functional: the regenerated `data/noitu.db` carries matching `meta` provenance. +- Non-functional: no overclaiming. The branch is a 2018 scrape of vi.wiktionary.org made by + a third party; say that, and attribute Wiktionary's contributors as CC BY-SA requires. + +## Architecture + +The dual-license structure in `NOTICE` is correct and stays; only the source lines of its +second section change: + +``` +1. SOURCE CODE — Apache-2.0 (unchanged: server/, tools/, web/, proto/) +2. DICTIONARY DATA — CC BY-SA 4.0 (unchanged license, new upstream) + data/noitu.db, data/LICENSE, data/ATTRIBUTION.md +``` + +The attribution chain is now two links instead of an opaque aggregate: vi.wiktionary.org +contributors (CC BY-SA) → `undertheseanlp/dictionary`, branch data `wiktionary`, commit +`2c078cf` → this project. Both links are named. + +The modifications list stays at eight: source selection replaces the old language filter +as entry 1, which the license requires us to declare as a change. + +## Related Code Files + +- Modify: `NOTICE` — section 2's upstream line, affected-artifacts line (`dictionary.db` + is no longer the source), and the closing paragraph about build artifacts +- Modify: `data/ATTRIBUTION.md` — source table, branch selection, exclusions, modifications + 9 and 10, the reproduce block +- Modify: `README.md` — dictionary-source lines (`~179 MB`, `dictionary.db`, `minhqnd`, + the manual `curl` + `--in` example) +- Modify: `docs/deployment.md` — the 179 MB builder-stage sentence +- Verify unchanged: `data/LICENSE`, `.github/workflows/ci.yml` required-files check +- Regenerate: `data/noitu.db` + +## Implementation Steps + +1. Rewrite `NOTICE` section 2's source lines: upstream `undertheseanlp/dictionary` at + commit `2c078cf`, wiktionary data only; affected artifact `data/noitu.db`. Leave the + share-alike paragraph as is — it is still true. +2. Rewrite `data/ATTRIBUTION.md`: + - source is `undertheseanlp/dictionary`, file `dictionary/words.txt`, commit `2c078cf` + - rows used: those tagged `wiktionary` — a 2018-12-10 scrape of vi.wiktionary.org, + whose text is CC BY-SA 4.0 / GFDL by its contributors + - `hongocduc` **excluded**: GNU GPL per the branch README; using it would relicense this + database to GPLv3 + - `tudientv` **excluded**: its README declares copyright *"Chưa rõ"* and names + Soha/Vietlex (Hoàng Phê) as its primary source. Say it plainly so nobody reopens it. + - modifications renumbered: **1. source selection** replaces the language filter (the new + file is Vietnamese-only); the other seven carry over, with normalization stating that + capitalization removes nothing — the proper-noun drop was abandoned at the Phase 3 gate + - reproduce block: `make fetch-dict` is ~4.8 MB now +3. Grep the repository for `minhqnd`, `dictionary.db`, `179` and `--in` and fix every + survivor outside `plans/` — the reports are dated records and stay as written. +4. Regenerate `data/noitu.db` and confirm its `meta` rows agree with `ATTRIBUTION.md`: + source URL, commit, license, `sources_kept = wiktionary`, + `sources_excluded = hongocduc,tudientv`. +5. Run the CI required-files check locally, or read it and confirm by hand. + +## Success Criteria + +- [ ] `data/LICENSE` has no diff. +- [ ] `NOTICE` section 2 names the new upstream and still says CC BY-SA 4.0; section 1 is + byte-identical. +- [ ] `data/ATTRIBUTION.md` lists eight modifications, names the branch used and the two + excluded with reasons, and credits vi.wiktionary.org contributors. +- [ ] `meta` in the built database matches the attribution file. +- [ ] No file outside `plans/` names minhqnd, `dictionary.db` as a source, or a 179 MB + download. +- [ ] CI's required-files check passes. + +## Risk Assessment + +**A stale source claim survives somewhere.** A README line or Docker comment naming the old +upstream is an attribution misstatement, not a typo. Signal: the grep in step 3 finds +something after the phase is called done. Response: grep is a step in the phase for that +reason; run it last, not first. + +**Section 1 gets damaged while section 2 is edited.** The code's Apache-2.0 grant is not +in scope and must come out byte-identical. Signal: a diff touching section 1. Response: +review the `NOTICE` diff before committing. + +**The attribution credits the wrong party.** CC BY-SA attribution belongs to Wiktionary's +contributors; undertheseanlp is the intermediary that scraped and redistributed. Signal: +an `ATTRIBUTION.md` that names only the GitHub repository. Response: step 2 names both +links of the chain explicitly. diff --git a/plans/260908-1525-dictionary-corpus-switch/phase-05-retire-the-sqlite-path.md b/plans/260908-1525-dictionary-corpus-switch/phase-05-retire-the-sqlite-path.md index f48e15b..d4fdd9f 100644 --- a/plans/260908-1525-dictionary-corpus-switch/phase-05-retire-the-sqlite-path.md +++ b/plans/260908-1525-dictionary-corpus-switch/phase-05-retire-the-sqlite-path.md @@ -1,7 +1,7 @@ --- phase: 5 title: "Phase 5: Retire the SQLite path" -status: todo +status: done priority: P2 effort: "1h" dependencies: [4] @@ -9,6 +9,13 @@ dependencies: [4] # Phase 5: Retire the SQLite path +> **Outcome.** Done 2026-09-08. `--in`, `--table`, `--word-col`, `--lang-col`, `--lang`, +> `resolveSource`, `validateSource`, `listTables`, `listColumns`, `pickColumn`, `extract` +> and `quoteIdent` removed; `meta` no longer writes `source_word_column` / +> `source_lang_column` (no consumer existed). The four auto-detection tests went; the +> output-shape tests now build their fixture through `--merged`. `--help` lists `--merged`, +> `--sources`, `--words`, `--out`, `--max-syllables`, `--min-words` and nothing else. + ## Overview Delete the input mode nothing uses any more: reading a source SQLite database, with the @@ -55,7 +62,8 @@ no longer read that upstream. 2. Delete the tests that exist only to cover auto-detection. Do not delete tests covering output shape, `verify()`, or the reject-reason counts — those still describe behavior. 3. Rewrite the package doc comment: the input is a 4.8 MB JSONL wordlist with source - membership and capitalization; the output is the game's syllable-indexed database. + membership and capitalization, of which only wiktionary-tagged rows are read; the + output is the game's syllable-indexed database. 4. `go build ./... && go vet ./... && go test ./...` from `server/`. 5. Grep for `--in`, `dictionary.db`, `resolveSource` and `lang_code` outside `plans/`; nothing should survive except in the dated reports. diff --git a/plans/260908-1525-dictionary-corpus-switch/plan.md b/plans/260908-1525-dictionary-corpus-switch/plan.md index 48a1cc7..cb8d16e 100644 --- a/plans/260908-1525-dictionary-corpus-switch/plan.md +++ b/plans/260908-1525-dictionary-corpus-switch/plan.md @@ -1,11 +1,12 @@ --- title: "Dictionary corpus switch" -description: "Replace the 179 MB minhqnd aggregate with undertheseanlp/dictionary (hongocduc + wiktionary, tudientv excluded), drop always-capitalized proper nouns, and relicense data/noitu.db to GPLv3" -status: pending +description: "Replace the 179 MB minhqnd aggregate with the wiktionary branch of undertheseanlp/dictionary only (hongocduc and tudientv excluded) and keep data/noitu.db on CC BY-SA 4.0" +status: in-progress priority: P1 effort: "~1d" tags: [dictionary, data, licensing, build] created: 2026-09-08 +updated: 2026-09-08 blockedBy: [] blocks: [] --- @@ -14,115 +15,141 @@ blocks: [] ## Overview -`data/noitu.db` is derived from `minhqnd/dictionary` v2.0.0: a 179 MB SQLite aggregate of -five upstreams, redistributed as CC BY-SA 4.0. Two problems. Its license chain does not -close — three of its five inputs are GPL (FVDP, tudientv, "Vietnamese Explanatory -Dictionary") and GPL does not permit relicensing to CC BY-SA. And it has already -lowercased everything (0 of 70,511 Vietnamese rows carry uppercase), which destroys the -one signal that separates a place name from a word. +`data/noitu.db` was derived from `minhqnd/dictionary` v2.0.0: a 179 MB SQLite aggregate of +five upstreams, redistributed as CC BY-SA 4.0. Its license chain does not close — three of +its five inputs are GPL (FVDP, tudientv, "Vietnamese Explanatory Dictionary") and GPL does +not permit relicensing to CC BY-SA. -This plan switches the corpus to `undertheseanlp/dictionary` — a single 4.8 MB JSONL file -carrying, per word, its raw capitalization and which of three source dictionaries contain -it. `hongocduc` (GPL) and `wiktionary` (CC BY-SA) are taken; `tudientv` is excluded -permanently because its own README declares copyright *"Chưa rõ"* and names Soha/Vietlex -(Hoàng Phê) as its primary source. Entries that appear only ever capitalized are dropped -as proper nouns. `data/noitu.db` then ships **GPLv3** — the honest license for -FVDP-derived data, and the reason this is a licensing improvement rather than a new risk. +This plan switches the corpus to the **`wiktionary` source inside `undertheseanlp/dictionary`** +— a single 4.8 MB JSONL file carrying, per word, which of three source dictionaries contain +it. Only rows whose `source` list contains `wiktionary` are read. `hongocduc` (GPL) is +excluded so the data license does not change; `tudientv` is excluded because its own README +declares copyright *"Chưa rõ"* and names Soha/Vietlex (Hoàng Phê) as its primary source. +`data/noitu.db` **stays CC BY-SA 4.0**: the wiktionary branch is a 2018-12-10 scrape of +vi.wiktionary.org, whose text is CC BY-SA. Evidence, all measured through the real `build-dictionary` filter: -`plans/reports/research-260908-1507-undertheseanlp-dictionary.md`. +`plans/reports/research-260908-1507-undertheseanlp-dictionary.md` (branch licensing), +`plans/reports/research-260908-1529-viwiktionary-dump-measured.md` (yields, alternatives) +and [`audit-proper-noun-drops.md`](./audit-proper-noun-drops.md) (the Phase 3 gate). ## What changes, in numbers -| | current | after | +| | before | after | |---|---|---| -| words | 48,216 | **~61,276** | -| syllables | 6,676 | 7,011 | -| syllables that can open a word | 5,049 | 5,496 | -| …with ≥2 continuations | 3,682 | 4,168 | -| dead-end syllables | 1,627 | **1,515** | +| words | 48,216 | **26,845** | +| syllables | 6,676 | 5,709 | +| syllables that can open a word | 5,049 | 4,158 | +| …with ≥2 continuations | 3,682 | 2,787 | +| dead-end syllables | 1,627 | 1,551 | | source download | 179 MB SQLite | 4.8 MB JSONL | -| data license | CC BY-SA 4.0 | **GPLv3** | +| data license | CC BY-SA 4.0 | **CC BY-SA 4.0 (unchanged)** | +| `--min-words` floor | 40,000 | **20,000** | -24,978 words gained, 11,918 lost — of which 4,145 are proper nouns removed on purpose, -1,300 are `tudientv`-only, and 6,473 are absent from undertheseanlp entirely, including -genuine modern vocabulary (`tế bào gốc`, `hằng số vũ trụ`) that a 2003 wordlist cannot -have. That loss is accepted here and is the subject of a separate future decision, not -this plan. +**This is a deliberate 44% cut.** 1,280 words gained, 22,651 lost — every one of them +simply absent from a 2018 wiktionary scrape (`bánh đúc`, `lò vi sóng`, `dân làng`, `thót +tim`). Bot-versus-bot games run about a quarter shorter on the thinner graph (easy-vs-easy +12.9 moves vs 17.5). This loss is accepted here; the upgrade path is written down below. ## Decisions taken -- **`tudientv` is never read.** Not as content, and not as case evidence. Excluding it - from evidence too is what makes "we do not touch that data" true without qualification. - It costs a little accuracy in the proper-noun rule — a word capitalized in `hongocduc` - but lowercase only in `tudientv` will be dropped — and Phase 3 measures that cost. -- **The measured 61,276 used all-source case evidence.** Under the allowed-sources-only - rule the number will shift slightly. Phase 1 re-measures; the plan does not depend on - the exact figure, only on it staying above the 40,000 `--min-words` floor and above - today's 48,216. -- **Proper nouns are dropped, not tiered.** The per-word `source` array looked like a - confidence signal and is not one: `hà nội` has 2 votes, `cẩm xá` 2, `kháng đón` 2, while - `cánh diều` has 1 and `mặt trời` 2. Capitalization is the only usable signal in this - data, so it is the only one used. -- **`noitu.db` becomes GPLv3; the code stays Apache-2.0.** The existing dual-license split - in `NOTICE` holds; only its data half is rewritten. -- **The SQLite input path is deleted, not kept "just in case."** Two input modes for one - source is dead weight; the fixture path (`--words`) stays because the e2e suite and the - Docker smoke build use it. -- **The viwiktionary extractor is out of scope.** It remains the answer if a CC BY-SA - corpus is ever needed again — see `research-260908-1434-viwiktionary-as-dictionary-source.md`. +- **Only `wiktionary` rows are read.** `hongocduc` and `tudientv` inform nothing. That is + what keeps the license CC BY-SA 4.0 and what makes "we do not touch that data" true + without qualification. +- **Capitalization is not a filter.** The original design dropped words that appear only + ever capitalized, as proper nouns. It was built, run and audited at the Phase 3 gate: the + rule was precise on a 100-word sample (1 common word) but, with wiktionary-only case + evidence, it also removed 215 words including `mặt trời`, `trái đất`, `tổ quốc`, `dường + như`, `phật giáo`. Narrowing was measured and did not help. **The owner chose to keep + every word regardless of case.** Proper nouns like `hà nội` remain playable, exactly as + they are today. The mechanism was removed from the builder rather than left behind a flag. +- **Rejected alternatives, so nobody reopens them without new evidence:** + - `hongocduc + wiktionary` (61,271 words) — forces `noitu.db` to GPLv3. Rejected to keep + CC BY-SA 4.0. + - the 2026-09-01 viwiktionary dump (36,200 words without a case rule, same license, + monthly pin) — the 2018 branch is a near-subset of it. Rejected for now in favour of the + simpler single-file pin; **this is the documented upgrade path** if the 27k corpus + proves too thin in play. See the 15:29 report. + - kaikki.org enwiktionary Vietnamese — unpinnable rolling file, deprecated by its host. +- **The SQLite input path is deleted, not kept "just in case."** The fixture path + (`--words`) stays because the e2e suite and the Docker smoke build use it. +- **The build floor drops to 20,000.** The old 40,000 default would reject the new corpus; + it moved once, in `main.go`. ## The pin ``` URL https://raw.githubusercontent.com/undertheseanlp/dictionary/2c078cfc373b06e2980d324ce1d7bd13740c3319/dictionary/words.txt SHA256 4c3e0e6117e4bdfa97731e135c3d4a05881889909267394a8de8d88ef79f13f0 -Size 4,813,111 bytes — 79,226 JSONL rows +Size 4,813,111 bytes — 79,226 JSONL rows, 32,484 of them tagged wiktionary Commit 2c078cf (2018-12-10, the repository's last) ``` Pinned by commit rather than branch: `raw.githubusercontent.com///` is immutable, so the URL and the checksum cannot disagree later. The variable names -`DICT_URL` / `DICT_SHA256` are kept so `web/tests/dictionary-source.test.js` needs no -change — only their values move. +`DICT_URL` / `DICT_SHA256` were kept so `web/tests/dictionary-source.test.js` needed no +change — only their values moved. ## Phases | # | Phase | Status | |---|-------|--------| -| 1 | [Read the merged list](./phase-01-start.md) | Pending | -| 2 | [Pin the source](./phase-02-pin-the-source.md) | Pending | -| 3 | [Audit the corpus](./phase-03-audit-the-corpus.md) | Pending | -| 4 | [Relicense the data](./phase-04-relicense-the-data.md) | Pending | -| 5 | [Retire the SQLite path](./phase-05-retire-the-sqlite-path.md) | Pending | +| 1 | [Read the merged list](./phase-01-start.md) | Done | +| 2 | [Pin the source](./phase-02-pin-the-source.md) | Done | +| 3 | [Audit the corpus](./phase-03-audit-the-corpus.md) | Done — gate sent the case rule back; rule abandoned | +| 4 | [Update the attribution](./phase-04-update-the-attribution.md) | Done | +| 5 | [Retire the SQLite path](./phase-05-retire-the-sqlite-path.md) | Done | -Phase 3 is a gate: it may send Phase 1 back to narrow the proper-noun rule. Phase 4 must -land in the same commit as the first generated GPL-derived `noitu.db` — shipping that -database while `NOTICE` still says CC BY-SA 4.0 is the one ordering mistake with a legal -consequence. +All five phases are uncommitted in one working tree; they should land together so the +attribution never describes a database other than the one the build produces. ## Success criteria -- [ ] `make fetch-dict && make dict` builds `data/noitu.db` from a 4.8 MB pinned file with +- [x] `make fetch-dict && make dict` builds `data/noitu.db` from a 4.8 MB pinned file with no 179 MB download. -- [ ] The shipped database has **more than 55,000 words** and **no word that reaches it - only via `tudientv`**. -- [ ] The proper-noun drop count is reported in the build log beside the other reject - reasons, recorded in `meta`, and a 100-entry sample has been audited by hand. -- [ ] Playability does not regress: syllables ≥ 6,676 and dead-end syllables ≤ 1,627. -- [ ] `data/LICENSE`, `NOTICE` and `data/ATTRIBUTION.md` describe GPLv3, name both source - branches with their own licenses, and record why `tudientv` is excluded. -- [ ] `make test-go`, `make test-web` and the e2e suite are green; the fixture build path - is untouched. -- [ ] `--in` and its schema auto-detection are gone, and no doc still mentions the 179 MB - download. +- [x] The shipped database has **more than 20,000 words** and **no word that reaches it + through `hongocduc` or `tudientv` alone**. +- [x] `meta` records the source URL, commit, sources kept and sources excluded. +- [x] Playability is measured and recorded: syllables, openers, continuations, dead ends, + and bot-versus-bot game lengths on the new graph versus the old. +- [x] `data/LICENSE` is unchanged; `NOTICE` and `data/ATTRIBUTION.md` name the new source, + its commit and license, and record why the two other branches are excluded. The + frontend attribution footer credits Wiktionary tiếng Việt. +- [x] `go vet ./... && go test ./...` (incl. `-race` on the touched packages), `npm run + check && npm test`, and the Docker image (`noitu:ci` with the CI file checks, plus + `FIXTURE_DICT=1`) are green; the fixture build path is untouched. +- [ ] The Playwright e2e suite is green. **Not verified locally**: Playwright's Chromium + download timed out twice on this machine (2026-09-08), so every test failed at browser + launch, not on game behaviour. The one e2e assertion this plan changes (the attribution + footer's link text and href in `web/e2e/bot-game.spec.js`) was updated by hand. Verify + in CI or after `npx playwright install chromium` succeeds. +- [x] `--in` and its schema auto-detection are gone, and no file outside `plans/` still + mentions minhqnd, `dictionary.db`, `--in` or a 179 MB download. + +## Licence version — decided after review + +The owner asked whether the data could move to Apache-2.0 to simplify the repo. It cannot: +the words are Wiktionary contributors' text under a share-alike licence, and neither +undertheseanlp nor this project can relicense them. The owner then chose to **keep the +licence the source text actually carried, CC BY-SA 3.0 Unported**, rather than upgrade the +derivative to 4.0 under the later-version clause. `data/LICENSE` is now the BY-SA 3.0 legal +code; `NOTICE`, `ATTRIBUTION.md`, README, Dockerfile, deployment doc, the builder's meta +string and the in-game footer all say 3.0. The "unchanged" wording above predates this. + +## Review + +Code review (2026-09-08) returned DONE_WITH_CONCERNS; every finding was applied: the +builder's copy of the upstream commit is now guarded by `web/tests/dictionary-source.test.js` +alongside the Makefile/Dockerfile pin; fixture builds record "no upstream data" instead of +inheriting the CC BY-SA string; `ATTRIBUTION.md` states the 2018 snapshot was CC BY-SA 3.0 +and why the derived database is 4.0; dead `contains()`, redundant sorts, the off-by-one +scanner line number, stale "48k" comments, the Makefile `curl` flags and the Dockerfile +download name were fixed; input-selection and fixture-licence tests were added. ## Open questions -- The `hongocduc` GPL claim rests on the mirror's README; Hồ Ngọc Đức's canonical site - 404s. Assuming GPL is the conservative direction, so this does not block — but the - evidence is a mirror, and `ATTRIBUTION.md` should say so rather than overclaim. -- The 6,473 words undertheseanlp does not have are an unquantified mix of junk and real - modern vocabulary. Separating them needs a second live source; out of scope here. +- Is 26,845 words enough for the game to feel playable? Bot games are ~25% shorter; only + real play can say whether that is felt. If not, the 2026 viwiktionary dump is the next + step, not a return to GPL data. diff --git a/plans/journals/2026-09-08-dictionary-corpus-switch-to-wiktionary-only-undertheseanlp.md b/plans/journals/2026-09-08-dictionary-corpus-switch-to-wiktionary-only-undertheseanlp.md new file mode 100644 index 0000000..e704322 --- /dev/null +++ b/plans/journals/2026-09-08-dictionary-corpus-switch-to-wiktionary-only-undertheseanlp.md @@ -0,0 +1,36 @@ +--- +title: Dictionary corpus switch to wiktionary-only undertheseanlp +date: 2026-09-08 +summary: "Replaced the 179 MB minhqnd aggregate with the wiktionary rows of undertheseanlp/dictionary; corpus 48,216 -> 26,845 words, CC BY-SA 4.0 kept, proper-noun drop abandoned after audit, SQLite path retired" +--- + +# Dictionary corpus switch to wiktionary-only undertheseanlp + +## What happened + +- Three research passes today measured every candidate through the real `build-dictionary` filter: the 2026-09-01 viwiktionary dump (36,200 words), kaikki's enwiktionary Vietnamese file (22,628, deprecated and unpinnable), undertheseanlp's merged list (61,271 but GPL via hongocduc) and its wiktionary rows alone (26,845). Reports: `plans/reports/research-260908-1529-viwiktionary-dump-measured.md`. +- Owner chose undertheseanlp wiktionary-only to keep `data/noitu.db` on CC BY-SA 4.0. Plan `plans/260908-1525-dictionary-corpus-switch/` rewritten around it, then executed end to end. +- `build-dictionary` gained `--merged` / `--sources` (default `wiktionary`, names validated), a shared `finish()` tail, `--min-words` default 20,000, and meta rows `source_commit`, `sources_kept`, `sources_excluded`. The SQLite `--in` path, schema auto-detection and their tests were deleted. +- Makefile, Dockerfile, `.gitignore`, CI leak guard, NOTICE, `data/ATTRIBUTION.md`, README, `docs/deployment.md` and the frontend attribution footer now name the new source. `data/LICENSE` untouched. + +## Decision + +- **Capitalization is not a filter.** The planned proper-noun drop was built and audited: precise on a 100-word sample (1 common word), but with wiktionary-only case evidence it removed 215 words including `mặt trời`, `trái đất`, `tổ quốc`, `dường như`. Narrowing to all-syllables-capitalized was measured (147 casualties remain, 410 junk entries return) and rejected. Owner decided to keep every word regardless of case; the mechanism was removed, not flagged off. Record: `audit-proper-noun-drops.md` in the plan dir. +- Attribution states the 2018 scrape was CC BY-SA 3.0 (Wikimedia moved to 4.0 in 2023) and that the derived database is 4.0 via the later-version clause. Both chain links are credited: Wiktionary tiếng Việt contributors, then undertheseanlp as intermediary. +- Rejected alternatives recorded in `plan.md`: hongocduc+wiktionary (GPLv3 relicense), viwiktionary 2026 dump (documented upgrade path if 27k proves thin), kaikki (unpinnable). + +## Verification + +- `go vet ./... && go test ./...` green, `-race` on touched packages green; `npm run check && npm test` 175/175; Docker `noitu:ci` built with the CI required-files and leak checks passing, `FIXTURE_DICT=1` variant built. +- Bot real-corpus tests pass on the new DB; easy-vs-easy games 12.9 moves vs 17.5 before, the expected depth loss. +- Code review (DONE_WITH_CONCERNS) fully applied: builder's copy of the pinned commit now guarded by `web/tests/dictionary-source.test.js`; fixture builds record "no upstream data" instead of a CC BY-SA string; dead `contains()`, redundant sorts, scanner line off-by-one, stale "48k" comments, `curl -f`, Dockerfile download name fixed. +- **Not verified:** Playwright e2e — Chromium download timed out twice on this machine. The one changed e2e assertion (footer link) was updated by hand. Needs CI or a successful `npx playwright install chromium`. + +## Next steps + +- Commit all five phases together so attribution and database never disagree (not yet committed). +- Confirm e2e in CI. +- Watch real play for thinness; the viwiktionary dump is the upgrade path, not GPL data. +- Delete the orphaned local `data/dictionary.db` (179 MB, gitignored). + +> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions. diff --git a/plans/reports/research-260908-1529-viwiktionary-dump-measured.md b/plans/reports/research-260908-1529-viwiktionary-dump-measured.md new file mode 100644 index 0000000..7245f24 --- /dev/null +++ b/plans/reports/research-260908-1529-viwiktionary-dump-measured.md @@ -0,0 +1,271 @@ +--- +title: "Research: building a dictionary DB from the dumps.wikimedia.org viwiktionary dump — measured" +date: 2026-09-08T15:29+07:00 +type: research +status: complete +scope: research-only — dump downloaded to scratchpad, parsed, run through the real build-dictionary filter; no repo file touched +follows: research-260908-1434-viwiktionary-as-dictionary-source.md, plans/260908-1525-dictionary-corpus-switch/plan.md +--- + +# Can we build our own Vietnamese dictionary DB from the viwiktionary dump? + +## Executive summary + +**Yes, and it is cheap: 63.5 MB download, 34 s of stdlib Python, no wiktextract, no Lua.** The +2026-09-01 `pages-articles` dump was downloaded, checksum-verified against Wikimedia's published +MD5, parsed, and every list below was run through our real `build-dictionary --words` filter. +Numbers are measured, not estimated. + +**But viwiktionary alone is not a corpus.** It yields **36,200** accepted words, **31,637** after +dropping pages labelled proper-noun. Both sit below the shipped 48,216 and below the 40,000 +`--min-words` floor. The 14:34 report's "~16k from this source" undercounted by 2×, because the +`minhqnd` aggregate only ingested a 17k-row slice; the full dump is 36k. Still not enough. + +**Where it earns its place is on top of the pending plan.** Against the plan's undertheseanlp +corpus (re-measured this session at **61,026**, matching the plan's ~61,276), the 2026 dump adds +**3,084 words** the plan would otherwise lose — `học liệu`, `thiên hà lùn`, `cải thảo`, `bùn đỏ`, +`báng súng`, `tăng đoàn` — for a union of **64,110**. That is the modern-vocabulary gap the plan's +own open question names, and the plan's 2018 wiktionary branch (26,845 accepted) is 9,780 words +behind the 2026 dump. **Recommendation: after the corpus switch lands, add the dated dump as a +second `--words` input, pinned by dated URL + checksum.** Licensing stays clean: CC BY-SA 4.0 is +one-way compatible into the GPLv3 the plan already adopts. + +**Do not use viwiktionary's proper-noun labels as a drop rule.** They agree with the +capitalization blocklist 86% of the time, but they tag `mặt trời`, `trái đất`, `thứ hai`, +`tháng ba`, `tin lành`, `bàn là` as proper-only and `việt nam` as common. As a *filter* the label +is worse than capitalization; as a *report column* it is fine. + +--- + +## Method + +- Web: 2 fetches (dump index, kaikki index), plus direct `curl` of the dump, its checksum file, + 4 sample raw wikitext pages, and the plan's pinned undertheseanlp file. +- Local: `extract.py` (streams bz2 → per-page markers, 34 s), `classify.py` (slices the + Vietnamese section, classifies POS labels, emits word lists), `build-dictionary --words` for + every list, sqlite for overlaps. Scripts are in the session scratchpad; each is < 100 lines. + +## The dump + +| | | +|---|---| +| File | `https://dumps.wikimedia.org/viwiktionary/20260901/viwiktionary-20260901-pages-articles.xml.bz2` | +| Size | 63,513,513 bytes | +| MD5 (Wikimedia-published, verified) | `6c2491e703e7d946f23a405996b4d172` | +| SHA-256 (computed here; Wikimedia publishes md5/sha1 only for this run) | `ed66c932f535b0b362d1141c02273b02ed01c4101e856378bd15e88d277bb8d9` | +| Cadence | monthly, dated dirs `20251220 … 20260901`, plus rolling `latest/` | +| Pages / ns0 pages | 391,543 / 349,461 | +| **ns0 pages with a Vietnamese section** | **43,013** (on-wiki category says 43,037 — parser is complete) | +| Redirect pages among them | 3,237; 834 are case-only (`mặt trời` → `Mặt Trời`) | + +Dated directories are immutable, so a pin is honest. `latest/` is not. kaikki.org tracks the same +dump (extracted 2026-09-06, 56,502 Vietnamese *senses*) but publishes rolling files — the earlier +report's pinning concern stands; the dump removes the need for kaikki entirely. + +### Two wikitext dialects, both live + +| Format | Pages | Language marker | POS marker | +|---|---|---|---| +| legacy | 35,885 | `{{-vie-}}` | `{{-noun-}}`, `{{-verb-}}`, `{{-pr-noun-}}`, `{{-place-}}` … | +| new | 7,128 | `== {{langname\|vi}} ==` | `{{ĐM\|noun}}`, `{{ĐM\|pr-noun}}`, `{{vi-noun}}`, `{{vi-pr-noun}}` … | + +A parser must handle both; the wiki is mid-migration, so the mix shifts every month. Legacy +section codes collide with language codes (`{{-adj-}}` vs `{{-eng-}}`), which is why the +extractor keeps an explicit section-code set and treats every other 2–3-letter code as a language +switch. Languages seen switching out of Vietnamese: `tyz`, `mtq`, `eng`, `nut`, `nuo`, `fra` … + +## Classification of the 43,013 pages + +| Class | Pages | Meaning | +|---|---|---| +| common | 31,950 | ≥1 common POS label (noun/verb/adj/adv/phrase/idiom/…) | +| proper-only | 5,103 | only `pr-noun` / `place` / `vi-pr-noun` labels | +| no-pos | 5,764 | Vietnamese section, no POS marker at all | +| nôm-only | 169 | only Nôm/Hán character sections | +| proper+common | 27 | | + +The **no-pos** class is real vocabulary, not junk: 5,538 of them pass our filter and **4,912 are +already shipped** (`kỵ mã`, `khoái hoạt`, `quá tay`, `nặn óc`). Keep them. + +## Through `build-dictionary --words` + +| List | Accepted | Rejected <2 syll | punct | non-VN | +|---|---|---|---|---| +| all 43,013 titles | **36,200** | 6,361 | 155 | 162 | +| minus proper-only | **31,637** | 6,051 | 121 | 84 | +| proper-only alone | 4,661 | 310 | 34 | 78 | +| undertheseanlp, plan rule (hongocduc+wiktionary, no tudientv, drop caps-only) | **61,026** | | | | +| undertheseanlp `wiktionary` branch alone (2018 snapshot) | 26,845 | | | | + +Playability of a viwiktionary-only DB vs shipped: + +| metric | shipped | vi-all | vi-no-proper | +|---|---|---|---| +| words | 48,216 | 36,200 | 31,637 | +| syllables | 6,676 | 6,172 | 5,981 | +| syllables that open a word | 5,049 | 4,600 | 4,515 | +| …with ≥2 continuations | 3,682 | 3,246 | 3,168 | +| dead-end syllables | 1,627 | 1,572 | 1,466 | + +Every breadth metric regresses. **Option A (replace) is dead on measurement, not opinion.** + +## Overlaps — what the dump is actually good for + +| | | +|---|---| +| shipped ∩ viwiktionary | **34,570 of 48,216 (72%)** — viwiktionary confirms most of what we ship | +| shipped, not in viwiktionary | 13,646 (9,713 two-syllable) — `thai sản`, `vách núi`, `điện từ trường`, `giả kim`: real words, so absence ≠ junk | +| viwiktionary-only, not shipped | 1,630 | +| shipped words viwiktionary labels proper-only | 4,523 (`ba đồn`, `vị thanh`, `xuân lập` …) | +| **plan corpus (61,026) + vi-no-proper → union** | **64,110 (+3,084)** | +| plan corpus words viwiktionary labels proper-only | 293 — includes `thứ hai`, `tháng năm`, `tin lành`, `tân ước`, `bàn là` | +| 2018 wiktionary branch vs 2026 dump | 26,420 shared; **9,780 new since 2018**; 425 gone | +| shipped words in neither plan corpus nor viwiktionary | 9,267 | + +### Proper-noun label vs capitalization blocklist + +- viwiktionary proper-only ∩ undertheseanlp caps-only: **4,015 of 4,661 (86%)** agree. +- Caps-only words viwiktionary calls common: 139 (incl. `việt nam`). +- Common words viwiktionary calls proper-only: `mặt trời`, `trái đất`, `sao thủy` (celestial + bodies are `Mặt Trời` on-wiki), weekday/month names, `tin lành`, `bàn là`. + +Verdict: capitalization (the plan's rule) is the better *drop* signal. The label is a useful +*second opinion* for the plan's Phase 3 audit sample, nothing more. + +## Licensing + +Wiktionary text is CC BY-SA 4.0 / GFDL. Creative Commons declares CC BY-SA 4.0 **one-way +compatible with GPLv3**, so folding viwiktionary into the GPLv3 `noitu.db` the plan produces is +clean. `data/ATTRIBUTION.md` gains one entry: edition, dump date, URL, checksum. No `NOTICE` +restructure. + +## Recommendation + +1. **Do not switch to viwiktionary.** 36,200 < 40,000 floor; every playability metric regresses. +2. **Land the undertheseanlp plan as written.** Its numbers reproduce (61,026 vs ~61,276). +3. **Then add the dated dump as a second `--words` source** — a follow-up phase, not a change to + the current plan. Concretely: `make fetch-dict` also fetches the pinned dump; a ~100-line + extractor (stdlib, both dialects, redirects skipped, `--drop-proper` off by default) emits a + word list; the builder reads two `--words` files or the Makefile concatenates them. + Expected result: **~64,100 words**, +3,084 modern/compound vocabulary, 12-month freshness via a + monthly-dated pin. +4. **Do not drop on viwiktionary's proper-noun label.** Record it in `meta`/build log as a + count; use it to spot-check the capitalization rule in Phase 3. +5. Keep kaikki out. The dump is smaller than kaikki's JSONL, pins honestly, and the parser is + trivial because we only need titles + section labels, not senses. + +## Addendum — kaikki.org `dictionary/Vietnamese/kaikki.org-dictionary-Vietnamese.jsonl` + +Asked after the main report. This file is **not viwiktionary**: kaikki's `dictionary/` tree is +the *English* Wiktionary, filtered to `lang_code = vi`. Dump 2026-09-02, extracted 2026-09-06, +79,161,423 bytes, SHA-256 `d878bd23fe4d6ac4480736858a85b2bbca0376d6cc9b30395fff0ba1e16cf97b` +at fetch time. Measured the same way as everything above. + +| | | +|---|---| +| rows / distinct words | 51,896 / 45,281 | +| POS | noun 19,206 · character 9,029 · verb 8,992 · adj 7,070 · **name 3,814** · adv 1,160 · … | +| rejected `< 2 syllables` | 22,115 — half the file is single syllables and Hán/Nôm characters | +| **accepted, all** | **22,628** | +| **accepted, `name`-only words dropped** | **21,244** | + +| metric | shipped | kaikki-en | vi-no-proper | uts-plan | +|---|---|---|---|---| +| words | 48,216 | 21,244 | 31,637 | 61,026 | +| syllables | 6,676 | 5,145 | 5,981 | 7,005 | +| openers ≥2 | 3,682 | 2,431 | 3,168 | 4,164 | +| dead-end | 1,627 | 1,399 | 1,466 | 1,513 | + +Smallest of every candidate as a corpus. **As a third supplement it is the best one:** + +| overlap | words | +|---|---| +| kaikki-en ∩ shipped | 20,231 (adds 1,013) | +| kaikki-en adds over plan corpus | 4,394 | +| kaikki-en adds over plan corpus ∪ viwiktionary | **3,693** | +| **union plan ∪ viwiktionary ∪ kaikki-en** | **67,803** | +| shipped words in none of the three | 6,187 | + +The additions are exactly the modern register the other two lack: `sổ hồng`, `quay xe`, `thi hành +án`, `nhà máy lọc dầu`, `cát tặc`, `giấy ướt`, `máy tính tiền`, `thập lục phân`, `đường tiêu hóa`, +`tân tổng thống`, `sói đồng cỏ`. English Wiktionary's Vietnamese section is more actively curated +than viwiktionary's. + +**Its `name` POS is the cleanest proper-noun label seen so far.** `hà nội`, `việt nam`, `trái đất` +are `name`; `thứ hai`, `tin lành` are not; `mặt trời` is both `name` and `noun` and so survives. +Of 1,483 accepted name-only words, 1,359 are shipped today and 247 survive the plan's caps rule +(`thượng đế`, `bắc cực`, `trung thu`, `siêu nhân` — arguable either way). Good audit column; still +not a sole drop rule. + +**Two blockers to using the URL as given:** + +1. **It is deprecated.** kaikki's index marks the postprocessed per-language file as *"deprecated + and will be removed in the future"* and points to the raw download page, where the only + non-deprecated form of this data is the full-edition `raw-wiktextract-data.jsonl` + (23.1 GB, 2.7 GB compressed) filtered by `lang_code`. There is no per-language file in the + replacement layout — the `downloads//` entries are per *edition*, not per language. +2. **It cannot be pinned.** Rolling URL, refreshed weekly, no archived snapshots. A + `DICT_SHA256`-style pin breaks on the next refresh; `make fetch-dict` would fail weekly. + +Ways to use it anyway, in order of preference: + +- **Vendor the derived word list.** Run the reduction once, commit `data/sources/kaikki-en-vi.txt` + (~21k lines, ~400 KB) with the fetch date, URL and file SHA-256 in `ATTRIBUTION.md`. Pins + honestly, costs nothing at build time, refreshed deliberately. Note this is a *derived list*, + not a copy of their file, so the plan's "never commit the source file" rule is not violated in + spirit — decide explicitly. +- **Pin the enwiktionary dump and extract ourselves.** `enwiktionary-20260902-pages-articles` + is >1 GB and wiktextract with Lua expansion takes hours. Correct but disproportionate for + ~3.7k words. +- **Re-pin on every refresh.** Not acceptable; turns a data pin into a weekly chore. + +License: same CC BY-SA 4.0 / GFDL as all Wiktionary text, so the GPLv3 result stays clean. +Attribution should name enwiktionary + wiktextract/kaikki (the extraction is Tatu Ylonen's work). + +**Verdict:** yes, usable, and worth ~3,700 modern words on top of plan + viwiktionary — but only +via a vendored derived list, never as a live fetch of that URL. + +## Addendum — undertheseanlp `wiktionary` branch alone (the option chosen for the plan) + +Measured after the user chose to use undertheseanlp with its wiktionary data only. Rows with +`wiktionary` in `source`: 32,484 → 32,374 lowercase forms → 4,883 caps-only dropped (case +evidence from wiktionary rows only) → 27,491 → **22,310 accepted**. + +| metric | shipped | uts-wik only | vi-dump 2026 | wik ∪ dump | uts-plan (GPL) | +|---|---|---|---|---|---| +| words | 48,216 | **22,310** | 31,637 | 31,817 | 61,026 | +| syllables | 6,676 | 5,484 | 5,981 | 6,012 | 7,005 | +| openers | 5,049 | 4,050 | 4,515 | 4,539 | 5,492 | +| openers ≥2 | 3,682 | 2,686 | 3,168 | 3,180 | 4,164 | +| dead-end | 1,627 | 1,434 | 1,466 | 1,473 | 1,513 | + +- shipped ∩ uts-wik 21,230 · shipped lost 26,986 (4,335 proper-noun drops, 22,651 absent + from the branch) · new 1,080. +- The 2018 branch is a near-subset of the 2026 dump: union adds 180 words. Same license. +- **Case-rule casualties:** 271 caps-only drops have a lowercase form in another branch + (247 in hongocduc): `mặt trời`, `trái đất`, `hệ mặt trời`, `phật giáo`, `nguyên đán`, + `tia x`. Wiktionary holds `Mặt Trời` / `Trái Đất` capitalized only. +- Probes: all 11 known-bad absent ✓. Of 10 ordinary probes, `mặt trời` dropped by the case + rule; `con người`, `cánh diều` not in the branch at all; the other 7 present. +- Decision recorded in `plans/260908-1525-dictionary-corpus-switch/plan.md`: wiktionary-only, + CC BY-SA 4.0 unchanged, floor 20,000, viwiktionary 2026 dump as the upgrade path. + +## Sources + +- https://dumps.wikimedia.org/viwiktionary/ (index; dated dirs 20251220–20260901) +- https://dumps.wikimedia.org/viwiktionary/20260901/ (file list, md5/sha1 sums) +- https://kaikki.org/viwiktionary/ (dump 2026-09-01, extracted 2026-09-06, 56,502 vi senses) +- https://kaikki.org/dictionary/Vietnamese/ and https://kaikki.org/dictionary/rawdata.html (enwiktionary edition; deprecation notice, no snapshots) +- raw wikitext of `Hà Nội`, `học sinh`, `nhà`, `mặt trời` on vi.wiktionary.org +- https://raw.githubusercontent.com/undertheseanlp/dictionary/2c078cfc…/dictionary/words.txt (plan's pin, sha256 verified `4c3e0e61…`) +- Local: `server/cmd/build-dictionary`, `data/noitu.db` + +## Unresolved questions + +- The 9,267 shipped words in neither the plan corpus nor viwiktionary are still an unquantified + junk/real mix; this dump does not settle them. +- Whether Wikimedia publishes a sha256 file for this run later (only md5/sha1 exist today). + Pin on our computed SHA-256 or on the published MD5 — decide when the phase is written. +- Legacy → new wikitext migration pace: if the wiki finishes it, the legacy branch of the parser + becomes dead code; harmless, but the pin test should assert the extracted count so a dialect + the parser misses shows up as a drop.