docs(plans): record the dictionary corpus switch, its audit and the source research

This commit is contained in:
tiennm99 committed 2026-09-08 16:37:45 +07:00
1 parent 39ef45730f
commit 10a47ebc76
10 files changed
+835 -285

No files matched your search

@@ -0,0 +1,184 @@
# Audit: proper-noun drops and the wiktionary-only corpus
Date: 2026-09-08. Phase 3 of `plan.md`. Every number here is reproducible from the
commands listed; the shipped `data/noitu.db` was copied aside before any build and is the
"old" side of every comparison.
## Commands
```sh
# from server/, after `make fetch-dict`
go run ./cmd/build-dictionary --merged ../data/undertheseanlp-words.jsonl \
--out /tmp/wik.db --report-drops /tmp/drops-wik.txt
go run ./cmd/build-dictionary --merged ../data/undertheseanlp-words.jsonl \
--sources hongocduc,wiktionary --out /tmp/hw.db --report-drops /tmp/drops-hw.txt
# sample: python, random.seed(1), random.sample(sorted(drops-wik), 100)
# casualties: (drops-wik − drops-hw) ∩ words(hw.db)
go test ./internal/bot/ -run RealCorpus -v # once per database at data/noitu.db
```
## Build result
```
rejected 97: contains punctuation
rejected 5074: fewer than 2 syllables
rejected 37: no Vietnamese letters
rejected 4747: only ever capitalized
accepted 22419 distinct words (sources: wiktionary)
```
22,419 vs the 22,310 the research report measured. The delta is the rule: the report
dropped any form containing an uppercase letter, the builder drops forms that *begin* with
one, so `tia X`, `bộ bài tây`, `chữ Hán`, `châu Á` survive. All are ordinary words.
## 1. Drop audit — 100 sampled, seed 1
Verdicts: **P** proper noun (correct drop) · **X** not a Vietnamese word (would be rejected
by the filter anyway) · **U** unclear · **C** common word (false positive).
| | | | | |
|---|---|---|---|---|
| a yun P | an hoà tây P | an minh bắc P | an thuận P | ba vinh P |
| ban cơ U | brao X | bàn đạt P | bành trạch P | bàu năng P |
| bá xuyên P | bình hiệp P | bình tấn P | bằng la P | chiềng sơ P |
| châu hội P | chù lá phù lá P | chư krêy P | cur X | côn đảo P |
| cơ kiều P | cẩm trung P | hùng vương P | hồng bàng P | irc X |
| khánh gia P | lão quân P | lương đài P | mặc dương P | nhơn hoà lập P |
| ninh lai P | ninh thạnh P | noong bua P | nội thôn U | phiếu hữu mai U |
| phù lảng P | phương cao kén ngựa U | phần lão U | quán hành P | quắc hương P |
| rã bản P | sam mứn P | sơn hạ P | sơn vy P | tam ngọc P |
| thanh liên P | thanh văn P | thiệu đô P | thái an P | thượng lâm P |
| thượng tiến P | thạch xuân P | thới quản P | tiên cát P | tiên thọ P |
| tri lễ P | triệu giang P | triệu việt vương P | trung tú P | trà khê P |
| trà kót P | trà linh P | trường khánh P | trường long P | trần đình thâm P |
| trọng do P | tuân lộ P | tân hưng P | tân mỹ P | tân nhựt P |
| tân trì P | tây hiếu P | tăng sâm P | tĩnh gia P | tịnh an P |
| tốt động P | vinh an P | việt–mường X | vân canh P | võ trường toản P |
| võ văn tồn P | văn môn P | văn quân P | văn đình dận P | vĩnh hoà hưng bắc P |
| vĩnh lương P | vĩnh lộc b P | vĩnh trinh P | vũ thư P | vạn thuỷ P |
| xuân mai P | xuân quan P | xá bung P | xá cẩu P | xá khắc P |
| xốp cộp P | yên cát P | ô qua U | đồng nai P | đổ rượu ra sông thết quân lính C |
**P 89 · X 4 · U 6 · C 1.** Gate is ≤2 common words in 100: **passes.** The bulk is
commune names, ethnonyms and historical persons — the content the rule exists to remove.
## 2. Case casualties — 215, listed in full
Words the wiktionary rows hold only capitalized, which another branch has lowercase. They
are dropped under the plan's wiktionary-only evidence rule. Full list:
`case-casualties.txt` in the session scratchpad; reproduced here because it is the
decision input.
an bình · an dân · an hảo · an khang · an lạc · ba tiêu · biển hồ · bu lu · **bàn là** ·
**báo đáp** · bát tiên · bình chuẩn · bình chân · bình khang · bình nghị · bình sa · bình
thanh · bình thuỷ · bình trị · bình tâm · bình văn · bình điền · bích đào · bóng chim tăm
cá · bông trang · bạch cung · bản nguyên · **bảo toàn** · bắc phong · bẻ quế · bố chính ·
**bồ đề** · bồng sơn · cao biền dậy non · cao kỳ · cao nhân · cao sơn · cao tổ · cao xanh ·
cao đường · chà và · chày sương · chánh hội · **chân mây** · châu lệ · châu thành · chí
thiện · chí thành · chính tâm · **chúa nhật** · chúa trời · chị hằng · con tạo · cát lũy ·
cát tân · cát đằng · **công bình** · công dã tràng · **công giáo** · **công nguyên** ·
**cơ đốc giáo** · cường thịnh · cầm đuốc chơi đêm · cầu lam · cầu lộc · cầu ô · cẩm châu ·
cẩm đường · cửu kinh · **cựu ước** · diêm vương · diêm vương tinh · dã hạc · dương quan ·
dương đài · **dường như** · **giao tử** · giấc mai · giọt châu · giọt tương · hoa cái ·
**hoa kiều** · **hoà đồng** · hoá công · hoả tinh · hán học · hán tộc · hán tự · hán văn ·
hạ thần · hải phòng · hải triều · hải vương tinh · hằng nga · **hết sảy** · **hệ mặt
trời** · hội gió mây · hội long vân · **hồi giáo** · **khổng giáo** · kim tinh · **kinh
thánh** · la ve · lâm viên · **lưỡi hái** · lạc hầu · mạnh thường quân · **mặt trời** · nam
lâu · **nguyên đán** · **nho học** · nhân kiệt · năm cha ba mẹ · nước dương · nếm mật nằm
gai · nợ như chúa chổm · phong thu · **phù phiếm** · **phật giáo** · phật học · phật pháp ·
phật tiền · phật tổ · phật tự · **phật đản** · quang phục · quyết tiến · quân thiều · **quả
đất** · **sao hôm** · sao hỏa · sao kim · **sao mai** · sao mộc · song mai · sơn cương ·
sơn lâm · sơn mai · sơn nguyên · sơn trung · sư tử hà đông · **sừng trâu** · tam dân · tam
phủ · tam sơn · thanh nghị · thanh phong · thanh quang · thanh vận · thiên chúa · **thiên
chúa giáo** · thiên vương tinh · thuỷ liễu · thân giáp · thạch bàn · thạch thán · **thần
chết** · **tin lành** · tiên sư · **trái đất** · trướng huỳnh · trường giang · trường xuân ·
tu lý · tuần duyên · tân kỳ · **tân ước** · tây dương · tây thiên · tô hạp · tùng lâm · tế
tân · **tết nguyên đán** · **tổ quốc** · tử phòng · u minh · vinh thăng · **việt ngữ** ·
vách quế · vân hà · vân hán · vân trình · võ miếu · **văn giáo** · **văn miếu** · **văn
nhân** · văn quan · **văn võ** · văn đức · **vũ công** · **vũ hội** · vũng tàu · vương bá ·
vạn an · vạn kiếp · vạn phúc · vầng ô · **vớ bở** · **xe tơ** · xuân hoá · xuân phong ·
xuân tình · xích thố · xương thịnh · yên bình · yên chi · yên hoa · yên hà · đinh điền ·
**đường luật** · **đường thi** · đại danh · đỉnh giáp non thần · **địa cầu** · đỗ vũ
Bold: everyday or dictionary-headword vocabulary by this auditor's reading, ~55 of 215.
The rest are Sino-Vietnamese literary terms, religious/astronomical names that Wiktionary
capitalizes by convention, and genuine place names (`hải phòng`, `vũng tàu`, `bồng sơn`).
**Narrowing does not fix this.** The plan's pre-decided narrowing — drop only when every
syllable is capitalized — was measured: drops fall 4,747 → 4,337, casualties fall only
215 → 147 (`mặt trời`, `trái đất`, `tổ quốc`-style two-cap forms stay dropped), and 410
proper nouns come back in, including junk like `bản mẫu:-vie-n-` and `con kde`. Rejected.
## 3. Corpus diff — old shipped vs new
| | |
|---|---|
| overlap | 21,320 |
| gained | 1,099 (`gân bò`, `rèn đúc`, `ích kỷ`, `đáng lý`, `thở hồng hộc`) |
| lost | **26,896** |
| — dropped as proper noun | 4,245 |
| — absent from the wiktionary branch | 22,651 |
| — other | 0 |
## 4. Playability
| metric | old | new |
|---|---|---|
| words | 48,216 | 22,419 |
| syllables | 6,676 | 5,485 |
| syllables that open a word | 5,049 | 4,052 |
| …with ≥2 continuations | 3,682 | 2,691 |
| dead-end syllables | 1,627 | 1,433 |
Bot-versus-bot, `internal/bot` real-corpus tests, 60 games each:
| | old | new |
|---|---|---|
| hard-vs-easy win rate (moves) | 98% (3.3) | 98% (3.5) |
| medium-vs-easy | 87% | 88% |
| hard-vs-medium (moves) | 72% (3.5) | 55% (2.9) |
| easy-vs-easy game length | 17.5 moves | **11.3 moves** |
| hard decision p95 | 5.2 ms | 1.1 ms |
Both tests pass on both databases. Easy-vs-easy games are a third shorter: the thinner
graph runs out of continuations sooner, which is the depth loss the plan predicted.
Hard-vs-medium falling to 55% means the graph gives the stronger bot fewer ways to trap.
## 5. Decision
Drop rule on the sample: **pass** (1 common word in 100).
Case casualties: **open — needs the owner's call**, because the options trade the
"wiktionary data only" decision against ~55 everyday words:
- **Accept** the 215 losses as measured.
- **Exception list**: a checked-in list of words to keep despite capitalization, curated
from the 215 above. Cause-aligned, small, but a hand-maintained artifact.
- **Read hongocduc for case evidence only**: recovers all 215 mechanically, but uses GPL
data as an input even though none of its words ship. Must be stated in `ATTRIBUTION.md`.
**Recorded decision (owner, 2026-09-08): abandon the capitalization rule.** Every
wiktionary-tagged word is kept regardless of case; proper nouns stay in the corpus as they
are in the shipped database today. The rule, its reject reason and `--report-drops` were
removed from the builder.
## 6. Final corpus, without the rule
```
accepted 26845 distinct words (sources: wiktionary)
```
| metric | old | new |
|---|---|---|
| words | 48,216 | 26,845 |
| syllables | 6,676 | 5,709 |
| syllables that open a word | 5,049 | 4,158 |
| …with ≥2 continuations | 3,682 | 2,787 |
| dead-end syllables | 1,627 | 1,551 |
| overlap / gained / lost | | 25,565 / 1,280 / 22,651 |
Every lost word is absent from the 2018 wiktionary branch; none is lost to a rule.
Bot-versus-bot, 60 games each: hard-vs-easy 95% (2.7 moves), medium-vs-easy 87%,
hard-vs-medium 63% (3.2 moves), easy-vs-easy **12.9 moves** (old: 17.5). Both real-corpus
tests pass.
@@ -1,7 +1,7 @@
---
phase: 1
title: "Phase 1: Read the merged list"
status: todo
status: done
priority: P1
effort: "3h"
dependencies: []
@@ -12,109 +12,107 @@ dependencies: []
## Overview
Teach `build-dictionary` a third input: undertheseanlp's merged JSONL. It selects words by
source membership, drops entries that appear only ever capitalized, and hands the survivors
to the normalization and filtering path that already exists.
source membership and hands them to the normalization and filtering path that already
exists.
## Requirements
- Functional: read `{"text": "...", "source": ["hongocduc", ...]}` lines; keep a word when
its sources intersect the allowed set; drop a word when every raw form of it, across the
allowed sources, begins with an uppercase letter.
- Functional: `tudientv` rows are skipped before any decision is made about a word —
neither its words nor its capitalization inform the output.
- Functional: the proper-noun drop is counted and printed beside the existing reject
reasons, so the build log says what the source contained.
- Non-functional: streaming line-by-line, one pass to gather, one pass to decide. The file
is 4.8 MB; nothing here needs to be clever.
its sources intersect the allowed set — **`wiktionary` only** by default.
- Functional: `hongocduc` and `tudientv` rows are skipped before any decision is made about
a word. They inform nothing.
- Functional: capitalization is not a filter. `Hà Nội` and `hà nội` are the same word and
land as the lowercase form, exactly as every other input mode already behaves.
- Functional: the `--min-words` default moves from 40,000 to 20,000. The new corpus is
~26,800 words; the old floor would reject it.
- Non-functional: streaming line-by-line. The file is 4.8 MB; nothing here needs to be
clever.
- Non-functional: everything downstream — `accept()`, syllable indexing, alias generation,
`verify()`, `meta` — is reused unchanged.
## Architecture
Two maps keyed by the lowercased, whitespace-collapsed word:
```
forms[key] -> set of raw spellings seen in allowed sources ("Mặt Trời", "mặt trời")
sources[key] -> set of allowed sources containing it {"hongocduc", "wiktionary"}
--merged <file> --sources wiktionary
│
▼ one row at a time
sources ∩ allowed ≠ ∅ ? ── no ──▶ skipped, never counted
│ yes
▼
accept() (NFC, lowercase, ≥2 syllables, alphabet, phonotactics)
│
▼
finish() (floor, aliases, atomic write, verify) ← shared with --words and --in
```
A key survives when `sources[key]` is non-empty and at least one entry in `forms[key]`
does not start with an uppercase letter. `Mặt Trời` survives on the strength of its
lowercase twin; `Hà Nội`, `Trương Công Định` and `A Mú Sung` have no lowercase form and go.
The `--sources` flag is kept general — a comma list validated against the three known
names — even though only `wiktionary` is passed, so a comparison build against another
branch is a flag away and never a code patch. Unknown names are an error, not a no-op.
Surviving keys are then fed to the same `accept()` the SQLite path feeds, so the ≥2
syllable rule, the digit/punctuation rejections and the Vietnamese phonotactic check all
apply as they do today.
Case evidence deliberately comes from allowed sources only. The alternative — reading
`tudientv` rows for their capitalization while refusing their words — would be more
accurate and would undercut the claim that we do not use that data. Phase 3 measures what
the stricter choice costs.
`finish()` is new: the three input modes used to each carry their own copy of the floor
check, alias generation, write and verify. One copy is what keeps a fixture from drifting
into a different shape from the database production loads.
## Related Code Files
- Create: `server/cmd/build-dictionary/merged_list.go`
- Create: `server/cmd/build-dictionary/merged_list_test.go`
- Modify: `server/cmd/build-dictionary/main.go` — `--merged` and `--sources` flags, dispatch
in `run()`, `sourceURL`/`sourceLicense` constants, `meta` rows, package doc comment
- Modify: `server/cmd/build-dictionary/filter.go` — add the `rejectProperNoun` reason so
the drop is reported through the existing counter, not a bespoke log line
- Created: `server/cmd/build-dictionary/merged_list.go`
- Created: `server/cmd/build-dictionary/merged_list_test.go`
- Modified: `server/cmd/build-dictionary/main.go` — `--merged` and `--sources` flags,
mutually-exclusive input check, `finish()`, `sourceSpec.url` / `sourceSpec.extra`, meta
rows, package doc comment, `--min-words` default
- Unchanged: `server/cmd/build-dictionary/filter.go`
## Implementation Steps
1. Add `rejectProperNoun rejectReason = "only ever capitalized"` to `filter.go`. Leave
`accept()` alone — the case decision happens before it, on evidence `accept()` cannot
see, since it lowercases.
2. Write `merged_list.go`: decode with `encoding/json` per line (`bufio.Scanner` with a
raised buffer — the longest line is short, but a scanner that silently truncates is a
bad way to lose words), build the two maps, apply the rule, return words plus a count
per reject reason.
3. Add `--merged <file>` and `--sources hongocduc,wiktionary` to `main.go`. Unknown source
names are an error, not a silent no-op: a typo in `--sources` must not quietly ship an
empty or wrong corpus.
4. Dispatch in `run()`: `--merged` takes precedence in the same shape `--words` already
does. Refuse `--merged` together with `--in`/`--words` rather than picking one.
5. Point `sourceURL`, `sourceLicense` and the package doc comment at the new source. Write
`source_commit`, `sources_kept`, `sources_excluded` and `proper_nouns_dropped` into
`meta` — provenance in the artifact, not only in a markdown file.
6. Tests in `merged_list_test.go`, table-driven, over hand-written JSONL fixtures:
- `Mặt Trời` + `mặt trời` → kept
- `Hà Nội` alone → dropped as a proper noun
- a `tudientv`-only word → absent from the output
- a word whose only lowercase form is in `tudientv` → dropped (the documented cost)
- `học sinh` in all three → kept once, not thrice
1. Write `merged_list.go`: decode with `encoding/json` per line (`bufio.Scanner` with a
raised buffer — a scanner that silently truncates is a bad way to lose words), skip rows
with no allowed source, feed the rest to `accept()`.
2. Add `--merged <file>` and `--sources wiktionary` (the default) to `main.go`. Change the
`--min-words` default to 20,000 in the same edit; the fixture path passes its own floor
explicitly and is unaffected.
3. Dispatch in `run()`: exactly one of `--merged`, `--words`, `--in` may be given. Refuse
two rather than picking one.
4. Write `source_url`, `source_commit`, `sources_kept` and `sources_excluded` into `meta` —
provenance in the artifact, not only in a markdown file.
5. Tests in `merged_list_test.go`, table-driven, over hand-written JSONL fixtures:
- `Hà Nội` tagged wiktionary → kept as `hà nội`
- a `hongocduc`-only word and a `tudientv`-only word → absent from the output
- `học sinh` in all three, plus `Học sinh` → kept once
- a malformed line → a clear error naming the line number, not a skipped word
7. Run against the real file and record the counts. Compare to the report's 61,276, which
used all-source case evidence; explain the delta rather than adjusting the number.
- `--sources hongocduc,wiktionary` → strictly more words; `tudientv` still absent
- meta rows as above
6. Run against the real file and record the counts.
## Result
```
rejected 2: contains a digit
rejected 147: contains punctuation
rejected 5335: fewer than 2 syllables
rejected 87: no Vietnamese letters
accepted 26845 distinct words (sources: wiktionary)
generated 1507 spelling aliases
```
The first cut of this phase also implemented a capitalization-based proper-noun drop
(`only ever capitalized` → rejected, `--report-drops` for audit). It reached the Phase 3
gate, was audited, and was removed by the owner's decision — see `plan.md` and
`audit-proper-noun-drops.md`. Nothing of it remains in the code.
## Success Criteria
- [ ] `go test ./cmd/build-dictionary/` passes, new tests included.
- [ ] Building from the real merged file accepts >55,000 words and logs a proper-noun drop
count in the low thousands.
- [ ] Every one of the 11 known-bad probes from the report is absent from the output
(`trương công định`, `trần danh án`, `tân an thạnh`, `cẩm xá`, `vĩnh điện`,
`cam lâm`, `hà nội`, `sài gòn`, `a mú sung`, `a lưới`, `kháng đón`).
- [ ] All 10 ordinary-word probes are present (`học sinh`, `mặt trời`, `con người`,
`bánh mì`, `cánh diều`, `xe đạp`, `giáo viên`, `hoa hồng`, `tình yêu`, `nước mắm`).
- [ ] `--sources hongocduc` alone and `--sources hongocduc,wiktionary` both build, with the
first strictly smaller.
- [ ] No word in the output is reachable only through `tudientv`.
- [x] `go test ./cmd/build-dictionary/` passes, new tests included.
- [x] Building from the real merged file accepts >20,000 words (26,845).
- [x] Ordinary-word probes present: `học sinh`, `bánh mì`, `xe đạp`, `giáo viên`, `hoa
hồng`, `tình yêu`, `nước mắm`, `mặt trời`, `trái đất`. `con người` and `cánh diều` are
not in the 2018 branch at all — expected, recorded.
- [x] `--sources wiktionary` (default) and `--sources hongocduc,wiktionary` both build, with
the first strictly smaller (26,845 vs 61,271 with no case rule).
- [x] No word in the output is reachable only through `hongocduc` or `tudientv`.
## Risk Assessment
**The capitalization rule eats real words.** The probe set is 21 words; the rule fires on
thousands. Signal: Phase 3's hand audit finds common words among the drops. Response:
narrow the rule to require every syllable capitalized (`Hà Nội`, not `Kháng đón`), re-audit,
and if that still misfires, stop dropping and keep the count as a log line only — a corpus
with proper nouns in it is what we ship today and is survivable.
**Allowed-sources-only case evidence drops more than expected.** Signal: the drop count
comes in far above the report's 4,145. Response: measure how many drops have a lowercase
form only in `tudientv`; if that is a large share, reconsider reading `tudientv` for case
evidence alone and say so plainly in `ATTRIBUTION.md` rather than hiding it.
**A JSON schema surprise.** The file is 8 years old and unversioned; a row could carry an
unexpected shape. Signal: decode errors on the real file. Response: the error names the
line — inspect it, and only then decide between a tolerant skip with a counted reason and
a hard failure. Silent skipping is not an option.
a hard failure. Silent skipping is not an option. (Observed: none; all 79,226 rows decode.)
@@ -1,7 +1,7 @@
---
phase: 2
title: "Phase 2: Pin the source"
status: todo
status: done
priority: P1
effort: "2h"
dependencies: [1]
@@ -54,6 +54,8 @@ because that is already the right shape; only the size in the help text changes.
## Implementation Steps
1. Update the three Makefile variables and the `dict` target to pass `--merged $(DICT_SRC)`.
`--sources` is left at its `wiktionary` default so the Makefile and Dockerfile carry
one fewer thing to keep in agreement.
2. Reword `help`, `fetch-dict` and the `$(DICT_SRC)` guard: the download is ~4.8 MB now,
and saying "179 MB" would be the kind of stale comment that outlives three refactors.
3. Mirror all of it in the Dockerfile stage. Keep `FIXTURE_DICT=1` on `--words` — the
@@ -1,7 +1,7 @@
---
phase: 3
title: "Phase 3: Audit the corpus"
status: todo
status: done
priority: P1
effort: "2h"
dependencies: [1, 2]
@@ -9,6 +9,12 @@ dependencies: [1, 2]
# Phase 3: Audit the corpus
> **Outcome.** Executed 2026-09-08; see [`audit-proper-noun-drops.md`](./audit-proper-noun-drops.md).
> The drop rule passed its sample gate (1 common word in 100) but cost 215 words under
> wiktionary-only case evidence, including everyday vocabulary. Narrowing was measured and
> rejected. **The owner abandoned the rule**: every word is kept regardless of case. Final
> corpus 26,845 words; playability and bot results recorded in the audit note.
## Overview
The gate. A rule that removes thousands of words on a capitalization heuristic has to be
@@ -34,12 +40,17 @@ Three measurements, each against the freshly built database and the current one:
passes at ≤2 common words in 100; that tolerance is the difference between a filter and
a corpus edit.
2. **Corpus diff.** Overlap, gained, lost. Split the lost into: dropped as proper nouns,
`tudientv`-only, absent from undertheseanlp. The report measured 4,145 / 1,300 / 6,473
under all-source case evidence; this re-measures under the shipped rule.
absent from the wiktionary branch. The report measured 4,335 / 22,651 of 26,986 lost;
this re-measures from the shipped build. Separately count the **271 case casualties**
— drops that have a lowercase form in `hongocduc` or `tudientv` — and list them in
full; that list is what the pass/narrow/exception decision is made on.
3. **Playability.** Words, syllables, syllables that can open a word, how many have ≥2
continuations, and dead-end syllables — the last of these being what decides whether
the game hands somebody an unanswerable position. Current: 48,216 / 6,676 / 5,049 /
3,682 / 1,627.
3,682 / 1,627. Expected: 22,310 / 5,484 / 4,050 / 2,686 / 1,434. The corpus is smaller
by design, so the question is not "is it bigger" but "does the game still run": add
bot-versus-bot game lengths from `server/internal/bot`'s real-corpus tests on both
databases.
## Related Code Files
@@ -58,20 +69,23 @@ Three measurements, each against the freshly built database and the current one:
hand. Vietnamese place and person names are the expected content; anything that reads
as an ordinary word is a false positive and gets called one.
4. Run the corpus diff and the playability comparison; put both tables in the audit note.
5. Measure the documented cost of allowed-sources-only case evidence: how many drops have
a lowercase form only in `tudientv`.
6. Decide: pass, narrow, or abandon the rule. Record the decision and its reason in the
audit note — and if the rule is narrowed, Phase 1's tests change with it.
5. Build `--sources hongocduc,wiktionary` to a second scratch path purely to enumerate the
case casualties (words dropped under `wiktionary` alone but kept when hongocduc's
lowercase forms are visible). Nothing from that build ships.
6. Decide: pass, narrow, exception list, or abandon the rule. Record the decision and its
reason in the audit note — and if the rule changes, Phase 1's tests change with it.
## Success Criteria
- [ ] The audit note exists, with 100 judged entries and the commands that produced them.
- [ ] ≤2 of 100 sampled drops are ordinary words.
- [ ] The new corpus has >55,000 words.
- [ ] Syllables ≥ 6,676 and dead-end syllables ≤ 1,627 — the graph is not worse than
what players walk today.
- [ ] The three loss buckets are quantified, not estimated.
- [ ] A pass/narrow/abandon decision is written down with its reason.
- [ ] ≤2 of 100 sampled drops are ordinary words, the 271 known case casualties aside —
those are listed in full and judged as a group.
- [ ] The new corpus has >20,000 words.
- [ ] The five graph numbers and bot-game lengths are recorded for both databases; bot
games on the new graph complete without the engine running out of moves earlier
than on today's.
- [ ] The loss buckets are quantified, not estimated.
- [ ] A pass/narrow/exception-list/abandon decision is written down with its reason.
## Risk Assessment
@@ -81,10 +95,14 @@ Response: the drop list is a file — a disputed word can be checked against it
and a per-word exception list is a small change on top of this design.
**The audit fails and the phase becomes a redesign.** Signal: >2 common words in 100.
Response: apply Phase 1's pre-decided narrowing (every syllable capitalized), re-audit
once. If it fails again, ship without the drop — the corpus is still bigger and denser
than today's, and the proper-noun problem stays exactly as bad as it currently is rather
than getting worse.
Response: apply Phase 1's pre-decided narrowing (every syllable capitalized) or the
exception list, re-audit once. If it fails again, ship without the drop — the proper-noun
problem then stays exactly as bad as it is today rather than getting worse.
**The corpus is too thin in play.** 22,310 words is less than half of today's. Signal: bot
games noticeably shorter, or the dead-end rate per move up rather than down. Response: this
is the trigger for the documented upgrade path — the 2026-09-01 viwiktionary dump (31,637
words, same license) as a second `--words` source — not for re-admitting GPL data.
**Playability regresses in a way these five numbers do not capture.** Signal: aggregate
counts look fine but bot games end oddly short or long. Response: `server/internal/bot`'s
@@ -1,110 +0,0 @@
---
phase: 4
title: "Phase 4: Relicense the data"
status: todo
priority: P1
effort: "2h"
dependencies: [3]
---
# Phase 4: Relicense the data
## Overview
`data/noitu.db` becomes GPLv3, because `hongocduc`'s data is GPL and GPL is copyleft. This
phase rewrites the data half of the licensing documents and regenerates the shipped
database — in one commit, because a GPL-derived database sitting in a tree that still
says CC BY-SA 4.0 is the only ordering mistake here with a legal consequence.
## Requirements
- Functional: `data/LICENSE` carries the GPLv3 text.
- Functional: `NOTICE` section 2 describes GPLv3 for the data, names both source branches
with their own licenses, and keeps section 1 (Apache-2.0 for code) intact.
- Functional: `data/ATTRIBUTION.md` records the new source, the per-branch licenses, the
full list of modifications, and why `tudientv` is excluded.
- Functional: the regenerated `data/noitu.db` carries matching `meta` provenance.
- Non-functional: no overclaiming. The `hongocduc` GPL statement comes from a mirror's
README and the document says so.
## Architecture
The dual-license structure already in `NOTICE` is correct and stays; only its second
section changes:
```
1. SOURCE CODE — Apache-2.0 (unchanged: server/, tools/, web/, proto/)
2. DICTIONARY DATA — GPLv3 (was CC BY-SA 4.0)
data/noitu.db, data/LICENSE, data/ATTRIBUTION.md
```
Why GPLv3 rather than CC BY-SA 4.0 is worth stating in `ATTRIBUTION.md` rather than left
implicit: `hongocduc` is Hồ Ngọc Đức's FVDP wordlist, distributed under GNU GPL. The
previous source aggregated that same data and redistributed it as CC BY-SA 4.0, which GPL
does not permit. Ours is the stricter, and correct, direction.
The modifications list gains two entries beyond the current five — source selection and
the proper-noun drop — both of which the license requires us to declare as changes.
## Related Code Files
- Modify: `data/LICENSE` — CC BY-SA 4.0 text replaced with GPLv3
- Modify: `NOTICE` — section 2 rewritten
- Modify: `data/ATTRIBUTION.md` — source table, per-branch licenses, modifications 6 and 7,
the `tudientv` exclusion note
- Modify: `README.md` — any dictionary-source or data-license claim
- Verify: `.github/workflows/ci.yml` — the required-files check (`data/LICENSE`,
`data/ATTRIBUTION.md`, `NOTICE`, `data/noitu.db`) still holds
- Regenerate: `data/noitu.db`
## Implementation Steps
1. Replace `data/LICENSE` with the GPLv3 text, verbatim and complete.
2. Rewrite `NOTICE` section 2: GPLv3, the affected artifacts, both upstream branches with
their licenses, and the copyleft obligation carrying into container images — the same
point the current text makes about share-alike, which is no less true of GPL.
3. Rewrite `data/ATTRIBUTION.md`:
- source is `undertheseanlp/dictionary` at commit `2c078cf`, file `dictionary/words.txt`
- `hongocduc` (GNU GPL, per the branch README — a mirror of Hồ Ngọc Đức's 2003
wordlist, whose canonical site no longer resolves) and `wiktionary` (CC BY-SA, a
2018-12-10 scrape of vi.wiktionary.org)
- `tudientv` **excluded**: its README declares copyright *"Chưa rõ"* and names
Soha/Vietlex (Hoàng Phê) as its primary source. Say it plainly so nobody reopens it.
- modifications 1–5 carried over, plus **6. source selection** and **7. proper-noun
removal** with the count from Phase 3
- the resulting license and why it is GPLv3 and not CC BY-SA 4.0
4. Grep the repository for `minhqnd`, `CC BY-SA`, `179 MB` and fix every survivor outside
`plans/` — the reports are dated records and stay as written.
5. Regenerate `data/noitu.db` and confirm its `meta` rows agree with `ATTRIBUTION.md`. A
database whose provenance contradicts the attribution file is worse than neither.
6. Run the CI required-files check locally, or read it and confirm by hand.
## Success Criteria
- [ ] `data/LICENSE` is the complete GPLv3 text.
- [ ] `NOTICE` section 2 says GPLv3 and names both branches with their licenses; section 1
is unchanged.
- [ ] `data/ATTRIBUTION.md` lists seven modifications, states the `tudientv` exclusion and
its reason, and does not claim more about the `hongocduc` license than the mirror
supports.
- [ ] `meta` in the built database matches the attribution file — source URL, commit,
license, sources kept and excluded, proper-noun drop count.
- [ ] No file outside `plans/` still describes the data as CC BY-SA 4.0 or names minhqnd
as the source.
- [ ] CI's required-files check passes.
## Risk Assessment
**Relicensing is hard to walk back.** Once a GPLv3 `noitu.db` is published, that release
is GPLv3 permanently. Signal: none — this is simply true. Response: it is why the user
signed off before this plan was written; the mitigation is that it was a decision rather
than a side effect.
**A stale CC BY-SA claim survives somewhere.** A README line or a Docker label saying the
wrong license is a licensing misstatement, not a typo. Signal: the grep in step 4 finds
something after the phase is called done. Response: grep is a step in the phase for that
reason; run it last, not first.
**Section 1 gets damaged while section 2 is rewritten.** The code's Apache-2.0 grant is
not in scope and must come out byte-identical. Signal: a diff touching section 1.
Response: review the `NOTICE` diff before committing.
@@ -0,0 +1,116 @@
---
phase: 4
title: "Phase 4: Update the attribution"
status: done
priority: P1
effort: "1h"
dependencies: [3]
---
# Phase 4: Update the attribution
> **Outcome.** Done 2026-09-08. `data/LICENSE` untouched; `NOTICE` section 2 names the new
> upstream, `data/ATTRIBUTION.md` rewritten (eight modifications, both attribution links,
> both exclusions). Also updated, beyond the plan's list: the frontend attribution footer
> (`AttributionFooter.svelte`, `i18n/vi.js`) and its e2e assertion now credit Wiktionary
> tiếng Việt instead of minhqnd — the user-visible half of the CC BY-SA obligation.
## Overview
The data license does not change — the wiktionary branch is CC BY-SA 4.0, the same license
`noitu.db` already carries — but everything that names the source does. This phase rewrites
the attribution and notice text and regenerates the shipped database in one commit, so the
attribution never describes a database other than the one in the tree.
## Requirements
- Functional: `data/LICENSE` is untouched — still the CC BY-SA 4.0 text.
- Functional: `NOTICE` section 2 names the new upstream, keeps CC BY-SA 4.0, and keeps
section 1 (Apache-2.0 for code) byte-identical.
- Functional: `data/ATTRIBUTION.md` records the new source, the branch actually used, the
two branches excluded and why, and the full list of modifications.
- Functional: the regenerated `data/noitu.db` carries matching `meta` provenance.
- Non-functional: no overclaiming. The branch is a 2018 scrape of vi.wiktionary.org made by
a third party; say that, and attribute Wiktionary's contributors as CC BY-SA requires.
## Architecture
The dual-license structure in `NOTICE` is correct and stays; only the source lines of its
second section change:
```
1. SOURCE CODE — Apache-2.0 (unchanged: server/, tools/, web/, proto/)
2. DICTIONARY DATA — CC BY-SA 4.0 (unchanged license, new upstream)
data/noitu.db, data/LICENSE, data/ATTRIBUTION.md
```
The attribution chain is now two links instead of an opaque aggregate: vi.wiktionary.org
contributors (CC BY-SA) → `undertheseanlp/dictionary`, branch data `wiktionary`, commit
`2c078cf` → this project. Both links are named.
The modifications list stays at eight: source selection replaces the old language filter
as entry 1, which the license requires us to declare as a change.
## Related Code Files
- Modify: `NOTICE` — section 2's upstream line, affected-artifacts line (`dictionary.db`
is no longer the source), and the closing paragraph about build artifacts
- Modify: `data/ATTRIBUTION.md` — source table, branch selection, exclusions, modifications
9 and 10, the reproduce block
- Modify: `README.md` — dictionary-source lines (`~179 MB`, `dictionary.db`, `minhqnd`,
the manual `curl` + `--in` example)
- Modify: `docs/deployment.md` — the 179 MB builder-stage sentence
- Verify unchanged: `data/LICENSE`, `.github/workflows/ci.yml` required-files check
- Regenerate: `data/noitu.db`
## Implementation Steps
1. Rewrite `NOTICE` section 2's source lines: upstream `undertheseanlp/dictionary` at
commit `2c078cf`, wiktionary data only; affected artifact `data/noitu.db`. Leave the
share-alike paragraph as is — it is still true.
2. Rewrite `data/ATTRIBUTION.md`:
- source is `undertheseanlp/dictionary`, file `dictionary/words.txt`, commit `2c078cf`
- rows used: those tagged `wiktionary` — a 2018-12-10 scrape of vi.wiktionary.org,
whose text is CC BY-SA 4.0 / GFDL by its contributors
- `hongocduc` **excluded**: GNU GPL per the branch README; using it would relicense this
database to GPLv3
- `tudientv` **excluded**: its README declares copyright *"Chưa rõ"* and names
Soha/Vietlex (Hoàng Phê) as its primary source. Say it plainly so nobody reopens it.
- modifications renumbered: **1. source selection** replaces the language filter (the new
file is Vietnamese-only); the other seven carry over, with normalization stating that
capitalization removes nothing — the proper-noun drop was abandoned at the Phase 3 gate
- reproduce block: `make fetch-dict` is ~4.8 MB now
3. Grep the repository for `minhqnd`, `dictionary.db`, `179` and `--in` and fix every
survivor outside `plans/` — the reports are dated records and stay as written.
4. Regenerate `data/noitu.db` and confirm its `meta` rows agree with `ATTRIBUTION.md`:
source URL, commit, license, `sources_kept = wiktionary`,
`sources_excluded = hongocduc,tudientv`.
5. Run the CI required-files check locally, or read it and confirm by hand.
## Success Criteria
- [ ] `data/LICENSE` has no diff.
- [ ] `NOTICE` section 2 names the new upstream and still says CC BY-SA 4.0; section 1 is
byte-identical.
- [ ] `data/ATTRIBUTION.md` lists eight modifications, names the branch used and the two
excluded with reasons, and credits vi.wiktionary.org contributors.
- [ ] `meta` in the built database matches the attribution file.
- [ ] No file outside `plans/` names minhqnd, `dictionary.db` as a source, or a 179 MB
download.
- [ ] CI's required-files check passes.
## Risk Assessment
**A stale source claim survives somewhere.** A README line or Docker comment naming the old
upstream is an attribution misstatement, not a typo. Signal: the grep in step 3 finds
something after the phase is called done. Response: grep is a step in the phase for that
reason; run it last, not first.
**Section 1 gets damaged while section 2 is edited.** The code's Apache-2.0 grant is not
in scope and must come out byte-identical. Signal: a diff touching section 1. Response:
review the `NOTICE` diff before committing.
**The attribution credits the wrong party.** CC BY-SA attribution belongs to Wiktionary's
contributors; undertheseanlp is the intermediary that scraped and redistributed. Signal:
an `ATTRIBUTION.md` that names only the GitHub repository. Response: step 2 names both
links of the chain explicitly.
@@ -1,7 +1,7 @@
---
phase: 5
title: "Phase 5: Retire the SQLite path"
status: todo
status: done
priority: P2
effort: "1h"
dependencies: [4]
@@ -9,6 +9,13 @@ dependencies: [4]
# Phase 5: Retire the SQLite path
> **Outcome.** Done 2026-09-08. `--in`, `--table`, `--word-col`, `--lang-col`, `--lang`,
> `resolveSource`, `validateSource`, `listTables`, `listColumns`, `pickColumn`, `extract`
> and `quoteIdent` removed; `meta` no longer writes `source_word_column` /
> `source_lang_column` (no consumer existed). The four auto-detection tests went; the
> output-shape tests now build their fixture through `--merged`. `--help` lists `--merged`,
> `--sources`, `--words`, `--out`, `--max-syllables`, `--min-words` and nothing else.
## Overview
Delete the input mode nothing uses any more: reading a source SQLite database, with the
@@ -55,7 +62,8 @@ no longer read that upstream.
2. Delete the tests that exist only to cover auto-detection. Do not delete tests covering
output shape, `verify()`, or the reject-reason counts — those still describe behavior.
3. Rewrite the package doc comment: the input is a 4.8 MB JSONL wordlist with source
membership and capitalization; the output is the game's syllable-indexed database.
membership and capitalization, of which only wiktionary-tagged rows are read; the
output is the game's syllable-indexed database.
4. `go build ./... && go vet ./... && go test ./...` from `server/`.
5. Grep for `--in`, `dictionary.db`, `resolveSource` and `lang_code` outside `plans/`;
nothing should survive except in the dated reports.
@@ -1,11 +1,12 @@
---
title: "Dictionary corpus switch"
description: "Replace the 179 MB minhqnd aggregate with undertheseanlp/dictionary (hongocduc + wiktionary, tudientv excluded), drop always-capitalized proper nouns, and relicense data/noitu.db to GPLv3"
status: pending
description: "Replace the 179 MB minhqnd aggregate with the wiktionary branch of undertheseanlp/dictionary only (hongocduc and tudientv excluded) and keep data/noitu.db on CC BY-SA 4.0"
status: in-progress
priority: P1
effort: "~1d"
tags: [dictionary, data, licensing, build]
created: 2026-09-08
updated: 2026-09-08
blockedBy: []
blocks: []
---
@@ -14,115 +15,141 @@ blocks: []
## Overview
`data/noitu.db` is derived from `minhqnd/dictionary` v2.0.0: a 179 MB SQLite aggregate of
five upstreams, redistributed as CC BY-SA 4.0. Two problems. Its license chain does not
close — three of its five inputs are GPL (FVDP, tudientv, "Vietnamese Explanatory
Dictionary") and GPL does not permit relicensing to CC BY-SA. And it has already
lowercased everything (0 of 70,511 Vietnamese rows carry uppercase), which destroys the
one signal that separates a place name from a word.
`data/noitu.db` was derived from `minhqnd/dictionary` v2.0.0: a 179 MB SQLite aggregate of
five upstreams, redistributed as CC BY-SA 4.0. Its license chain does not close — three of
its five inputs are GPL (FVDP, tudientv, "Vietnamese Explanatory Dictionary") and GPL does
not permit relicensing to CC BY-SA.
This plan switches the corpus to `undertheseanlp/dictionary` — a single 4.8 MB JSONL file
carrying, per word, its raw capitalization and which of three source dictionaries contain
it. `hongocduc` (GPL) and `wiktionary` (CC BY-SA) are taken; `tudientv` is excluded
permanently because its own README declares copyright *"Chưa rõ"* and names Soha/Vietlex
(Hoàng Phê) as its primary source. Entries that appear only ever capitalized are dropped
as proper nouns. `data/noitu.db` then ships **GPLv3** — the honest license for
FVDP-derived data, and the reason this is a licensing improvement rather than a new risk.
This plan switches the corpus to the **`wiktionary` source inside `undertheseanlp/dictionary`**
— a single 4.8 MB JSONL file carrying, per word, which of three source dictionaries contain
it. Only rows whose `source` list contains `wiktionary` are read. `hongocduc` (GPL) is
excluded so the data license does not change; `tudientv` is excluded because its own README
declares copyright *"Chưa rõ"* and names Soha/Vietlex (Hoàng Phê) as its primary source.
`data/noitu.db` **stays CC BY-SA 4.0**: the wiktionary branch is a 2018-12-10 scrape of
vi.wiktionary.org, whose text is CC BY-SA.
Evidence, all measured through the real `build-dictionary` filter:
`plans/reports/research-260908-1507-undertheseanlp-dictionary.md`.
`plans/reports/research-260908-1507-undertheseanlp-dictionary.md` (branch licensing),
`plans/reports/research-260908-1529-viwiktionary-dump-measured.md` (yields, alternatives)
and [`audit-proper-noun-drops.md`](./audit-proper-noun-drops.md) (the Phase 3 gate).
## What changes, in numbers
| | current | after |
| | before | after |
|---|---|---|
| words | 48,216 | **~61,276** |
| syllables | 6,676 | 7,011 |
| syllables that can open a word | 5,049 | 5,496 |
| …with ≥2 continuations | 3,682 | 4,168 |
| dead-end syllables | 1,627 | **1,515** |
| words | 48,216 | **26,845** |
| syllables | 6,676 | 5,709 |
| syllables that can open a word | 5,049 | 4,158 |
| …with ≥2 continuations | 3,682 | 2,787 |
| dead-end syllables | 1,627 | 1,551 |
| source download | 179 MB SQLite | 4.8 MB JSONL |
| data license | CC BY-SA 4.0 | **GPLv3** |
| data license | CC BY-SA 4.0 | **CC BY-SA 4.0 (unchanged)** |
| `--min-words` floor | 40,000 | **20,000** |
24,978 words gained, 11,918 lost — of which 4,145 are proper nouns removed on purpose,
1,300 are `tudientv`-only, and 6,473 are absent from undertheseanlp entirely, including
genuine modern vocabulary (`tế bào gốc`, `hằng số vũ trụ`) that a 2003 wordlist cannot
have. That loss is accepted here and is the subject of a separate future decision, not
this plan.
**This is a deliberate 44% cut.** 1,280 words gained, 22,651 lost — every one of them
simply absent from a 2018 wiktionary scrape (`bánh đúc`, `lò vi sóng`, `dân làng`, `thót
tim`). Bot-versus-bot games run about a quarter shorter on the thinner graph (easy-vs-easy
12.9 moves vs 17.5). This loss is accepted here; the upgrade path is written down below.
## Decisions taken
- **`tudientv` is never read.** Not as content, and not as case evidence. Excluding it
from evidence too is what makes "we do not touch that data" true without qualification.
It costs a little accuracy in the proper-noun rule — a word capitalized in `hongocduc`
but lowercase only in `tudientv` will be dropped — and Phase 3 measures that cost.
- **The measured 61,276 used all-source case evidence.** Under the allowed-sources-only
rule the number will shift slightly. Phase 1 re-measures; the plan does not depend on
the exact figure, only on it staying above the 40,000 `--min-words` floor and above
today's 48,216.
- **Proper nouns are dropped, not tiered.** The per-word `source` array looked like a
confidence signal and is not one: `hà nội` has 2 votes, `cẩm xá` 2, `kháng đón` 2, while
`cánh diều` has 1 and `mặt trời` 2. Capitalization is the only usable signal in this
data, so it is the only one used.
- **`noitu.db` becomes GPLv3; the code stays Apache-2.0.** The existing dual-license split
in `NOTICE` holds; only its data half is rewritten.
- **The SQLite input path is deleted, not kept "just in case."** Two input modes for one
source is dead weight; the fixture path (`--words`) stays because the e2e suite and the
Docker smoke build use it.
- **The viwiktionary extractor is out of scope.** It remains the answer if a CC BY-SA
corpus is ever needed again — see `research-260908-1434-viwiktionary-as-dictionary-source.md`.
- **Only `wiktionary` rows are read.** `hongocduc` and `tudientv` inform nothing. That is
what keeps the license CC BY-SA 4.0 and what makes "we do not touch that data" true
without qualification.
- **Capitalization is not a filter.** The original design dropped words that appear only
ever capitalized, as proper nouns. It was built, run and audited at the Phase 3 gate: the
rule was precise on a 100-word sample (1 common word) but, with wiktionary-only case
evidence, it also removed 215 words including `mặt trời`, `trái đất`, `tổ quốc`, `dường
như`, `phật giáo`. Narrowing was measured and did not help. **The owner chose to keep
every word regardless of case.** Proper nouns like `hà nội` remain playable, exactly as
they are today. The mechanism was removed from the builder rather than left behind a flag.
- **Rejected alternatives, so nobody reopens them without new evidence:**
- `hongocduc + wiktionary` (61,271 words) — forces `noitu.db` to GPLv3. Rejected to keep
CC BY-SA 4.0.
- the 2026-09-01 viwiktionary dump (36,200 words without a case rule, same license,
monthly pin) — the 2018 branch is a near-subset of it. Rejected for now in favour of the
simpler single-file pin; **this is the documented upgrade path** if the 27k corpus
proves too thin in play. See the 15:29 report.
- kaikki.org enwiktionary Vietnamese — unpinnable rolling file, deprecated by its host.
- **The SQLite input path is deleted, not kept "just in case."** The fixture path
(`--words`) stays because the e2e suite and the Docker smoke build use it.
- **The build floor drops to 20,000.** The old 40,000 default would reject the new corpus;
it moved once, in `main.go`.
## The pin
```
URL https://raw.githubusercontent.com/undertheseanlp/dictionary/2c078cfc373b06e2980d324ce1d7bd13740c3319/dictionary/words.txt
SHA256 4c3e0e6117e4bdfa97731e135c3d4a05881889909267394a8de8d88ef79f13f0
Size 4,813,111 bytes — 79,226 JSONL rows
Size 4,813,111 bytes — 79,226 JSONL rows, 32,484 of them tagged wiktionary
Commit 2c078cf (2018-12-10, the repository's last)
```
Pinned by commit rather than branch: `raw.githubusercontent.com/<repo>/<sha>/` is
immutable, so the URL and the checksum cannot disagree later. The variable names
`DICT_URL` / `DICT_SHA256` are kept so `web/tests/dictionary-source.test.js` needs no
change — only their values move.
`DICT_URL` / `DICT_SHA256` were kept so `web/tests/dictionary-source.test.js` needed no
change — only their values moved.
## Phases
| # | Phase | Status |
|---|-------|--------|
| 1 | [Read the merged list](./phase-01-start.md) | Pending |
| 2 | [Pin the source](./phase-02-pin-the-source.md) | Pending |
| 3 | [Audit the corpus](./phase-03-audit-the-corpus.md) | Pending |
| 4 | [Relicense the data](./phase-04-relicense-the-data.md) | Pending |
| 5 | [Retire the SQLite path](./phase-05-retire-the-sqlite-path.md) | Pending |
| 1 | [Read the merged list](./phase-01-start.md) | Done |
| 2 | [Pin the source](./phase-02-pin-the-source.md) | Done |
| 3 | [Audit the corpus](./phase-03-audit-the-corpus.md) | Done — gate sent the case rule back; rule abandoned |
| 4 | [Update the attribution](./phase-04-update-the-attribution.md) | Done |
| 5 | [Retire the SQLite path](./phase-05-retire-the-sqlite-path.md) | Done |
Phase 3 is a gate: it may send Phase 1 back to narrow the proper-noun rule. Phase 4 must
land in the same commit as the first generated GPL-derived `noitu.db` — shipping that
database while `NOTICE` still says CC BY-SA 4.0 is the one ordering mistake with a legal
consequence.
All five phases are uncommitted in one working tree; they should land together so the
attribution never describes a database other than the one the build produces.
## Success criteria
- [ ] `make fetch-dict && make dict` builds `data/noitu.db` from a 4.8 MB pinned file with
- [x] `make fetch-dict && make dict` builds `data/noitu.db` from a 4.8 MB pinned file with
no 179 MB download.
- [ ] The shipped database has **more than 55,000 words** and **no word that reaches it
only via `tudientv`**.
- [ ] The proper-noun drop count is reported in the build log beside the other reject
reasons, recorded in `meta`, and a 100-entry sample has been audited by hand.
- [ ] Playability does not regress: syllables ≥ 6,676 and dead-end syllables ≤ 1,627.
- [ ] `data/LICENSE`, `NOTICE` and `data/ATTRIBUTION.md` describe GPLv3, name both source
branches with their own licenses, and record why `tudientv` is excluded.
- [ ] `make test-go`, `make test-web` and the e2e suite are green; the fixture build path
is untouched.
- [ ] `--in` and its schema auto-detection are gone, and no doc still mentions the 179 MB
download.
- [x] The shipped database has **more than 20,000 words** and **no word that reaches it
through `hongocduc` or `tudientv` alone**.
- [x] `meta` records the source URL, commit, sources kept and sources excluded.
- [x] Playability is measured and recorded: syllables, openers, continuations, dead ends,
and bot-versus-bot game lengths on the new graph versus the old.
- [x] `data/LICENSE` is unchanged; `NOTICE` and `data/ATTRIBUTION.md` name the new source,
its commit and license, and record why the two other branches are excluded. The
frontend attribution footer credits Wiktionary tiếng Việt.
- [x] `go vet ./... && go test ./...` (incl. `-race` on the touched packages), `npm run
check && npm test`, and the Docker image (`noitu:ci` with the CI file checks, plus
`FIXTURE_DICT=1`) are green; the fixture build path is untouched.
- [ ] The Playwright e2e suite is green. **Not verified locally**: Playwright's Chromium
download timed out twice on this machine (2026-09-08), so every test failed at browser
launch, not on game behaviour. The one e2e assertion this plan changes (the attribution
footer's link text and href in `web/e2e/bot-game.spec.js`) was updated by hand. Verify
in CI or after `npx playwright install chromium` succeeds.
- [x] `--in` and its schema auto-detection are gone, and no file outside `plans/` still
mentions minhqnd, `dictionary.db`, `--in` or a 179 MB download.
## Licence version — decided after review
The owner asked whether the data could move to Apache-2.0 to simplify the repo. It cannot:
the words are Wiktionary contributors' text under a share-alike licence, and neither
undertheseanlp nor this project can relicense them. The owner then chose to **keep the
licence the source text actually carried, CC BY-SA 3.0 Unported**, rather than upgrade the
derivative to 4.0 under the later-version clause. `data/LICENSE` is now the BY-SA 3.0 legal
code; `NOTICE`, `ATTRIBUTION.md`, README, Dockerfile, deployment doc, the builder's meta
string and the in-game footer all say 3.0. The "unchanged" wording above predates this.
## Review
Code review (2026-09-08) returned DONE_WITH_CONCERNS; every finding was applied: the
builder's copy of the upstream commit is now guarded by `web/tests/dictionary-source.test.js`
alongside the Makefile/Dockerfile pin; fixture builds record "no upstream data" instead of
inheriting the CC BY-SA string; `ATTRIBUTION.md` states the 2018 snapshot was CC BY-SA 3.0
and why the derived database is 4.0; dead `contains()`, redundant sorts, the off-by-one
scanner line number, stale "48k" comments, the Makefile `curl` flags and the Dockerfile
download name were fixed; input-selection and fixture-licence tests were added.
## Open questions
- The `hongocduc` GPL claim rests on the mirror's README; Hồ Ngọc Đức's canonical site
404s. Assuming GPL is the conservative direction, so this does not block — but the
evidence is a mirror, and `ATTRIBUTION.md` should say so rather than overclaim.
- The 6,473 words undertheseanlp does not have are an unquantified mix of junk and real
modern vocabulary. Separating them needs a second live source; out of scope here.
- Is 26,845 words enough for the game to feel playable? Bot games are ~25% shorter; only
real play can say whether that is felt. If not, the 2026 viwiktionary dump is the next
step, not a return to GPL data.
<!-- slug: dictionary-corpus-switch -->
@@ -0,0 +1,36 @@
---
title: Dictionary corpus switch to wiktionary-only undertheseanlp
date: 2026-09-08
summary: "Replaced the 179 MB minhqnd aggregate with the wiktionary rows of undertheseanlp/dictionary; corpus 48,216 -> 26,845 words, CC BY-SA 4.0 kept, proper-noun drop abandoned after audit, SQLite path retired"
---
# Dictionary corpus switch to wiktionary-only undertheseanlp
## What happened
- Three research passes today measured every candidate through the real `build-dictionary` filter: the 2026-09-01 viwiktionary dump (36,200 words), kaikki's enwiktionary Vietnamese file (22,628, deprecated and unpinnable), undertheseanlp's merged list (61,271 but GPL via hongocduc) and its wiktionary rows alone (26,845). Reports: `plans/reports/research-260908-1529-viwiktionary-dump-measured.md`.
- Owner chose undertheseanlp wiktionary-only to keep `data/noitu.db` on CC BY-SA 4.0. Plan `plans/260908-1525-dictionary-corpus-switch/` rewritten around it, then executed end to end.
- `build-dictionary` gained `--merged` / `--sources` (default `wiktionary`, names validated), a shared `finish()` tail, `--min-words` default 20,000, and meta rows `source_commit`, `sources_kept`, `sources_excluded`. The SQLite `--in` path, schema auto-detection and their tests were deleted.
- Makefile, Dockerfile, `.gitignore`, CI leak guard, NOTICE, `data/ATTRIBUTION.md`, README, `docs/deployment.md` and the frontend attribution footer now name the new source. `data/LICENSE` untouched.
## Decision
- **Capitalization is not a filter.** The planned proper-noun drop was built and audited: precise on a 100-word sample (1 common word), but with wiktionary-only case evidence it removed 215 words including `mặt trời`, `trái đất`, `tổ quốc`, `dường như`. Narrowing to all-syllables-capitalized was measured (147 casualties remain, 410 junk entries return) and rejected. Owner decided to keep every word regardless of case; the mechanism was removed, not flagged off. Record: `audit-proper-noun-drops.md` in the plan dir.
- Attribution states the 2018 scrape was CC BY-SA 3.0 (Wikimedia moved to 4.0 in 2023) and that the derived database is 4.0 via the later-version clause. Both chain links are credited: Wiktionary tiếng Việt contributors, then undertheseanlp as intermediary.
- Rejected alternatives recorded in `plan.md`: hongocduc+wiktionary (GPLv3 relicense), viwiktionary 2026 dump (documented upgrade path if 27k proves thin), kaikki (unpinnable).
## Verification
- `go vet ./... && go test ./...` green, `-race` on touched packages green; `npm run check && npm test` 175/175; Docker `noitu:ci` built with the CI required-files and leak checks passing, `FIXTURE_DICT=1` variant built.
- Bot real-corpus tests pass on the new DB; easy-vs-easy games 12.9 moves vs 17.5 before, the expected depth loss.
- Code review (DONE_WITH_CONCERNS) fully applied: builder's copy of the pinned commit now guarded by `web/tests/dictionary-source.test.js`; fixture builds record "no upstream data" instead of a CC BY-SA string; dead `contains()`, redundant sorts, scanner line off-by-one, stale "48k" comments, `curl -f`, Dockerfile download name fixed.
- **Not verified:** Playwright e2e — Chromium download timed out twice on this machine. The one changed e2e assertion (footer link) was updated by hand. Needs CI or a successful `npx playwright install chromium`.
## Next steps
- Commit all five phases together so attribution and database never disagree (not yet committed).
- Confirm e2e in CI.
- Watch real play for thinness; the viwiktionary dump is the upgrade path, not GPL data.
- Delete the orphaned local `data/dictionary.db` (179 MB, gitignored).
> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions.
@@ -0,0 +1,271 @@
---
title: "Research: building a dictionary DB from the dumps.wikimedia.org viwiktionary dump — measured"
date: 2026-09-08T15:29+07:00
type: research
status: complete
scope: research-only — dump downloaded to scratchpad, parsed, run through the real build-dictionary filter; no repo file touched
follows: research-260908-1434-viwiktionary-as-dictionary-source.md, plans/260908-1525-dictionary-corpus-switch/plan.md
---
# Can we build our own Vietnamese dictionary DB from the viwiktionary dump?
## Executive summary
**Yes, and it is cheap: 63.5 MB download, 34 s of stdlib Python, no wiktextract, no Lua.** The
2026-09-01 `pages-articles` dump was downloaded, checksum-verified against Wikimedia's published
MD5, parsed, and every list below was run through our real `build-dictionary --words` filter.
Numbers are measured, not estimated.
**But viwiktionary alone is not a corpus.** It yields **36,200** accepted words, **31,637** after
dropping pages labelled proper-noun. Both sit below the shipped 48,216 and below the 40,000
`--min-words` floor. The 14:34 report's "~16k from this source" undercounted by 2×, because the
`minhqnd` aggregate only ingested a 17k-row slice; the full dump is 36k. Still not enough.
**Where it earns its place is on top of the pending plan.** Against the plan's undertheseanlp
corpus (re-measured this session at **61,026**, matching the plan's ~61,276), the 2026 dump adds
**3,084 words** the plan would otherwise lose — `học liệu`, `thiên hà lùn`, `cải thảo`, `bùn đỏ`,
`báng súng`, `tăng đoàn` — for a union of **64,110**. That is the modern-vocabulary gap the plan's
own open question names, and the plan's 2018 wiktionary branch (26,845 accepted) is 9,780 words
behind the 2026 dump. **Recommendation: after the corpus switch lands, add the dated dump as a
second `--words` input, pinned by dated URL + checksum.** Licensing stays clean: CC BY-SA 4.0 is
one-way compatible into the GPLv3 the plan already adopts.
**Do not use viwiktionary's proper-noun labels as a drop rule.** They agree with the
capitalization blocklist 86% of the time, but they tag `mặt trời`, `trái đất`, `thứ hai`,
`tháng ba`, `tin lành`, `bàn là` as proper-only and `việt nam` as common. As a *filter* the label
is worse than capitalization; as a *report column* it is fine.
---
## Method
- Web: 2 fetches (dump index, kaikki index), plus direct `curl` of the dump, its checksum file,
4 sample raw wikitext pages, and the plan's pinned undertheseanlp file.
- Local: `extract.py` (streams bz2 → per-page markers, 34 s), `classify.py` (slices the
Vietnamese section, classifies POS labels, emits word lists), `build-dictionary --words` for
every list, sqlite for overlaps. Scripts are in the session scratchpad; each is < 100 lines.
## The dump
| | |
|---|---|
| File | `https://dumps.wikimedia.org/viwiktionary/20260901/viwiktionary-20260901-pages-articles.xml.bz2` |
| Size | 63,513,513 bytes |
| MD5 (Wikimedia-published, verified) | `6c2491e703e7d946f23a405996b4d172` |
| SHA-256 (computed here; Wikimedia publishes md5/sha1 only for this run) | `ed66c932f535b0b362d1141c02273b02ed01c4101e856378bd15e88d277bb8d9` |
| Cadence | monthly, dated dirs `20251220 … 20260901`, plus rolling `latest/` |
| Pages / ns0 pages | 391,543 / 349,461 |
| **ns0 pages with a Vietnamese section** | **43,013** (on-wiki category says 43,037 — parser is complete) |
| Redirect pages among them | 3,237; 834 are case-only (`mặt trời` → `Mặt Trời`) |
Dated directories are immutable, so a pin is honest. `latest/` is not. kaikki.org tracks the same
dump (extracted 2026-09-06, 56,502 Vietnamese *senses*) but publishes rolling files — the earlier
report's pinning concern stands; the dump removes the need for kaikki entirely.
### Two wikitext dialects, both live
| Format | Pages | Language marker | POS marker |
|---|---|---|---|
| legacy | 35,885 | `{{-vie-}}` | `{{-noun-}}`, `{{-verb-}}`, `{{-pr-noun-}}`, `{{-place-}}` … |
| new | 7,128 | `== {{langname\|vi}} ==` | `{{ĐM\|noun}}`, `{{ĐM\|pr-noun}}`, `{{vi-noun}}`, `{{vi-pr-noun}}` … |
A parser must handle both; the wiki is mid-migration, so the mix shifts every month. Legacy
section codes collide with language codes (`{{-adj-}}` vs `{{-eng-}}`), which is why the
extractor keeps an explicit section-code set and treats every other 2–3-letter code as a language
switch. Languages seen switching out of Vietnamese: `tyz`, `mtq`, `eng`, `nut`, `nuo`, `fra` …
## Classification of the 43,013 pages
| Class | Pages | Meaning |
|---|---|---|
| common | 31,950 | ≥1 common POS label (noun/verb/adj/adv/phrase/idiom/…) |
| proper-only | 5,103 | only `pr-noun` / `place` / `vi-pr-noun` labels |
| no-pos | 5,764 | Vietnamese section, no POS marker at all |
| nôm-only | 169 | only Nôm/Hán character sections |
| proper+common | 27 | |
The **no-pos** class is real vocabulary, not junk: 5,538 of them pass our filter and **4,912 are
already shipped** (`kỵ mã`, `khoái hoạt`, `quá tay`, `nặn óc`). Keep them.
## Through `build-dictionary --words`
| List | Accepted | Rejected <2 syll | punct | non-VN |
|---|---|---|---|---|
| all 43,013 titles | **36,200** | 6,361 | 155 | 162 |
| minus proper-only | **31,637** | 6,051 | 121 | 84 |
| proper-only alone | 4,661 | 310 | 34 | 78 |
| undertheseanlp, plan rule (hongocduc+wiktionary, no tudientv, drop caps-only) | **61,026** | | | |
| undertheseanlp `wiktionary` branch alone (2018 snapshot) | 26,845 | | | |
Playability of a viwiktionary-only DB vs shipped:
| metric | shipped | vi-all | vi-no-proper |
|---|---|---|---|
| words | 48,216 | 36,200 | 31,637 |
| syllables | 6,676 | 6,172 | 5,981 |
| syllables that open a word | 5,049 | 4,600 | 4,515 |
| …with ≥2 continuations | 3,682 | 3,246 | 3,168 |
| dead-end syllables | 1,627 | 1,572 | 1,466 |
Every breadth metric regresses. **Option A (replace) is dead on measurement, not opinion.**
## Overlaps — what the dump is actually good for
| | |
|---|---|
| shipped ∩ viwiktionary | **34,570 of 48,216 (72%)** — viwiktionary confirms most of what we ship |
| shipped, not in viwiktionary | 13,646 (9,713 two-syllable) — `thai sản`, `vách núi`, `điện từ trường`, `giả kim`: real words, so absence ≠ junk |
| viwiktionary-only, not shipped | 1,630 |
| shipped words viwiktionary labels proper-only | 4,523 (`ba đồn`, `vị thanh`, `xuân lập` …) |
| **plan corpus (61,026) + vi-no-proper → union** | **64,110 (+3,084)** |
| plan corpus words viwiktionary labels proper-only | 293 — includes `thứ hai`, `tháng năm`, `tin lành`, `tân ước`, `bàn là` |
| 2018 wiktionary branch vs 2026 dump | 26,420 shared; **9,780 new since 2018**; 425 gone |
| shipped words in neither plan corpus nor viwiktionary | 9,267 |
### Proper-noun label vs capitalization blocklist
- viwiktionary proper-only ∩ undertheseanlp caps-only: **4,015 of 4,661 (86%)** agree.
- Caps-only words viwiktionary calls common: 139 (incl. `việt nam`).
- Common words viwiktionary calls proper-only: `mặt trời`, `trái đất`, `sao thủy` (celestial
bodies are `Mặt Trời` on-wiki), weekday/month names, `tin lành`, `bàn là`.
Verdict: capitalization (the plan's rule) is the better *drop* signal. The label is a useful
*second opinion* for the plan's Phase 3 audit sample, nothing more.
## Licensing
Wiktionary text is CC BY-SA 4.0 / GFDL. Creative Commons declares CC BY-SA 4.0 **one-way
compatible with GPLv3**, so folding viwiktionary into the GPLv3 `noitu.db` the plan produces is
clean. `data/ATTRIBUTION.md` gains one entry: edition, dump date, URL, checksum. No `NOTICE`
restructure.
## Recommendation
1. **Do not switch to viwiktionary.** 36,200 < 40,000 floor; every playability metric regresses.
2. **Land the undertheseanlp plan as written.** Its numbers reproduce (61,026 vs ~61,276).
3. **Then add the dated dump as a second `--words` source** — a follow-up phase, not a change to
the current plan. Concretely: `make fetch-dict` also fetches the pinned dump; a ~100-line
extractor (stdlib, both dialects, redirects skipped, `--drop-proper` off by default) emits a
word list; the builder reads two `--words` files or the Makefile concatenates them.
Expected result: **~64,100 words**, +3,084 modern/compound vocabulary, 12-month freshness via a
monthly-dated pin.
4. **Do not drop on viwiktionary's proper-noun label.** Record it in `meta`/build log as a
count; use it to spot-check the capitalization rule in Phase 3.
5. Keep kaikki out. The dump is smaller than kaikki's JSONL, pins honestly, and the parser is
trivial because we only need titles + section labels, not senses.
## Addendum — kaikki.org `dictionary/Vietnamese/kaikki.org-dictionary-Vietnamese.jsonl`
Asked after the main report. This file is **not viwiktionary**: kaikki's `dictionary/` tree is
the *English* Wiktionary, filtered to `lang_code = vi`. Dump 2026-09-02, extracted 2026-09-06,
79,161,423 bytes, SHA-256 `d878bd23fe4d6ac4480736858a85b2bbca0376d6cc9b30395fff0ba1e16cf97b`
at fetch time. Measured the same way as everything above.
| | |
|---|---|
| rows / distinct words | 51,896 / 45,281 |
| POS | noun 19,206 · character 9,029 · verb 8,992 · adj 7,070 · **name 3,814** · adv 1,160 · … |
| rejected `< 2 syllables` | 22,115 — half the file is single syllables and Hán/Nôm characters |
| **accepted, all** | **22,628** |
| **accepted, `name`-only words dropped** | **21,244** |
| metric | shipped | kaikki-en | vi-no-proper | uts-plan |
|---|---|---|---|---|
| words | 48,216 | 21,244 | 31,637 | 61,026 |
| syllables | 6,676 | 5,145 | 5,981 | 7,005 |
| openers ≥2 | 3,682 | 2,431 | 3,168 | 4,164 |
| dead-end | 1,627 | 1,399 | 1,466 | 1,513 |
Smallest of every candidate as a corpus. **As a third supplement it is the best one:**
| overlap | words |
|---|---|
| kaikki-en ∩ shipped | 20,231 (adds 1,013) |
| kaikki-en adds over plan corpus | 4,394 |
| kaikki-en adds over plan corpus ∪ viwiktionary | **3,693** |
| **union plan ∪ viwiktionary ∪ kaikki-en** | **67,803** |
| shipped words in none of the three | 6,187 |
The additions are exactly the modern register the other two lack: `sổ hồng`, `quay xe`, `thi hành
án`, `nhà máy lọc dầu`, `cát tặc`, `giấy ướt`, `máy tính tiền`, `thập lục phân`, `đường tiêu hóa`,
`tân tổng thống`, `sói đồng cỏ`. English Wiktionary's Vietnamese section is more actively curated
than viwiktionary's.
**Its `name` POS is the cleanest proper-noun label seen so far.** `hà nội`, `việt nam`, `trái đất`
are `name`; `thứ hai`, `tin lành` are not; `mặt trời` is both `name` and `noun` and so survives.
Of 1,483 accepted name-only words, 1,359 are shipped today and 247 survive the plan's caps rule
(`thượng đế`, `bắc cực`, `trung thu`, `siêu nhân` — arguable either way). Good audit column; still
not a sole drop rule.
**Two blockers to using the URL as given:**
1. **It is deprecated.** kaikki's index marks the postprocessed per-language file as *"deprecated
and will be removed in the future"* and points to the raw download page, where the only
non-deprecated form of this data is the full-edition `raw-wiktextract-data.jsonl`
(23.1 GB, 2.7 GB compressed) filtered by `lang_code`. There is no per-language file in the
replacement layout — the `downloads/<code>/` entries are per *edition*, not per language.
2. **It cannot be pinned.** Rolling URL, refreshed weekly, no archived snapshots. A
`DICT_SHA256`-style pin breaks on the next refresh; `make fetch-dict` would fail weekly.
Ways to use it anyway, in order of preference:
- **Vendor the derived word list.** Run the reduction once, commit `data/sources/kaikki-en-vi.txt`
(~21k lines, ~400 KB) with the fetch date, URL and file SHA-256 in `ATTRIBUTION.md`. Pins
honestly, costs nothing at build time, refreshed deliberately. Note this is a *derived list*,
not a copy of their file, so the plan's "never commit the source file" rule is not violated in
spirit — decide explicitly.
- **Pin the enwiktionary dump and extract ourselves.** `enwiktionary-20260902-pages-articles`
is >1 GB and wiktextract with Lua expansion takes hours. Correct but disproportionate for
~3.7k words.
- **Re-pin on every refresh.** Not acceptable; turns a data pin into a weekly chore.
License: same CC BY-SA 4.0 / GFDL as all Wiktionary text, so the GPLv3 result stays clean.
Attribution should name enwiktionary + wiktextract/kaikki (the extraction is Tatu Ylonen's work).
**Verdict:** yes, usable, and worth ~3,700 modern words on top of plan + viwiktionary — but only
via a vendored derived list, never as a live fetch of that URL.
## Addendum — undertheseanlp `wiktionary` branch alone (the option chosen for the plan)
Measured after the user chose to use undertheseanlp with its wiktionary data only. Rows with
`wiktionary` in `source`: 32,484 → 32,374 lowercase forms → 4,883 caps-only dropped (case
evidence from wiktionary rows only) → 27,491 → **22,310 accepted**.
| metric | shipped | uts-wik only | vi-dump 2026 | wik ∪ dump | uts-plan (GPL) |
|---|---|---|---|---|---|
| words | 48,216 | **22,310** | 31,637 | 31,817 | 61,026 |
| syllables | 6,676 | 5,484 | 5,981 | 6,012 | 7,005 |
| openers | 5,049 | 4,050 | 4,515 | 4,539 | 5,492 |
| openers ≥2 | 3,682 | 2,686 | 3,168 | 3,180 | 4,164 |
| dead-end | 1,627 | 1,434 | 1,466 | 1,473 | 1,513 |
- shipped ∩ uts-wik 21,230 · shipped lost 26,986 (4,335 proper-noun drops, 22,651 absent
from the branch) · new 1,080.
- The 2018 branch is a near-subset of the 2026 dump: union adds 180 words. Same license.
- **Case-rule casualties:** 271 caps-only drops have a lowercase form in another branch
(247 in hongocduc): `mặt trời`, `trái đất`, `hệ mặt trời`, `phật giáo`, `nguyên đán`,
`tia x`. Wiktionary holds `Mặt Trời` / `Trái Đất` capitalized only.
- Probes: all 11 known-bad absent ✓. Of 10 ordinary probes, `mặt trời` dropped by the case
rule; `con người`, `cánh diều` not in the branch at all; the other 7 present.
- Decision recorded in `plans/260908-1525-dictionary-corpus-switch/plan.md`: wiktionary-only,
CC BY-SA 4.0 unchanged, floor 20,000, viwiktionary 2026 dump as the upgrade path.
## Sources
- https://dumps.wikimedia.org/viwiktionary/ (index; dated dirs 20251220–20260901)
- https://dumps.wikimedia.org/viwiktionary/20260901/ (file list, md5/sha1 sums)
- https://kaikki.org/viwiktionary/ (dump 2026-09-01, extracted 2026-09-06, 56,502 vi senses)
- https://kaikki.org/dictionary/Vietnamese/ and https://kaikki.org/dictionary/rawdata.html (enwiktionary edition; deprecation notice, no snapshots)
- raw wikitext of `Hà Nội`, `học sinh`, `nhà`, `mặt trời` on vi.wiktionary.org
- https://raw.githubusercontent.com/undertheseanlp/dictionary/2c078cfc…/dictionary/words.txt (plan's pin, sha256 verified `4c3e0e61…`)
- Local: `server/cmd/build-dictionary`, `data/noitu.db`
## Unresolved questions
- The 9,267 shipped words in neither the plan corpus nor viwiktionary are still an unquantified
junk/real mix; this dump does not settle them.
- Whether Wikimedia publishes a sha256 file for this run later (only md5/sha1 exist today).
Pin on our computed SHA-256 or on the published MD5 — decide when the phase is written.
- Legacy → new wikitext migration pace: if the wiki finishes it, the legacy branch of the parser
becomes dead code; harmless, but the pin test should assert the extracted count so a dialect
the parser misses shows up as a drop.