mirror of
https://github.com/tiennm99/noitu.git
synced 2026-10-11 03:13:45 +00:00
docs(plans): record the dictionary corpus switch, its audit and the source research
This commit is contained in:
1 parent
39ef45730f
commit
10a47ebc76
10 files changed
+835
-285
No files matched your search
@@ -0,0 +1,184 @@
|
||||
# Audit: proper-noun drops and the wiktionary-only corpus
|
||||
|
||||
Date: 2026-09-08. Phase 3 of `plan.md`. Every number here is reproducible from the
|
||||
commands listed; the shipped `data/noitu.db` was copied aside before any build and is the
|
||||
"old" side of every comparison.
|
||||
|
||||
## Commands
|
||||
|
||||
```sh
|
||||
# from server/, after `make fetch-dict`
|
||||
go run ./cmd/build-dictionary --merged ../data/undertheseanlp-words.jsonl \
|
||||
--out /tmp/wik.db --report-drops /tmp/drops-wik.txt
|
||||
go run ./cmd/build-dictionary --merged ../data/undertheseanlp-words.jsonl \
|
||||
--sources hongocduc,wiktionary --out /tmp/hw.db --report-drops /tmp/drops-hw.txt
|
||||
# sample: python, random.seed(1), random.sample(sorted(drops-wik), 100)
|
||||
# casualties: (drops-wik − drops-hw) ∩ words(hw.db)
|
||||
go test ./internal/bot/ -run RealCorpus -v # once per database at data/noitu.db
|
||||
```
|
||||
|
||||
## Build result
|
||||
|
||||
```
|
||||
rejected 97: contains punctuation
|
||||
rejected 5074: fewer than 2 syllables
|
||||
rejected 37: no Vietnamese letters
|
||||
rejected 4747: only ever capitalized
|
||||
accepted 22419 distinct words (sources: wiktionary)
|
||||
```
|
||||
|
||||
22,419 vs the 22,310 the research report measured. The delta is the rule: the report
|
||||
dropped any form containing an uppercase letter, the builder drops forms that *begin* with
|
||||
one, so `tia X`, `bộ bài tây`, `chữ Hán`, `châu Á` survive. All are ordinary words.
|
||||
|
||||
## 1. Drop audit — 100 sampled, seed 1
|
||||
|
||||
Verdicts: **P** proper noun (correct drop) · **X** not a Vietnamese word (would be rejected
|
||||
by the filter anyway) · **U** unclear · **C** common word (false positive).
|
||||
|
||||
| | | | | |
|
||||
|---|---|---|---|---|
|
||||
| a yun P | an hoà tây P | an minh bắc P | an thuận P | ba vinh P |
|
||||
| ban cơ U | brao X | bàn đạt P | bành trạch P | bàu năng P |
|
||||
| bá xuyên P | bình hiệp P | bình tấn P | bằng la P | chiềng sơ P |
|
||||
| châu hội P | chù lá phù lá P | chư krêy P | cur X | côn đảo P |
|
||||
| cơ kiều P | cẩm trung P | hùng vương P | hồng bàng P | irc X |
|
||||
| khánh gia P | lão quân P | lương đài P | mặc dương P | nhơn hoà lập P |
|
||||
| ninh lai P | ninh thạnh P | noong bua P | nội thôn U | phiếu hữu mai U |
|
||||
| phù lảng P | phương cao kén ngựa U | phần lão U | quán hành P | quắc hương P |
|
||||
| rã bản P | sam mứn P | sơn hạ P | sơn vy P | tam ngọc P |
|
||||
| thanh liên P | thanh văn P | thiệu đô P | thái an P | thượng lâm P |
|
||||
| thượng tiến P | thạch xuân P | thới quản P | tiên cát P | tiên thọ P |
|
||||
| tri lễ P | triệu giang P | triệu việt vương P | trung tú P | trà khê P |
|
||||
| trà kót P | trà linh P | trường khánh P | trường long P | trần đình thâm P |
|
||||
| trọng do P | tuân lộ P | tân hưng P | tân mỹ P | tân nhựt P |
|
||||
| tân trì P | tây hiếu P | tăng sâm P | tĩnh gia P | tịnh an P |
|
||||
| tốt động P | vinh an P | việt–mường X | vân canh P | võ trường toản P |
|
||||
| võ văn tồn P | văn môn P | văn quân P | văn đình dận P | vĩnh hoà hưng bắc P |
|
||||
| vĩnh lương P | vĩnh lộc b P | vĩnh trinh P | vũ thư P | vạn thuỷ P |
|
||||
| xuân mai P | xuân quan P | xá bung P | xá cẩu P | xá khắc P |
|
||||
| xốp cộp P | yên cát P | ô qua U | đồng nai P | đổ rượu ra sông thết quân lính C |
|
||||
|
||||
**P 89 · X 4 · U 6 · C 1.** Gate is ≤2 common words in 100: **passes.** The bulk is
|
||||
commune names, ethnonyms and historical persons — the content the rule exists to remove.
|
||||
|
||||
## 2. Case casualties — 215, listed in full
|
||||
|
||||
Words the wiktionary rows hold only capitalized, which another branch has lowercase. They
|
||||
are dropped under the plan's wiktionary-only evidence rule. Full list:
|
||||
`case-casualties.txt` in the session scratchpad; reproduced here because it is the
|
||||
decision input.
|
||||
|
||||
an bình · an dân · an hảo · an khang · an lạc · ba tiêu · biển hồ · bu lu · **bàn là** ·
|
||||
**báo đáp** · bát tiên · bình chuẩn · bình chân · bình khang · bình nghị · bình sa · bình
|
||||
thanh · bình thuỷ · bình trị · bình tâm · bình văn · bình điền · bích đào · bóng chim tăm
|
||||
cá · bông trang · bạch cung · bản nguyên · **bảo toàn** · bắc phong · bẻ quế · bố chính ·
|
||||
**bồ đề** · bồng sơn · cao biền dậy non · cao kỳ · cao nhân · cao sơn · cao tổ · cao xanh ·
|
||||
cao đường · chà và · chày sương · chánh hội · **chân mây** · châu lệ · châu thành · chí
|
||||
thiện · chí thành · chính tâm · **chúa nhật** · chúa trời · chị hằng · con tạo · cát lũy ·
|
||||
cát tân · cát đằng · **công bình** · công dã tràng · **công giáo** · **công nguyên** ·
|
||||
**cơ đốc giáo** · cường thịnh · cầm đuốc chơi đêm · cầu lam · cầu lộc · cầu ô · cẩm châu ·
|
||||
cẩm đường · cửu kinh · **cựu ước** · diêm vương · diêm vương tinh · dã hạc · dương quan ·
|
||||
dương đài · **dường như** · **giao tử** · giấc mai · giọt châu · giọt tương · hoa cái ·
|
||||
**hoa kiều** · **hoà đồng** · hoá công · hoả tinh · hán học · hán tộc · hán tự · hán văn ·
|
||||
hạ thần · hải phòng · hải triều · hải vương tinh · hằng nga · **hết sảy** · **hệ mặt
|
||||
trời** · hội gió mây · hội long vân · **hồi giáo** · **khổng giáo** · kim tinh · **kinh
|
||||
thánh** · la ve · lâm viên · **lưỡi hái** · lạc hầu · mạnh thường quân · **mặt trời** · nam
|
||||
lâu · **nguyên đán** · **nho học** · nhân kiệt · năm cha ba mẹ · nước dương · nếm mật nằm
|
||||
gai · nợ như chúa chổm · phong thu · **phù phiếm** · **phật giáo** · phật học · phật pháp ·
|
||||
phật tiền · phật tổ · phật tự · **phật đản** · quang phục · quyết tiến · quân thiều · **quả
|
||||
đất** · **sao hôm** · sao hỏa · sao kim · **sao mai** · sao mộc · song mai · sơn cương ·
|
||||
sơn lâm · sơn mai · sơn nguyên · sơn trung · sư tử hà đông · **sừng trâu** · tam dân · tam
|
||||
phủ · tam sơn · thanh nghị · thanh phong · thanh quang · thanh vận · thiên chúa · **thiên
|
||||
chúa giáo** · thiên vương tinh · thuỷ liễu · thân giáp · thạch bàn · thạch thán · **thần
|
||||
chết** · **tin lành** · tiên sư · **trái đất** · trướng huỳnh · trường giang · trường xuân ·
|
||||
tu lý · tuần duyên · tân kỳ · **tân ước** · tây dương · tây thiên · tô hạp · tùng lâm · tế
|
||||
tân · **tết nguyên đán** · **tổ quốc** · tử phòng · u minh · vinh thăng · **việt ngữ** ·
|
||||
vách quế · vân hà · vân hán · vân trình · võ miếu · **văn giáo** · **văn miếu** · **văn
|
||||
nhân** · văn quan · **văn võ** · văn đức · **vũ công** · **vũ hội** · vũng tàu · vương bá ·
|
||||
vạn an · vạn kiếp · vạn phúc · vầng ô · **vớ bở** · **xe tơ** · xuân hoá · xuân phong ·
|
||||
xuân tình · xích thố · xương thịnh · yên bình · yên chi · yên hoa · yên hà · đinh điền ·
|
||||
**đường luật** · **đường thi** · đại danh · đỉnh giáp non thần · **địa cầu** · đỗ vũ
|
||||
|
||||
Bold: everyday or dictionary-headword vocabulary by this auditor's reading, ~55 of 215.
|
||||
The rest are Sino-Vietnamese literary terms, religious/astronomical names that Wiktionary
|
||||
capitalizes by convention, and genuine place names (`hải phòng`, `vũng tàu`, `bồng sơn`).
|
||||
|
||||
**Narrowing does not fix this.** The plan's pre-decided narrowing — drop only when every
|
||||
syllable is capitalized — was measured: drops fall 4,747 → 4,337, casualties fall only
|
||||
215 → 147 (`mặt trời`, `trái đất`, `tổ quốc`-style two-cap forms stay dropped), and 410
|
||||
proper nouns come back in, including junk like `bản mẫu:-vie-n-` and `con kde`. Rejected.
|
||||
|
||||
## 3. Corpus diff — old shipped vs new
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| overlap | 21,320 |
|
||||
| gained | 1,099 (`gân bò`, `rèn đúc`, `ích kỷ`, `đáng lý`, `thở hồng hộc`) |
|
||||
| lost | **26,896** |
|
||||
| — dropped as proper noun | 4,245 |
|
||||
| — absent from the wiktionary branch | 22,651 |
|
||||
| — other | 0 |
|
||||
|
||||
## 4. Playability
|
||||
|
||||
| metric | old | new |
|
||||
|---|---|---|
|
||||
| words | 48,216 | 22,419 |
|
||||
| syllables | 6,676 | 5,485 |
|
||||
| syllables that open a word | 5,049 | 4,052 |
|
||||
| …with ≥2 continuations | 3,682 | 2,691 |
|
||||
| dead-end syllables | 1,627 | 1,433 |
|
||||
|
||||
Bot-versus-bot, `internal/bot` real-corpus tests, 60 games each:
|
||||
|
||||
| | old | new |
|
||||
|---|---|---|
|
||||
| hard-vs-easy win rate (moves) | 98% (3.3) | 98% (3.5) |
|
||||
| medium-vs-easy | 87% | 88% |
|
||||
| hard-vs-medium (moves) | 72% (3.5) | 55% (2.9) |
|
||||
| easy-vs-easy game length | 17.5 moves | **11.3 moves** |
|
||||
| hard decision p95 | 5.2 ms | 1.1 ms |
|
||||
|
||||
Both tests pass on both databases. Easy-vs-easy games are a third shorter: the thinner
|
||||
graph runs out of continuations sooner, which is the depth loss the plan predicted.
|
||||
Hard-vs-medium falling to 55% means the graph gives the stronger bot fewer ways to trap.
|
||||
|
||||
## 5. Decision
|
||||
|
||||
Drop rule on the sample: **pass** (1 common word in 100).
|
||||
|
||||
Case casualties: **open — needs the owner's call**, because the options trade the
|
||||
"wiktionary data only" decision against ~55 everyday words:
|
||||
|
||||
- **Accept** the 215 losses as measured.
|
||||
- **Exception list**: a checked-in list of words to keep despite capitalization, curated
|
||||
from the 215 above. Cause-aligned, small, but a hand-maintained artifact.
|
||||
- **Read hongocduc for case evidence only**: recovers all 215 mechanically, but uses GPL
|
||||
data as an input even though none of its words ship. Must be stated in `ATTRIBUTION.md`.
|
||||
|
||||
**Recorded decision (owner, 2026-09-08): abandon the capitalization rule.** Every
|
||||
wiktionary-tagged word is kept regardless of case; proper nouns stay in the corpus as they
|
||||
are in the shipped database today. The rule, its reject reason and `--report-drops` were
|
||||
removed from the builder.
|
||||
|
||||
## 6. Final corpus, without the rule
|
||||
|
||||
```
|
||||
accepted 26845 distinct words (sources: wiktionary)
|
||||
```
|
||||
|
||||
| metric | old | new |
|
||||
|---|---|---|
|
||||
| words | 48,216 | 26,845 |
|
||||
| syllables | 6,676 | 5,709 |
|
||||
| syllables that open a word | 5,049 | 4,158 |
|
||||
| …with ≥2 continuations | 3,682 | 2,787 |
|
||||
| dead-end syllables | 1,627 | 1,551 |
|
||||
| overlap / gained / lost | | 25,565 / 1,280 / 22,651 |
|
||||
|
||||
Every lost word is absent from the 2018 wiktionary branch; none is lost to a rule.
|
||||
|
||||
Bot-versus-bot, 60 games each: hard-vs-easy 95% (2.7 moves), medium-vs-easy 87%,
|
||||
hard-vs-medium 63% (3.2 moves), easy-vs-easy **12.9 moves** (old: 17.5). Both real-corpus
|
||||
tests pass.
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 1
|
||||
title: "Phase 1: Read the merged list"
|
||||
status: todo
|
||||
status: done
|
||||
priority: P1
|
||||
effort: "3h"
|
||||
dependencies: []
|
||||
@@ -12,109 +12,107 @@ dependencies: []
|
||||
## Overview
|
||||
|
||||
Teach `build-dictionary` a third input: undertheseanlp's merged JSONL. It selects words by
|
||||
source membership, drops entries that appear only ever capitalized, and hands the survivors
|
||||
to the normalization and filtering path that already exists.
|
||||
source membership and hands them to the normalization and filtering path that already
|
||||
exists.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: read `{"text": "...", "source": ["hongocduc", ...]}` lines; keep a word when
|
||||
its sources intersect the allowed set; drop a word when every raw form of it, across the
|
||||
allowed sources, begins with an uppercase letter.
|
||||
- Functional: `tudientv` rows are skipped before any decision is made about a word —
|
||||
neither its words nor its capitalization inform the output.
|
||||
- Functional: the proper-noun drop is counted and printed beside the existing reject
|
||||
reasons, so the build log says what the source contained.
|
||||
- Non-functional: streaming line-by-line, one pass to gather, one pass to decide. The file
|
||||
is 4.8 MB; nothing here needs to be clever.
|
||||
its sources intersect the allowed set — **`wiktionary` only** by default.
|
||||
- Functional: `hongocduc` and `tudientv` rows are skipped before any decision is made about
|
||||
a word. They inform nothing.
|
||||
- Functional: capitalization is not a filter. `Hà Nội` and `hà nội` are the same word and
|
||||
land as the lowercase form, exactly as every other input mode already behaves.
|
||||
- Functional: the `--min-words` default moves from 40,000 to 20,000. The new corpus is
|
||||
~26,800 words; the old floor would reject it.
|
||||
- Non-functional: streaming line-by-line. The file is 4.8 MB; nothing here needs to be
|
||||
clever.
|
||||
- Non-functional: everything downstream — `accept()`, syllable indexing, alias generation,
|
||||
`verify()`, `meta` — is reused unchanged.
|
||||
|
||||
## Architecture
|
||||
|
||||
Two maps keyed by the lowercased, whitespace-collapsed word:
|
||||
|
||||
```
|
||||
forms[key] -> set of raw spellings seen in allowed sources ("Mặt Trời", "mặt trời")
|
||||
sources[key] -> set of allowed sources containing it {"hongocduc", "wiktionary"}
|
||||
--merged <file> --sources wiktionary
|
||||
│
|
||||
▼ one row at a time
|
||||
sources ∩ allowed ≠ ∅ ? ── no ──▶ skipped, never counted
|
||||
│ yes
|
||||
▼
|
||||
accept() (NFC, lowercase, ≥2 syllables, alphabet, phonotactics)
|
||||
│
|
||||
▼
|
||||
finish() (floor, aliases, atomic write, verify) ← shared with --words and --in
|
||||
```
|
||||
|
||||
A key survives when `sources[key]` is non-empty and at least one entry in `forms[key]`
|
||||
does not start with an uppercase letter. `Mặt Trời` survives on the strength of its
|
||||
lowercase twin; `Hà Nội`, `Trương Công Định` and `A Mú Sung` have no lowercase form and go.
|
||||
The `--sources` flag is kept general — a comma list validated against the three known
|
||||
names — even though only `wiktionary` is passed, so a comparison build against another
|
||||
branch is a flag away and never a code patch. Unknown names are an error, not a no-op.
|
||||
|
||||
Surviving keys are then fed to the same `accept()` the SQLite path feeds, so the ≥2
|
||||
syllable rule, the digit/punctuation rejections and the Vietnamese phonotactic check all
|
||||
apply as they do today.
|
||||
|
||||
Case evidence deliberately comes from allowed sources only. The alternative — reading
|
||||
`tudientv` rows for their capitalization while refusing their words — would be more
|
||||
accurate and would undercut the claim that we do not use that data. Phase 3 measures what
|
||||
the stricter choice costs.
|
||||
`finish()` is new: the three input modes used to each carry their own copy of the floor
|
||||
check, alias generation, write and verify. One copy is what keeps a fixture from drifting
|
||||
into a different shape from the database production loads.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `server/cmd/build-dictionary/merged_list.go`
|
||||
- Create: `server/cmd/build-dictionary/merged_list_test.go`
|
||||
- Modify: `server/cmd/build-dictionary/main.go` — `--merged` and `--sources` flags, dispatch
|
||||
in `run()`, `sourceURL`/`sourceLicense` constants, `meta` rows, package doc comment
|
||||
- Modify: `server/cmd/build-dictionary/filter.go` — add the `rejectProperNoun` reason so
|
||||
the drop is reported through the existing counter, not a bespoke log line
|
||||
- Created: `server/cmd/build-dictionary/merged_list.go`
|
||||
- Created: `server/cmd/build-dictionary/merged_list_test.go`
|
||||
- Modified: `server/cmd/build-dictionary/main.go` — `--merged` and `--sources` flags,
|
||||
mutually-exclusive input check, `finish()`, `sourceSpec.url` / `sourceSpec.extra`, meta
|
||||
rows, package doc comment, `--min-words` default
|
||||
- Unchanged: `server/cmd/build-dictionary/filter.go`
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Add `rejectProperNoun rejectReason = "only ever capitalized"` to `filter.go`. Leave
|
||||
`accept()` alone — the case decision happens before it, on evidence `accept()` cannot
|
||||
see, since it lowercases.
|
||||
2. Write `merged_list.go`: decode with `encoding/json` per line (`bufio.Scanner` with a
|
||||
raised buffer — the longest line is short, but a scanner that silently truncates is a
|
||||
bad way to lose words), build the two maps, apply the rule, return words plus a count
|
||||
per reject reason.
|
||||
3. Add `--merged <file>` and `--sources hongocduc,wiktionary` to `main.go`. Unknown source
|
||||
names are an error, not a silent no-op: a typo in `--sources` must not quietly ship an
|
||||
empty or wrong corpus.
|
||||
4. Dispatch in `run()`: `--merged` takes precedence in the same shape `--words` already
|
||||
does. Refuse `--merged` together with `--in`/`--words` rather than picking one.
|
||||
5. Point `sourceURL`, `sourceLicense` and the package doc comment at the new source. Write
|
||||
`source_commit`, `sources_kept`, `sources_excluded` and `proper_nouns_dropped` into
|
||||
`meta` — provenance in the artifact, not only in a markdown file.
|
||||
6. Tests in `merged_list_test.go`, table-driven, over hand-written JSONL fixtures:
|
||||
- `Mặt Trời` + `mặt trời` → kept
|
||||
- `Hà Nội` alone → dropped as a proper noun
|
||||
- a `tudientv`-only word → absent from the output
|
||||
- a word whose only lowercase form is in `tudientv` → dropped (the documented cost)
|
||||
- `học sinh` in all three → kept once, not thrice
|
||||
1. Write `merged_list.go`: decode with `encoding/json` per line (`bufio.Scanner` with a
|
||||
raised buffer — a scanner that silently truncates is a bad way to lose words), skip rows
|
||||
with no allowed source, feed the rest to `accept()`.
|
||||
2. Add `--merged <file>` and `--sources wiktionary` (the default) to `main.go`. Change the
|
||||
`--min-words` default to 20,000 in the same edit; the fixture path passes its own floor
|
||||
explicitly and is unaffected.
|
||||
3. Dispatch in `run()`: exactly one of `--merged`, `--words`, `--in` may be given. Refuse
|
||||
two rather than picking one.
|
||||
4. Write `source_url`, `source_commit`, `sources_kept` and `sources_excluded` into `meta` —
|
||||
provenance in the artifact, not only in a markdown file.
|
||||
5. Tests in `merged_list_test.go`, table-driven, over hand-written JSONL fixtures:
|
||||
- `Hà Nội` tagged wiktionary → kept as `hà nội`
|
||||
- a `hongocduc`-only word and a `tudientv`-only word → absent from the output
|
||||
- `học sinh` in all three, plus `Học sinh` → kept once
|
||||
- a malformed line → a clear error naming the line number, not a skipped word
|
||||
7. Run against the real file and record the counts. Compare to the report's 61,276, which
|
||||
used all-source case evidence; explain the delta rather than adjusting the number.
|
||||
- `--sources hongocduc,wiktionary` → strictly more words; `tudientv` still absent
|
||||
- meta rows as above
|
||||
6. Run against the real file and record the counts.
|
||||
|
||||
## Result
|
||||
|
||||
```
|
||||
rejected 2: contains a digit
|
||||
rejected 147: contains punctuation
|
||||
rejected 5335: fewer than 2 syllables
|
||||
rejected 87: no Vietnamese letters
|
||||
accepted 26845 distinct words (sources: wiktionary)
|
||||
generated 1507 spelling aliases
|
||||
```
|
||||
|
||||
The first cut of this phase also implemented a capitalization-based proper-noun drop
|
||||
(`only ever capitalized` → rejected, `--report-drops` for audit). It reached the Phase 3
|
||||
gate, was audited, and was removed by the owner's decision — see `plan.md` and
|
||||
`audit-proper-noun-drops.md`. Nothing of it remains in the code.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] `go test ./cmd/build-dictionary/` passes, new tests included.
|
||||
- [ ] Building from the real merged file accepts >55,000 words and logs a proper-noun drop
|
||||
count in the low thousands.
|
||||
- [ ] Every one of the 11 known-bad probes from the report is absent from the output
|
||||
(`trương công định`, `trần danh án`, `tân an thạnh`, `cẩm xá`, `vĩnh điện`,
|
||||
`cam lâm`, `hà nội`, `sài gòn`, `a mú sung`, `a lưới`, `kháng đón`).
|
||||
- [ ] All 10 ordinary-word probes are present (`học sinh`, `mặt trời`, `con người`,
|
||||
`bánh mì`, `cánh diều`, `xe đạp`, `giáo viên`, `hoa hồng`, `tình yêu`, `nước mắm`).
|
||||
- [ ] `--sources hongocduc` alone and `--sources hongocduc,wiktionary` both build, with the
|
||||
first strictly smaller.
|
||||
- [ ] No word in the output is reachable only through `tudientv`.
|
||||
- [x] `go test ./cmd/build-dictionary/` passes, new tests included.
|
||||
- [x] Building from the real merged file accepts >20,000 words (26,845).
|
||||
- [x] Ordinary-word probes present: `học sinh`, `bánh mì`, `xe đạp`, `giáo viên`, `hoa
|
||||
hồng`, `tình yêu`, `nước mắm`, `mặt trời`, `trái đất`. `con người` and `cánh diều` are
|
||||
not in the 2018 branch at all — expected, recorded.
|
||||
- [x] `--sources wiktionary` (default) and `--sources hongocduc,wiktionary` both build, with
|
||||
the first strictly smaller (26,845 vs 61,271 with no case rule).
|
||||
- [x] No word in the output is reachable only through `hongocduc` or `tudientv`.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
**The capitalization rule eats real words.** The probe set is 21 words; the rule fires on
|
||||
thousands. Signal: Phase 3's hand audit finds common words among the drops. Response:
|
||||
narrow the rule to require every syllable capitalized (`Hà Nội`, not `Kháng đón`), re-audit,
|
||||
and if that still misfires, stop dropping and keep the count as a log line only — a corpus
|
||||
with proper nouns in it is what we ship today and is survivable.
|
||||
|
||||
**Allowed-sources-only case evidence drops more than expected.** Signal: the drop count
|
||||
comes in far above the report's 4,145. Response: measure how many drops have a lowercase
|
||||
form only in `tudientv`; if that is a large share, reconsider reading `tudientv` for case
|
||||
evidence alone and say so plainly in `ATTRIBUTION.md` rather than hiding it.
|
||||
|
||||
**A JSON schema surprise.** The file is 8 years old and unversioned; a row could carry an
|
||||
unexpected shape. Signal: decode errors on the real file. Response: the error names the
|
||||
line — inspect it, and only then decide between a tolerant skip with a counted reason and
|
||||
a hard failure. Silent skipping is not an option.
|
||||
a hard failure. Silent skipping is not an option. (Observed: none; all 79,226 rows decode.)
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 2
|
||||
title: "Phase 2: Pin the source"
|
||||
status: todo
|
||||
status: done
|
||||
priority: P1
|
||||
effort: "2h"
|
||||
dependencies: [1]
|
||||
@@ -54,6 +54,8 @@ because that is already the right shape; only the size in the help text changes.
|
||||
## Implementation Steps
|
||||
|
||||
1. Update the three Makefile variables and the `dict` target to pass `--merged $(DICT_SRC)`.
|
||||
`--sources` is left at its `wiktionary` default so the Makefile and Dockerfile carry
|
||||
one fewer thing to keep in agreement.
|
||||
2. Reword `help`, `fetch-dict` and the `$(DICT_SRC)` guard: the download is ~4.8 MB now,
|
||||
and saying "179 MB" would be the kind of stale comment that outlives three refactors.
|
||||
3. Mirror all of it in the Dockerfile stage. Keep `FIXTURE_DICT=1` on `--words` — the
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 3
|
||||
title: "Phase 3: Audit the corpus"
|
||||
status: todo
|
||||
status: done
|
||||
priority: P1
|
||||
effort: "2h"
|
||||
dependencies: [1, 2]
|
||||
@@ -9,6 +9,12 @@ dependencies: [1, 2]
|
||||
|
||||
# Phase 3: Audit the corpus
|
||||
|
||||
> **Outcome.** Executed 2026-09-08; see [`audit-proper-noun-drops.md`](./audit-proper-noun-drops.md).
|
||||
> The drop rule passed its sample gate (1 common word in 100) but cost 215 words under
|
||||
> wiktionary-only case evidence, including everyday vocabulary. Narrowing was measured and
|
||||
> rejected. **The owner abandoned the rule**: every word is kept regardless of case. Final
|
||||
> corpus 26,845 words; playability and bot results recorded in the audit note.
|
||||
|
||||
## Overview
|
||||
|
||||
The gate. A rule that removes thousands of words on a capitalization heuristic has to be
|
||||
@@ -34,12 +40,17 @@ Three measurements, each against the freshly built database and the current one:
|
||||
passes at ≤2 common words in 100; that tolerance is the difference between a filter and
|
||||
a corpus edit.
|
||||
2. **Corpus diff.** Overlap, gained, lost. Split the lost into: dropped as proper nouns,
|
||||
`tudientv`-only, absent from undertheseanlp. The report measured 4,145 / 1,300 / 6,473
|
||||
under all-source case evidence; this re-measures under the shipped rule.
|
||||
absent from the wiktionary branch. The report measured 4,335 / 22,651 of 26,986 lost;
|
||||
this re-measures from the shipped build. Separately count the **271 case casualties**
|
||||
— drops that have a lowercase form in `hongocduc` or `tudientv` — and list them in
|
||||
full; that list is what the pass/narrow/exception decision is made on.
|
||||
3. **Playability.** Words, syllables, syllables that can open a word, how many have ≥2
|
||||
continuations, and dead-end syllables — the last of these being what decides whether
|
||||
the game hands somebody an unanswerable position. Current: 48,216 / 6,676 / 5,049 /
|
||||
3,682 / 1,627.
|
||||
3,682 / 1,627. Expected: 22,310 / 5,484 / 4,050 / 2,686 / 1,434. The corpus is smaller
|
||||
by design, so the question is not "is it bigger" but "does the game still run": add
|
||||
bot-versus-bot game lengths from `server/internal/bot`'s real-corpus tests on both
|
||||
databases.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
@@ -58,20 +69,23 @@ Three measurements, each against the freshly built database and the current one:
|
||||
hand. Vietnamese place and person names are the expected content; anything that reads
|
||||
as an ordinary word is a false positive and gets called one.
|
||||
4. Run the corpus diff and the playability comparison; put both tables in the audit note.
|
||||
5. Measure the documented cost of allowed-sources-only case evidence: how many drops have
|
||||
a lowercase form only in `tudientv`.
|
||||
6. Decide: pass, narrow, or abandon the rule. Record the decision and its reason in the
|
||||
audit note — and if the rule is narrowed, Phase 1's tests change with it.
|
||||
5. Build `--sources hongocduc,wiktionary` to a second scratch path purely to enumerate the
|
||||
case casualties (words dropped under `wiktionary` alone but kept when hongocduc's
|
||||
lowercase forms are visible). Nothing from that build ships.
|
||||
6. Decide: pass, narrow, exception list, or abandon the rule. Record the decision and its
|
||||
reason in the audit note — and if the rule changes, Phase 1's tests change with it.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] The audit note exists, with 100 judged entries and the commands that produced them.
|
||||
- [ ] ≤2 of 100 sampled drops are ordinary words.
|
||||
- [ ] The new corpus has >55,000 words.
|
||||
- [ ] Syllables ≥ 6,676 and dead-end syllables ≤ 1,627 — the graph is not worse than
|
||||
what players walk today.
|
||||
- [ ] The three loss buckets are quantified, not estimated.
|
||||
- [ ] A pass/narrow/abandon decision is written down with its reason.
|
||||
- [ ] ≤2 of 100 sampled drops are ordinary words, the 271 known case casualties aside —
|
||||
those are listed in full and judged as a group.
|
||||
- [ ] The new corpus has >20,000 words.
|
||||
- [ ] The five graph numbers and bot-game lengths are recorded for both databases; bot
|
||||
games on the new graph complete without the engine running out of moves earlier
|
||||
than on today's.
|
||||
- [ ] The loss buckets are quantified, not estimated.
|
||||
- [ ] A pass/narrow/exception-list/abandon decision is written down with its reason.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
@@ -81,10 +95,14 @@ Response: the drop list is a file — a disputed word can be checked against it
|
||||
and a per-word exception list is a small change on top of this design.
|
||||
|
||||
**The audit fails and the phase becomes a redesign.** Signal: >2 common words in 100.
|
||||
Response: apply Phase 1's pre-decided narrowing (every syllable capitalized), re-audit
|
||||
once. If it fails again, ship without the drop — the corpus is still bigger and denser
|
||||
than today's, and the proper-noun problem stays exactly as bad as it currently is rather
|
||||
than getting worse.
|
||||
Response: apply Phase 1's pre-decided narrowing (every syllable capitalized) or the
|
||||
exception list, re-audit once. If it fails again, ship without the drop — the proper-noun
|
||||
problem then stays exactly as bad as it is today rather than getting worse.
|
||||
|
||||
**The corpus is too thin in play.** 22,310 words is less than half of today's. Signal: bot
|
||||
games noticeably shorter, or the dead-end rate per move up rather than down. Response: this
|
||||
is the trigger for the documented upgrade path — the 2026-09-01 viwiktionary dump (31,637
|
||||
words, same license) as a second `--words` source — not for re-admitting GPL data.
|
||||
|
||||
**Playability regresses in a way these five numbers do not capture.** Signal: aggregate
|
||||
counts look fine but bot games end oddly short or long. Response: `server/internal/bot`'s
|
||||
|
||||
@@ -1,110 +0,0 @@
|
||||
---
|
||||
phase: 4
|
||||
title: "Phase 4: Relicense the data"
|
||||
status: todo
|
||||
priority: P1
|
||||
effort: "2h"
|
||||
dependencies: [3]
|
||||
---
|
||||
|
||||
# Phase 4: Relicense the data
|
||||
|
||||
## Overview
|
||||
|
||||
`data/noitu.db` becomes GPLv3, because `hongocduc`'s data is GPL and GPL is copyleft. This
|
||||
phase rewrites the data half of the licensing documents and regenerates the shipped
|
||||
database — in one commit, because a GPL-derived database sitting in a tree that still
|
||||
says CC BY-SA 4.0 is the only ordering mistake here with a legal consequence.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: `data/LICENSE` carries the GPLv3 text.
|
||||
- Functional: `NOTICE` section 2 describes GPLv3 for the data, names both source branches
|
||||
with their own licenses, and keeps section 1 (Apache-2.0 for code) intact.
|
||||
- Functional: `data/ATTRIBUTION.md` records the new source, the per-branch licenses, the
|
||||
full list of modifications, and why `tudientv` is excluded.
|
||||
- Functional: the regenerated `data/noitu.db` carries matching `meta` provenance.
|
||||
- Non-functional: no overclaiming. The `hongocduc` GPL statement comes from a mirror's
|
||||
README and the document says so.
|
||||
|
||||
## Architecture
|
||||
|
||||
The dual-license structure already in `NOTICE` is correct and stays; only its second
|
||||
section changes:
|
||||
|
||||
```
|
||||
1. SOURCE CODE — Apache-2.0 (unchanged: server/, tools/, web/, proto/)
|
||||
2. DICTIONARY DATA — GPLv3 (was CC BY-SA 4.0)
|
||||
data/noitu.db, data/LICENSE, data/ATTRIBUTION.md
|
||||
```
|
||||
|
||||
Why GPLv3 rather than CC BY-SA 4.0 is worth stating in `ATTRIBUTION.md` rather than left
|
||||
implicit: `hongocduc` is Hồ Ngọc Đức's FVDP wordlist, distributed under GNU GPL. The
|
||||
previous source aggregated that same data and redistributed it as CC BY-SA 4.0, which GPL
|
||||
does not permit. Ours is the stricter, and correct, direction.
|
||||
|
||||
The modifications list gains two entries beyond the current five — source selection and
|
||||
the proper-noun drop — both of which the license requires us to declare as changes.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Modify: `data/LICENSE` — CC BY-SA 4.0 text replaced with GPLv3
|
||||
- Modify: `NOTICE` — section 2 rewritten
|
||||
- Modify: `data/ATTRIBUTION.md` — source table, per-branch licenses, modifications 6 and 7,
|
||||
the `tudientv` exclusion note
|
||||
- Modify: `README.md` — any dictionary-source or data-license claim
|
||||
- Verify: `.github/workflows/ci.yml` — the required-files check (`data/LICENSE`,
|
||||
`data/ATTRIBUTION.md`, `NOTICE`, `data/noitu.db`) still holds
|
||||
- Regenerate: `data/noitu.db`
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Replace `data/LICENSE` with the GPLv3 text, verbatim and complete.
|
||||
2. Rewrite `NOTICE` section 2: GPLv3, the affected artifacts, both upstream branches with
|
||||
their licenses, and the copyleft obligation carrying into container images — the same
|
||||
point the current text makes about share-alike, which is no less true of GPL.
|
||||
3. Rewrite `data/ATTRIBUTION.md`:
|
||||
- source is `undertheseanlp/dictionary` at commit `2c078cf`, file `dictionary/words.txt`
|
||||
- `hongocduc` (GNU GPL, per the branch README — a mirror of Hồ Ngọc Đức's 2003
|
||||
wordlist, whose canonical site no longer resolves) and `wiktionary` (CC BY-SA, a
|
||||
2018-12-10 scrape of vi.wiktionary.org)
|
||||
- `tudientv` **excluded**: its README declares copyright *"Chưa rõ"* and names
|
||||
Soha/Vietlex (Hoàng Phê) as its primary source. Say it plainly so nobody reopens it.
|
||||
- modifications 1–5 carried over, plus **6. source selection** and **7. proper-noun
|
||||
removal** with the count from Phase 3
|
||||
- the resulting license and why it is GPLv3 and not CC BY-SA 4.0
|
||||
4. Grep the repository for `minhqnd`, `CC BY-SA`, `179 MB` and fix every survivor outside
|
||||
`plans/` — the reports are dated records and stay as written.
|
||||
5. Regenerate `data/noitu.db` and confirm its `meta` rows agree with `ATTRIBUTION.md`. A
|
||||
database whose provenance contradicts the attribution file is worse than neither.
|
||||
6. Run the CI required-files check locally, or read it and confirm by hand.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] `data/LICENSE` is the complete GPLv3 text.
|
||||
- [ ] `NOTICE` section 2 says GPLv3 and names both branches with their licenses; section 1
|
||||
is unchanged.
|
||||
- [ ] `data/ATTRIBUTION.md` lists seven modifications, states the `tudientv` exclusion and
|
||||
its reason, and does not claim more about the `hongocduc` license than the mirror
|
||||
supports.
|
||||
- [ ] `meta` in the built database matches the attribution file — source URL, commit,
|
||||
license, sources kept and excluded, proper-noun drop count.
|
||||
- [ ] No file outside `plans/` still describes the data as CC BY-SA 4.0 or names minhqnd
|
||||
as the source.
|
||||
- [ ] CI's required-files check passes.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
**Relicensing is hard to walk back.** Once a GPLv3 `noitu.db` is published, that release
|
||||
is GPLv3 permanently. Signal: none — this is simply true. Response: it is why the user
|
||||
signed off before this plan was written; the mitigation is that it was a decision rather
|
||||
than a side effect.
|
||||
|
||||
**A stale CC BY-SA claim survives somewhere.** A README line or a Docker label saying the
|
||||
wrong license is a licensing misstatement, not a typo. Signal: the grep in step 4 finds
|
||||
something after the phase is called done. Response: grep is a step in the phase for that
|
||||
reason; run it last, not first.
|
||||
|
||||
**Section 1 gets damaged while section 2 is rewritten.** The code's Apache-2.0 grant is
|
||||
not in scope and must come out byte-identical. Signal: a diff touching section 1.
|
||||
Response: review the `NOTICE` diff before committing.
|
||||
@@ -0,0 +1,116 @@
|
||||
---
|
||||
phase: 4
|
||||
title: "Phase 4: Update the attribution"
|
||||
status: done
|
||||
priority: P1
|
||||
effort: "1h"
|
||||
dependencies: [3]
|
||||
---
|
||||
|
||||
# Phase 4: Update the attribution
|
||||
|
||||
> **Outcome.** Done 2026-09-08. `data/LICENSE` untouched; `NOTICE` section 2 names the new
|
||||
> upstream, `data/ATTRIBUTION.md` rewritten (eight modifications, both attribution links,
|
||||
> both exclusions). Also updated, beyond the plan's list: the frontend attribution footer
|
||||
> (`AttributionFooter.svelte`, `i18n/vi.js`) and its e2e assertion now credit Wiktionary
|
||||
> tiếng Việt instead of minhqnd — the user-visible half of the CC BY-SA obligation.
|
||||
|
||||
## Overview
|
||||
|
||||
The data license does not change — the wiktionary branch is CC BY-SA 4.0, the same license
|
||||
`noitu.db` already carries — but everything that names the source does. This phase rewrites
|
||||
the attribution and notice text and regenerates the shipped database in one commit, so the
|
||||
attribution never describes a database other than the one in the tree.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: `data/LICENSE` is untouched — still the CC BY-SA 4.0 text.
|
||||
- Functional: `NOTICE` section 2 names the new upstream, keeps CC BY-SA 4.0, and keeps
|
||||
section 1 (Apache-2.0 for code) byte-identical.
|
||||
- Functional: `data/ATTRIBUTION.md` records the new source, the branch actually used, the
|
||||
two branches excluded and why, and the full list of modifications.
|
||||
- Functional: the regenerated `data/noitu.db` carries matching `meta` provenance.
|
||||
- Non-functional: no overclaiming. The branch is a 2018 scrape of vi.wiktionary.org made by
|
||||
a third party; say that, and attribute Wiktionary's contributors as CC BY-SA requires.
|
||||
|
||||
## Architecture
|
||||
|
||||
The dual-license structure in `NOTICE` is correct and stays; only the source lines of its
|
||||
second section change:
|
||||
|
||||
```
|
||||
1. SOURCE CODE — Apache-2.0 (unchanged: server/, tools/, web/, proto/)
|
||||
2. DICTIONARY DATA — CC BY-SA 4.0 (unchanged license, new upstream)
|
||||
data/noitu.db, data/LICENSE, data/ATTRIBUTION.md
|
||||
```
|
||||
|
||||
The attribution chain is now two links instead of an opaque aggregate: vi.wiktionary.org
|
||||
contributors (CC BY-SA) → `undertheseanlp/dictionary`, branch data `wiktionary`, commit
|
||||
`2c078cf` → this project. Both links are named.
|
||||
|
||||
The modifications list stays at eight: source selection replaces the old language filter
|
||||
as entry 1, which the license requires us to declare as a change.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Modify: `NOTICE` — section 2's upstream line, affected-artifacts line (`dictionary.db`
|
||||
is no longer the source), and the closing paragraph about build artifacts
|
||||
- Modify: `data/ATTRIBUTION.md` — source table, branch selection, exclusions, modifications
|
||||
9 and 10, the reproduce block
|
||||
- Modify: `README.md` — dictionary-source lines (`~179 MB`, `dictionary.db`, `minhqnd`,
|
||||
the manual `curl` + `--in` example)
|
||||
- Modify: `docs/deployment.md` — the 179 MB builder-stage sentence
|
||||
- Verify unchanged: `data/LICENSE`, `.github/workflows/ci.yml` required-files check
|
||||
- Regenerate: `data/noitu.db`
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Rewrite `NOTICE` section 2's source lines: upstream `undertheseanlp/dictionary` at
|
||||
commit `2c078cf`, wiktionary data only; affected artifact `data/noitu.db`. Leave the
|
||||
share-alike paragraph as is — it is still true.
|
||||
2. Rewrite `data/ATTRIBUTION.md`:
|
||||
- source is `undertheseanlp/dictionary`, file `dictionary/words.txt`, commit `2c078cf`
|
||||
- rows used: those tagged `wiktionary` — a 2018-12-10 scrape of vi.wiktionary.org,
|
||||
whose text is CC BY-SA 4.0 / GFDL by its contributors
|
||||
- `hongocduc` **excluded**: GNU GPL per the branch README; using it would relicense this
|
||||
database to GPLv3
|
||||
- `tudientv` **excluded**: its README declares copyright *"Chưa rõ"* and names
|
||||
Soha/Vietlex (Hoàng Phê) as its primary source. Say it plainly so nobody reopens it.
|
||||
- modifications renumbered: **1. source selection** replaces the language filter (the new
|
||||
file is Vietnamese-only); the other seven carry over, with normalization stating that
|
||||
capitalization removes nothing — the proper-noun drop was abandoned at the Phase 3 gate
|
||||
- reproduce block: `make fetch-dict` is ~4.8 MB now
|
||||
3. Grep the repository for `minhqnd`, `dictionary.db`, `179` and `--in` and fix every
|
||||
survivor outside `plans/` — the reports are dated records and stay as written.
|
||||
4. Regenerate `data/noitu.db` and confirm its `meta` rows agree with `ATTRIBUTION.md`:
|
||||
source URL, commit, license, `sources_kept = wiktionary`,
|
||||
`sources_excluded = hongocduc,tudientv`.
|
||||
5. Run the CI required-files check locally, or read it and confirm by hand.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] `data/LICENSE` has no diff.
|
||||
- [ ] `NOTICE` section 2 names the new upstream and still says CC BY-SA 4.0; section 1 is
|
||||
byte-identical.
|
||||
- [ ] `data/ATTRIBUTION.md` lists eight modifications, names the branch used and the two
|
||||
excluded with reasons, and credits vi.wiktionary.org contributors.
|
||||
- [ ] `meta` in the built database matches the attribution file.
|
||||
- [ ] No file outside `plans/` names minhqnd, `dictionary.db` as a source, or a 179 MB
|
||||
download.
|
||||
- [ ] CI's required-files check passes.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
**A stale source claim survives somewhere.** A README line or Docker comment naming the old
|
||||
upstream is an attribution misstatement, not a typo. Signal: the grep in step 3 finds
|
||||
something after the phase is called done. Response: grep is a step in the phase for that
|
||||
reason; run it last, not first.
|
||||
|
||||
**Section 1 gets damaged while section 2 is edited.** The code's Apache-2.0 grant is not
|
||||
in scope and must come out byte-identical. Signal: a diff touching section 1. Response:
|
||||
review the `NOTICE` diff before committing.
|
||||
|
||||
**The attribution credits the wrong party.** CC BY-SA attribution belongs to Wiktionary's
|
||||
contributors; undertheseanlp is the intermediary that scraped and redistributed. Signal:
|
||||
an `ATTRIBUTION.md` that names only the GitHub repository. Response: step 2 names both
|
||||
links of the chain explicitly.
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 5
|
||||
title: "Phase 5: Retire the SQLite path"
|
||||
status: todo
|
||||
status: done
|
||||
priority: P2
|
||||
effort: "1h"
|
||||
dependencies: [4]
|
||||
@@ -9,6 +9,13 @@ dependencies: [4]
|
||||
|
||||
# Phase 5: Retire the SQLite path
|
||||
|
||||
> **Outcome.** Done 2026-09-08. `--in`, `--table`, `--word-col`, `--lang-col`, `--lang`,
|
||||
> `resolveSource`, `validateSource`, `listTables`, `listColumns`, `pickColumn`, `extract`
|
||||
> and `quoteIdent` removed; `meta` no longer writes `source_word_column` /
|
||||
> `source_lang_column` (no consumer existed). The four auto-detection tests went; the
|
||||
> output-shape tests now build their fixture through `--merged`. `--help` lists `--merged`,
|
||||
> `--sources`, `--words`, `--out`, `--max-syllables`, `--min-words` and nothing else.
|
||||
|
||||
## Overview
|
||||
|
||||
Delete the input mode nothing uses any more: reading a source SQLite database, with the
|
||||
@@ -55,7 +62,8 @@ no longer read that upstream.
|
||||
2. Delete the tests that exist only to cover auto-detection. Do not delete tests covering
|
||||
output shape, `verify()`, or the reject-reason counts — those still describe behavior.
|
||||
3. Rewrite the package doc comment: the input is a 4.8 MB JSONL wordlist with source
|
||||
membership and capitalization; the output is the game's syllable-indexed database.
|
||||
membership and capitalization, of which only wiktionary-tagged rows are read; the
|
||||
output is the game's syllable-indexed database.
|
||||
4. `go build ./... && go vet ./... && go test ./...` from `server/`.
|
||||
5. Grep for `--in`, `dictionary.db`, `resolveSource` and `lang_code` outside `plans/`;
|
||||
nothing should survive except in the dated reports.
|
||||
|
||||
@@ -1,11 +1,12 @@
|
||||
---
|
||||
title: "Dictionary corpus switch"
|
||||
description: "Replace the 179 MB minhqnd aggregate with undertheseanlp/dictionary (hongocduc + wiktionary, tudientv excluded), drop always-capitalized proper nouns, and relicense data/noitu.db to GPLv3"
|
||||
status: pending
|
||||
description: "Replace the 179 MB minhqnd aggregate with the wiktionary branch of undertheseanlp/dictionary only (hongocduc and tudientv excluded) and keep data/noitu.db on CC BY-SA 4.0"
|
||||
status: in-progress
|
||||
priority: P1
|
||||
effort: "~1d"
|
||||
tags: [dictionary, data, licensing, build]
|
||||
created: 2026-09-08
|
||||
updated: 2026-09-08
|
||||
blockedBy: []
|
||||
blocks: []
|
||||
---
|
||||
@@ -14,115 +15,141 @@ blocks: []
|
||||
|
||||
## Overview
|
||||
|
||||
`data/noitu.db` is derived from `minhqnd/dictionary` v2.0.0: a 179 MB SQLite aggregate of
|
||||
five upstreams, redistributed as CC BY-SA 4.0. Two problems. Its license chain does not
|
||||
close — three of its five inputs are GPL (FVDP, tudientv, "Vietnamese Explanatory
|
||||
Dictionary") and GPL does not permit relicensing to CC BY-SA. And it has already
|
||||
lowercased everything (0 of 70,511 Vietnamese rows carry uppercase), which destroys the
|
||||
one signal that separates a place name from a word.
|
||||
`data/noitu.db` was derived from `minhqnd/dictionary` v2.0.0: a 179 MB SQLite aggregate of
|
||||
five upstreams, redistributed as CC BY-SA 4.0. Its license chain does not close — three of
|
||||
its five inputs are GPL (FVDP, tudientv, "Vietnamese Explanatory Dictionary") and GPL does
|
||||
not permit relicensing to CC BY-SA.
|
||||
|
||||
This plan switches the corpus to `undertheseanlp/dictionary` — a single 4.8 MB JSONL file
|
||||
carrying, per word, its raw capitalization and which of three source dictionaries contain
|
||||
it. `hongocduc` (GPL) and `wiktionary` (CC BY-SA) are taken; `tudientv` is excluded
|
||||
permanently because its own README declares copyright *"Chưa rõ"* and names Soha/Vietlex
|
||||
(Hoàng Phê) as its primary source. Entries that appear only ever capitalized are dropped
|
||||
as proper nouns. `data/noitu.db` then ships **GPLv3** — the honest license for
|
||||
FVDP-derived data, and the reason this is a licensing improvement rather than a new risk.
|
||||
This plan switches the corpus to the **`wiktionary` source inside `undertheseanlp/dictionary`**
|
||||
— a single 4.8 MB JSONL file carrying, per word, which of three source dictionaries contain
|
||||
it. Only rows whose `source` list contains `wiktionary` are read. `hongocduc` (GPL) is
|
||||
excluded so the data license does not change; `tudientv` is excluded because its own README
|
||||
declares copyright *"Chưa rõ"* and names Soha/Vietlex (Hoàng Phê) as its primary source.
|
||||
`data/noitu.db` **stays CC BY-SA 4.0**: the wiktionary branch is a 2018-12-10 scrape of
|
||||
vi.wiktionary.org, whose text is CC BY-SA.
|
||||
|
||||
Evidence, all measured through the real `build-dictionary` filter:
|
||||
`plans/reports/research-260908-1507-undertheseanlp-dictionary.md`.
|
||||
`plans/reports/research-260908-1507-undertheseanlp-dictionary.md` (branch licensing),
|
||||
`plans/reports/research-260908-1529-viwiktionary-dump-measured.md` (yields, alternatives)
|
||||
and [`audit-proper-noun-drops.md`](./audit-proper-noun-drops.md) (the Phase 3 gate).
|
||||
|
||||
## What changes, in numbers
|
||||
|
||||
| | current | after |
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| words | 48,216 | **~61,276** |
|
||||
| syllables | 6,676 | 7,011 |
|
||||
| syllables that can open a word | 5,049 | 5,496 |
|
||||
| …with ≥2 continuations | 3,682 | 4,168 |
|
||||
| dead-end syllables | 1,627 | **1,515** |
|
||||
| words | 48,216 | **26,845** |
|
||||
| syllables | 6,676 | 5,709 |
|
||||
| syllables that can open a word | 5,049 | 4,158 |
|
||||
| …with ≥2 continuations | 3,682 | 2,787 |
|
||||
| dead-end syllables | 1,627 | 1,551 |
|
||||
| source download | 179 MB SQLite | 4.8 MB JSONL |
|
||||
| data license | CC BY-SA 4.0 | **GPLv3** |
|
||||
| data license | CC BY-SA 4.0 | **CC BY-SA 4.0 (unchanged)** |
|
||||
| `--min-words` floor | 40,000 | **20,000** |
|
||||
|
||||
24,978 words gained, 11,918 lost — of which 4,145 are proper nouns removed on purpose,
|
||||
1,300 are `tudientv`-only, and 6,473 are absent from undertheseanlp entirely, including
|
||||
genuine modern vocabulary (`tế bào gốc`, `hằng số vũ trụ`) that a 2003 wordlist cannot
|
||||
have. That loss is accepted here and is the subject of a separate future decision, not
|
||||
this plan.
|
||||
**This is a deliberate 44% cut.** 1,280 words gained, 22,651 lost — every one of them
|
||||
simply absent from a 2018 wiktionary scrape (`bánh đúc`, `lò vi sóng`, `dân làng`, `thót
|
||||
tim`). Bot-versus-bot games run about a quarter shorter on the thinner graph (easy-vs-easy
|
||||
12.9 moves vs 17.5). This loss is accepted here; the upgrade path is written down below.
|
||||
|
||||
## Decisions taken
|
||||
|
||||
- **`tudientv` is never read.** Not as content, and not as case evidence. Excluding it
|
||||
from evidence too is what makes "we do not touch that data" true without qualification.
|
||||
It costs a little accuracy in the proper-noun rule — a word capitalized in `hongocduc`
|
||||
but lowercase only in `tudientv` will be dropped — and Phase 3 measures that cost.
|
||||
- **The measured 61,276 used all-source case evidence.** Under the allowed-sources-only
|
||||
rule the number will shift slightly. Phase 1 re-measures; the plan does not depend on
|
||||
the exact figure, only on it staying above the 40,000 `--min-words` floor and above
|
||||
today's 48,216.
|
||||
- **Proper nouns are dropped, not tiered.** The per-word `source` array looked like a
|
||||
confidence signal and is not one: `hà nội` has 2 votes, `cẩm xá` 2, `kháng đón` 2, while
|
||||
`cánh diều` has 1 and `mặt trời` 2. Capitalization is the only usable signal in this
|
||||
data, so it is the only one used.
|
||||
- **`noitu.db` becomes GPLv3; the code stays Apache-2.0.** The existing dual-license split
|
||||
in `NOTICE` holds; only its data half is rewritten.
|
||||
- **The SQLite input path is deleted, not kept "just in case."** Two input modes for one
|
||||
source is dead weight; the fixture path (`--words`) stays because the e2e suite and the
|
||||
Docker smoke build use it.
|
||||
- **The viwiktionary extractor is out of scope.** It remains the answer if a CC BY-SA
|
||||
corpus is ever needed again — see `research-260908-1434-viwiktionary-as-dictionary-source.md`.
|
||||
- **Only `wiktionary` rows are read.** `hongocduc` and `tudientv` inform nothing. That is
|
||||
what keeps the license CC BY-SA 4.0 and what makes "we do not touch that data" true
|
||||
without qualification.
|
||||
- **Capitalization is not a filter.** The original design dropped words that appear only
|
||||
ever capitalized, as proper nouns. It was built, run and audited at the Phase 3 gate: the
|
||||
rule was precise on a 100-word sample (1 common word) but, with wiktionary-only case
|
||||
evidence, it also removed 215 words including `mặt trời`, `trái đất`, `tổ quốc`, `dường
|
||||
như`, `phật giáo`. Narrowing was measured and did not help. **The owner chose to keep
|
||||
every word regardless of case.** Proper nouns like `hà nội` remain playable, exactly as
|
||||
they are today. The mechanism was removed from the builder rather than left behind a flag.
|
||||
- **Rejected alternatives, so nobody reopens them without new evidence:**
|
||||
- `hongocduc + wiktionary` (61,271 words) — forces `noitu.db` to GPLv3. Rejected to keep
|
||||
CC BY-SA 4.0.
|
||||
- the 2026-09-01 viwiktionary dump (36,200 words without a case rule, same license,
|
||||
monthly pin) — the 2018 branch is a near-subset of it. Rejected for now in favour of the
|
||||
simpler single-file pin; **this is the documented upgrade path** if the 27k corpus
|
||||
proves too thin in play. See the 15:29 report.
|
||||
- kaikki.org enwiktionary Vietnamese — unpinnable rolling file, deprecated by its host.
|
||||
- **The SQLite input path is deleted, not kept "just in case."** The fixture path
|
||||
(`--words`) stays because the e2e suite and the Docker smoke build use it.
|
||||
- **The build floor drops to 20,000.** The old 40,000 default would reject the new corpus;
|
||||
it moved once, in `main.go`.
|
||||
|
||||
## The pin
|
||||
|
||||
```
|
||||
URL https://raw.githubusercontent.com/undertheseanlp/dictionary/2c078cfc373b06e2980d324ce1d7bd13740c3319/dictionary/words.txt
|
||||
SHA256 4c3e0e6117e4bdfa97731e135c3d4a05881889909267394a8de8d88ef79f13f0
|
||||
Size 4,813,111 bytes — 79,226 JSONL rows
|
||||
Size 4,813,111 bytes — 79,226 JSONL rows, 32,484 of them tagged wiktionary
|
||||
Commit 2c078cf (2018-12-10, the repository's last)
|
||||
```
|
||||
|
||||
Pinned by commit rather than branch: `raw.githubusercontent.com/<repo>/<sha>/` is
|
||||
immutable, so the URL and the checksum cannot disagree later. The variable names
|
||||
`DICT_URL` / `DICT_SHA256` are kept so `web/tests/dictionary-source.test.js` needs no
|
||||
change — only their values move.
|
||||
`DICT_URL` / `DICT_SHA256` were kept so `web/tests/dictionary-source.test.js` needed no
|
||||
change — only their values moved.
|
||||
|
||||
## Phases
|
||||
|
||||
| # | Phase | Status |
|
||||
|---|-------|--------|
|
||||
| 1 | [Read the merged list](./phase-01-start.md) | Pending |
|
||||
| 2 | [Pin the source](./phase-02-pin-the-source.md) | Pending |
|
||||
| 3 | [Audit the corpus](./phase-03-audit-the-corpus.md) | Pending |
|
||||
| 4 | [Relicense the data](./phase-04-relicense-the-data.md) | Pending |
|
||||
| 5 | [Retire the SQLite path](./phase-05-retire-the-sqlite-path.md) | Pending |
|
||||
| 1 | [Read the merged list](./phase-01-start.md) | Done |
|
||||
| 2 | [Pin the source](./phase-02-pin-the-source.md) | Done |
|
||||
| 3 | [Audit the corpus](./phase-03-audit-the-corpus.md) | Done — gate sent the case rule back; rule abandoned |
|
||||
| 4 | [Update the attribution](./phase-04-update-the-attribution.md) | Done |
|
||||
| 5 | [Retire the SQLite path](./phase-05-retire-the-sqlite-path.md) | Done |
|
||||
|
||||
Phase 3 is a gate: it may send Phase 1 back to narrow the proper-noun rule. Phase 4 must
|
||||
land in the same commit as the first generated GPL-derived `noitu.db` — shipping that
|
||||
database while `NOTICE` still says CC BY-SA 4.0 is the one ordering mistake with a legal
|
||||
consequence.
|
||||
All five phases are uncommitted in one working tree; they should land together so the
|
||||
attribution never describes a database other than the one the build produces.
|
||||
|
||||
## Success criteria
|
||||
|
||||
- [ ] `make fetch-dict && make dict` builds `data/noitu.db` from a 4.8 MB pinned file with
|
||||
- [x] `make fetch-dict && make dict` builds `data/noitu.db` from a 4.8 MB pinned file with
|
||||
no 179 MB download.
|
||||
- [ ] The shipped database has **more than 55,000 words** and **no word that reaches it
|
||||
only via `tudientv`**.
|
||||
- [ ] The proper-noun drop count is reported in the build log beside the other reject
|
||||
reasons, recorded in `meta`, and a 100-entry sample has been audited by hand.
|
||||
- [ ] Playability does not regress: syllables ≥ 6,676 and dead-end syllables ≤ 1,627.
|
||||
- [ ] `data/LICENSE`, `NOTICE` and `data/ATTRIBUTION.md` describe GPLv3, name both source
|
||||
branches with their own licenses, and record why `tudientv` is excluded.
|
||||
- [ ] `make test-go`, `make test-web` and the e2e suite are green; the fixture build path
|
||||
is untouched.
|
||||
- [ ] `--in` and its schema auto-detection are gone, and no doc still mentions the 179 MB
|
||||
download.
|
||||
- [x] The shipped database has **more than 20,000 words** and **no word that reaches it
|
||||
through `hongocduc` or `tudientv` alone**.
|
||||
- [x] `meta` records the source URL, commit, sources kept and sources excluded.
|
||||
- [x] Playability is measured and recorded: syllables, openers, continuations, dead ends,
|
||||
and bot-versus-bot game lengths on the new graph versus the old.
|
||||
- [x] `data/LICENSE` is unchanged; `NOTICE` and `data/ATTRIBUTION.md` name the new source,
|
||||
its commit and license, and record why the two other branches are excluded. The
|
||||
frontend attribution footer credits Wiktionary tiếng Việt.
|
||||
- [x] `go vet ./... && go test ./...` (incl. `-race` on the touched packages), `npm run
|
||||
check && npm test`, and the Docker image (`noitu:ci` with the CI file checks, plus
|
||||
`FIXTURE_DICT=1`) are green; the fixture build path is untouched.
|
||||
- [ ] The Playwright e2e suite is green. **Not verified locally**: Playwright's Chromium
|
||||
download timed out twice on this machine (2026-09-08), so every test failed at browser
|
||||
launch, not on game behaviour. The one e2e assertion this plan changes (the attribution
|
||||
footer's link text and href in `web/e2e/bot-game.spec.js`) was updated by hand. Verify
|
||||
in CI or after `npx playwright install chromium` succeeds.
|
||||
- [x] `--in` and its schema auto-detection are gone, and no file outside `plans/` still
|
||||
mentions minhqnd, `dictionary.db`, `--in` or a 179 MB download.
|
||||
|
||||
## Licence version — decided after review
|
||||
|
||||
The owner asked whether the data could move to Apache-2.0 to simplify the repo. It cannot:
|
||||
the words are Wiktionary contributors' text under a share-alike licence, and neither
|
||||
undertheseanlp nor this project can relicense them. The owner then chose to **keep the
|
||||
licence the source text actually carried, CC BY-SA 3.0 Unported**, rather than upgrade the
|
||||
derivative to 4.0 under the later-version clause. `data/LICENSE` is now the BY-SA 3.0 legal
|
||||
code; `NOTICE`, `ATTRIBUTION.md`, README, Dockerfile, deployment doc, the builder's meta
|
||||
string and the in-game footer all say 3.0. The "unchanged" wording above predates this.
|
||||
|
||||
## Review
|
||||
|
||||
Code review (2026-09-08) returned DONE_WITH_CONCERNS; every finding was applied: the
|
||||
builder's copy of the upstream commit is now guarded by `web/tests/dictionary-source.test.js`
|
||||
alongside the Makefile/Dockerfile pin; fixture builds record "no upstream data" instead of
|
||||
inheriting the CC BY-SA string; `ATTRIBUTION.md` states the 2018 snapshot was CC BY-SA 3.0
|
||||
and why the derived database is 4.0; dead `contains()`, redundant sorts, the off-by-one
|
||||
scanner line number, stale "48k" comments, the Makefile `curl` flags and the Dockerfile
|
||||
download name were fixed; input-selection and fixture-licence tests were added.
|
||||
|
||||
## Open questions
|
||||
|
||||
- The `hongocduc` GPL claim rests on the mirror's README; Hồ Ngọc Đức's canonical site
|
||||
404s. Assuming GPL is the conservative direction, so this does not block — but the
|
||||
evidence is a mirror, and `ATTRIBUTION.md` should say so rather than overclaim.
|
||||
- The 6,473 words undertheseanlp does not have are an unquantified mix of junk and real
|
||||
modern vocabulary. Separating them needs a second live source; out of scope here.
|
||||
- Is 26,845 words enough for the game to feel playable? Bot games are ~25% shorter; only
|
||||
real play can say whether that is felt. If not, the 2026 viwiktionary dump is the next
|
||||
step, not a return to GPL data.
|
||||
|
||||
<!-- slug: dictionary-corpus-switch -->
|
||||
+36
@@ -0,0 +1,36 @@
|
||||
---
|
||||
title: Dictionary corpus switch to wiktionary-only undertheseanlp
|
||||
date: 2026-09-08
|
||||
summary: "Replaced the 179 MB minhqnd aggregate with the wiktionary rows of undertheseanlp/dictionary; corpus 48,216 -> 26,845 words, CC BY-SA 4.0 kept, proper-noun drop abandoned after audit, SQLite path retired"
|
||||
---
|
||||
|
||||
# Dictionary corpus switch to wiktionary-only undertheseanlp
|
||||
|
||||
## What happened
|
||||
|
||||
- Three research passes today measured every candidate through the real `build-dictionary` filter: the 2026-09-01 viwiktionary dump (36,200 words), kaikki's enwiktionary Vietnamese file (22,628, deprecated and unpinnable), undertheseanlp's merged list (61,271 but GPL via hongocduc) and its wiktionary rows alone (26,845). Reports: `plans/reports/research-260908-1529-viwiktionary-dump-measured.md`.
|
||||
- Owner chose undertheseanlp wiktionary-only to keep `data/noitu.db` on CC BY-SA 4.0. Plan `plans/260908-1525-dictionary-corpus-switch/` rewritten around it, then executed end to end.
|
||||
- `build-dictionary` gained `--merged` / `--sources` (default `wiktionary`, names validated), a shared `finish()` tail, `--min-words` default 20,000, and meta rows `source_commit`, `sources_kept`, `sources_excluded`. The SQLite `--in` path, schema auto-detection and their tests were deleted.
|
||||
- Makefile, Dockerfile, `.gitignore`, CI leak guard, NOTICE, `data/ATTRIBUTION.md`, README, `docs/deployment.md` and the frontend attribution footer now name the new source. `data/LICENSE` untouched.
|
||||
|
||||
## Decision
|
||||
|
||||
- **Capitalization is not a filter.** The planned proper-noun drop was built and audited: precise on a 100-word sample (1 common word), but with wiktionary-only case evidence it removed 215 words including `mặt trời`, `trái đất`, `tổ quốc`, `dường như`. Narrowing to all-syllables-capitalized was measured (147 casualties remain, 410 junk entries return) and rejected. Owner decided to keep every word regardless of case; the mechanism was removed, not flagged off. Record: `audit-proper-noun-drops.md` in the plan dir.
|
||||
- Attribution states the 2018 scrape was CC BY-SA 3.0 (Wikimedia moved to 4.0 in 2023) and that the derived database is 4.0 via the later-version clause. Both chain links are credited: Wiktionary tiếng Việt contributors, then undertheseanlp as intermediary.
|
||||
- Rejected alternatives recorded in `plan.md`: hongocduc+wiktionary (GPLv3 relicense), viwiktionary 2026 dump (documented upgrade path if 27k proves thin), kaikki (unpinnable).
|
||||
|
||||
## Verification
|
||||
|
||||
- `go vet ./... && go test ./...` green, `-race` on touched packages green; `npm run check && npm test` 175/175; Docker `noitu:ci` built with the CI required-files and leak checks passing, `FIXTURE_DICT=1` variant built.
|
||||
- Bot real-corpus tests pass on the new DB; easy-vs-easy games 12.9 moves vs 17.5 before, the expected depth loss.
|
||||
- Code review (DONE_WITH_CONCERNS) fully applied: builder's copy of the pinned commit now guarded by `web/tests/dictionary-source.test.js`; fixture builds record "no upstream data" instead of a CC BY-SA string; dead `contains()`, redundant sorts, scanner line off-by-one, stale "48k" comments, `curl -f`, Dockerfile download name fixed.
|
||||
- **Not verified:** Playwright e2e — Chromium download timed out twice on this machine. The one changed e2e assertion (footer link) was updated by hand. Needs CI or a successful `npx playwright install chromium`.
|
||||
|
||||
## Next steps
|
||||
|
||||
- Commit all five phases together so attribution and database never disagree (not yet committed).
|
||||
- Confirm e2e in CI.
|
||||
- Watch real play for thinness; the viwiktionary dump is the upgrade path, not GPL data.
|
||||
- Delete the orphaned local `data/dictionary.db` (179 MB, gitignored).
|
||||
|
||||
> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions.
|
||||
@@ -0,0 +1,271 @@
|
||||
---
|
||||
title: "Research: building a dictionary DB from the dumps.wikimedia.org viwiktionary dump — measured"
|
||||
date: 2026-09-08T15:29+07:00
|
||||
type: research
|
||||
status: complete
|
||||
scope: research-only — dump downloaded to scratchpad, parsed, run through the real build-dictionary filter; no repo file touched
|
||||
follows: research-260908-1434-viwiktionary-as-dictionary-source.md, plans/260908-1525-dictionary-corpus-switch/plan.md
|
||||
---
|
||||
|
||||
# Can we build our own Vietnamese dictionary DB from the viwiktionary dump?
|
||||
|
||||
## Executive summary
|
||||
|
||||
**Yes, and it is cheap: 63.5 MB download, 34 s of stdlib Python, no wiktextract, no Lua.** The
|
||||
2026-09-01 `pages-articles` dump was downloaded, checksum-verified against Wikimedia's published
|
||||
MD5, parsed, and every list below was run through our real `build-dictionary --words` filter.
|
||||
Numbers are measured, not estimated.
|
||||
|
||||
**But viwiktionary alone is not a corpus.** It yields **36,200** accepted words, **31,637** after
|
||||
dropping pages labelled proper-noun. Both sit below the shipped 48,216 and below the 40,000
|
||||
`--min-words` floor. The 14:34 report's "~16k from this source" undercounted by 2×, because the
|
||||
`minhqnd` aggregate only ingested a 17k-row slice; the full dump is 36k. Still not enough.
|
||||
|
||||
**Where it earns its place is on top of the pending plan.** Against the plan's undertheseanlp
|
||||
corpus (re-measured this session at **61,026**, matching the plan's ~61,276), the 2026 dump adds
|
||||
**3,084 words** the plan would otherwise lose — `học liệu`, `thiên hà lùn`, `cải thảo`, `bùn đỏ`,
|
||||
`báng súng`, `tăng đoàn` — for a union of **64,110**. That is the modern-vocabulary gap the plan's
|
||||
own open question names, and the plan's 2018 wiktionary branch (26,845 accepted) is 9,780 words
|
||||
behind the 2026 dump. **Recommendation: after the corpus switch lands, add the dated dump as a
|
||||
second `--words` input, pinned by dated URL + checksum.** Licensing stays clean: CC BY-SA 4.0 is
|
||||
one-way compatible into the GPLv3 the plan already adopts.
|
||||
|
||||
**Do not use viwiktionary's proper-noun labels as a drop rule.** They agree with the
|
||||
capitalization blocklist 86% of the time, but they tag `mặt trời`, `trái đất`, `thứ hai`,
|
||||
`tháng ba`, `tin lành`, `bàn là` as proper-only and `việt nam` as common. As a *filter* the label
|
||||
is worse than capitalization; as a *report column* it is fine.
|
||||
|
||||
---
|
||||
|
||||
## Method
|
||||
|
||||
- Web: 2 fetches (dump index, kaikki index), plus direct `curl` of the dump, its checksum file,
|
||||
4 sample raw wikitext pages, and the plan's pinned undertheseanlp file.
|
||||
- Local: `extract.py` (streams bz2 → per-page markers, 34 s), `classify.py` (slices the
|
||||
Vietnamese section, classifies POS labels, emits word lists), `build-dictionary --words` for
|
||||
every list, sqlite for overlaps. Scripts are in the session scratchpad; each is < 100 lines.
|
||||
|
||||
## The dump
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| File | `https://dumps.wikimedia.org/viwiktionary/20260901/viwiktionary-20260901-pages-articles.xml.bz2` |
|
||||
| Size | 63,513,513 bytes |
|
||||
| MD5 (Wikimedia-published, verified) | `6c2491e703e7d946f23a405996b4d172` |
|
||||
| SHA-256 (computed here; Wikimedia publishes md5/sha1 only for this run) | `ed66c932f535b0b362d1141c02273b02ed01c4101e856378bd15e88d277bb8d9` |
|
||||
| Cadence | monthly, dated dirs `20251220 … 20260901`, plus rolling `latest/` |
|
||||
| Pages / ns0 pages | 391,543 / 349,461 |
|
||||
| **ns0 pages with a Vietnamese section** | **43,013** (on-wiki category says 43,037 — parser is complete) |
|
||||
| Redirect pages among them | 3,237; 834 are case-only (`mặt trời` → `Mặt Trời`) |
|
||||
|
||||
Dated directories are immutable, so a pin is honest. `latest/` is not. kaikki.org tracks the same
|
||||
dump (extracted 2026-09-06, 56,502 Vietnamese *senses*) but publishes rolling files — the earlier
|
||||
report's pinning concern stands; the dump removes the need for kaikki entirely.
|
||||
|
||||
### Two wikitext dialects, both live
|
||||
|
||||
| Format | Pages | Language marker | POS marker |
|
||||
|---|---|---|---|
|
||||
| legacy | 35,885 | `{{-vie-}}` | `{{-noun-}}`, `{{-verb-}}`, `{{-pr-noun-}}`, `{{-place-}}` … |
|
||||
| new | 7,128 | `== {{langname\|vi}} ==` | `{{ĐM\|noun}}`, `{{ĐM\|pr-noun}}`, `{{vi-noun}}`, `{{vi-pr-noun}}` … |
|
||||
|
||||
A parser must handle both; the wiki is mid-migration, so the mix shifts every month. Legacy
|
||||
section codes collide with language codes (`{{-adj-}}` vs `{{-eng-}}`), which is why the
|
||||
extractor keeps an explicit section-code set and treats every other 2–3-letter code as a language
|
||||
switch. Languages seen switching out of Vietnamese: `tyz`, `mtq`, `eng`, `nut`, `nuo`, `fra` …
|
||||
|
||||
## Classification of the 43,013 pages
|
||||
|
||||
| Class | Pages | Meaning |
|
||||
|---|---|---|
|
||||
| common | 31,950 | ≥1 common POS label (noun/verb/adj/adv/phrase/idiom/…) |
|
||||
| proper-only | 5,103 | only `pr-noun` / `place` / `vi-pr-noun` labels |
|
||||
| no-pos | 5,764 | Vietnamese section, no POS marker at all |
|
||||
| nôm-only | 169 | only Nôm/Hán character sections |
|
||||
| proper+common | 27 | |
|
||||
|
||||
The **no-pos** class is real vocabulary, not junk: 5,538 of them pass our filter and **4,912 are
|
||||
already shipped** (`kỵ mã`, `khoái hoạt`, `quá tay`, `nặn óc`). Keep them.
|
||||
|
||||
## Through `build-dictionary --words`
|
||||
|
||||
| List | Accepted | Rejected <2 syll | punct | non-VN |
|
||||
|---|---|---|---|---|
|
||||
| all 43,013 titles | **36,200** | 6,361 | 155 | 162 |
|
||||
| minus proper-only | **31,637** | 6,051 | 121 | 84 |
|
||||
| proper-only alone | 4,661 | 310 | 34 | 78 |
|
||||
| undertheseanlp, plan rule (hongocduc+wiktionary, no tudientv, drop caps-only) | **61,026** | | | |
|
||||
| undertheseanlp `wiktionary` branch alone (2018 snapshot) | 26,845 | | | |
|
||||
|
||||
Playability of a viwiktionary-only DB vs shipped:
|
||||
|
||||
| metric | shipped | vi-all | vi-no-proper |
|
||||
|---|---|---|---|
|
||||
| words | 48,216 | 36,200 | 31,637 |
|
||||
| syllables | 6,676 | 6,172 | 5,981 |
|
||||
| syllables that open a word | 5,049 | 4,600 | 4,515 |
|
||||
| …with ≥2 continuations | 3,682 | 3,246 | 3,168 |
|
||||
| dead-end syllables | 1,627 | 1,572 | 1,466 |
|
||||
|
||||
Every breadth metric regresses. **Option A (replace) is dead on measurement, not opinion.**
|
||||
|
||||
## Overlaps — what the dump is actually good for
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| shipped ∩ viwiktionary | **34,570 of 48,216 (72%)** — viwiktionary confirms most of what we ship |
|
||||
| shipped, not in viwiktionary | 13,646 (9,713 two-syllable) — `thai sản`, `vách núi`, `điện từ trường`, `giả kim`: real words, so absence ≠ junk |
|
||||
| viwiktionary-only, not shipped | 1,630 |
|
||||
| shipped words viwiktionary labels proper-only | 4,523 (`ba đồn`, `vị thanh`, `xuân lập` …) |
|
||||
| **plan corpus (61,026) + vi-no-proper → union** | **64,110 (+3,084)** |
|
||||
| plan corpus words viwiktionary labels proper-only | 293 — includes `thứ hai`, `tháng năm`, `tin lành`, `tân ước`, `bàn là` |
|
||||
| 2018 wiktionary branch vs 2026 dump | 26,420 shared; **9,780 new since 2018**; 425 gone |
|
||||
| shipped words in neither plan corpus nor viwiktionary | 9,267 |
|
||||
|
||||
### Proper-noun label vs capitalization blocklist
|
||||
|
||||
- viwiktionary proper-only ∩ undertheseanlp caps-only: **4,015 of 4,661 (86%)** agree.
|
||||
- Caps-only words viwiktionary calls common: 139 (incl. `việt nam`).
|
||||
- Common words viwiktionary calls proper-only: `mặt trời`, `trái đất`, `sao thủy` (celestial
|
||||
bodies are `Mặt Trời` on-wiki), weekday/month names, `tin lành`, `bàn là`.
|
||||
|
||||
Verdict: capitalization (the plan's rule) is the better *drop* signal. The label is a useful
|
||||
*second opinion* for the plan's Phase 3 audit sample, nothing more.
|
||||
|
||||
## Licensing
|
||||
|
||||
Wiktionary text is CC BY-SA 4.0 / GFDL. Creative Commons declares CC BY-SA 4.0 **one-way
|
||||
compatible with GPLv3**, so folding viwiktionary into the GPLv3 `noitu.db` the plan produces is
|
||||
clean. `data/ATTRIBUTION.md` gains one entry: edition, dump date, URL, checksum. No `NOTICE`
|
||||
restructure.
|
||||
|
||||
## Recommendation
|
||||
|
||||
1. **Do not switch to viwiktionary.** 36,200 < 40,000 floor; every playability metric regresses.
|
||||
2. **Land the undertheseanlp plan as written.** Its numbers reproduce (61,026 vs ~61,276).
|
||||
3. **Then add the dated dump as a second `--words` source** — a follow-up phase, not a change to
|
||||
the current plan. Concretely: `make fetch-dict` also fetches the pinned dump; a ~100-line
|
||||
extractor (stdlib, both dialects, redirects skipped, `--drop-proper` off by default) emits a
|
||||
word list; the builder reads two `--words` files or the Makefile concatenates them.
|
||||
Expected result: **~64,100 words**, +3,084 modern/compound vocabulary, 12-month freshness via a
|
||||
monthly-dated pin.
|
||||
4. **Do not drop on viwiktionary's proper-noun label.** Record it in `meta`/build log as a
|
||||
count; use it to spot-check the capitalization rule in Phase 3.
|
||||
5. Keep kaikki out. The dump is smaller than kaikki's JSONL, pins honestly, and the parser is
|
||||
trivial because we only need titles + section labels, not senses.
|
||||
|
||||
## Addendum — kaikki.org `dictionary/Vietnamese/kaikki.org-dictionary-Vietnamese.jsonl`
|
||||
|
||||
Asked after the main report. This file is **not viwiktionary**: kaikki's `dictionary/` tree is
|
||||
the *English* Wiktionary, filtered to `lang_code = vi`. Dump 2026-09-02, extracted 2026-09-06,
|
||||
79,161,423 bytes, SHA-256 `d878bd23fe4d6ac4480736858a85b2bbca0376d6cc9b30395fff0ba1e16cf97b`
|
||||
at fetch time. Measured the same way as everything above.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| rows / distinct words | 51,896 / 45,281 |
|
||||
| POS | noun 19,206 · character 9,029 · verb 8,992 · adj 7,070 · **name 3,814** · adv 1,160 · … |
|
||||
| rejected `< 2 syllables` | 22,115 — half the file is single syllables and Hán/Nôm characters |
|
||||
| **accepted, all** | **22,628** |
|
||||
| **accepted, `name`-only words dropped** | **21,244** |
|
||||
|
||||
| metric | shipped | kaikki-en | vi-no-proper | uts-plan |
|
||||
|---|---|---|---|---|
|
||||
| words | 48,216 | 21,244 | 31,637 | 61,026 |
|
||||
| syllables | 6,676 | 5,145 | 5,981 | 7,005 |
|
||||
| openers ≥2 | 3,682 | 2,431 | 3,168 | 4,164 |
|
||||
| dead-end | 1,627 | 1,399 | 1,466 | 1,513 |
|
||||
|
||||
Smallest of every candidate as a corpus. **As a third supplement it is the best one:**
|
||||
|
||||
| overlap | words |
|
||||
|---|---|
|
||||
| kaikki-en ∩ shipped | 20,231 (adds 1,013) |
|
||||
| kaikki-en adds over plan corpus | 4,394 |
|
||||
| kaikki-en adds over plan corpus ∪ viwiktionary | **3,693** |
|
||||
| **union plan ∪ viwiktionary ∪ kaikki-en** | **67,803** |
|
||||
| shipped words in none of the three | 6,187 |
|
||||
|
||||
The additions are exactly the modern register the other two lack: `sổ hồng`, `quay xe`, `thi hành
|
||||
án`, `nhà máy lọc dầu`, `cát tặc`, `giấy ướt`, `máy tính tiền`, `thập lục phân`, `đường tiêu hóa`,
|
||||
`tân tổng thống`, `sói đồng cỏ`. English Wiktionary's Vietnamese section is more actively curated
|
||||
than viwiktionary's.
|
||||
|
||||
**Its `name` POS is the cleanest proper-noun label seen so far.** `hà nội`, `việt nam`, `trái đất`
|
||||
are `name`; `thứ hai`, `tin lành` are not; `mặt trời` is both `name` and `noun` and so survives.
|
||||
Of 1,483 accepted name-only words, 1,359 are shipped today and 247 survive the plan's caps rule
|
||||
(`thượng đế`, `bắc cực`, `trung thu`, `siêu nhân` — arguable either way). Good audit column; still
|
||||
not a sole drop rule.
|
||||
|
||||
**Two blockers to using the URL as given:**
|
||||
|
||||
1. **It is deprecated.** kaikki's index marks the postprocessed per-language file as *"deprecated
|
||||
and will be removed in the future"* and points to the raw download page, where the only
|
||||
non-deprecated form of this data is the full-edition `raw-wiktextract-data.jsonl`
|
||||
(23.1 GB, 2.7 GB compressed) filtered by `lang_code`. There is no per-language file in the
|
||||
replacement layout — the `downloads/<code>/` entries are per *edition*, not per language.
|
||||
2. **It cannot be pinned.** Rolling URL, refreshed weekly, no archived snapshots. A
|
||||
`DICT_SHA256`-style pin breaks on the next refresh; `make fetch-dict` would fail weekly.
|
||||
|
||||
Ways to use it anyway, in order of preference:
|
||||
|
||||
- **Vendor the derived word list.** Run the reduction once, commit `data/sources/kaikki-en-vi.txt`
|
||||
(~21k lines, ~400 KB) with the fetch date, URL and file SHA-256 in `ATTRIBUTION.md`. Pins
|
||||
honestly, costs nothing at build time, refreshed deliberately. Note this is a *derived list*,
|
||||
not a copy of their file, so the plan's "never commit the source file" rule is not violated in
|
||||
spirit — decide explicitly.
|
||||
- **Pin the enwiktionary dump and extract ourselves.** `enwiktionary-20260902-pages-articles`
|
||||
is >1 GB and wiktextract with Lua expansion takes hours. Correct but disproportionate for
|
||||
~3.7k words.
|
||||
- **Re-pin on every refresh.** Not acceptable; turns a data pin into a weekly chore.
|
||||
|
||||
License: same CC BY-SA 4.0 / GFDL as all Wiktionary text, so the GPLv3 result stays clean.
|
||||
Attribution should name enwiktionary + wiktextract/kaikki (the extraction is Tatu Ylonen's work).
|
||||
|
||||
**Verdict:** yes, usable, and worth ~3,700 modern words on top of plan + viwiktionary — but only
|
||||
via a vendored derived list, never as a live fetch of that URL.
|
||||
|
||||
## Addendum — undertheseanlp `wiktionary` branch alone (the option chosen for the plan)
|
||||
|
||||
Measured after the user chose to use undertheseanlp with its wiktionary data only. Rows with
|
||||
`wiktionary` in `source`: 32,484 → 32,374 lowercase forms → 4,883 caps-only dropped (case
|
||||
evidence from wiktionary rows only) → 27,491 → **22,310 accepted**.
|
||||
|
||||
| metric | shipped | uts-wik only | vi-dump 2026 | wik ∪ dump | uts-plan (GPL) |
|
||||
|---|---|---|---|---|---|
|
||||
| words | 48,216 | **22,310** | 31,637 | 31,817 | 61,026 |
|
||||
| syllables | 6,676 | 5,484 | 5,981 | 6,012 | 7,005 |
|
||||
| openers | 5,049 | 4,050 | 4,515 | 4,539 | 5,492 |
|
||||
| openers ≥2 | 3,682 | 2,686 | 3,168 | 3,180 | 4,164 |
|
||||
| dead-end | 1,627 | 1,434 | 1,466 | 1,473 | 1,513 |
|
||||
|
||||
- shipped ∩ uts-wik 21,230 · shipped lost 26,986 (4,335 proper-noun drops, 22,651 absent
|
||||
from the branch) · new 1,080.
|
||||
- The 2018 branch is a near-subset of the 2026 dump: union adds 180 words. Same license.
|
||||
- **Case-rule casualties:** 271 caps-only drops have a lowercase form in another branch
|
||||
(247 in hongocduc): `mặt trời`, `trái đất`, `hệ mặt trời`, `phật giáo`, `nguyên đán`,
|
||||
`tia x`. Wiktionary holds `Mặt Trời` / `Trái Đất` capitalized only.
|
||||
- Probes: all 11 known-bad absent ✓. Of 10 ordinary probes, `mặt trời` dropped by the case
|
||||
rule; `con người`, `cánh diều` not in the branch at all; the other 7 present.
|
||||
- Decision recorded in `plans/260908-1525-dictionary-corpus-switch/plan.md`: wiktionary-only,
|
||||
CC BY-SA 4.0 unchanged, floor 20,000, viwiktionary 2026 dump as the upgrade path.
|
||||
|
||||
## Sources
|
||||
|
||||
- https://dumps.wikimedia.org/viwiktionary/ (index; dated dirs 20251220–20260901)
|
||||
- https://dumps.wikimedia.org/viwiktionary/20260901/ (file list, md5/sha1 sums)
|
||||
- https://kaikki.org/viwiktionary/ (dump 2026-09-01, extracted 2026-09-06, 56,502 vi senses)
|
||||
- https://kaikki.org/dictionary/Vietnamese/ and https://kaikki.org/dictionary/rawdata.html (enwiktionary edition; deprecation notice, no snapshots)
|
||||
- raw wikitext of `Hà Nội`, `học sinh`, `nhà`, `mặt trời` on vi.wiktionary.org
|
||||
- https://raw.githubusercontent.com/undertheseanlp/dictionary/2c078cfc…/dictionary/words.txt (plan's pin, sha256 verified `4c3e0e61…`)
|
||||
- Local: `server/cmd/build-dictionary`, `data/noitu.db`
|
||||
|
||||
## Unresolved questions
|
||||
|
||||
- The 9,267 shipped words in neither the plan corpus nor viwiktionary are still an unquantified
|
||||
junk/real mix; this dump does not settle them.
|
||||
- Whether Wikimedia publishes a sha256 file for this run later (only md5/sha1 exist today).
|
||||
Pin on our computed SHA-256 or on the published MD5 — decide when the phase is written.
|
||||
- Legacy → new wikitext migration pace: if the wiki finishes it, the legacy branch of the parser
|
||||
becomes dead code; harmless, but the pin test should assert the extracted count so a dialect
|
||||
the parser misses shows up as a drop.
|
||||
Reference in new issue
Block a user