docs(plans): record the kaikki corpus plan, measurement and review

This commit is contained in:
tiennm99 committed 2026-09-08 17:17:58 +07:00
1 parent 0a5c5b06fe
commit ff5627336d
9 files changed
+688 -1

No files matched your search

@@ -1,7 +1,7 @@
---
title: "Dictionary corpus switch"
description: "Replace the 179 MB minhqnd aggregate with the wiktionary branch of undertheseanlp/dictionary only (hongocduc and tudientv excluded) and keep data/noitu.db on CC BY-SA 4.0"
status: in-progress
status: completed
priority: P1
effort: "~1d"
tags: [dictionary, data, licensing, build]
@@ -0,0 +1,97 @@
# Measurement: first kaikki build vs the undertheseanlp corpus
Date: 2026-09-08. Phase 3 of `plan.md`. The pre-switch `data/noitu.db` (26,845 words) was
copied aside before the first kaikki build and is the "old" side of every table.
## The file measured
```
source_url https://kaikki.org/viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl
source_sha256 51ddc2fbd73cd7e2a468ccff3c31200928e7200cc58622f528ba3eb84c7a49f0
source_rows 44564
source_fetched_at 2026-09-08T10:04:08Z (kaikki's export of 2026-09-06, dump 2026-09-01)
size 62,288,004 bytes; longest line 61,552 bytes
```
Build log:
```
rejected 4: contains a digit
rejected 136: contains punctuation
rejected 8348: fewer than 2 syllables
rejected 160: no Vietnamese letters
parts of speech: noun 16219, verb 9355, adj 6698, name 5180, unknown 4652, adv 1024,
phrase 392, proverb 322, character 198, intj 175, pron 100, conj 70, prep 48, num 45,
abbrev 38, particle 18, romanization 14, prefix 13, suffix 2, det 1
accepted 34813 distinct words
generated 1664 spelling aliases (552 skipped as ambiguous or already real words)
```
## Commands
```sh
cp data/noitu.db /tmp/old.db # before the switch
make fetch-dict && make dict # or the raw curl + go run lines in README
cd server && go test ./internal/bot/ -run RealCorpus -v -count=1 # once per database at data/noitu.db
# graph numbers and diff: sqlite, see the queries below
```
```sql
SELECT COUNT(*) FROM words;
SELECT COUNT(*) FROM syllables;
SELECT COUNT(DISTINCT first) FROM words;
SELECT COUNT(*) FROM (SELECT first FROM words GROUP BY first HAVING COUNT(*) >= 2);
SELECT COUNT(*) FROM syllables WHERE out_degree = 0;
```
## Playability
| metric | old (undertheseanlp) | new (kaikki) |
|---|---|---|
| words | 26,845 | **34,813** |
| syllables | 5,709 | 6,081 |
| syllables that open a word | 4,158 | 4,487 |
| …with ≥2 continuations | 2,787 | 3,163 |
| dead-end syllables | 1,551 | 1,594 (27.2% → 26.2% of syllables) |
Bot-versus-bot, 60 games each, same run:
| | old | new |
|---|---|---|
| hard-vs-easy win rate (moves) | 98% (3.6) | 97% (3.2) |
| medium-vs-easy | 87% | 95% |
| hard-vs-medium (moves) | 57% (3.1) | 67% (3.1) |
| easy-vs-easy game length | 14.4 moves | 14.4 moves |
| hard decision p95 | 1.5 ms | 3.1 ms |
Both real-corpus tests pass on both databases. Game length is unchanged; the stronger bots
win more often on the denser graph, which is the expected direction.
## Corpus diff
| | |
|---|---|
| shared | 25,392 |
| gained | **9,421** — `công nghiệp hóa`, `bóng râm`, `bội thu`, `đồng dao`, `bỏ ngỏ`, `đắt hàng`, `ấn độ giáo`, `chạy làng`, `béo nung núc`, `nhìn muốn rụng trứng` |
| lost | **1,453** — see below |
### The 1,453 lost words
Sample of 25 (seed 7): `trịnh tuệ`, `tuân khanh`, `qua quít`, `cúc pha`, `mành mành`,
`cây bài chặt`, `trảm phong`, `kéo cánh`, `ngọt lự`, `rưng rức`, `làu nhàu`, `trong lúc`,
`khí hư`, `trần ích tắc`, `nháo nhác`, `trần cảnh`, `óc ách`, `lườm lườm`, `hỗn luân`,
`trụi lủi`, `tục tác`, `lảu nhảu`, `phát sầu`, `hãng hàng không quốc gia việt nam`, `trói ké`.
Two kinds. Roughly a fifth are historical person names (`trần ích tắc`, `trần cảnh`, `trịnh
tuệ`) whose pages have since been deleted or moved. The rest are real vocabulary, heavily
reduplicatives (`rưng rức`, `nháo nhác`, `óc ách`, `làu nhàu`, `lườm lườm`), that the 2018
scrape had and the current wiktextract export does not. Most likely those pages still exist
but in a markup the `vi` extractor does not yet parse — the raw 2026 dump has 36,200
words to kaikki's 34,813, a 1,387-word gap of the same order. **Not acted on:** the plan
does not merge sources, and the net is +7,968 words. If the reduplicatives matter, the
option is a second `--words` input reduced from the raw dump, as the previous plan's
research described.
## Verdict
Every number above today's; nothing regresses. Corpus accepted as built.
@@ -0,0 +1,118 @@
---
phase: 1
title: "Phase 1: Read the kaikki file"
status: completed
priority: P1
effort: "3h"
dependencies: []
---
# Phase 1: Read the kaikki file
## Overview
Replace the undertheseanlp reader in `build-dictionary` with one for kaikki's wiktextract
JSONL, hashing the input as it streams so the database records exactly which bytes it was
built from.
## Requirements
- Functional: `--kaikki <file>` reads one JSON object per line, keeps rows whose
`lang_code` is `vi`, and feeds `word` to the existing `accept()`.
- Functional: rows with any other `lang_code` are counted and rejected, not silently
skipped — the file is vi-only today, and a change in that would be worth seeing in the log.
- Functional: `pos` is tallied per value and printed with the reject counts. It never filters.
- Functional: the SHA-256 of the file and the row count are computed during the single
streaming pass and written to `meta` as `source_sha256` and `source_rows`, with
`source_fetched_at` taken from the file's modification time.
- Functional: `--merged` and `--sources` are removed along with `merged_list.go` and its
tests. Exactly one of `--kaikki` / `--words` must be given, same rule as today.
- Functional: the `--min-words` default moves from 20,000 to 30,000.
- Non-functional: a malformed line is an error naming the line, as today. Lines can be
long — a row carries every sense and translation — so the scanner buffer must allow
several megabytes; measure the longest line on the real file and set the cap above it
with headroom, or switch to `bufio.Reader.ReadBytes('\n')` which has no cap.
- Non-functional: `finish()`, `accept()`, aliases, `write`, `verify` unchanged.
## Architecture
```
--kaikki <file>
│ one line at a time, bytes also fed to sha256.New()
▼
json.Unmarshal → {word, lang_code, pos}
│
├─ lang_code != "vi" → rejects["not Vietnamese-language entry"]++
▼
accept(word) (NFC, lowercase, ≥2 syllables, alphabet, phonotactics)
▼
finish() (floor 30,000, aliases, atomic write, verify)
```
Decode only the three fields via a struct; `encoding/json` ignores the rest, so the 62 MB
of senses and translations cost I/O but not memory.
`sourceSpec.extra` already exists for mode-specific meta rows; the kaikki provenance uses it:
```
source_url the kaikki URL (constant kaikkiSourceURL)
source_sha256 hex of the streamed bytes
source_rows lines decoded
source_fetched_at file mtime, RFC 3339 UTC
source_license CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
```
## Related Code Files
- Create: `server/cmd/build-dictionary/kaikki_list.go`
- Create: `server/cmd/build-dictionary/kaikki_list_test.go`
- Delete: `server/cmd/build-dictionary/merged_list.go`, `merged_list_test.go`
- Modify: `server/cmd/build-dictionary/main.go` — flags, dispatch, doc comment, `--min-words`
default, `mergedProvenance` call site
- Modify: `server/cmd/build-dictionary/main_test.go` — `fixtureSource`/`defaultRows` emit
kaikki-shaped rows (`{"word": ..., "pos": ..., "lang_code": "vi"}`); the "excluded source"
row becomes a `lang_code: "en"` row
## Implementation Steps
1. Write `kaikki_list.go`: `kaikkiRow{Word, Pos, LangCode}`, `readKaikkiList(path, maxSyllables)`
returning words, rejects, a `pos` tally, and a `kaikkiProvenance` (sha256, rows, mtime).
Wrap the file in `io.TeeReader` into `sha256.New()` so hashing is free.
2. Add `rejectNotVietnamese rejectReason = "not a Vietnamese-language entry"` to `filter.go`.
3. In `main.go`: replace `merged`/`sources` with `kaikki` in `config`, flags and `run()`;
`--min-words` default 30,000; `runFromKaikkiList` logs rejects, the POS tally
(sorted, one line) and the accepted count, then calls `finish()` with the provenance spec.
4. Delete the merged reader and its tests. Rewrite `main_test.go` fixtures to kaikki rows.
5. Tests in `kaikki_list_test.go`: a `vi` row kept; an `en` row rejected and counted;
`Hà Nội` kept as `hà nội`; a word listed twice with different `pos` kept once; a malformed
line names its number; `meta` carries a 64-hex `source_sha256` equal to `sha256sum` of the
fixture file, the right `source_rows`, and no `source_commit`/`sources_*` keys.
6. Run against the real file (`scratchpad/kk-vi-edition.jsonl` from this session, or a fresh
download) and record the counts in phase 3.
## Success Criteria
- [x] `go test ./cmd/build-dictionary/` passes; `merged_list*.go` no longer exist.
- [x] Real file builds >30,000 words (measured 34,813) and logs a POS tally.
- [x] `build-dictionary --help` lists `--kaikki`, `--words`, `--out`, `--max-syllables`,
`--min-words` and nothing else.
- [x] `meta.source_sha256` of a build equals `sha256sum` of the input file.
- [x] The fixture path (`--words ... --min-words 150`) is byte-for-byte unaffected in
behaviour: same 205 words from `testdata/fixture-words.txt`.
## Risk Assessment
**A row exceeds the scanner buffer.** Signal: `bufio.Scanner: token too long` on the real
file. Response: pre-decided — use `bufio.Reader` with no line cap rather than guessing a
buffer size; measure the longest line once and note it in the phase result.
**kaikki changes the row shape.** `word` and `lang_code` are wiktextract's stable core fields
and unlikely to move. Signal: zero accepted words → the 30,000 floor fails the build with a
clear message. Response: read the new shape, adjust the struct; the failure mode is loud by
design.
**The file arrives partially.** `curl -f` catches HTTP errors but not a truncated body, and
a body cut exactly on a line boundary parses cleanly. Signal: a malformed final line (error
names it) or a word count under the 30,000 floor (13.8% headroom on 34,813). Response: the
fetch writes to a `.part` name and renames only on success (review finding, Phase 2), so an
interrupted download is never left for the next `make dict`; the floor covers the rest.
@@ -0,0 +1,105 @@
---
phase: 2
title: "Phase 2: Fetch the newest file"
status: completed
priority: P1
effort: "2h"
dependencies: [1]
---
# Phase 2: Fetch the newest file
## Overview
Point the Makefile and the Docker dictionary stage at kaikki's rolling URL with no checksum,
keep the two build files in agreement, and make sure a failed download cannot pass as a
success anywhere.
## Requirements
- Functional: `make fetch-dict` downloads the kaikki file to `data/kaikki-viwiktionary-vi.jsonl`
with `curl -fL` and fails on any HTTP error; `make dict` builds from it with `--kaikki`.
- Functional: `verify-dict` and `DICT_SHA256` are removed from the Makefile and Dockerfile;
nothing in the repository claims a checksum for this file.
- Functional: the Docker `dict` stage downloads the same URL and runs the same command;
`FIXTURE_DICT=1` keeps working untouched.
- Functional: `web/tests/dictionary-source.test.js` asserts the Makefile and Dockerfile URLs
agree, that the URL is the kaikki viwiktionary `Tiếng Việt` file, and that the builder's
`kaikkiSourceURL` constant is the same URL. The three checksum assertions go.
- Functional: the CI leak guard greps for `kaikki.*\.jsonl$`; `.gitignore` ignores the new
file name.
- Non-functional: no resume flag (`-C -`) — a rolling file can change between attempts, and a
resumed download would splice two versions together.
## Architecture
```make
DICT_URL := https://kaikki.org/viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl
DICT_SRC := data/kaikki-viwiktionary-vi.jsonl
DICT_OUT := data/noitu.db
fetch-dict:
@mkdir -p data
curl -fL -o $(DICT_SRC).part $(DICT_URL) && mv $(DICT_SRC).part $(DICT_SRC)
dict: $(DICT_SRC)
cd server && go run ./cmd/build-dictionary --kaikki ../$(DICT_SRC) --out ../$(DICT_OUT)
```
The Dockerfile mirrors it with `ARG DICT_URL` only. The URL must stay percent-encoded in
both files: Make and `sh` would otherwise split on the space in `Tiếng Việt`.
Correctness now rests on three checks instead of a checksum: `curl -f` plus an atomic
`.part` rename (HTTP errors and interrupted downloads), the JSON decoder (shape, and
truncation that lands mid-line), the 30,000 floor (content, including a truncation that
lands on a line boundary).
## Related Code Files
- Modify: `Makefile` — variables, `help` text, `fetch-dict`, remove `verify-dict` and its
`.PHONY` entry, `dict` flags
- Modify: `Dockerfile` — `ARG DICT_URL`, remove `ARG DICT_SHA256` and the `sha256sum` line,
download file name, `--kaikki`; stage comment (no longer "pinned … checked by digest")
- Modify: `web/tests/dictionary-source.test.js` — drop checksum tests, keep URL agreement,
add the kaikki-path and builder-constant assertions
- Modify: `.github/workflows/ci.yml` — leak-guard pattern and comment
- Modify: `.gitignore` — `data/kaikki-viwiktionary-vi.jsonl` replaces the undertheseanlp entry
- Modify: `README.md` — Make targets table (`verify-dict` row gone), the manual `curl` and
`build-dictionary` lines
## Implementation Steps
1. Update the Makefile variables and targets; delete `verify-dict`.
2. Mirror in the Dockerfile; rewrite the stage comment to say the file is fetched fresh and
identified by the SHA-256 the builder records.
3. Rewrite the pin test as described. Run `npx vitest run tests/dictionary-source.test.js`.
4. Update the CI leak guard, `.gitignore` and the README lines.
5. Run `make fetch-dict && make dict` (or the raw commands on Windows) from a clean `data/`.
6. Build the image with and without `FIXTURE_DICT=1`; run the CI file checks against the
real one and confirm no `*.jsonl` is inside.
7. Simulate a bad download once: point `DICT_URL` at a 404 and confirm `curl -f` fails the
target; feed the builder a truncated copy of the file and confirm the build fails.
## Success Criteria
- [x] `make fetch-dict` fetches ~62 MB and `make dict` builds >30,000 words.
- [x] `grep -rn DICT_SHA256 --exclude-dir=plans .` finds nothing; `make verify-dict` is not a
target.
- [x] `web/tests/dictionary-source.test.js` passes and fails if either URL is edited alone.
- [x] Both image variants build; the real one passes the CI file checks; no `.jsonl` inside.
- [x] A 404 URL and a truncated file each fail the build with a readable message.
## Risk Assessment
**kaikki.org is a single volunteer-run host.** An outage breaks `make fetch-dict` and the
release image build until it returns. Signal: `curl: (22)` on fetch. Response: wait, or
build from a previously fetched local copy (`make dict` needs only the file); if outages
recur, keep a copy as a release asset and point `DICT_URL` at it — the open question in
`plan.md`.
**Two builds ship different words.** By design. Signal: none needed. Response: `meta`
identifies the bytes; the audit note in phase 3 records what this first build contained.
**The space in the URL.** Signal: `curl` fetching `https://kaikki.org/viwiktionary/Ti%E1%BA%BFng`
and failing. Response: the URL stays percent-encoded in every file, and the pin test checks
that the string contains `Ti%E1%BA%BFng%20Vi%E1%BB%87t`.
@@ -0,0 +1,81 @@
---
phase: 3
title: "Phase 3: Measure the corpus"
status: completed
priority: P1
effort: "1h"
dependencies: [1, 2]
---
# Phase 3: Measure the corpus
## Overview
Record what the first kaikki build contains against today's database: the graph numbers, the
words gained and lost, the POS mix, and bot-versus-bot behaviour. Measurement, not a gate —
no rule removes words any more, so there is nothing to audit by hand.
## Requirements
- Functional: a note in this plan directory with the numbers below and the commands that
produced them, plus the `source_sha256` of the file measured so the note is tied to bytes.
- Functional: a diff against the current `data/noitu.db` (copied aside before Phase 2
overwrites it): shared, gained, lost, with 25 random samples of each.
- Functional: playability on both databases: words, syllables, openers, openers with ≥2
continuations, dead ends, and the `internal/bot` real-corpus test output.
- Non-functional: every number reproducible from a written command.
## Architecture
Pre-measured this session on kaikki's 2026-09-06 file (sha256 `51ddc2fbd73cd7e2…`); the
phase re-measures on whatever the build fetched and explains any delta:
| | today | kaikki-vi (2026-09-06) |
|---|---|---|
| words | 26,845 | 34,813 |
| syllables | 5,709 | 6,081 |
| openers ≥2 continuations | 2,787 | 3,163 |
| dead-end syllables | 1,551 | 1,594 |
| shared / gained / lost | | 25,392 / 9,421 / 1,453 |
POS of the 41,507 distinct words: noun 16,219 · verb 9,355 · adj 6,698 · name 5,180 ·
unknown 4,652 · adv 1,024 · phrase 392 · proverb 322 · character 198. `name` and `unknown`
are kept; the log prints the tally so a future decision has the numbers.
## Related Code Files
- Create: `plans/260908-1653-kaikki-viwiktionary-corpus/measurement.md`
- Read only: the pre-switch `data/noitu.db` copy, the new build, `server/internal/bot`
real-corpus tests
## Implementation Steps
1. Before Phase 2 overwrites it, copy today's `data/noitu.db` to a scratch path.
2. After the first `make dict`, record `meta.source_sha256`, `source_rows`,
`source_fetched_at` and the build log's reject and POS lines.
3. Compute the diff and the five graph numbers for both databases (sqlite, one script; keep
the script text in the note).
4. Run `go test ./internal/bot/ -run RealCorpus -v` against both databases (swap the file
under `data/noitu.db`, restore afterwards) and record the game-length and win-rate lines.
5. Look at the 1,453 lost words: they are words in the 2018 scrape that Wiktionary has since
deleted or renamed. Sample 25, note what they look like (expected: misspellings, moved
pages, deleted junk). No action unless the sample is mostly real vocabulary.
## Success Criteria
- [x] `measurement.md` exists with the tables, samples, bot lines, commands and the source
SHA-256.
- [x] Words > 30,000; syllables, openers and ≥2-continuation counts all above today's.
- [x] Bot real-corpus tests pass on the new database; easy-vs-easy game length is not
shorter than today's 12.9 moves.
- [x] The lost-word sample is characterised in one paragraph.
## Risk Assessment
**The fetched file differs from the one measured today.** Certain over time. Signal: counts
off from the table above. Response: record the new numbers and the SHA-256; the deltas are
the point of the note, not a failure.
**Dead ends rise slightly (1,551 → 1,594).** More words bring more rare final syllables.
Signal: already known. Response: the ratio of dead ends to syllables falls (27.2% → 26.2%),
so the graph is denser, not sparser; record it and move on.
@@ -0,0 +1,94 @@
---
phase: 4
title: "Phase 4: Relicense to 4.0"
status: completed
priority: P1
effort: "1h"
dependencies: [3]
---
# Phase 4: Relicense to 4.0
## Overview
Current Wiktionary text is CC BY-SA 4.0, so the derived database is too. Restore the 4.0
legal code, rewrite the attribution chain for a source with no commit to cite, and update
every place that says 3.0 — in the same commit as the first kaikki-built database.
## Requirements
- Functional: `data/LICENSE` is the CC BY-SA 4.0 legal code, byte-identical to the file this
repository shipped before commit `39ef457` (`git show e19b083:data/LICENSE`).
- Functional: `NOTICE` section 2 says 4.0 and names Wiktionary tiếng Việt as the original
work and kaikki.org/wiktextract as the extraction; section 1 byte-identical.
- Functional: `data/ATTRIBUTION.md` records the source URL, that the file is refreshed
weekly and unpinned, that the exact bytes of any build are identified by
`meta.source_sha256`, the three attribution links, and the modifications list.
- Functional: the in-game footer links to the 4.0 deed and says 4.0; the e2e assertion
matches.
- Non-functional: no overclaiming about dates — the attribution names the dump date only as
"the Wikimedia dump kaikki extracted at fetch time; see `source_fetched_at`", not a fixed
date that goes stale.
## Architecture
Attribution chain, all three named:
```
Wiktionary tiếng Việt contributors authors, CC BY-SA 4.0
→ wiktextract / kaikki.org (Tatu Ylonen) extraction; cite LREC 2022 as kaikki requests
→ this project filter + index, data/noitu.db, CC BY-SA 4.0
```
Modifications list (renumbered): 1 language selection (`lang_code = vi`), 2 length filter,
3 content filter, 4 normalization, 5 spelling aliases, 6 added columns and tables,
7 deduplication, 8 dropped fields (senses, translations, POS, categories — everything but
the word form).
## Related Code Files
- Modify: `data/LICENSE` — restore 4.0 text from git history
- Modify: `NOTICE` — section 2 heading, source lines, share-alike sentence
- Modify: `data/ATTRIBUTION.md` — full rewrite of Source, Modifications 1 and 8, reproduce
block (no checksum step)
- Modify: `README.md` — licence table row, dictionary-source paragraph
- Modify: `Dockerfile`, `docs/deployment.md`, `.github/workflows/ci.yml` — 3.0 → 4.0 in comments
- Modify: `server/cmd/build-dictionary/main.go` doc comment, `server/cmd/noitu-server/main.go`
and `server/internal/dictionary/store.go` comments, `store_test.go` seed strings
- Modify: `web/src/lib/components/AttributionFooter.svelte` (deed URL, comment),
`web/src/lib/i18n/vi.js` (`attributionLicense`), `web/e2e/bot-game.spec.js` (regex)
- Verify unchanged: `.github/workflows/ci.yml` required-files check
## Implementation Steps
1. `git show e19b083:data/LICENSE > data/LICENSE`; confirm the first lines read
"Attribution-ShareAlike 4.0 International".
2. Rewrite `NOTICE` section 2 source lines; leave the share-alike paragraph's meaning, change
the version.
3. Rewrite `data/ATTRIBUTION.md` as described. State plainly: "The file is not pinned. Each
build fetches kaikki's current export; the SHA-256 and row count of the file a given
database was built from are recorded in its `meta` table."
4. `grep -rn "3\.0\|by-sa/3" --exclude-dir=plans .` and fix every hit; then the same for
`undertheseanlp`, `DICT_SHA256`, `--merged`, `--sources`.
5. Regenerate `data/noitu.db`; confirm `meta.source_license` says 4.0 and matches the files.
6. Run the full verification set from the previous plan: Go vet/test, web check/test, both
image builds with the CI file checks.
## Success Criteria
- [x] `data/LICENSE` is the 4.0 text; `git diff e19b083 -- data/LICENSE` is empty.
- [x] No file outside `plans/` says CC BY-SA 3.0 or names undertheseanlp.
- [x] `NOTICE` section 1 has no diff.
- [x] `ATTRIBUTION.md` names all three links and states the unpinned-fetch policy.
- [x] `meta.source_license` in the built database says 4.0.
- [x] All verification suites green; the commit contains both the licence files and the
builder change so no intermediate tree mixes 3.0 text with 4.0 data.
## Risk Assessment
**A stale 3.0 or undertheseanlp claim survives.** Signal: the grep in step 4 finds one after
the phase is called done. Response: grep is a step for that reason; run it last.
**The footer wording.** `giấy phép CC BY-SA 3.0` → `4.0` is one string, but the e2e regex
and the deed URL must change with it. Signal: e2e footer test fails in CI. Response: all
three edits are listed above; do them together.
@@ -0,0 +1,142 @@
---
title: "Kaikki viwiktionary corpus"
description: "Replace the 2018 undertheseanlp wiktionary rows with kaikki.org's current extraction of Wiktionary tiếng Việt, fetched unpinned at build time, and move the data licence to CC BY-SA 4.0"
status: completed
priority: P1
effort: "~1d"
tags: [dictionary, data, licensing, build]
created: 2026-09-08
blockedBy: []
blocks: []
---
# Kaikki viwiktionary corpus
## Overview
`data/noitu.db` is built from the `wiktionary` rows of `undertheseanlp/dictionary`: a
2018-12-10 scrape of vi.wiktionary.org, 26,845 words after filtering, pinned by commit and
SHA-256 (plan `260908-1525-dictionary-corpus-switch`, completed today). It is the same
website eight years staler than it needs to be.
This plan switches the source to **kaikki.org's extraction of Wiktionary tiếng Việt** — the
`Tiếng Việt` language file of the `viwiktionary` edition, produced by wiktextract from the
monthly Wikimedia dump (currently 2026-09-01) and refreshed roughly weekly. Same authors,
same website, current text. Measured through the real `build-dictionary` filter this
session: **34,813 words**, a near-superset of today's corpus (25,392 shared, 9,421 gained,
1,453 lost) and 96% of the raw dump (36,200).
Two owner decisions shape the plan:
- **viwiktionary only.** kaikki's English-Wiktionary Vietnamese file (21,244 words, marked
deprecated by kaikki) is not used. The game is Vietnamese; one edition, one attribution.
- **No pin — fetch the newest file every build.** kaikki publishes rolling files with no
archived snapshots, so a checksum pin would break weekly. The owner chose freshness over
reproducibility. The build therefore has to *fail loudly* on a bad download (HTTP error,
malformed JSON, too few words) and *record what it actually got* (SHA-256 and row count in
`meta`) so any shipped database can still be traced to exact bytes.
Consequence: the current text of Wiktionary is CC BY-SA **4.0** (Wikimedia moved from 3.0
on 2023-06-29), so `data/noitu.db` returns to 4.0. The 3.0 chosen this afternoon only ever
applied to the 2018 snapshot.
Evidence: `plans/reports/research-260908-1529-viwiktionary-dump-measured.md` (dump and
kaikki-en measurements) plus this session's measurement of the kaikki-vi file, recorded in
[phase 3](./phase-03-measure-the-corpus.md).
## What changes, in numbers
| | today | after |
|---|---|---|
| words | 26,845 | **34,813** (30,231 if `name` POS were dropped — it is not) |
| syllables | 5,709 | 6,081 |
| syllables with ≥2 continuations | 2,787 | 3,163 |
| dead-end syllables | 1,551 | 1,594 |
| source | 4.8 MB JSONL, 2018, pinned | **62 MB JSONL, current, unpinned** |
| data licence | CC BY-SA 3.0 | **CC BY-SA 4.0** |
| `--min-words` floor | 20,000 | **30,000** |
| build reproducible | yes | **no** — traceable via `meta.source_sha256` |
## The source
```
URL https://kaikki.org/viwiktionary/Tiếng Việt/kaikki.org-dictionary-TiếngViệt.jsonl
(percent-encoded: /viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl)
Size 62,288,004 bytes on 2026-09-06 — 44,564 rows, 41,507 distinct words, all lang_code "vi"
Row {"word": "trở thành", "pos": "verb", "lang_code": "vi", "lang": "Tiếng Việt", "senses": [...], ...}
```
Only `word` and `lang_code` are read. `pos` is counted for the build log but never filters:
the owner's earlier decision that capitalization does not remove words carries over to the
`name` POS (5,138 name-only words). 5,475 words carry uppercase and 229 lowercase forms
occur twice; `accept()` lowercases and dedupes as it always has.
## Decisions taken
- **One corpus input mode.** `--merged`/`--sources` (undertheseanlp) is replaced by
`--kaikki`, not kept beside it. Two readers for one database is dead weight, as the SQLite
path was. The fixture path `--words` stays.
- **The floor moves to 30,000.** 34,813 measured; a truncated or reshaped download must
fail, and 30,000 is the highest round number with a comfortable margin.
- **Provenance is measured, not asserted.** The reader hashes the file as it streams it and
writes `source_sha256`, `source_rows` and `source_fetched_at` into `meta`. There is no
commit to record, so this is what "which bytes" means from now on.
- **Attribution names three links.** Wiktionary tiếng Việt contributors (authors, CC BY-SA
4.0), wiktextract/kaikki.org by Tatu Ylonen (extraction; kaikki asks for the LREC 2022
citation), and this project. No GitHub intermediary any more.
- **`verify-dict` goes.** There is nothing to verify against. The
Makefile/Dockerfile agreement test keeps guarding the URL and drops the checksum clauses.
## Phases
| # | Phase | Status |
|---|-------|--------|
| 1 | [Read the kaikki file](./phase-01-start.md) | Pending |
| 2 | [Fetch the newest file](./phase-02-fetch-the-newest-file.md) | Pending |
| 3 | [Measure the corpus](./phase-03-measure-the-corpus.md) | Pending |
| 4 | [Relicense to 4.0](./phase-04-relicense-to-4-0.md) | Pending |
Phase 4 lands in the same commit as the first kaikki-built `noitu.db`: a 4.0-derived
database in a tree whose `NOTICE` says 3.0 is the one ordering mistake with a legal
consequence. Phase 3 is measurement, not a gate — there is no drop rule left to audit.
## Success criteria
- [x] `make fetch-dict && make dict` downloads kaikki's current file and builds
`data/noitu.db` with **more than 30,000 words**; a truncated or non-JSON download fails
the build with a message naming the cause.
- [x] `meta` records `source_url`, `source_sha256`, `source_rows`, `source_fetched_at` and
`source_license = CC BY-SA 4.0`; no `source_commit`, `sources_kept` or
`sources_excluded` remain.
- [x] `data/LICENSE` is the CC BY-SA 4.0 legal code; `NOTICE`, `data/ATTRIBUTION.md`, README,
Dockerfile, deployment doc, the builder's licence string and the in-game footer all say
4.0 and credit Wiktionary tiếng Việt contributors and wiktextract/kaikki.org.
- [x] Playability measured and recorded against today's database, including bot game
lengths; nothing regresses below today's numbers.
- [x] `go vet ./... && go test ./...`, `npm run check && npm test`, the Docker image with the
CI file checks and `FIXTURE_DICT=1` are green. e2e: not runnable on this machine (Playwright Chromium download fails); the one changed assertion was updated by hand, CI verifies.
- [x] No file outside `plans/` names undertheseanlp, `--merged`, `--sources` or `DICT_SHA256`.
## Review
Code review (2026-09-08) returned DONE_WITH_CONCERNS; every finding was applied: `fetch-dict`
downloads to a `.part` name and renames on success so an interrupted fetch never feeds the
next `make dict`; a failed `Stat` is an error rather than a zero `source_fetched_at`;
`builder_version` bumped to 3 for the changed meta contract; `.gitignore` covers
`data/*.jsonl`; `ATTRIBUTION.md` states kaikki's redistribution terms; bare JSON literals
are malformed rather than "foreign-language"; the non-EOF read-error line number is exact;
the CI leak guard matches any `.jsonl`; byte-exact reader tests cover a missing final
newline, an HTML error page, a mid-line cut and a bare literal. The reviewer's residual
note stands: with no pin, the build trusts kaikki.org over TLS, and the recorded SHA-256 is
the compensating control; blast radius is game content only.
## Open questions
- kaikki's weekly refresh means two builds a week apart can differ. Accepted by the owner;
`source_sha256` in `meta` is the answer to "which words did this image ship". Should a
release ever need to be rebuilt bit-for-bit, the fallback is to keep the fetched file as a
release asset — a Makefile variable change, not a redesign.
- The `unknown` POS bucket (4,616 words such as `tình báo`, `trọc phú`, `hội quán`) is real
vocabulary the vi extractor could not label. Kept, like everything else.
<!-- slug: kaikki-viwiktionary-corpus -->
@@ -0,0 +1,25 @@
---
title: Planned the kaikki viwiktionary corpus switch
date: 2026-09-08
summary: "Plan to replace the 2018 undertheseanlp wiktionary rows with kaikki.org's current Wiktionary tiếng Việt extraction, fetched unpinned; 26,845 -> 34,813 words, data licence to CC BY-SA 4.0"
---
# Planned the kaikki viwiktionary corpus switch
## What happened
- Measured kaikki's `viwiktionary` edition Vietnamese file through the real filter: 34,813 words (25,392 shared with today, 9,421 gained, 1,453 lost), a 96% subset of the raw dump. kaikki's enwiktionary Vietnamese file is 21,244 and deprecated.
- Owner chose the viwiktionary file only, and to fetch kaikki's newest file each build with no checksum pin.
- Plan written: `plans/260908-1653-kaikki-viwiktionary-corpus/` — four phases: kaikki reader with streamed SHA-256 into `meta`, unpinned fetch with loud failure modes, measurement note, relicense to CC BY-SA 4.0.
- Previous plan `260908-1525-dictionary-corpus-switch` marked completed.
## Decision
- Freshness over reproducibility: two builds a week apart may differ; `meta.source_sha256` identifies the bytes. Fallback if bit-for-bit rebuilds are ever needed: keep the fetched file as a release asset.
- The 3.0 licence chosen earlier today only applied to the 2018 snapshot; current Wiktionary text is 4.0, so the database returns to 4.0.
## Next steps
- Validate or cook the plan.
> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions.
@@ -0,0 +1,25 @@
---
title: "Switched the corpus to kaikki's Wiktionary tiếng Việt export"
date: 2026-09-08
summary: "Replaced the pinned 2018 undertheseanlp rows with kaikki.org's current wiktextract export fetched unpinned; 26,845 -> 34,813 words; provenance by streamed SHA-256; data licence back to CC BY-SA 4.0"
---
# Switched the corpus to kaikki's Wiktionary tiếng Việt export
## What happened
- `build-dictionary` reads kaikki's JSONL with `--kaikki`, hashing the stream and writing `source_sha256`, `source_rows`, `source_fetched_at` into `meta`; `--merged`/`--sources` and the undertheseanlp reader are gone; `--min-words` floor 30,000; `builder_version` 3.
- Makefile and Dockerfile fetch the rolling kaikki URL with `curl -fL`, no checksum, atomic `.part` rename; `verify-dict` and `DICT_SHA256` removed; the web test guards Makefile == Dockerfile == builder constant and the viwiktionary path.
- Data licence back to CC BY-SA 4.0 (current Wiktionary text); `data/LICENSE` restored from git history; NOTICE/ATTRIBUTION/README/footer credit Wiktionary tiếng Việt contributors and wiktextract/kaikki.org with the LREC citation.
- Measured: 34,813 words (+9,421 / -1,453 vs before), all graph metrics up, bot game length unchanged. Lost words are mostly reduplicatives the vi extractor misses and deleted person-name pages; recorded in `measurement.md`.
- Verified: Go vet/test, web check/test (174), both Docker variants with CI file checks, 404 and truncated-download failure modes. e2e not runnable locally (Chromium download fails), CI covers it.
## Decision
- Owner: viwiktionary edition only; fetch newest, no pin. Two builds a week apart may differ; the SHA-256 in `meta` identifies the bytes. Fallback if reproducibility is ever needed: keep a fetched copy as a release asset.
## Next steps
- Commit (not yet committed); watch CI's e2e job.
> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions.