diff --git a/plans/260908-1525-dictionary-corpus-switch/plan.md b/plans/260908-1525-dictionary-corpus-switch/plan.md index cb8d16e..f234306 100644 --- a/plans/260908-1525-dictionary-corpus-switch/plan.md +++ b/plans/260908-1525-dictionary-corpus-switch/plan.md @@ -1,7 +1,7 @@ --- title: "Dictionary corpus switch" description: "Replace the 179 MB minhqnd aggregate with the wiktionary branch of undertheseanlp/dictionary only (hongocduc and tudientv excluded) and keep data/noitu.db on CC BY-SA 4.0" -status: in-progress +status: completed priority: P1 effort: "~1d" tags: [dictionary, data, licensing, build] diff --git a/plans/260908-1653-kaikki-viwiktionary-corpus/measurement.md b/plans/260908-1653-kaikki-viwiktionary-corpus/measurement.md new file mode 100644 index 0000000..738c4a7 --- /dev/null +++ b/plans/260908-1653-kaikki-viwiktionary-corpus/measurement.md @@ -0,0 +1,97 @@ +# Measurement: first kaikki build vs the undertheseanlp corpus + +Date: 2026-09-08. Phase 3 of `plan.md`. The pre-switch `data/noitu.db` (26,845 words) was +copied aside before the first kaikki build and is the "old" side of every table. + +## The file measured + +``` +source_url https://kaikki.org/viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl +source_sha256 51ddc2fbd73cd7e2a468ccff3c31200928e7200cc58622f528ba3eb84c7a49f0 +source_rows 44564 +source_fetched_at 2026-09-08T10:04:08Z (kaikki's export of 2026-09-06, dump 2026-09-01) +size 62,288,004 bytes; longest line 61,552 bytes +``` + +Build log: + +``` +rejected 4: contains a digit +rejected 136: contains punctuation +rejected 8348: fewer than 2 syllables +rejected 160: no Vietnamese letters +parts of speech: noun 16219, verb 9355, adj 6698, name 5180, unknown 4652, adv 1024, + phrase 392, proverb 322, character 198, intj 175, pron 100, conj 70, prep 48, num 45, + abbrev 38, particle 18, romanization 14, prefix 13, suffix 2, det 1 +accepted 34813 distinct words +generated 1664 spelling aliases (552 skipped as ambiguous or already real words) +``` + +## Commands + +```sh +cp data/noitu.db /tmp/old.db # before the switch +make fetch-dict && make dict # or the raw curl + go run lines in README +cd server && go test ./internal/bot/ -run RealCorpus -v -count=1 # once per database at data/noitu.db +# graph numbers and diff: sqlite, see the queries below +``` + +```sql +SELECT COUNT(*) FROM words; +SELECT COUNT(*) FROM syllables; +SELECT COUNT(DISTINCT first) FROM words; +SELECT COUNT(*) FROM (SELECT first FROM words GROUP BY first HAVING COUNT(*) >= 2); +SELECT COUNT(*) FROM syllables WHERE out_degree = 0; +``` + +## Playability + +| metric | old (undertheseanlp) | new (kaikki) | +|---|---|---| +| words | 26,845 | **34,813** | +| syllables | 5,709 | 6,081 | +| syllables that open a word | 4,158 | 4,487 | +| …with ≥2 continuations | 2,787 | 3,163 | +| dead-end syllables | 1,551 | 1,594 (27.2% → 26.2% of syllables) | + +Bot-versus-bot, 60 games each, same run: + +| | old | new | +|---|---|---| +| hard-vs-easy win rate (moves) | 98% (3.6) | 97% (3.2) | +| medium-vs-easy | 87% | 95% | +| hard-vs-medium (moves) | 57% (3.1) | 67% (3.1) | +| easy-vs-easy game length | 14.4 moves | 14.4 moves | +| hard decision p95 | 1.5 ms | 3.1 ms | + +Both real-corpus tests pass on both databases. Game length is unchanged; the stronger bots +win more often on the denser graph, which is the expected direction. + +## Corpus diff + +| | | +|---|---| +| shared | 25,392 | +| gained | **9,421** — `công nghiệp hóa`, `bóng râm`, `bội thu`, `đồng dao`, `bỏ ngỏ`, `đắt hàng`, `ấn độ giáo`, `chạy làng`, `béo nung núc`, `nhìn muốn rụng trứng` | +| lost | **1,453** — see below | + +### The 1,453 lost words + +Sample of 25 (seed 7): `trịnh tuệ`, `tuân khanh`, `qua quít`, `cúc pha`, `mành mành`, +`cây bài chặt`, `trảm phong`, `kéo cánh`, `ngọt lự`, `rưng rức`, `làu nhàu`, `trong lúc`, +`khí hư`, `trần ích tắc`, `nháo nhác`, `trần cảnh`, `óc ách`, `lườm lườm`, `hỗn luân`, +`trụi lủi`, `tục tác`, `lảu nhảu`, `phát sầu`, `hãng hàng không quốc gia việt nam`, `trói ké`. + +Two kinds. Roughly a fifth are historical person names (`trần ích tắc`, `trần cảnh`, `trịnh +tuệ`) whose pages have since been deleted or moved. The rest are real vocabulary, heavily +reduplicatives (`rưng rức`, `nháo nhác`, `óc ách`, `làu nhàu`, `lườm lườm`), that the 2018 +scrape had and the current wiktextract export does not. Most likely those pages still exist +but in a markup the `vi` extractor does not yet parse — the raw 2026 dump has 36,200 +words to kaikki's 34,813, a 1,387-word gap of the same order. **Not acted on:** the plan +does not merge sources, and the net is +7,968 words. If the reduplicatives matter, the +option is a second `--words` input reduced from the raw dump, as the previous plan's +research described. + +## Verdict + +Every number above today's; nothing regresses. Corpus accepted as built. diff --git a/plans/260908-1653-kaikki-viwiktionary-corpus/phase-01-start.md b/plans/260908-1653-kaikki-viwiktionary-corpus/phase-01-start.md new file mode 100644 index 0000000..983ea53 --- /dev/null +++ b/plans/260908-1653-kaikki-viwiktionary-corpus/phase-01-start.md @@ -0,0 +1,118 @@ +--- +phase: 1 +title: "Phase 1: Read the kaikki file" +status: completed +priority: P1 +effort: "3h" +dependencies: [] +--- + +# Phase 1: Read the kaikki file + +## Overview + +Replace the undertheseanlp reader in `build-dictionary` with one for kaikki's wiktextract +JSONL, hashing the input as it streams so the database records exactly which bytes it was +built from. + +## Requirements + +- Functional: `--kaikki ` reads one JSON object per line, keeps rows whose + `lang_code` is `vi`, and feeds `word` to the existing `accept()`. +- Functional: rows with any other `lang_code` are counted and rejected, not silently + skipped — the file is vi-only today, and a change in that would be worth seeing in the log. +- Functional: `pos` is tallied per value and printed with the reject counts. It never filters. +- Functional: the SHA-256 of the file and the row count are computed during the single + streaming pass and written to `meta` as `source_sha256` and `source_rows`, with + `source_fetched_at` taken from the file's modification time. +- Functional: `--merged` and `--sources` are removed along with `merged_list.go` and its + tests. Exactly one of `--kaikki` / `--words` must be given, same rule as today. +- Functional: the `--min-words` default moves from 20,000 to 30,000. +- Non-functional: a malformed line is an error naming the line, as today. Lines can be + long — a row carries every sense and translation — so the scanner buffer must allow + several megabytes; measure the longest line on the real file and set the cap above it + with headroom, or switch to `bufio.Reader.ReadBytes('\n')` which has no cap. +- Non-functional: `finish()`, `accept()`, aliases, `write`, `verify` unchanged. + +## Architecture + +``` +--kaikki + │ one line at a time, bytes also fed to sha256.New() + ▼ + json.Unmarshal → {word, lang_code, pos} + │ + ├─ lang_code != "vi" → rejects["not Vietnamese-language entry"]++ + ▼ + accept(word) (NFC, lowercase, ≥2 syllables, alphabet, phonotactics) + ▼ + finish() (floor 30,000, aliases, atomic write, verify) +``` + +Decode only the three fields via a struct; `encoding/json` ignores the rest, so the 62 MB +of senses and translations cost I/O but not memory. + +`sourceSpec.extra` already exists for mode-specific meta rows; the kaikki provenance uses it: + +``` +source_url the kaikki URL (constant kaikkiSourceURL) +source_sha256 hex of the streamed bytes +source_rows lines decoded +source_fetched_at file mtime, RFC 3339 UTC +source_license CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) +``` + +## Related Code Files + +- Create: `server/cmd/build-dictionary/kaikki_list.go` +- Create: `server/cmd/build-dictionary/kaikki_list_test.go` +- Delete: `server/cmd/build-dictionary/merged_list.go`, `merged_list_test.go` +- Modify: `server/cmd/build-dictionary/main.go` — flags, dispatch, doc comment, `--min-words` + default, `mergedProvenance` call site +- Modify: `server/cmd/build-dictionary/main_test.go` — `fixtureSource`/`defaultRows` emit + kaikki-shaped rows (`{"word": ..., "pos": ..., "lang_code": "vi"}`); the "excluded source" + row becomes a `lang_code: "en"` row + +## Implementation Steps + +1. Write `kaikki_list.go`: `kaikkiRow{Word, Pos, LangCode}`, `readKaikkiList(path, maxSyllables)` + returning words, rejects, a `pos` tally, and a `kaikkiProvenance` (sha256, rows, mtime). + Wrap the file in `io.TeeReader` into `sha256.New()` so hashing is free. +2. Add `rejectNotVietnamese rejectReason = "not a Vietnamese-language entry"` to `filter.go`. +3. In `main.go`: replace `merged`/`sources` with `kaikki` in `config`, flags and `run()`; + `--min-words` default 30,000; `runFromKaikkiList` logs rejects, the POS tally + (sorted, one line) and the accepted count, then calls `finish()` with the provenance spec. +4. Delete the merged reader and its tests. Rewrite `main_test.go` fixtures to kaikki rows. +5. Tests in `kaikki_list_test.go`: a `vi` row kept; an `en` row rejected and counted; + `Hà Nội` kept as `hà nội`; a word listed twice with different `pos` kept once; a malformed + line names its number; `meta` carries a 64-hex `source_sha256` equal to `sha256sum` of the + fixture file, the right `source_rows`, and no `source_commit`/`sources_*` keys. +6. Run against the real file (`scratchpad/kk-vi-edition.jsonl` from this session, or a fresh + download) and record the counts in phase 3. + +## Success Criteria + +- [x] `go test ./cmd/build-dictionary/` passes; `merged_list*.go` no longer exist. +- [x] Real file builds >30,000 words (measured 34,813) and logs a POS tally. +- [x] `build-dictionary --help` lists `--kaikki`, `--words`, `--out`, `--max-syllables`, + `--min-words` and nothing else. +- [x] `meta.source_sha256` of a build equals `sha256sum` of the input file. +- [x] The fixture path (`--words ... --min-words 150`) is byte-for-byte unaffected in + behaviour: same 205 words from `testdata/fixture-words.txt`. + +## Risk Assessment + +**A row exceeds the scanner buffer.** Signal: `bufio.Scanner: token too long` on the real +file. Response: pre-decided — use `bufio.Reader` with no line cap rather than guessing a +buffer size; measure the longest line once and note it in the phase result. + +**kaikki changes the row shape.** `word` and `lang_code` are wiktextract's stable core fields +and unlikely to move. Signal: zero accepted words → the 30,000 floor fails the build with a +clear message. Response: read the new shape, adjust the struct; the failure mode is loud by +design. + +**The file arrives partially.** `curl -f` catches HTTP errors but not a truncated body, and +a body cut exactly on a line boundary parses cleanly. Signal: a malformed final line (error +names it) or a word count under the 30,000 floor (13.8% headroom on 34,813). Response: the +fetch writes to a `.part` name and renames only on success (review finding, Phase 2), so an +interrupted download is never left for the next `make dict`; the floor covers the rest. diff --git a/plans/260908-1653-kaikki-viwiktionary-corpus/phase-02-fetch-the-newest-file.md b/plans/260908-1653-kaikki-viwiktionary-corpus/phase-02-fetch-the-newest-file.md new file mode 100644 index 0000000..9752ac8 --- /dev/null +++ b/plans/260908-1653-kaikki-viwiktionary-corpus/phase-02-fetch-the-newest-file.md @@ -0,0 +1,105 @@ +--- +phase: 2 +title: "Phase 2: Fetch the newest file" +status: completed +priority: P1 +effort: "2h" +dependencies: [1] +--- + +# Phase 2: Fetch the newest file + +## Overview + +Point the Makefile and the Docker dictionary stage at kaikki's rolling URL with no checksum, +keep the two build files in agreement, and make sure a failed download cannot pass as a +success anywhere. + +## Requirements + +- Functional: `make fetch-dict` downloads the kaikki file to `data/kaikki-viwiktionary-vi.jsonl` + with `curl -fL` and fails on any HTTP error; `make dict` builds from it with `--kaikki`. +- Functional: `verify-dict` and `DICT_SHA256` are removed from the Makefile and Dockerfile; + nothing in the repository claims a checksum for this file. +- Functional: the Docker `dict` stage downloads the same URL and runs the same command; + `FIXTURE_DICT=1` keeps working untouched. +- Functional: `web/tests/dictionary-source.test.js` asserts the Makefile and Dockerfile URLs + agree, that the URL is the kaikki viwiktionary `Tiếng Việt` file, and that the builder's + `kaikkiSourceURL` constant is the same URL. The three checksum assertions go. +- Functional: the CI leak guard greps for `kaikki.*\.jsonl$`; `.gitignore` ignores the new + file name. +- Non-functional: no resume flag (`-C -`) — a rolling file can change between attempts, and a + resumed download would splice two versions together. + +## Architecture + +```make +DICT_URL := https://kaikki.org/viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl +DICT_SRC := data/kaikki-viwiktionary-vi.jsonl +DICT_OUT := data/noitu.db + +fetch-dict: + @mkdir -p data + curl -fL -o $(DICT_SRC).part $(DICT_URL) && mv $(DICT_SRC).part $(DICT_SRC) + +dict: $(DICT_SRC) + cd server && go run ./cmd/build-dictionary --kaikki ../$(DICT_SRC) --out ../$(DICT_OUT) +``` + +The Dockerfile mirrors it with `ARG DICT_URL` only. The URL must stay percent-encoded in +both files: Make and `sh` would otherwise split on the space in `Tiếng Việt`. + +Correctness now rests on three checks instead of a checksum: `curl -f` plus an atomic +`.part` rename (HTTP errors and interrupted downloads), the JSON decoder (shape, and +truncation that lands mid-line), the 30,000 floor (content, including a truncation that +lands on a line boundary). + +## Related Code Files + +- Modify: `Makefile` — variables, `help` text, `fetch-dict`, remove `verify-dict` and its + `.PHONY` entry, `dict` flags +- Modify: `Dockerfile` — `ARG DICT_URL`, remove `ARG DICT_SHA256` and the `sha256sum` line, + download file name, `--kaikki`; stage comment (no longer "pinned … checked by digest") +- Modify: `web/tests/dictionary-source.test.js` — drop checksum tests, keep URL agreement, + add the kaikki-path and builder-constant assertions +- Modify: `.github/workflows/ci.yml` — leak-guard pattern and comment +- Modify: `.gitignore` — `data/kaikki-viwiktionary-vi.jsonl` replaces the undertheseanlp entry +- Modify: `README.md` — Make targets table (`verify-dict` row gone), the manual `curl` and + `build-dictionary` lines + +## Implementation Steps + +1. Update the Makefile variables and targets; delete `verify-dict`. +2. Mirror in the Dockerfile; rewrite the stage comment to say the file is fetched fresh and + identified by the SHA-256 the builder records. +3. Rewrite the pin test as described. Run `npx vitest run tests/dictionary-source.test.js`. +4. Update the CI leak guard, `.gitignore` and the README lines. +5. Run `make fetch-dict && make dict` (or the raw commands on Windows) from a clean `data/`. +6. Build the image with and without `FIXTURE_DICT=1`; run the CI file checks against the + real one and confirm no `*.jsonl` is inside. +7. Simulate a bad download once: point `DICT_URL` at a 404 and confirm `curl -f` fails the + target; feed the builder a truncated copy of the file and confirm the build fails. + +## Success Criteria + +- [x] `make fetch-dict` fetches ~62 MB and `make dict` builds >30,000 words. +- [x] `grep -rn DICT_SHA256 --exclude-dir=plans .` finds nothing; `make verify-dict` is not a + target. +- [x] `web/tests/dictionary-source.test.js` passes and fails if either URL is edited alone. +- [x] Both image variants build; the real one passes the CI file checks; no `.jsonl` inside. +- [x] A 404 URL and a truncated file each fail the build with a readable message. + +## Risk Assessment + +**kaikki.org is a single volunteer-run host.** An outage breaks `make fetch-dict` and the +release image build until it returns. Signal: `curl: (22)` on fetch. Response: wait, or +build from a previously fetched local copy (`make dict` needs only the file); if outages +recur, keep a copy as a release asset and point `DICT_URL` at it — the open question in +`plan.md`. + +**Two builds ship different words.** By design. Signal: none needed. Response: `meta` +identifies the bytes; the audit note in phase 3 records what this first build contained. + +**The space in the URL.** Signal: `curl` fetching `https://kaikki.org/viwiktionary/Ti%E1%BA%BFng` +and failing. Response: the URL stays percent-encoded in every file, and the pin test checks +that the string contains `Ti%E1%BA%BFng%20Vi%E1%BB%87t`. diff --git a/plans/260908-1653-kaikki-viwiktionary-corpus/phase-03-measure-the-corpus.md b/plans/260908-1653-kaikki-viwiktionary-corpus/phase-03-measure-the-corpus.md new file mode 100644 index 0000000..89f7812 --- /dev/null +++ b/plans/260908-1653-kaikki-viwiktionary-corpus/phase-03-measure-the-corpus.md @@ -0,0 +1,81 @@ +--- +phase: 3 +title: "Phase 3: Measure the corpus" +status: completed +priority: P1 +effort: "1h" +dependencies: [1, 2] +--- + +# Phase 3: Measure the corpus + +## Overview + +Record what the first kaikki build contains against today's database: the graph numbers, the +words gained and lost, the POS mix, and bot-versus-bot behaviour. Measurement, not a gate — +no rule removes words any more, so there is nothing to audit by hand. + +## Requirements + +- Functional: a note in this plan directory with the numbers below and the commands that + produced them, plus the `source_sha256` of the file measured so the note is tied to bytes. +- Functional: a diff against the current `data/noitu.db` (copied aside before Phase 2 + overwrites it): shared, gained, lost, with 25 random samples of each. +- Functional: playability on both databases: words, syllables, openers, openers with ≥2 + continuations, dead ends, and the `internal/bot` real-corpus test output. +- Non-functional: every number reproducible from a written command. + +## Architecture + +Pre-measured this session on kaikki's 2026-09-06 file (sha256 `51ddc2fbd73cd7e2…`); the +phase re-measures on whatever the build fetched and explains any delta: + +| | today | kaikki-vi (2026-09-06) | +|---|---|---| +| words | 26,845 | 34,813 | +| syllables | 5,709 | 6,081 | +| openers ≥2 continuations | 2,787 | 3,163 | +| dead-end syllables | 1,551 | 1,594 | +| shared / gained / lost | | 25,392 / 9,421 / 1,453 | + +POS of the 41,507 distinct words: noun 16,219 · verb 9,355 · adj 6,698 · name 5,180 · +unknown 4,652 · adv 1,024 · phrase 392 · proverb 322 · character 198. `name` and `unknown` +are kept; the log prints the tally so a future decision has the numbers. + +## Related Code Files + +- Create: `plans/260908-1653-kaikki-viwiktionary-corpus/measurement.md` +- Read only: the pre-switch `data/noitu.db` copy, the new build, `server/internal/bot` + real-corpus tests + +## Implementation Steps + +1. Before Phase 2 overwrites it, copy today's `data/noitu.db` to a scratch path. +2. After the first `make dict`, record `meta.source_sha256`, `source_rows`, + `source_fetched_at` and the build log's reject and POS lines. +3. Compute the diff and the five graph numbers for both databases (sqlite, one script; keep + the script text in the note). +4. Run `go test ./internal/bot/ -run RealCorpus -v` against both databases (swap the file + under `data/noitu.db`, restore afterwards) and record the game-length and win-rate lines. +5. Look at the 1,453 lost words: they are words in the 2018 scrape that Wiktionary has since + deleted or renamed. Sample 25, note what they look like (expected: misspellings, moved + pages, deleted junk). No action unless the sample is mostly real vocabulary. + +## Success Criteria + +- [x] `measurement.md` exists with the tables, samples, bot lines, commands and the source + SHA-256. +- [x] Words > 30,000; syllables, openers and ≥2-continuation counts all above today's. +- [x] Bot real-corpus tests pass on the new database; easy-vs-easy game length is not + shorter than today's 12.9 moves. +- [x] The lost-word sample is characterised in one paragraph. + +## Risk Assessment + +**The fetched file differs from the one measured today.** Certain over time. Signal: counts +off from the table above. Response: record the new numbers and the SHA-256; the deltas are +the point of the note, not a failure. + +**Dead ends rise slightly (1,551 → 1,594).** More words bring more rare final syllables. +Signal: already known. Response: the ratio of dead ends to syllables falls (27.2% → 26.2%), +so the graph is denser, not sparser; record it and move on. diff --git a/plans/260908-1653-kaikki-viwiktionary-corpus/phase-04-relicense-to-4-0.md b/plans/260908-1653-kaikki-viwiktionary-corpus/phase-04-relicense-to-4-0.md new file mode 100644 index 0000000..77a7b4c --- /dev/null +++ b/plans/260908-1653-kaikki-viwiktionary-corpus/phase-04-relicense-to-4-0.md @@ -0,0 +1,94 @@ +--- +phase: 4 +title: "Phase 4: Relicense to 4.0" +status: completed +priority: P1 +effort: "1h" +dependencies: [3] +--- + +# Phase 4: Relicense to 4.0 + +## Overview + +Current Wiktionary text is CC BY-SA 4.0, so the derived database is too. Restore the 4.0 +legal code, rewrite the attribution chain for a source with no commit to cite, and update +every place that says 3.0 — in the same commit as the first kaikki-built database. + +## Requirements + +- Functional: `data/LICENSE` is the CC BY-SA 4.0 legal code, byte-identical to the file this + repository shipped before commit `39ef457` (`git show e19b083:data/LICENSE`). +- Functional: `NOTICE` section 2 says 4.0 and names Wiktionary tiếng Việt as the original + work and kaikki.org/wiktextract as the extraction; section 1 byte-identical. +- Functional: `data/ATTRIBUTION.md` records the source URL, that the file is refreshed + weekly and unpinned, that the exact bytes of any build are identified by + `meta.source_sha256`, the three attribution links, and the modifications list. +- Functional: the in-game footer links to the 4.0 deed and says 4.0; the e2e assertion + matches. +- Non-functional: no overclaiming about dates — the attribution names the dump date only as + "the Wikimedia dump kaikki extracted at fetch time; see `source_fetched_at`", not a fixed + date that goes stale. + +## Architecture + +Attribution chain, all three named: + +``` +Wiktionary tiếng Việt contributors authors, CC BY-SA 4.0 + → wiktextract / kaikki.org (Tatu Ylonen) extraction; cite LREC 2022 as kaikki requests + → this project filter + index, data/noitu.db, CC BY-SA 4.0 +``` + +Modifications list (renumbered): 1 language selection (`lang_code = vi`), 2 length filter, +3 content filter, 4 normalization, 5 spelling aliases, 6 added columns and tables, +7 deduplication, 8 dropped fields (senses, translations, POS, categories — everything but +the word form). + +## Related Code Files + +- Modify: `data/LICENSE` — restore 4.0 text from git history +- Modify: `NOTICE` — section 2 heading, source lines, share-alike sentence +- Modify: `data/ATTRIBUTION.md` — full rewrite of Source, Modifications 1 and 8, reproduce + block (no checksum step) +- Modify: `README.md` — licence table row, dictionary-source paragraph +- Modify: `Dockerfile`, `docs/deployment.md`, `.github/workflows/ci.yml` — 3.0 → 4.0 in comments +- Modify: `server/cmd/build-dictionary/main.go` doc comment, `server/cmd/noitu-server/main.go` + and `server/internal/dictionary/store.go` comments, `store_test.go` seed strings +- Modify: `web/src/lib/components/AttributionFooter.svelte` (deed URL, comment), + `web/src/lib/i18n/vi.js` (`attributionLicense`), `web/e2e/bot-game.spec.js` (regex) +- Verify unchanged: `.github/workflows/ci.yml` required-files check + +## Implementation Steps + +1. `git show e19b083:data/LICENSE > data/LICENSE`; confirm the first lines read + "Attribution-ShareAlike 4.0 International". +2. Rewrite `NOTICE` section 2 source lines; leave the share-alike paragraph's meaning, change + the version. +3. Rewrite `data/ATTRIBUTION.md` as described. State plainly: "The file is not pinned. Each + build fetches kaikki's current export; the SHA-256 and row count of the file a given + database was built from are recorded in its `meta` table." +4. `grep -rn "3\.0\|by-sa/3" --exclude-dir=plans .` and fix every hit; then the same for + `undertheseanlp`, `DICT_SHA256`, `--merged`, `--sources`. +5. Regenerate `data/noitu.db`; confirm `meta.source_license` says 4.0 and matches the files. +6. Run the full verification set from the previous plan: Go vet/test, web check/test, both + image builds with the CI file checks. + +## Success Criteria + +- [x] `data/LICENSE` is the 4.0 text; `git diff e19b083 -- data/LICENSE` is empty. +- [x] No file outside `plans/` says CC BY-SA 3.0 or names undertheseanlp. +- [x] `NOTICE` section 1 has no diff. +- [x] `ATTRIBUTION.md` names all three links and states the unpinned-fetch policy. +- [x] `meta.source_license` in the built database says 4.0. +- [x] All verification suites green; the commit contains both the licence files and the + builder change so no intermediate tree mixes 3.0 text with 4.0 data. + +## Risk Assessment + +**A stale 3.0 or undertheseanlp claim survives.** Signal: the grep in step 4 finds one after +the phase is called done. Response: grep is a step for that reason; run it last. + +**The footer wording.** `giấy phép CC BY-SA 3.0` → `4.0` is one string, but the e2e regex +and the deed URL must change with it. Signal: e2e footer test fails in CI. Response: all +three edits are listed above; do them together. diff --git a/plans/260908-1653-kaikki-viwiktionary-corpus/plan.md b/plans/260908-1653-kaikki-viwiktionary-corpus/plan.md new file mode 100644 index 0000000..17f5f89 --- /dev/null +++ b/plans/260908-1653-kaikki-viwiktionary-corpus/plan.md @@ -0,0 +1,142 @@ +--- +title: "Kaikki viwiktionary corpus" +description: "Replace the 2018 undertheseanlp wiktionary rows with kaikki.org's current extraction of Wiktionary tiếng Việt, fetched unpinned at build time, and move the data licence to CC BY-SA 4.0" +status: completed +priority: P1 +effort: "~1d" +tags: [dictionary, data, licensing, build] +created: 2026-09-08 +blockedBy: [] +blocks: [] +--- + +# Kaikki viwiktionary corpus + +## Overview + +`data/noitu.db` is built from the `wiktionary` rows of `undertheseanlp/dictionary`: a +2018-12-10 scrape of vi.wiktionary.org, 26,845 words after filtering, pinned by commit and +SHA-256 (plan `260908-1525-dictionary-corpus-switch`, completed today). It is the same +website eight years staler than it needs to be. + +This plan switches the source to **kaikki.org's extraction of Wiktionary tiếng Việt** — the +`Tiếng Việt` language file of the `viwiktionary` edition, produced by wiktextract from the +monthly Wikimedia dump (currently 2026-09-01) and refreshed roughly weekly. Same authors, +same website, current text. Measured through the real `build-dictionary` filter this +session: **34,813 words**, a near-superset of today's corpus (25,392 shared, 9,421 gained, +1,453 lost) and 96% of the raw dump (36,200). + +Two owner decisions shape the plan: + +- **viwiktionary only.** kaikki's English-Wiktionary Vietnamese file (21,244 words, marked + deprecated by kaikki) is not used. The game is Vietnamese; one edition, one attribution. +- **No pin — fetch the newest file every build.** kaikki publishes rolling files with no + archived snapshots, so a checksum pin would break weekly. The owner chose freshness over + reproducibility. The build therefore has to *fail loudly* on a bad download (HTTP error, + malformed JSON, too few words) and *record what it actually got* (SHA-256 and row count in + `meta`) so any shipped database can still be traced to exact bytes. + +Consequence: the current text of Wiktionary is CC BY-SA **4.0** (Wikimedia moved from 3.0 +on 2023-06-29), so `data/noitu.db` returns to 4.0. The 3.0 chosen this afternoon only ever +applied to the 2018 snapshot. + +Evidence: `plans/reports/research-260908-1529-viwiktionary-dump-measured.md` (dump and +kaikki-en measurements) plus this session's measurement of the kaikki-vi file, recorded in +[phase 3](./phase-03-measure-the-corpus.md). + +## What changes, in numbers + +| | today | after | +|---|---|---| +| words | 26,845 | **34,813** (30,231 if `name` POS were dropped — it is not) | +| syllables | 5,709 | 6,081 | +| syllables with ≥2 continuations | 2,787 | 3,163 | +| dead-end syllables | 1,551 | 1,594 | +| source | 4.8 MB JSONL, 2018, pinned | **62 MB JSONL, current, unpinned** | +| data licence | CC BY-SA 3.0 | **CC BY-SA 4.0** | +| `--min-words` floor | 20,000 | **30,000** | +| build reproducible | yes | **no** — traceable via `meta.source_sha256` | + +## The source + +``` +URL https://kaikki.org/viwiktionary/Tiếng Việt/kaikki.org-dictionary-TiếngViệt.jsonl + (percent-encoded: /viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl) +Size 62,288,004 bytes on 2026-09-06 — 44,564 rows, 41,507 distinct words, all lang_code "vi" +Row {"word": "trở thành", "pos": "verb", "lang_code": "vi", "lang": "Tiếng Việt", "senses": [...], ...} +``` + +Only `word` and `lang_code` are read. `pos` is counted for the build log but never filters: +the owner's earlier decision that capitalization does not remove words carries over to the +`name` POS (5,138 name-only words). 5,475 words carry uppercase and 229 lowercase forms +occur twice; `accept()` lowercases and dedupes as it always has. + +## Decisions taken + +- **One corpus input mode.** `--merged`/`--sources` (undertheseanlp) is replaced by + `--kaikki`, not kept beside it. Two readers for one database is dead weight, as the SQLite + path was. The fixture path `--words` stays. +- **The floor moves to 30,000.** 34,813 measured; a truncated or reshaped download must + fail, and 30,000 is the highest round number with a comfortable margin. +- **Provenance is measured, not asserted.** The reader hashes the file as it streams it and + writes `source_sha256`, `source_rows` and `source_fetched_at` into `meta`. There is no + commit to record, so this is what "which bytes" means from now on. +- **Attribution names three links.** Wiktionary tiếng Việt contributors (authors, CC BY-SA + 4.0), wiktextract/kaikki.org by Tatu Ylonen (extraction; kaikki asks for the LREC 2022 + citation), and this project. No GitHub intermediary any more. +- **`verify-dict` goes.** There is nothing to verify against. The + Makefile/Dockerfile agreement test keeps guarding the URL and drops the checksum clauses. + +## Phases + +| # | Phase | Status | +|---|-------|--------| +| 1 | [Read the kaikki file](./phase-01-start.md) | Pending | +| 2 | [Fetch the newest file](./phase-02-fetch-the-newest-file.md) | Pending | +| 3 | [Measure the corpus](./phase-03-measure-the-corpus.md) | Pending | +| 4 | [Relicense to 4.0](./phase-04-relicense-to-4-0.md) | Pending | + +Phase 4 lands in the same commit as the first kaikki-built `noitu.db`: a 4.0-derived +database in a tree whose `NOTICE` says 3.0 is the one ordering mistake with a legal +consequence. Phase 3 is measurement, not a gate — there is no drop rule left to audit. + +## Success criteria + +- [x] `make fetch-dict && make dict` downloads kaikki's current file and builds + `data/noitu.db` with **more than 30,000 words**; a truncated or non-JSON download fails + the build with a message naming the cause. +- [x] `meta` records `source_url`, `source_sha256`, `source_rows`, `source_fetched_at` and + `source_license = CC BY-SA 4.0`; no `source_commit`, `sources_kept` or + `sources_excluded` remain. +- [x] `data/LICENSE` is the CC BY-SA 4.0 legal code; `NOTICE`, `data/ATTRIBUTION.md`, README, + Dockerfile, deployment doc, the builder's licence string and the in-game footer all say + 4.0 and credit Wiktionary tiếng Việt contributors and wiktextract/kaikki.org. +- [x] Playability measured and recorded against today's database, including bot game + lengths; nothing regresses below today's numbers. +- [x] `go vet ./... && go test ./...`, `npm run check && npm test`, the Docker image with the + CI file checks and `FIXTURE_DICT=1` are green. e2e: not runnable on this machine (Playwright Chromium download fails); the one changed assertion was updated by hand, CI verifies. +- [x] No file outside `plans/` names undertheseanlp, `--merged`, `--sources` or `DICT_SHA256`. + +## Review + +Code review (2026-09-08) returned DONE_WITH_CONCERNS; every finding was applied: `fetch-dict` +downloads to a `.part` name and renames on success so an interrupted fetch never feeds the +next `make dict`; a failed `Stat` is an error rather than a zero `source_fetched_at`; +`builder_version` bumped to 3 for the changed meta contract; `.gitignore` covers +`data/*.jsonl`; `ATTRIBUTION.md` states kaikki's redistribution terms; bare JSON literals +are malformed rather than "foreign-language"; the non-EOF read-error line number is exact; +the CI leak guard matches any `.jsonl`; byte-exact reader tests cover a missing final +newline, an HTML error page, a mid-line cut and a bare literal. The reviewer's residual +note stands: with no pin, the build trusts kaikki.org over TLS, and the recorded SHA-256 is +the compensating control; blast radius is game content only. + +## Open questions + +- kaikki's weekly refresh means two builds a week apart can differ. Accepted by the owner; + `source_sha256` in `meta` is the answer to "which words did this image ship". Should a + release ever need to be rebuilt bit-for-bit, the fallback is to keep the fetched file as a + release asset — a Makefile variable change, not a redesign. +- The `unknown` POS bucket (4,616 words such as `tình báo`, `trọc phú`, `hội quán`) is real + vocabulary the vi extractor could not label. Kept, like everything else. + + diff --git a/plans/journals/2026-09-08-planned-the-kaikki-viwiktionary-corpus-switch.md b/plans/journals/2026-09-08-planned-the-kaikki-viwiktionary-corpus-switch.md new file mode 100644 index 0000000..b942115 --- /dev/null +++ b/plans/journals/2026-09-08-planned-the-kaikki-viwiktionary-corpus-switch.md @@ -0,0 +1,25 @@ +--- +title: Planned the kaikki viwiktionary corpus switch +date: 2026-09-08 +summary: "Plan to replace the 2018 undertheseanlp wiktionary rows with kaikki.org's current Wiktionary tiếng Việt extraction, fetched unpinned; 26,845 -> 34,813 words, data licence to CC BY-SA 4.0" +--- + +# Planned the kaikki viwiktionary corpus switch + +## What happened + +- Measured kaikki's `viwiktionary` edition Vietnamese file through the real filter: 34,813 words (25,392 shared with today, 9,421 gained, 1,453 lost), a 96% subset of the raw dump. kaikki's enwiktionary Vietnamese file is 21,244 and deprecated. +- Owner chose the viwiktionary file only, and to fetch kaikki's newest file each build with no checksum pin. +- Plan written: `plans/260908-1653-kaikki-viwiktionary-corpus/` — four phases: kaikki reader with streamed SHA-256 into `meta`, unpinned fetch with loud failure modes, measurement note, relicense to CC BY-SA 4.0. +- Previous plan `260908-1525-dictionary-corpus-switch` marked completed. + +## Decision + +- Freshness over reproducibility: two builds a week apart may differ; `meta.source_sha256` identifies the bytes. Fallback if bit-for-bit rebuilds are ever needed: keep the fetched file as a release asset. +- The 3.0 licence chosen earlier today only applied to the 2018 snapshot; current Wiktionary text is 4.0, so the database returns to 4.0. + +## Next steps + +- Validate or cook the plan. + +> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions. diff --git a/plans/journals/2026-09-08-switched-the-corpus-to-kaikkis-wiktionary-ting-vit-export.md b/plans/journals/2026-09-08-switched-the-corpus-to-kaikkis-wiktionary-ting-vit-export.md new file mode 100644 index 0000000..2262cdc --- /dev/null +++ b/plans/journals/2026-09-08-switched-the-corpus-to-kaikkis-wiktionary-ting-vit-export.md @@ -0,0 +1,25 @@ +--- +title: "Switched the corpus to kaikki's Wiktionary tiếng Việt export" +date: 2026-09-08 +summary: "Replaced the pinned 2018 undertheseanlp rows with kaikki.org's current wiktextract export fetched unpinned; 26,845 -> 34,813 words; provenance by streamed SHA-256; data licence back to CC BY-SA 4.0" +--- + +# Switched the corpus to kaikki's Wiktionary tiếng Việt export + +## What happened + +- `build-dictionary` reads kaikki's JSONL with `--kaikki`, hashing the stream and writing `source_sha256`, `source_rows`, `source_fetched_at` into `meta`; `--merged`/`--sources` and the undertheseanlp reader are gone; `--min-words` floor 30,000; `builder_version` 3. +- Makefile and Dockerfile fetch the rolling kaikki URL with `curl -fL`, no checksum, atomic `.part` rename; `verify-dict` and `DICT_SHA256` removed; the web test guards Makefile == Dockerfile == builder constant and the viwiktionary path. +- Data licence back to CC BY-SA 4.0 (current Wiktionary text); `data/LICENSE` restored from git history; NOTICE/ATTRIBUTION/README/footer credit Wiktionary tiếng Việt contributors and wiktextract/kaikki.org with the LREC citation. +- Measured: 34,813 words (+9,421 / -1,453 vs before), all graph metrics up, bot game length unchanged. Lost words are mostly reduplicatives the vi extractor misses and deleted person-name pages; recorded in `measurement.md`. +- Verified: Go vet/test, web check/test (174), both Docker variants with CI file checks, 404 and truncated-download failure modes. e2e not runnable locally (Chromium download fails), CI covers it. + +## Decision + +- Owner: viwiktionary edition only; fetch newest, no pin. Two builds a week apart may differ; the SHA-256 in `meta` identifies the bytes. Fallback if reproducibility is ever needed: keep a fetched copy as a release asset. + +## Next steps + +- Commit (not yet committed); watch CI's e2e job. + +> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions.