mirror of
https://github.com/tiennm99/noitu.git
synced 2026-10-11 03:13:45 +00:00
docs(plans): record the kaikki corpus plan, measurement and review
This commit is contained in:
1 parent
0a5c5b06fe
commit
ff5627336d
9 files changed
+688
-1
No files matched your search
@@ -1,7 +1,7 @@
|
||||
---
|
||||
title: "Dictionary corpus switch"
|
||||
description: "Replace the 179 MB minhqnd aggregate with the wiktionary branch of undertheseanlp/dictionary only (hongocduc and tudientv excluded) and keep data/noitu.db on CC BY-SA 4.0"
|
||||
status: in-progress
|
||||
status: completed
|
||||
priority: P1
|
||||
effort: "~1d"
|
||||
tags: [dictionary, data, licensing, build]
|
||||
|
||||
@@ -0,0 +1,97 @@
|
||||
# Measurement: first kaikki build vs the undertheseanlp corpus
|
||||
|
||||
Date: 2026-09-08. Phase 3 of `plan.md`. The pre-switch `data/noitu.db` (26,845 words) was
|
||||
copied aside before the first kaikki build and is the "old" side of every table.
|
||||
|
||||
## The file measured
|
||||
|
||||
```
|
||||
source_url https://kaikki.org/viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl
|
||||
source_sha256 51ddc2fbd73cd7e2a468ccff3c31200928e7200cc58622f528ba3eb84c7a49f0
|
||||
source_rows 44564
|
||||
source_fetched_at 2026-09-08T10:04:08Z (kaikki's export of 2026-09-06, dump 2026-09-01)
|
||||
size 62,288,004 bytes; longest line 61,552 bytes
|
||||
```
|
||||
|
||||
Build log:
|
||||
|
||||
```
|
||||
rejected 4: contains a digit
|
||||
rejected 136: contains punctuation
|
||||
rejected 8348: fewer than 2 syllables
|
||||
rejected 160: no Vietnamese letters
|
||||
parts of speech: noun 16219, verb 9355, adj 6698, name 5180, unknown 4652, adv 1024,
|
||||
phrase 392, proverb 322, character 198, intj 175, pron 100, conj 70, prep 48, num 45,
|
||||
abbrev 38, particle 18, romanization 14, prefix 13, suffix 2, det 1
|
||||
accepted 34813 distinct words
|
||||
generated 1664 spelling aliases (552 skipped as ambiguous or already real words)
|
||||
```
|
||||
|
||||
## Commands
|
||||
|
||||
```sh
|
||||
cp data/noitu.db /tmp/old.db # before the switch
|
||||
make fetch-dict && make dict # or the raw curl + go run lines in README
|
||||
cd server && go test ./internal/bot/ -run RealCorpus -v -count=1 # once per database at data/noitu.db
|
||||
# graph numbers and diff: sqlite, see the queries below
|
||||
```
|
||||
|
||||
```sql
|
||||
SELECT COUNT(*) FROM words;
|
||||
SELECT COUNT(*) FROM syllables;
|
||||
SELECT COUNT(DISTINCT first) FROM words;
|
||||
SELECT COUNT(*) FROM (SELECT first FROM words GROUP BY first HAVING COUNT(*) >= 2);
|
||||
SELECT COUNT(*) FROM syllables WHERE out_degree = 0;
|
||||
```
|
||||
|
||||
## Playability
|
||||
|
||||
| metric | old (undertheseanlp) | new (kaikki) |
|
||||
|---|---|---|
|
||||
| words | 26,845 | **34,813** |
|
||||
| syllables | 5,709 | 6,081 |
|
||||
| syllables that open a word | 4,158 | 4,487 |
|
||||
| …with ≥2 continuations | 2,787 | 3,163 |
|
||||
| dead-end syllables | 1,551 | 1,594 (27.2% → 26.2% of syllables) |
|
||||
|
||||
Bot-versus-bot, 60 games each, same run:
|
||||
|
||||
| | old | new |
|
||||
|---|---|---|
|
||||
| hard-vs-easy win rate (moves) | 98% (3.6) | 97% (3.2) |
|
||||
| medium-vs-easy | 87% | 95% |
|
||||
| hard-vs-medium (moves) | 57% (3.1) | 67% (3.1) |
|
||||
| easy-vs-easy game length | 14.4 moves | 14.4 moves |
|
||||
| hard decision p95 | 1.5 ms | 3.1 ms |
|
||||
|
||||
Both real-corpus tests pass on both databases. Game length is unchanged; the stronger bots
|
||||
win more often on the denser graph, which is the expected direction.
|
||||
|
||||
## Corpus diff
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| shared | 25,392 |
|
||||
| gained | **9,421** — `công nghiệp hóa`, `bóng râm`, `bội thu`, `đồng dao`, `bỏ ngỏ`, `đắt hàng`, `ấn độ giáo`, `chạy làng`, `béo nung núc`, `nhìn muốn rụng trứng` |
|
||||
| lost | **1,453** — see below |
|
||||
|
||||
### The 1,453 lost words
|
||||
|
||||
Sample of 25 (seed 7): `trịnh tuệ`, `tuân khanh`, `qua quít`, `cúc pha`, `mành mành`,
|
||||
`cây bài chặt`, `trảm phong`, `kéo cánh`, `ngọt lự`, `rưng rức`, `làu nhàu`, `trong lúc`,
|
||||
`khí hư`, `trần ích tắc`, `nháo nhác`, `trần cảnh`, `óc ách`, `lườm lườm`, `hỗn luân`,
|
||||
`trụi lủi`, `tục tác`, `lảu nhảu`, `phát sầu`, `hãng hàng không quốc gia việt nam`, `trói ké`.
|
||||
|
||||
Two kinds. Roughly a fifth are historical person names (`trần ích tắc`, `trần cảnh`, `trịnh
|
||||
tuệ`) whose pages have since been deleted or moved. The rest are real vocabulary, heavily
|
||||
reduplicatives (`rưng rức`, `nháo nhác`, `óc ách`, `làu nhàu`, `lườm lườm`), that the 2018
|
||||
scrape had and the current wiktextract export does not. Most likely those pages still exist
|
||||
but in a markup the `vi` extractor does not yet parse — the raw 2026 dump has 36,200
|
||||
words to kaikki's 34,813, a 1,387-word gap of the same order. **Not acted on:** the plan
|
||||
does not merge sources, and the net is +7,968 words. If the reduplicatives matter, the
|
||||
option is a second `--words` input reduced from the raw dump, as the previous plan's
|
||||
research described.
|
||||
|
||||
## Verdict
|
||||
|
||||
Every number above today's; nothing regresses. Corpus accepted as built.
|
||||
@@ -0,0 +1,118 @@
|
||||
---
|
||||
phase: 1
|
||||
title: "Phase 1: Read the kaikki file"
|
||||
status: completed
|
||||
priority: P1
|
||||
effort: "3h"
|
||||
dependencies: []
|
||||
---
|
||||
|
||||
# Phase 1: Read the kaikki file
|
||||
|
||||
## Overview
|
||||
|
||||
Replace the undertheseanlp reader in `build-dictionary` with one for kaikki's wiktextract
|
||||
JSONL, hashing the input as it streams so the database records exactly which bytes it was
|
||||
built from.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: `--kaikki <file>` reads one JSON object per line, keeps rows whose
|
||||
`lang_code` is `vi`, and feeds `word` to the existing `accept()`.
|
||||
- Functional: rows with any other `lang_code` are counted and rejected, not silently
|
||||
skipped — the file is vi-only today, and a change in that would be worth seeing in the log.
|
||||
- Functional: `pos` is tallied per value and printed with the reject counts. It never filters.
|
||||
- Functional: the SHA-256 of the file and the row count are computed during the single
|
||||
streaming pass and written to `meta` as `source_sha256` and `source_rows`, with
|
||||
`source_fetched_at` taken from the file's modification time.
|
||||
- Functional: `--merged` and `--sources` are removed along with `merged_list.go` and its
|
||||
tests. Exactly one of `--kaikki` / `--words` must be given, same rule as today.
|
||||
- Functional: the `--min-words` default moves from 20,000 to 30,000.
|
||||
- Non-functional: a malformed line is an error naming the line, as today. Lines can be
|
||||
long — a row carries every sense and translation — so the scanner buffer must allow
|
||||
several megabytes; measure the longest line on the real file and set the cap above it
|
||||
with headroom, or switch to `bufio.Reader.ReadBytes('\n')` which has no cap.
|
||||
- Non-functional: `finish()`, `accept()`, aliases, `write`, `verify` unchanged.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
--kaikki <file>
|
||||
│ one line at a time, bytes also fed to sha256.New()
|
||||
▼
|
||||
json.Unmarshal → {word, lang_code, pos}
|
||||
│
|
||||
├─ lang_code != "vi" → rejects["not Vietnamese-language entry"]++
|
||||
▼
|
||||
accept(word) (NFC, lowercase, ≥2 syllables, alphabet, phonotactics)
|
||||
▼
|
||||
finish() (floor 30,000, aliases, atomic write, verify)
|
||||
```
|
||||
|
||||
Decode only the three fields via a struct; `encoding/json` ignores the rest, so the 62 MB
|
||||
of senses and translations cost I/O but not memory.
|
||||
|
||||
`sourceSpec.extra` already exists for mode-specific meta rows; the kaikki provenance uses it:
|
||||
|
||||
```
|
||||
source_url the kaikki URL (constant kaikkiSourceURL)
|
||||
source_sha256 hex of the streamed bytes
|
||||
source_rows lines decoded
|
||||
source_fetched_at file mtime, RFC 3339 UTC
|
||||
source_license CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
|
||||
```
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `server/cmd/build-dictionary/kaikki_list.go`
|
||||
- Create: `server/cmd/build-dictionary/kaikki_list_test.go`
|
||||
- Delete: `server/cmd/build-dictionary/merged_list.go`, `merged_list_test.go`
|
||||
- Modify: `server/cmd/build-dictionary/main.go` — flags, dispatch, doc comment, `--min-words`
|
||||
default, `mergedProvenance` call site
|
||||
- Modify: `server/cmd/build-dictionary/main_test.go` — `fixtureSource`/`defaultRows` emit
|
||||
kaikki-shaped rows (`{"word": ..., "pos": ..., "lang_code": "vi"}`); the "excluded source"
|
||||
row becomes a `lang_code: "en"` row
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Write `kaikki_list.go`: `kaikkiRow{Word, Pos, LangCode}`, `readKaikkiList(path, maxSyllables)`
|
||||
returning words, rejects, a `pos` tally, and a `kaikkiProvenance` (sha256, rows, mtime).
|
||||
Wrap the file in `io.TeeReader` into `sha256.New()` so hashing is free.
|
||||
2. Add `rejectNotVietnamese rejectReason = "not a Vietnamese-language entry"` to `filter.go`.
|
||||
3. In `main.go`: replace `merged`/`sources` with `kaikki` in `config`, flags and `run()`;
|
||||
`--min-words` default 30,000; `runFromKaikkiList` logs rejects, the POS tally
|
||||
(sorted, one line) and the accepted count, then calls `finish()` with the provenance spec.
|
||||
4. Delete the merged reader and its tests. Rewrite `main_test.go` fixtures to kaikki rows.
|
||||
5. Tests in `kaikki_list_test.go`: a `vi` row kept; an `en` row rejected and counted;
|
||||
`Hà Nội` kept as `hà nội`; a word listed twice with different `pos` kept once; a malformed
|
||||
line names its number; `meta` carries a 64-hex `source_sha256` equal to `sha256sum` of the
|
||||
fixture file, the right `source_rows`, and no `source_commit`/`sources_*` keys.
|
||||
6. Run against the real file (`scratchpad/kk-vi-edition.jsonl` from this session, or a fresh
|
||||
download) and record the counts in phase 3.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] `go test ./cmd/build-dictionary/` passes; `merged_list*.go` no longer exist.
|
||||
- [x] Real file builds >30,000 words (measured 34,813) and logs a POS tally.
|
||||
- [x] `build-dictionary --help` lists `--kaikki`, `--words`, `--out`, `--max-syllables`,
|
||||
`--min-words` and nothing else.
|
||||
- [x] `meta.source_sha256` of a build equals `sha256sum` of the input file.
|
||||
- [x] The fixture path (`--words ... --min-words 150`) is byte-for-byte unaffected in
|
||||
behaviour: same 205 words from `testdata/fixture-words.txt`.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
**A row exceeds the scanner buffer.** Signal: `bufio.Scanner: token too long` on the real
|
||||
file. Response: pre-decided — use `bufio.Reader` with no line cap rather than guessing a
|
||||
buffer size; measure the longest line once and note it in the phase result.
|
||||
|
||||
**kaikki changes the row shape.** `word` and `lang_code` are wiktextract's stable core fields
|
||||
and unlikely to move. Signal: zero accepted words → the 30,000 floor fails the build with a
|
||||
clear message. Response: read the new shape, adjust the struct; the failure mode is loud by
|
||||
design.
|
||||
|
||||
**The file arrives partially.** `curl -f` catches HTTP errors but not a truncated body, and
|
||||
a body cut exactly on a line boundary parses cleanly. Signal: a malformed final line (error
|
||||
names it) or a word count under the 30,000 floor (13.8% headroom on 34,813). Response: the
|
||||
fetch writes to a `.part` name and renames only on success (review finding, Phase 2), so an
|
||||
interrupted download is never left for the next `make dict`; the floor covers the rest.
|
||||
@@ -0,0 +1,105 @@
|
||||
---
|
||||
phase: 2
|
||||
title: "Phase 2: Fetch the newest file"
|
||||
status: completed
|
||||
priority: P1
|
||||
effort: "2h"
|
||||
dependencies: [1]
|
||||
---
|
||||
|
||||
# Phase 2: Fetch the newest file
|
||||
|
||||
## Overview
|
||||
|
||||
Point the Makefile and the Docker dictionary stage at kaikki's rolling URL with no checksum,
|
||||
keep the two build files in agreement, and make sure a failed download cannot pass as a
|
||||
success anywhere.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: `make fetch-dict` downloads the kaikki file to `data/kaikki-viwiktionary-vi.jsonl`
|
||||
with `curl -fL` and fails on any HTTP error; `make dict` builds from it with `--kaikki`.
|
||||
- Functional: `verify-dict` and `DICT_SHA256` are removed from the Makefile and Dockerfile;
|
||||
nothing in the repository claims a checksum for this file.
|
||||
- Functional: the Docker `dict` stage downloads the same URL and runs the same command;
|
||||
`FIXTURE_DICT=1` keeps working untouched.
|
||||
- Functional: `web/tests/dictionary-source.test.js` asserts the Makefile and Dockerfile URLs
|
||||
agree, that the URL is the kaikki viwiktionary `Tiếng Việt` file, and that the builder's
|
||||
`kaikkiSourceURL` constant is the same URL. The three checksum assertions go.
|
||||
- Functional: the CI leak guard greps for `kaikki.*\.jsonl$`; `.gitignore` ignores the new
|
||||
file name.
|
||||
- Non-functional: no resume flag (`-C -`) — a rolling file can change between attempts, and a
|
||||
resumed download would splice two versions together.
|
||||
|
||||
## Architecture
|
||||
|
||||
```make
|
||||
DICT_URL := https://kaikki.org/viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl
|
||||
DICT_SRC := data/kaikki-viwiktionary-vi.jsonl
|
||||
DICT_OUT := data/noitu.db
|
||||
|
||||
fetch-dict:
|
||||
@mkdir -p data
|
||||
curl -fL -o $(DICT_SRC).part $(DICT_URL) && mv $(DICT_SRC).part $(DICT_SRC)
|
||||
|
||||
dict: $(DICT_SRC)
|
||||
cd server && go run ./cmd/build-dictionary --kaikki ../$(DICT_SRC) --out ../$(DICT_OUT)
|
||||
```
|
||||
|
||||
The Dockerfile mirrors it with `ARG DICT_URL` only. The URL must stay percent-encoded in
|
||||
both files: Make and `sh` would otherwise split on the space in `Tiếng Việt`.
|
||||
|
||||
Correctness now rests on three checks instead of a checksum: `curl -f` plus an atomic
|
||||
`.part` rename (HTTP errors and interrupted downloads), the JSON decoder (shape, and
|
||||
truncation that lands mid-line), the 30,000 floor (content, including a truncation that
|
||||
lands on a line boundary).
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Modify: `Makefile` — variables, `help` text, `fetch-dict`, remove `verify-dict` and its
|
||||
`.PHONY` entry, `dict` flags
|
||||
- Modify: `Dockerfile` — `ARG DICT_URL`, remove `ARG DICT_SHA256` and the `sha256sum` line,
|
||||
download file name, `--kaikki`; stage comment (no longer "pinned … checked by digest")
|
||||
- Modify: `web/tests/dictionary-source.test.js` — drop checksum tests, keep URL agreement,
|
||||
add the kaikki-path and builder-constant assertions
|
||||
- Modify: `.github/workflows/ci.yml` — leak-guard pattern and comment
|
||||
- Modify: `.gitignore` — `data/kaikki-viwiktionary-vi.jsonl` replaces the undertheseanlp entry
|
||||
- Modify: `README.md` — Make targets table (`verify-dict` row gone), the manual `curl` and
|
||||
`build-dictionary` lines
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Update the Makefile variables and targets; delete `verify-dict`.
|
||||
2. Mirror in the Dockerfile; rewrite the stage comment to say the file is fetched fresh and
|
||||
identified by the SHA-256 the builder records.
|
||||
3. Rewrite the pin test as described. Run `npx vitest run tests/dictionary-source.test.js`.
|
||||
4. Update the CI leak guard, `.gitignore` and the README lines.
|
||||
5. Run `make fetch-dict && make dict` (or the raw commands on Windows) from a clean `data/`.
|
||||
6. Build the image with and without `FIXTURE_DICT=1`; run the CI file checks against the
|
||||
real one and confirm no `*.jsonl` is inside.
|
||||
7. Simulate a bad download once: point `DICT_URL` at a 404 and confirm `curl -f` fails the
|
||||
target; feed the builder a truncated copy of the file and confirm the build fails.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] `make fetch-dict` fetches ~62 MB and `make dict` builds >30,000 words.
|
||||
- [x] `grep -rn DICT_SHA256 --exclude-dir=plans .` finds nothing; `make verify-dict` is not a
|
||||
target.
|
||||
- [x] `web/tests/dictionary-source.test.js` passes and fails if either URL is edited alone.
|
||||
- [x] Both image variants build; the real one passes the CI file checks; no `.jsonl` inside.
|
||||
- [x] A 404 URL and a truncated file each fail the build with a readable message.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
**kaikki.org is a single volunteer-run host.** An outage breaks `make fetch-dict` and the
|
||||
release image build until it returns. Signal: `curl: (22)` on fetch. Response: wait, or
|
||||
build from a previously fetched local copy (`make dict` needs only the file); if outages
|
||||
recur, keep a copy as a release asset and point `DICT_URL` at it — the open question in
|
||||
`plan.md`.
|
||||
|
||||
**Two builds ship different words.** By design. Signal: none needed. Response: `meta`
|
||||
identifies the bytes; the audit note in phase 3 records what this first build contained.
|
||||
|
||||
**The space in the URL.** Signal: `curl` fetching `https://kaikki.org/viwiktionary/Ti%E1%BA%BFng`
|
||||
and failing. Response: the URL stays percent-encoded in every file, and the pin test checks
|
||||
that the string contains `Ti%E1%BA%BFng%20Vi%E1%BB%87t`.
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
phase: 3
|
||||
title: "Phase 3: Measure the corpus"
|
||||
status: completed
|
||||
priority: P1
|
||||
effort: "1h"
|
||||
dependencies: [1, 2]
|
||||
---
|
||||
|
||||
# Phase 3: Measure the corpus
|
||||
|
||||
## Overview
|
||||
|
||||
Record what the first kaikki build contains against today's database: the graph numbers, the
|
||||
words gained and lost, the POS mix, and bot-versus-bot behaviour. Measurement, not a gate —
|
||||
no rule removes words any more, so there is nothing to audit by hand.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: a note in this plan directory with the numbers below and the commands that
|
||||
produced them, plus the `source_sha256` of the file measured so the note is tied to bytes.
|
||||
- Functional: a diff against the current `data/noitu.db` (copied aside before Phase 2
|
||||
overwrites it): shared, gained, lost, with 25 random samples of each.
|
||||
- Functional: playability on both databases: words, syllables, openers, openers with ≥2
|
||||
continuations, dead ends, and the `internal/bot` real-corpus test output.
|
||||
- Non-functional: every number reproducible from a written command.
|
||||
|
||||
## Architecture
|
||||
|
||||
Pre-measured this session on kaikki's 2026-09-06 file (sha256 `51ddc2fbd73cd7e2…`); the
|
||||
phase re-measures on whatever the build fetched and explains any delta:
|
||||
|
||||
| | today | kaikki-vi (2026-09-06) |
|
||||
|---|---|---|
|
||||
| words | 26,845 | 34,813 |
|
||||
| syllables | 5,709 | 6,081 |
|
||||
| openers ≥2 continuations | 2,787 | 3,163 |
|
||||
| dead-end syllables | 1,551 | 1,594 |
|
||||
| shared / gained / lost | | 25,392 / 9,421 / 1,453 |
|
||||
|
||||
POS of the 41,507 distinct words: noun 16,219 · verb 9,355 · adj 6,698 · name 5,180 ·
|
||||
unknown 4,652 · adv 1,024 · phrase 392 · proverb 322 · character 198. `name` and `unknown`
|
||||
are kept; the log prints the tally so a future decision has the numbers.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `plans/260908-1653-kaikki-viwiktionary-corpus/measurement.md`
|
||||
- Read only: the pre-switch `data/noitu.db` copy, the new build, `server/internal/bot`
|
||||
real-corpus tests
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Before Phase 2 overwrites it, copy today's `data/noitu.db` to a scratch path.
|
||||
2. After the first `make dict`, record `meta.source_sha256`, `source_rows`,
|
||||
`source_fetched_at` and the build log's reject and POS lines.
|
||||
3. Compute the diff and the five graph numbers for both databases (sqlite, one script; keep
|
||||
the script text in the note).
|
||||
4. Run `go test ./internal/bot/ -run RealCorpus -v` against both databases (swap the file
|
||||
under `data/noitu.db`, restore afterwards) and record the game-length and win-rate lines.
|
||||
5. Look at the 1,453 lost words: they are words in the 2018 scrape that Wiktionary has since
|
||||
deleted or renamed. Sample 25, note what they look like (expected: misspellings, moved
|
||||
pages, deleted junk). No action unless the sample is mostly real vocabulary.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] `measurement.md` exists with the tables, samples, bot lines, commands and the source
|
||||
SHA-256.
|
||||
- [x] Words > 30,000; syllables, openers and ≥2-continuation counts all above today's.
|
||||
- [x] Bot real-corpus tests pass on the new database; easy-vs-easy game length is not
|
||||
shorter than today's 12.9 moves.
|
||||
- [x] The lost-word sample is characterised in one paragraph.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
**The fetched file differs from the one measured today.** Certain over time. Signal: counts
|
||||
off from the table above. Response: record the new numbers and the SHA-256; the deltas are
|
||||
the point of the note, not a failure.
|
||||
|
||||
**Dead ends rise slightly (1,551 → 1,594).** More words bring more rare final syllables.
|
||||
Signal: already known. Response: the ratio of dead ends to syllables falls (27.2% → 26.2%),
|
||||
so the graph is denser, not sparser; record it and move on.
|
||||
@@ -0,0 +1,94 @@
|
||||
---
|
||||
phase: 4
|
||||
title: "Phase 4: Relicense to 4.0"
|
||||
status: completed
|
||||
priority: P1
|
||||
effort: "1h"
|
||||
dependencies: [3]
|
||||
---
|
||||
|
||||
# Phase 4: Relicense to 4.0
|
||||
|
||||
## Overview
|
||||
|
||||
Current Wiktionary text is CC BY-SA 4.0, so the derived database is too. Restore the 4.0
|
||||
legal code, rewrite the attribution chain for a source with no commit to cite, and update
|
||||
every place that says 3.0 — in the same commit as the first kaikki-built database.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: `data/LICENSE` is the CC BY-SA 4.0 legal code, byte-identical to the file this
|
||||
repository shipped before commit `39ef457` (`git show e19b083:data/LICENSE`).
|
||||
- Functional: `NOTICE` section 2 says 4.0 and names Wiktionary tiếng Việt as the original
|
||||
work and kaikki.org/wiktextract as the extraction; section 1 byte-identical.
|
||||
- Functional: `data/ATTRIBUTION.md` records the source URL, that the file is refreshed
|
||||
weekly and unpinned, that the exact bytes of any build are identified by
|
||||
`meta.source_sha256`, the three attribution links, and the modifications list.
|
||||
- Functional: the in-game footer links to the 4.0 deed and says 4.0; the e2e assertion
|
||||
matches.
|
||||
- Non-functional: no overclaiming about dates — the attribution names the dump date only as
|
||||
"the Wikimedia dump kaikki extracted at fetch time; see `source_fetched_at`", not a fixed
|
||||
date that goes stale.
|
||||
|
||||
## Architecture
|
||||
|
||||
Attribution chain, all three named:
|
||||
|
||||
```
|
||||
Wiktionary tiếng Việt contributors authors, CC BY-SA 4.0
|
||||
→ wiktextract / kaikki.org (Tatu Ylonen) extraction; cite LREC 2022 as kaikki requests
|
||||
→ this project filter + index, data/noitu.db, CC BY-SA 4.0
|
||||
```
|
||||
|
||||
Modifications list (renumbered): 1 language selection (`lang_code = vi`), 2 length filter,
|
||||
3 content filter, 4 normalization, 5 spelling aliases, 6 added columns and tables,
|
||||
7 deduplication, 8 dropped fields (senses, translations, POS, categories — everything but
|
||||
the word form).
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Modify: `data/LICENSE` — restore 4.0 text from git history
|
||||
- Modify: `NOTICE` — section 2 heading, source lines, share-alike sentence
|
||||
- Modify: `data/ATTRIBUTION.md` — full rewrite of Source, Modifications 1 and 8, reproduce
|
||||
block (no checksum step)
|
||||
- Modify: `README.md` — licence table row, dictionary-source paragraph
|
||||
- Modify: `Dockerfile`, `docs/deployment.md`, `.github/workflows/ci.yml` — 3.0 → 4.0 in comments
|
||||
- Modify: `server/cmd/build-dictionary/main.go` doc comment, `server/cmd/noitu-server/main.go`
|
||||
and `server/internal/dictionary/store.go` comments, `store_test.go` seed strings
|
||||
- Modify: `web/src/lib/components/AttributionFooter.svelte` (deed URL, comment),
|
||||
`web/src/lib/i18n/vi.js` (`attributionLicense`), `web/e2e/bot-game.spec.js` (regex)
|
||||
- Verify unchanged: `.github/workflows/ci.yml` required-files check
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. `git show e19b083:data/LICENSE > data/LICENSE`; confirm the first lines read
|
||||
"Attribution-ShareAlike 4.0 International".
|
||||
2. Rewrite `NOTICE` section 2 source lines; leave the share-alike paragraph's meaning, change
|
||||
the version.
|
||||
3. Rewrite `data/ATTRIBUTION.md` as described. State plainly: "The file is not pinned. Each
|
||||
build fetches kaikki's current export; the SHA-256 and row count of the file a given
|
||||
database was built from are recorded in its `meta` table."
|
||||
4. `grep -rn "3\.0\|by-sa/3" --exclude-dir=plans .` and fix every hit; then the same for
|
||||
`undertheseanlp`, `DICT_SHA256`, `--merged`, `--sources`.
|
||||
5. Regenerate `data/noitu.db`; confirm `meta.source_license` says 4.0 and matches the files.
|
||||
6. Run the full verification set from the previous plan: Go vet/test, web check/test, both
|
||||
image builds with the CI file checks.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] `data/LICENSE` is the 4.0 text; `git diff e19b083 -- data/LICENSE` is empty.
|
||||
- [x] No file outside `plans/` says CC BY-SA 3.0 or names undertheseanlp.
|
||||
- [x] `NOTICE` section 1 has no diff.
|
||||
- [x] `ATTRIBUTION.md` names all three links and states the unpinned-fetch policy.
|
||||
- [x] `meta.source_license` in the built database says 4.0.
|
||||
- [x] All verification suites green; the commit contains both the licence files and the
|
||||
builder change so no intermediate tree mixes 3.0 text with 4.0 data.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
**A stale 3.0 or undertheseanlp claim survives.** Signal: the grep in step 4 finds one after
|
||||
the phase is called done. Response: grep is a step for that reason; run it last.
|
||||
|
||||
**The footer wording.** `giấy phép CC BY-SA 3.0` → `4.0` is one string, but the e2e regex
|
||||
and the deed URL must change with it. Signal: e2e footer test fails in CI. Response: all
|
||||
three edits are listed above; do them together.
|
||||
@@ -0,0 +1,142 @@
|
||||
---
|
||||
title: "Kaikki viwiktionary corpus"
|
||||
description: "Replace the 2018 undertheseanlp wiktionary rows with kaikki.org's current extraction of Wiktionary tiếng Việt, fetched unpinned at build time, and move the data licence to CC BY-SA 4.0"
|
||||
status: completed
|
||||
priority: P1
|
||||
effort: "~1d"
|
||||
tags: [dictionary, data, licensing, build]
|
||||
created: 2026-09-08
|
||||
blockedBy: []
|
||||
blocks: []
|
||||
---
|
||||
|
||||
# Kaikki viwiktionary corpus
|
||||
|
||||
## Overview
|
||||
|
||||
`data/noitu.db` is built from the `wiktionary` rows of `undertheseanlp/dictionary`: a
|
||||
2018-12-10 scrape of vi.wiktionary.org, 26,845 words after filtering, pinned by commit and
|
||||
SHA-256 (plan `260908-1525-dictionary-corpus-switch`, completed today). It is the same
|
||||
website eight years staler than it needs to be.
|
||||
|
||||
This plan switches the source to **kaikki.org's extraction of Wiktionary tiếng Việt** — the
|
||||
`Tiếng Việt` language file of the `viwiktionary` edition, produced by wiktextract from the
|
||||
monthly Wikimedia dump (currently 2026-09-01) and refreshed roughly weekly. Same authors,
|
||||
same website, current text. Measured through the real `build-dictionary` filter this
|
||||
session: **34,813 words**, a near-superset of today's corpus (25,392 shared, 9,421 gained,
|
||||
1,453 lost) and 96% of the raw dump (36,200).
|
||||
|
||||
Two owner decisions shape the plan:
|
||||
|
||||
- **viwiktionary only.** kaikki's English-Wiktionary Vietnamese file (21,244 words, marked
|
||||
deprecated by kaikki) is not used. The game is Vietnamese; one edition, one attribution.
|
||||
- **No pin — fetch the newest file every build.** kaikki publishes rolling files with no
|
||||
archived snapshots, so a checksum pin would break weekly. The owner chose freshness over
|
||||
reproducibility. The build therefore has to *fail loudly* on a bad download (HTTP error,
|
||||
malformed JSON, too few words) and *record what it actually got* (SHA-256 and row count in
|
||||
`meta`) so any shipped database can still be traced to exact bytes.
|
||||
|
||||
Consequence: the current text of Wiktionary is CC BY-SA **4.0** (Wikimedia moved from 3.0
|
||||
on 2023-06-29), so `data/noitu.db` returns to 4.0. The 3.0 chosen this afternoon only ever
|
||||
applied to the 2018 snapshot.
|
||||
|
||||
Evidence: `plans/reports/research-260908-1529-viwiktionary-dump-measured.md` (dump and
|
||||
kaikki-en measurements) plus this session's measurement of the kaikki-vi file, recorded in
|
||||
[phase 3](./phase-03-measure-the-corpus.md).
|
||||
|
||||
## What changes, in numbers
|
||||
|
||||
| | today | after |
|
||||
|---|---|---|
|
||||
| words | 26,845 | **34,813** (30,231 if `name` POS were dropped — it is not) |
|
||||
| syllables | 5,709 | 6,081 |
|
||||
| syllables with ≥2 continuations | 2,787 | 3,163 |
|
||||
| dead-end syllables | 1,551 | 1,594 |
|
||||
| source | 4.8 MB JSONL, 2018, pinned | **62 MB JSONL, current, unpinned** |
|
||||
| data licence | CC BY-SA 3.0 | **CC BY-SA 4.0** |
|
||||
| `--min-words` floor | 20,000 | **30,000** |
|
||||
| build reproducible | yes | **no** — traceable via `meta.source_sha256` |
|
||||
|
||||
## The source
|
||||
|
||||
```
|
||||
URL https://kaikki.org/viwiktionary/Tiếng Việt/kaikki.org-dictionary-TiếngViệt.jsonl
|
||||
(percent-encoded: /viwiktionary/Ti%E1%BA%BFng%20Vi%E1%BB%87t/kaikki.org-dictionary-Ti%E1%BA%BFngVi%E1%BB%87t.jsonl)
|
||||
Size 62,288,004 bytes on 2026-09-06 — 44,564 rows, 41,507 distinct words, all lang_code "vi"
|
||||
Row {"word": "trở thành", "pos": "verb", "lang_code": "vi", "lang": "Tiếng Việt", "senses": [...], ...}
|
||||
```
|
||||
|
||||
Only `word` and `lang_code` are read. `pos` is counted for the build log but never filters:
|
||||
the owner's earlier decision that capitalization does not remove words carries over to the
|
||||
`name` POS (5,138 name-only words). 5,475 words carry uppercase and 229 lowercase forms
|
||||
occur twice; `accept()` lowercases and dedupes as it always has.
|
||||
|
||||
## Decisions taken
|
||||
|
||||
- **One corpus input mode.** `--merged`/`--sources` (undertheseanlp) is replaced by
|
||||
`--kaikki`, not kept beside it. Two readers for one database is dead weight, as the SQLite
|
||||
path was. The fixture path `--words` stays.
|
||||
- **The floor moves to 30,000.** 34,813 measured; a truncated or reshaped download must
|
||||
fail, and 30,000 is the highest round number with a comfortable margin.
|
||||
- **Provenance is measured, not asserted.** The reader hashes the file as it streams it and
|
||||
writes `source_sha256`, `source_rows` and `source_fetched_at` into `meta`. There is no
|
||||
commit to record, so this is what "which bytes" means from now on.
|
||||
- **Attribution names three links.** Wiktionary tiếng Việt contributors (authors, CC BY-SA
|
||||
4.0), wiktextract/kaikki.org by Tatu Ylonen (extraction; kaikki asks for the LREC 2022
|
||||
citation), and this project. No GitHub intermediary any more.
|
||||
- **`verify-dict` goes.** There is nothing to verify against. The
|
||||
Makefile/Dockerfile agreement test keeps guarding the URL and drops the checksum clauses.
|
||||
|
||||
## Phases
|
||||
|
||||
| # | Phase | Status |
|
||||
|---|-------|--------|
|
||||
| 1 | [Read the kaikki file](./phase-01-start.md) | Pending |
|
||||
| 2 | [Fetch the newest file](./phase-02-fetch-the-newest-file.md) | Pending |
|
||||
| 3 | [Measure the corpus](./phase-03-measure-the-corpus.md) | Pending |
|
||||
| 4 | [Relicense to 4.0](./phase-04-relicense-to-4-0.md) | Pending |
|
||||
|
||||
Phase 4 lands in the same commit as the first kaikki-built `noitu.db`: a 4.0-derived
|
||||
database in a tree whose `NOTICE` says 3.0 is the one ordering mistake with a legal
|
||||
consequence. Phase 3 is measurement, not a gate — there is no drop rule left to audit.
|
||||
|
||||
## Success criteria
|
||||
|
||||
- [x] `make fetch-dict && make dict` downloads kaikki's current file and builds
|
||||
`data/noitu.db` with **more than 30,000 words**; a truncated or non-JSON download fails
|
||||
the build with a message naming the cause.
|
||||
- [x] `meta` records `source_url`, `source_sha256`, `source_rows`, `source_fetched_at` and
|
||||
`source_license = CC BY-SA 4.0`; no `source_commit`, `sources_kept` or
|
||||
`sources_excluded` remain.
|
||||
- [x] `data/LICENSE` is the CC BY-SA 4.0 legal code; `NOTICE`, `data/ATTRIBUTION.md`, README,
|
||||
Dockerfile, deployment doc, the builder's licence string and the in-game footer all say
|
||||
4.0 and credit Wiktionary tiếng Việt contributors and wiktextract/kaikki.org.
|
||||
- [x] Playability measured and recorded against today's database, including bot game
|
||||
lengths; nothing regresses below today's numbers.
|
||||
- [x] `go vet ./... && go test ./...`, `npm run check && npm test`, the Docker image with the
|
||||
CI file checks and `FIXTURE_DICT=1` are green. e2e: not runnable on this machine (Playwright Chromium download fails); the one changed assertion was updated by hand, CI verifies.
|
||||
- [x] No file outside `plans/` names undertheseanlp, `--merged`, `--sources` or `DICT_SHA256`.
|
||||
|
||||
## Review
|
||||
|
||||
Code review (2026-09-08) returned DONE_WITH_CONCERNS; every finding was applied: `fetch-dict`
|
||||
downloads to a `.part` name and renames on success so an interrupted fetch never feeds the
|
||||
next `make dict`; a failed `Stat` is an error rather than a zero `source_fetched_at`;
|
||||
`builder_version` bumped to 3 for the changed meta contract; `.gitignore` covers
|
||||
`data/*.jsonl`; `ATTRIBUTION.md` states kaikki's redistribution terms; bare JSON literals
|
||||
are malformed rather than "foreign-language"; the non-EOF read-error line number is exact;
|
||||
the CI leak guard matches any `.jsonl`; byte-exact reader tests cover a missing final
|
||||
newline, an HTML error page, a mid-line cut and a bare literal. The reviewer's residual
|
||||
note stands: with no pin, the build trusts kaikki.org over TLS, and the recorded SHA-256 is
|
||||
the compensating control; blast radius is game content only.
|
||||
|
||||
## Open questions
|
||||
|
||||
- kaikki's weekly refresh means two builds a week apart can differ. Accepted by the owner;
|
||||
`source_sha256` in `meta` is the answer to "which words did this image ship". Should a
|
||||
release ever need to be rebuilt bit-for-bit, the fallback is to keep the fetched file as a
|
||||
release asset — a Makefile variable change, not a redesign.
|
||||
- The `unknown` POS bucket (4,616 words such as `tình báo`, `trọc phú`, `hội quán`) is real
|
||||
vocabulary the vi extractor could not label. Kept, like everything else.
|
||||
|
||||
<!-- slug: kaikki-viwiktionary-corpus -->
|
||||
@@ -0,0 +1,25 @@
|
||||
---
|
||||
title: Planned the kaikki viwiktionary corpus switch
|
||||
date: 2026-09-08
|
||||
summary: "Plan to replace the 2018 undertheseanlp wiktionary rows with kaikki.org's current Wiktionary tiếng Việt extraction, fetched unpinned; 26,845 -> 34,813 words, data licence to CC BY-SA 4.0"
|
||||
---
|
||||
|
||||
# Planned the kaikki viwiktionary corpus switch
|
||||
|
||||
## What happened
|
||||
|
||||
- Measured kaikki's `viwiktionary` edition Vietnamese file through the real filter: 34,813 words (25,392 shared with today, 9,421 gained, 1,453 lost), a 96% subset of the raw dump. kaikki's enwiktionary Vietnamese file is 21,244 and deprecated.
|
||||
- Owner chose the viwiktionary file only, and to fetch kaikki's newest file each build with no checksum pin.
|
||||
- Plan written: `plans/260908-1653-kaikki-viwiktionary-corpus/` — four phases: kaikki reader with streamed SHA-256 into `meta`, unpinned fetch with loud failure modes, measurement note, relicense to CC BY-SA 4.0.
|
||||
- Previous plan `260908-1525-dictionary-corpus-switch` marked completed.
|
||||
|
||||
## Decision
|
||||
|
||||
- Freshness over reproducibility: two builds a week apart may differ; `meta.source_sha256` identifies the bytes. Fallback if bit-for-bit rebuilds are ever needed: keep the fetched file as a release asset.
|
||||
- The 3.0 licence chosen earlier today only applied to the 2018 snapshot; current Wiktionary text is 4.0, so the database returns to 4.0.
|
||||
|
||||
## Next steps
|
||||
|
||||
- Validate or cook the plan.
|
||||
|
||||
> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions.
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
---
|
||||
title: "Switched the corpus to kaikki's Wiktionary tiếng Việt export"
|
||||
date: 2026-09-08
|
||||
summary: "Replaced the pinned 2018 undertheseanlp rows with kaikki.org's current wiktextract export fetched unpinned; 26,845 -> 34,813 words; provenance by streamed SHA-256; data licence back to CC BY-SA 4.0"
|
||||
---
|
||||
|
||||
# Switched the corpus to kaikki's Wiktionary tiếng Việt export
|
||||
|
||||
## What happened
|
||||
|
||||
- `build-dictionary` reads kaikki's JSONL with `--kaikki`, hashing the stream and writing `source_sha256`, `source_rows`, `source_fetched_at` into `meta`; `--merged`/`--sources` and the undertheseanlp reader are gone; `--min-words` floor 30,000; `builder_version` 3.
|
||||
- Makefile and Dockerfile fetch the rolling kaikki URL with `curl -fL`, no checksum, atomic `.part` rename; `verify-dict` and `DICT_SHA256` removed; the web test guards Makefile == Dockerfile == builder constant and the viwiktionary path.
|
||||
- Data licence back to CC BY-SA 4.0 (current Wiktionary text); `data/LICENSE` restored from git history; NOTICE/ATTRIBUTION/README/footer credit Wiktionary tiếng Việt contributors and wiktextract/kaikki.org with the LREC citation.
|
||||
- Measured: 34,813 words (+9,421 / -1,453 vs before), all graph metrics up, bot game length unchanged. Lost words are mostly reduplicatives the vi extractor misses and deleted person-name pages; recorded in `measurement.md`.
|
||||
- Verified: Go vet/test, web check/test (174), both Docker variants with CI file checks, 404 and truncated-download failure modes. e2e not runnable locally (Chromium download fails), CI covers it.
|
||||
|
||||
## Decision
|
||||
|
||||
- Owner: viwiktionary edition only; fetch newest, no pin. Two builds a week apart may differ; the SHA-256 in `meta` identifies the bytes. Fallback if reproducibility is ever needed: keep a fetched copy as a release asset.
|
||||
|
||||
## Next steps
|
||||
|
||||
- Commit (not yet committed); watch CI's e2e job.
|
||||
|
||||
> Historical work record — not durable authority. Prefer docs/specs/ADRs for current decisions.
|
||||
Reference in new issue
Block a user