5.8 KiB
phase, title, status, priority, effort, dependencies
| phase | title | status | priority | effort | dependencies |
|---|---|---|---|---|---|
| 1 | Phase 1: Read the kaikki file | completed | P1 | 3h |
Phase 1: Read the kaikki file
Overview
Replace the undertheseanlp reader in build-dictionary with one for kaikki's wiktextract
JSONL, hashing the input as it streams so the database records exactly which bytes it was
built from.
Requirements
- Functional:
--kaikki <file>reads one JSON object per line, keeps rows whoselang_codeisvi, and feedswordto the existingaccept(). - Functional: rows with any other
lang_codeare counted and rejected, not silently skipped — the file is vi-only today, and a change in that would be worth seeing in the log. - Functional:
posis tallied per value and printed with the reject counts. It never filters. - Functional: the SHA-256 of the file and the row count are computed during the single
streaming pass and written to
metaassource_sha256andsource_rows, withsource_fetched_attaken from the file's modification time. - Functional:
--mergedand--sourcesare removed along withmerged_list.goand its tests. Exactly one of--kaikki/--wordsmust be given, same rule as today. - Functional: the
--min-wordsdefault moves from 20,000 to 30,000. - Non-functional: a malformed line is an error naming the line, as today. Lines can be
long — a row carries every sense and translation — so the scanner buffer must allow
several megabytes; measure the longest line on the real file and set the cap above it
with headroom, or switch to
bufio.Reader.ReadBytes('\n')which has no cap. - Non-functional:
finish(),accept(), aliases,write,verifyunchanged.
Architecture
--kaikki <file>
│ one line at a time, bytes also fed to sha256.New()
▼
json.Unmarshal → {word, lang_code, pos}
│
├─ lang_code != "vi" → rejects["not Vietnamese-language entry"]++
▼
accept(word) (NFC, lowercase, ≥2 syllables, alphabet, phonotactics)
▼
finish() (floor 30,000, aliases, atomic write, verify)
Decode only the three fields via a struct; encoding/json ignores the rest, so the 62 MB
of senses and translations cost I/O but not memory.
sourceSpec.extra already exists for mode-specific meta rows; the kaikki provenance uses it:
source_url the kaikki URL (constant kaikkiSourceURL)
source_sha256 hex of the streamed bytes
source_rows lines decoded
source_fetched_at file mtime, RFC 3339 UTC
source_license CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Related Code Files
- Create:
server/cmd/build-dictionary/kaikki_list.go - Create:
server/cmd/build-dictionary/kaikki_list_test.go - Delete:
server/cmd/build-dictionary/merged_list.go,merged_list_test.go - Modify:
server/cmd/build-dictionary/main.go— flags, dispatch, doc comment,--min-wordsdefault,mergedProvenancecall site - Modify:
server/cmd/build-dictionary/main_test.go—fixtureSource/defaultRowsemit kaikki-shaped rows ({"word": ..., "pos": ..., "lang_code": "vi"}); the "excluded source" row becomes alang_code: "en"row
Implementation Steps
- Write
kaikki_list.go:kaikkiRow{Word, Pos, LangCode},readKaikkiList(path, maxSyllables)returning words, rejects, apostally, and akaikkiProvenance(sha256, rows, mtime). Wrap the file inio.TeeReaderintosha256.New()so hashing is free. - Add
rejectNotVietnamese rejectReason = "not a Vietnamese-language entry"tofilter.go. - In
main.go: replacemerged/sourceswithkaikkiinconfig, flags andrun();--min-wordsdefault 30,000;runFromKaikkiListlogs rejects, the POS tally (sorted, one line) and the accepted count, then callsfinish()with the provenance spec. - Delete the merged reader and its tests. Rewrite
main_test.gofixtures to kaikki rows. - Tests in
kaikki_list_test.go: avirow kept; anenrow rejected and counted;Hà Nộikept ashà nội; a word listed twice with differentposkept once; a malformed line names its number;metacarries a 64-hexsource_sha256equal tosha256sumof the fixture file, the rightsource_rows, and nosource_commit/sources_*keys. - Run against the real file (
scratchpad/kk-vi-edition.jsonlfrom this session, or a fresh download) and record the counts in phase 3.
Success Criteria
go test ./cmd/build-dictionary/passes;merged_list*.gono longer exist.- Real file builds >30,000 words (measured 34,813) and logs a POS tally.
build-dictionary --helplists--kaikki,--words,--out,--max-syllables,--min-wordsand nothing else.meta.source_sha256of a build equalssha256sumof the input file.- The fixture path (
--words ... --min-words 150) is byte-for-byte unaffected in behaviour: same 205 words fromtestdata/fixture-words.txt.
Risk Assessment
A row exceeds the scanner buffer. Signal: bufio.Scanner: token too long on the real
file. Response: pre-decided — use bufio.Reader with no line cap rather than guessing a
buffer size; measure the longest line once and note it in the phase result.
kaikki changes the row shape. word and lang_code are wiktextract's stable core fields
and unlikely to move. Signal: zero accepted words → the 30,000 floor fails the build with a
clear message. Response: read the new shape, adjust the struct; the failure mode is loud by
design.
The file arrives partially. curl -f catches HTTP errors but not a truncated body, and
a body cut exactly on a line boundary parses cleanly. Signal: a malformed final line (error
names it) or a word count under the 30,000 floor (13.8% headroom on 34,813). Response: the
fetch writes to a .part name and renames only on success (review finding, Phase 2), so an
interrupted download is never left for the next make dict; the floor covers the rest.