reader for both wikitext dialects; `meanings(word, ord, pos, gloss)` table; `meaning_count`/`words_with_meaning`/`source_pages` in meta, `source_rows` gone, builder_version 5; `--dump`/`--min-pages` replace `--kaikki`; attribution names the dump and the definition excerpts; 36,200 words, 96.9% with a meaning, every kaikki word kept.
Replace the pinned 2018 undertheseanlp wordlist with kaikki.org's current
wiktextract export of the Vietnamese Wiktionary, read with --kaikki. The
same authors and website, eight years fresher: 34,813 words instead of
26,845, with every graph metric up and bot game length unchanged.
The export is fetched fresh for every build and is not pinned, by the
owner's decision: kaikki keeps no dated snapshots, so a checksum would
break weekly. The builder therefore hashes the file as it streams it and
records source_sha256, source_rows and source_fetched_at in meta; the
fetch downloads to a .part name and renames on success; a truncated or
non-JSON body fails the build, and the --min-words floor rises to 30,000.
DICT_SHA256 and verify-dict are gone; the Makefile, Dockerfile and the
builder's URL constant are held in agreement by a test.
Current Wiktionary text is CC BY-SA 4.0, so the data licence returns to
4.0: data/LICENSE is restored, and NOTICE, ATTRIBUTION, README, the image
docs and the in-game footer credit Wiktionary tiếng Việt's contributors
and wiktextract/kaikki.org. builder_version becomes 3 for the changed
meta contract.
Replace the 179 MB minhqnd SQLite aggregate with the 4.8 MB
undertheseanlp/dictionary JSONL, pinned by commit and SHA-256, reading
only rows tagged "wiktionary". The two other wordlists in that file are
never read: hongocduc is GPL and would force a relicense, tudientv is an
unlicensed derivative of a commercial dictionary.
build-dictionary gains --merged and --sources (names validated, default
wiktionary), a shared finish() tail, and meta rows for the source commit
and the sources kept and excluded. The SQLite --in path, its schema
auto-detection and their tests are removed. Fixture builds now record
that they carry no upstream data instead of inheriting a licence string.
The --min-words floor moves from 40,000 to 20,000; the corpus is 26,845
words, down from 48,216, all of the loss being words absent from the
2018 Wiktionary scrape. Capitalization is not a filter.
The data licence follows the source text: CC BY-SA 3.0 Unported, which
is what vi.wiktionary.org carried in 2018. LICENSE, NOTICE, ATTRIBUTION,
the README, the image docs, the builder's meta string and the in-game
footer all name Wiktionary tiếng Việt's contributors as the authors and
undertheseanlp as the intermediary. The Makefile/Dockerfile pin test now
also checks the commit the builder stamps into the database.
One distroless image of about 25 MB carries the binary, the built frontend and
the derived dictionary. The 179 MB upstream release is downloaded in a builder
stage and never reaches the final image; the derived wordlist is copied in as
its own layer alongside its licence, attribution and notice, because CC BY-SA
4.0 applies wherever that data is distributed and an image is distribution.
FIXTURE_DICT=1 builds the same Dockerfile against the checked-in word sample,
so the image is built and smoke-tested on every push rather than only at
release. An image built only at release time is an image that breaks at release
time.
CI runs the Go suite under race detection, the frontend type check and tests,
the browser suite, and the image with its licence assertions. The wire contract
keeps its own workflow; the test steps it duplicated were removed from it.
docs/deployment.md covers configuration, the reverse-proxy settings that each
break the game in a way that looks like something else, and what a restart
costs.