Commit Graph
6 Commits
Author SHA1 Message Date
tiennm99 80f216d56b feat(dict): build the corpus and word meanings from the Wikimedia viwiktionary dump
reader for both wikitext dialects; `meanings(word, ord, pos, gloss)` table; `meaning_count`/`words_with_meaning`/`source_pages` in meta, `source_rows` gone, builder_version 5; `--dump`/`--min-pages` replace `--kaikki`; attribution names the dump and the definition excerpts; 36,200 words, 96.9% with a meaning, every kaikki word kept.
2026-09-08 22:48:26 +07:00
tiennm99 557de1af94 refactor: remove unused knobs, strings and stale text
Drop what nothing uses any more: the --max-syllables flag with its
sourceSpec field, meta row and reject reason; the store's clearRejection;
the --gap CSS token; three orphaned i18n strings; two exports with no
importer; the preview npm script; the proto.yml "baseline exists" step
that has been unconditionally true since the schema landed on main; and
two root-level .gitignore entries for paths Playwright never writes.

Correct text that outlived its subject: NOTICE named a tools/ directory
that never existed, the README still pointed at the first plan and
carried a migration note for retired sources, and a few comments still
said "release" for an upstream that is a weekly export. The pin test
header now says three copies and also holds README and ATTRIBUTION to
the same URL; the Make-targets table lists clean and help.

Wire fixtures regenerated so client_hello uses the ProtocolVersion
constant and server_game_started carries the 30s turn limit.
builder_version becomes 4 for the dropped meta row.
2026-09-08 18:01:37 +07:00
tiennm99 0a5c5b06fe feat(dict): derive the corpus from kaikki.org's Wiktionary tiếng Việt export
Replace the pinned 2018 undertheseanlp wordlist with kaikki.org's current
wiktextract export of the Vietnamese Wiktionary, read with --kaikki. The
same authors and website, eight years fresher: 34,813 words instead of
26,845, with every graph metric up and bot game length unchanged.

The export is fetched fresh for every build and is not pinned, by the
owner's decision: kaikki keeps no dated snapshots, so a checksum would
break weekly. The builder therefore hashes the file as it streams it and
records source_sha256, source_rows and source_fetched_at in meta; the
fetch downloads to a .part name and renames on success; a truncated or
non-JSON body fails the build, and the --min-words floor rises to 30,000.
DICT_SHA256 and verify-dict are gone; the Makefile, Dockerfile and the
builder's URL constant are held in agreement by a test.

Current Wiktionary text is CC BY-SA 4.0, so the data licence returns to
4.0: data/LICENSE is restored, and NOTICE, ATTRIBUTION, README, the image
docs and the in-game footer credit Wiktionary tiếng Việt's contributors
and wiktextract/kaikki.org. builder_version becomes 3 for the changed
meta contract.
2026-09-08 17:17:58 +07:00
tiennm99 39ef45730f feat(dict): derive the corpus from undertheseanlp's wiktionary rows
Replace the 179 MB minhqnd SQLite aggregate with the 4.8 MB
undertheseanlp/dictionary JSONL, pinned by commit and SHA-256, reading
only rows tagged "wiktionary". The two other wordlists in that file are
never read: hongocduc is GPL and would force a relicense, tudientv is an
unlicensed derivative of a commercial dictionary.

build-dictionary gains --merged and --sources (names validated, default
wiktionary), a shared finish() tail, and meta rows for the source commit
and the sources kept and excluded. The SQLite --in path, its schema
auto-detection and their tests are removed. Fixture builds now record
that they carry no upstream data instead of inheriting a licence string.
The --min-words floor moves from 40,000 to 20,000; the corpus is 26,845
words, down from 48,216, all of the loss being words absent from the
2018 Wiktionary scrape. Capitalization is not a filter.

The data licence follows the source text: CC BY-SA 3.0 Unported, which
is what vi.wiktionary.org carried in 2018. LICENSE, NOTICE, ATTRIBUTION,
the README, the image docs, the builder's meta string and the in-game
footer all name Wiktionary tiếng Việt's contributors as the authors and
undertheseanlp as the intermediary. The Makefile/Dockerfile pin test now
also checks the commit the builder stamps into the database.
2026-09-08 16:37:34 +07:00
tiennm99 b2cad42b0c build: package the game as a container image and wire CI
One distroless image of about 25 MB carries the binary, the built frontend and
the derived dictionary. The 179 MB upstream release is downloaded in a builder
stage and never reaches the final image; the derived wordlist is copied in as
its own layer alongside its licence, attribution and notice, because CC BY-SA
4.0 applies wherever that data is distributed and an image is distribution.

FIXTURE_DICT=1 builds the same Dockerfile against the checked-in word sample,
so the image is built and smoke-tested on every push rather than only at
release. An image built only at release time is an image that breaks at release
time.

CI runs the Go suite under race detection, the frontend type check and tests,
the browser suite, and the image with its licence assertions. The wire contract
keeps its own workflow; the test steps it duplicated were removed from it.

docs/deployment.md covers configuration, the reverse-proxy settings that each
break the game in a way that looks like something else, and what a restart
costs.
2026-09-05 14:18:18 +07:00
tiennm99 4b4afea679 feat(dictionary): build Vietnamese wordlist from upstream dictionary
Reduce the 179 MB minhqnd/dictionary SQLite release to a ~3 MB game
wordlist: 48,216 Vietnamese words of two or more syllables, indexed by
first and last syllable with an out-degree table for dead-end detection.

Source schema is auto-detected rather than hardcoded, since it is someone
else's release artifact; explicit flags override it and are validated
against the real tables, because SQLite silently reads an unknown
double-quoted column as a string literal.

Accept a word only if every syllable fits Vietnamese phonotactics. An
alphabet check is not enough: "credit card" and "come out" use only
letters Vietnamese has, and the multilingual source tags them as
Vietnamese. Onset matching backtracks so the gi digraph does not swallow
the nucleus of common words like "gi", "gin" and "gi" (rust).

Record accepted spelling variants in an alias table rather than solving
tone placement at runtime. Tone shifting applies only to open oa/oe/uy
syllables, since "hoan" and "hoai" have a single correct spelling, and
"qu" is a consonant onset. The i/y alternation uses an onset allowlist
plus explicit pairs, because it is lexical rather than productive.
Variants are generated as a cross product over syllables so a word with
two variable syllables still offers the fully modern spelling.

Build to a temporary file and rename only after commit, so a failed run
cannot leave an empty database where a good one was, then re-open the
result and verify its invariants on disk.

Licensing: the derived data is CC BY-SA 4.0 and stays a separate artifact
from the Apache-2.0 code, loaded at runtime and never embedded. Ships
NOTICE, data/LICENSE and an attribution file recording every change.
2026-09-04 16:25:43 +07:00