Load the whole dictionary into maps at Open and close the database before
Open returns. The plan called for per-lookup SQLite, but a round-trip
benchmarked at 55us against 8.9ns for a map hit, and the hard bot in a
later phase explores hundreds of candidates inside a 150ms budget. The
in-memory form is also simpler: no connection pool, no prepared
statements, no tail latency. Costs ~70ms and ~7.8MB at startup.
Resolve returns the canonical word, never the spelling the player typed.
Canonicalization moves either end: about half the aliases differ in the
last syllable and more than a third in the first, so "sy hai" resolves to
"si hai". FirstSyllable and LastSyllable report the canonical's ends, and
the engine must chain on those or it will reject legal moves.
WordsStartingWith yields an iterator rather than the backing slice. A
caller could otherwise sort, shuffle or append into dictionary state:
verified that a write landed in the store, that most buckets have spare
capacity for append to scribble into, and that concurrent callers race.
Open validates what it loaded against the builder's recorded word count,
cross-checks every out-degree against the words actually indexed, and
rejects orphan aliases. A truncated database otherwise opens cleanly and
the server starts, rejects every word, and fails every room creation.
RandomOpeningWord picks from a pre-sorted slice by binary search instead
of rebuilding a filtered copy per call, cutting room creation from 374us
and 720KB to 18ns and no allocation.
Escape the database path when building the URI: SQLite reads # as a
fragment delimiter, so an unescaped path opens a different file and
reports a misleading schema error.
The store does not log. A library writing to the global logger fights
structured logging later, and the caller has WordCount, AliasCount and
License to state the CC BY-SA attribution itself.
Reduce the 179 MB minhqnd/dictionary SQLite release to a ~3 MB game
wordlist: 48,216 Vietnamese words of two or more syllables, indexed by
first and last syllable with an out-degree table for dead-end detection.
Source schema is auto-detected rather than hardcoded, since it is someone
else's release artifact; explicit flags override it and are validated
against the real tables, because SQLite silently reads an unknown
double-quoted column as a string literal.
Accept a word only if every syllable fits Vietnamese phonotactics. An
alphabet check is not enough: "credit card" and "come out" use only
letters Vietnamese has, and the multilingual source tags them as
Vietnamese. Onset matching backtracks so the gi digraph does not swallow
the nucleus of common words like "gi", "gin" and "gi" (rust).
Record accepted spelling variants in an alias table rather than solving
tone placement at runtime. Tone shifting applies only to open oa/oe/uy
syllables, since "hoan" and "hoai" have a single correct spelling, and
"qu" is a consonant onset. The i/y alternation uses an onset allowlist
plus explicit pairs, because it is lexical rather than productive.
Variants are generated as a cross product over syllables so a word with
two variable syllables still offers the fully modern spelling.
Build to a temporary file and rename only after commit, so a failed run
cannot leave an empty database where a good one was, then re-open the
result and verify its invariants on disk.
Licensing: the derived data is CC BY-SA 4.0 and stays a separate artifact
from the Apache-2.0 code, loaded at runtime and never embedded. Ships
NOTICE, data/LICENSE and an attribution file recording every change.