proto/noitu/v1/game.proto is the single source of truth for every WebSocket message. buf generates Go types into server/gen and JavaScript types into web/src/lib/proto; both trees are committed so building needs no codegen toolchain. The Go suite emits binary fixtures into proto/testdata and the JavaScript suite decodes the same bytes, so the two generated clients are checked against one artifact rather than against each other's assumptions. CI lints the schema, rejects breaking changes against main, and fails when the committed generated trees drift from the schema. game.NumRejectReasons and game.NumEndReasons let the mapping tests prove every engine reason has a wire value without guessing where the enum ends.
8.7 KiB
title, status, phase, priority, effort, dependencies
| title | status | phase | priority | effort | dependencies |
|---|---|---|---|---|---|
| Phase 1: Foundations and Data Pipeline | done | 1 | P1 | 3d |
Phase 1: Foundations and Data Pipeline
Overview
Stand up the repo skeleton and produce data/noitu.db — the derived, normalized Vietnamese
wordlist of words with at least 2 syllables — from the upstream minhqnd/dictionary
SQLite release, with full CC BY-SA 4.0 compliance. Nothing downstream can be built or
tested without this data.
Requirements
Functional
- Reproducible build: upstream
dictionary.db→data/noitu.db, runnable by any contributor - Output contains only Vietnamese (
lang_code = 'vi') entries of 2 or more space-separated syllables - All words NFC-normalized and lowercased
- Tone-placement aliases resolved (
hoà→hòa,thuý→thúy,quí→quý, …) - Precomputed
first/lastsyllable columns and out-degree table for O(1) engine lookups - Build fails loudly if entry count falls below a floor (40,000) or any row has fewer than 2 syllables
--max-syllablesflag available (default: no cap) so a phrase cap can be applied later without a code change
Non-functional
- Output DB is a few MB, opens read-only, no writes at runtime
- License artifacts are correct and consistent across all five locations
Architecture
Source: github.com/minhqnd/dictionary release
v2.0.0 → dictionary.db,
179 MB (verified via the GitHub API; not in git; 357k+ entries, 1,500+ language pairs).
Code MIT, data CC BY-SA 4.0.
We do not republish a derived dictionary — every build downloads this asset and derives
locally. See plan.md → Data Distribution for the CI and Docker consequences.
Derived schema (data/noitu.db):
CREATE TABLE words (
word TEXT PRIMARY KEY, -- normalized, e.g. "pháp luật"
first TEXT NOT NULL, -- first syllable, "pháp"
last TEXT NOT NULL, -- last syllable, "luật"
syllables INTEGER NOT NULL -- 2, 3, 4, …
) WITHOUT ROWID;
CREATE INDEX idx_words_first ON words(first);
-- Out-degree per syllable: how many words start with it. Drives bot heuristics
-- and instant dead-end detection.
CREATE TABLE syllables (
syllable TEXT PRIMARY KEY,
out_degree INTEGER NOT NULL
) WITHOUT ROWID;
-- Accepted spelling variants that map to a canonical word. Built offline so the
-- runtime never has to reason about Vietnamese tone placement.
CREATE TABLE aliases (
variant TEXT PRIMARY KEY,
canonical TEXT NOT NULL REFERENCES words(word)
) WITHOUT ROWID;
CREATE TABLE meta (key TEXT PRIMARY KEY, value TEXT);
-- rows: source_url, source_license, built_at, word_count, builder_version
Why an alias table instead of a runtime normalizer: Vietnamese old-style vs new-style
tone placement (hoà vs hòa) produces genuinely different codepoint sequences. A runtime
algorithm to re-place tone marks is fiddly and easy to get subtly wrong; generating both
variants once at build time is data, cheap to fix, and trivially testable. KISS.
Pipeline:
dictionary.db ──► filter lang_code='vi'
──► NFC normalize (golang.org/x/text/unicode/norm) + lowercase + collapse spaces
──► keep len(strings.Fields(w)) >= 2 (and <= --max-syllables when set)
──► drop entries containing digits, latin-only tokens, or punctuation
──► dedupe
──► generate tone-placement variants → aliases
──► compute first/last/syllables + out_degree
──► write noitu.db + meta
──► assert: count >= 40000, every row >= 2 syllables, no orphan aliases
Note on the graph: allowing 3+ syllable words does not change the graph model — an edge
still runs first ──word──► last; it may simply span more syllables in between. Out-degree,
dead-end detection, and the engine's lookups are unaffected.
Related Code Files
- Create:
server/cmd/build-dictionary/main.go— CLI:--in dictionary.db --out data/noitu.db [--max-syllables N] - Create:
server/cmd/build-dictionary/filter.go— vi + ≥2-syllable + junk filters - Create:
server/cmd/build-dictionary/aliases.go— tone-placement variant generation - Create:
server/cmd/build-dictionary/main_test.go— filter/alias unit tests with fixtures - Create:
data/LICENSE— CC BY-SA 4.0 full text - Create:
data/ATTRIBUTION.md - Create:
NOTICE - Create:
README.md(replace stub) — quickstart, architecture, License section - Create:
.gitignore—data/*.db(covers both the 179 MB upstream and the derived DB),web/node_modules,web/build, server binaries - Create:
Makefile(orTaskfile.yml) —make fetch-dict,make dict,make server,make web,make test - Create:
server/go.mod—module github.com/tiennm99dev/noitu/server, Go 1.25+ (raised by modernc.org/sqlite; local toolchain is 1.26.5) - Modify: none (repo is empty apart from
LICENSE+README.md)
Implementation Steps
.gitignore,Makefile,server/go.mod(go mod init), directory skeleton per plan layout.- Write
data/LICENSE(verbatim CC BY-SA 4.0) anddata/ATTRIBUTION.mdnaming: source repo + release URL, upstream sources (Wiktionary,vntk/dictionary), license URL, and the explicit modification list fromplan.md. - Write
NOTICE+ README License section stating the split: Apache-2.0 code / CC BY-SA 4.0 data, and that the two are distributed as separate artifacts. server/cmd/build-dictionary: open upstream DB read-only viamodernc.org/sqlite, streamviwords.- Normalization:
norm.NFC,strings.ToLower,strings.Fields→ join with single space. - Filters: at least 2 fields (and ≤
--max-syllableswhen set); reject any token containing a digit, ASCII-only letters, or punctuation. - Alias generation: for each canonical word, emit old-style tone-placement variants for the
oa/oe/uynuclei; skip a variant if it collides with a different canonical word (log collisions). - Compute
first,last,syllables,out_degree; write output DB in one transaction; writemetarows. - Assertions: word count ≥ 40,000; 100% of rows ≥ 2 syllables; every alias resolves; fail non-zero otherwise.
make fetch-dictdownloads the 179 MB asset from the pinned release URL intodata/(resumable, checksum-logged);make dictderives from it. Both documented in the README as one-time setup.- Unit tests on fixtures: normalization, ≥2-syllable filter (including a 3- and a 4-syllable entry accepted and a 1-syllable rejected), junk rejection, alias generation, collision handling.
Success Criteria
make fetch-dict && make dictproducesdata/noitu.dbfrom the upstream releasesqlite3 data/noitu.db "SELECT COUNT(*) FROM words"≥ 40,000- Zero rows where
wordhas fewer than 2 syllables; 3- and 4-syllable entries present hoà,thuý,quíeach resolve throughaliasesto a canonical word present inwordsmetarecords source URL, license, build timestamp, word countdata/LICENSE,data/ATTRIBUTION.md,NOTICE, README license section all present and mutually consistentgo test ./cmd/...greendata/*.dbis git-ignored
Risk Assessment
| Risk | Signal | Response |
|---|---|---|
| Upstream schema differs from README-documented shape | Query errors on first run | Inspect actual tables with .schema first; adapt the extraction query — the pipeline stages after extraction are schema-independent |
| Fewer than 40k clean ≥2-syllable entries | Count assertion fails | Union with Viet74K filtered to ≥2 syllables; add it as a second --in source and extend ATTRIBUTION.md |
| 179 MB download is slow or the release URL moves | make fetch-dict fails or stalls |
URL is pinned to release v2.0.0, download is resumable, and the derived DB is cached locally so the fetch is a one-time cost per contributor |
| Allowing 3+ syllables admits multi-word phrases that are not really words | Playtest complaints | --max-syllables flag already present; applying a cap is a rebuild, not a code change |
| Upstream data contains proper nouns / non-words that make the game feel wrong | Playtest complaints in later phases | Add a denylist file consumed by the builder; regenerate — no code change needed |
| Alias generation creates false accepts (a variant that is a different real word) | Collision log non-empty | Collisions are skipped by design and logged; review the log before release |
| CC BY-SA obligations misread | License review | Attribution + share-alike on the data artifact only; code untouched. NOTICE states the boundary explicitly |