6.0 KiB
phase, title, status, priority, effort, dependencies
| phase | title | status | priority | effort | dependencies | |
|---|---|---|---|---|---|---|
| 3 | Phase 3: Meanings in the database | completed | P1 | 4h |
|
Phase 3: Meanings in the database
Overview
Persist the senses phase 1 extracts in a meanings table, load them in dictionary.Store
behind a Meanings(word) method, and let the fixture word list carry meanings so tests and
the e2e suite have some.
Requirements
- Functional: schema gains
CREATE TABLE meanings ( word TEXT NOT NULL, ord INTEGER NOT NULL, pos TEXT NOT NULL, -- Vietnamese part-of-speech label, '' when the heading was unmapped gloss TEXT NOT NULL, PRIMARY KEY (word, ord) ) WITHOUT ROWID;ordis the 0-based sense order from the page. No foreign key pragma;verify()checks the join instead, as it does for aliases. (Validation session 1:poscolumn.) - Functional:
finish()takes the meanings map besidewords;writeToinserts them in the same transaction.metagainsmeaning_count(rows) andwords_with_meaning. - Functional:
verify()adds: no meaning row whose word is missing fromwords; no empty gloss; no gloss over 200 runes;words_with_meaningis at least 60% ofword_countfor a--dumpbuild (fixture builds are exempt — the check is passed a flag, or keyed on the source table prefix). - Functional:
dictionary.Storeloadsmeaningsintomap[string][]Senseordered byord, wheretype Sense struct { Pos, Gloss string }is exported from thedictionarypackage, and exposesMeanings(word string) []Sensereturning nil for a word with none.Resolvefirst: the caller passes a canonical word, as it does forFirstSyllable. Startup log gains the meanings count next towords. - Functional:
--wordsaccepts an optional meaning column:word<TAB>sense<TAB>sense…, where a sense ispos|glossor justgloss(the pipe never survives stripping, so it is a safe separator). Lines without a tab have no meaning.testdata/fixture-words.txtgains hand-written meanings for the words the e2e suite plays (phase 5 says which) and for at least half the list, so the fixture database exercises the same store paths as the real one. - Non-functional:
Storememory grows by the text size, a few MB. Startup time unchanged in practice (one more ordered scan). - Non-functional:
builder_version = "5"(set in phase 1; this is the contract it names).
Architecture
build-dictionary dictionary.Store
words map[string]entry ─┐ words map[string]wordInfo
meanings map[string][]sense ─┼─► sqlite ─► meanings map[string][]Sense
aliases map[string]string ─┘ Meanings(word) []Sense
The store's validate() already cross-checks word_count; add the same for
meaning_count so a database whose meanings table was truncated on disk is refused at
startup rather than served silently without meanings.
Related Code Files
- Modify:
server/cmd/build-dictionary/main.go—finish,write,writeTo,verify,runFromWordList(tab-separated senses),sourceSpecunchanged - Modify:
server/cmd/build-dictionary/main_test.go— meanings written and verified; a fixture list with tabs; averifyfailure on an orphan meaning row - Modify:
server/internal/dictionary/store.go—meaningsfield,loadMeanings,Meanings(),validatecount check,MeaningCount() - Modify:
server/internal/dictionary/store_test.go— hand-built databases gain the table;Meaningsreturns ordered senses and nil; a mismatchedmeaning_countis refused - Modify:
server/cmd/noitu-server/main.go— startup log line - Modify:
testdata/fixture-words.txt— meanings column; header comment documents the format - Modify:
web/e2e/fixture-dictionary.js— it reads the list one word per line (fixture-dictionary.js:15-19); cut each line at the first tab before trimming, or the e2e graph gains words that are reallyword<TAB>sensestrings (validation session 1)
Implementation Steps
- Schema, insert, meta and
verifyin the builder; tests first for the orphan-row and empty-gloss failures. runFromWordListtab parsing; a test that a line with two tabs yields two senses in order,danh từ|…splits into label and gloss, a cell without a pipe has an empty label, and a line without a tab yields none. Updatefixture-dictionary.jsin the same step.- Store: load, method, validate; tests.
- Fixture list: add meanings. Rebuild
data/fixture.db(make fixture-dict) and start the server against it; the log shows the count. - Rebuild the real database from phase 1's dump and confirm
verifypasses the 60% rule.
Success Criteria
go test ./cmd/build-dictionary ./internal/dictionary -racegreen.sqlite3 data/noitu.db 'SELECT COUNT(*) FROM meanings'is above 40,000 andwords_with_meaningabove 60% of words (phase 6 records the actual).store.Meanings("học sinh")on the real database returns a non-empty ordered slice whose first sense hasPos == "danh từ"; on a word with no#line it returns nil.npm run test:e2estill builds its word graph from the fixture list correctly (no tab-carrying "words").- The fixture database has meanings for the e2e words and the server starts on it.
Risk Assessment
The 60% rule is a guess. Research counted 5,764 of 43,013 pages with no POS marker, and
some of those have no # line either; the true coverage is unknown until phase 1 runs.
Signal: verify fails on the real dump with a coverage in the 50s. Response: measure, set
the floor a comfortable margin under the measurement, and record why in the code comment.
The rule exists to catch a stripper that suddenly returns nothing, not to demand quality.
Fixture meanings drift from fixture words. A word renamed in the list loses its meaning silently. Signal: an e2e assertion on a meaning fails. Response: the builder's fixture path logs words without meaning; keep the list short and hand-checked.