docs(plans): research, brainstorm and health-scan reports with synthesis brief

This commit is contained in:
tiennm99 committed 2026-09-21 00:30:10 +07:00
1 parent dd3b46397f
commit a5052fca3f
4 files changed
+896

No files matched your search

@@ -0,0 +1,287 @@
---
title: "noitu — improvement directions"
date: 2026-09-21
type: brainstorm
status: advisory, read-only
scope: gameplay/bot, product, UX, dictionary, architecture/ops, DX, docs
---
# noitu — improvement directions
Method: read `README.md`, `docs/deployment.md`, `proto/noitu/v1/game.proto`, `server/internal/{game,bot,dictionary,vietnamese,wsapi}`, `server/cmd/{noitu-server,build-dictionary}`, `web/src` (stores, routes, components, `i18n/vi.js`), `Makefile`, `Dockerfile`, both CI workflows, the three 2026-09-10 UI/UX reports, the cleanup review, and every journal tail. Everything below cites a file I read. Speculation is labelled.
## Baseline: what is already true (so nothing here repeats it)
The three 2026-09-10 UI/UX reports are **mostly delivered** — verified in the tree, not assumed:
- Tokens: `--border-strong`, `--accent: #12692f`, `--space-1..8`, `--text-1..9`, `--radius-pill` all exist in `web/src/app.css:1-130`. The three AA failures are closed.
- Clock: `mine` / `stalled` / `spoken` derivations in `CountdownRing.svelte:14-51` — the false-panic red on the bot's turn, the stalled-socket clock and the screen-reader time marks are all done.
- `interactive-widget=resizes-content` in `app.html:13-16`; skip link at `+layout.svelte:18`; per-route `<svelte:head>` on all three routes; `reconnectNow()` at `ws/connection.svelte.js:53`; clipboard fallback plus visible invite URL at `RoomCodePanel.svelte:50,104`; `unready: 'Bỏ sẵn sàng'` (`vi.js:120`); lobby-local error block (`Lobby.svelte:192-203`); invite-path nickname gate (`online/+page.svelte:93,378`); `RoomCodePanel compact` kept in the post-game lobby (`Lobby.svelte:88-94`).
Two findings from those reports are **still open** and are folded into C1 below: the chat toggle still sits under the whole chain (`online/+page.svelte:362` renders `ChatPanel` after `GameBoard`), and it is only collapsible while `playing && !wide`, so lobby chat has no unread badge.
The cleanup review (`code-review-260908-2302`) is fully applied; `deadcode`/`staticcheck` were clean at `f00d0ef` and nothing since has added a `TODO`.
---
## A. Gameplay, rules and bot
### A1. The bot plays to survive; the human plays for points
**Problem.** `game/engine.go:265-307` prices a word on chain length, syllable count, speed and link rarity. The bot never sees any of it: `bot.Board` (`bot/bot.go:50-55`) exposes only `LegalMoves`, `Used`, `WordsStartingWith`, `LastSyllable`, and `hard.Choose` maximises "how few replies the opponent has" (`strategy_hard.go:137-166`). So Hard plays a *strangling* game while the player is scored on a *fluent* one. The ladder is also coarse — three constants (`hardKillRate 0.85`, `searchDepth 4`, `branchCap 12`) and one latency test (`realcorpus_test.go:149-202`, worst case must stay under 150ms).
**Options.** (a) Retune the three constants and add a fourth "Chuyên gia" difficulty at higher depth — cheapest, but the fourth tier is the same personality with more nodes. (b) Give the search a small points term, so tighter *and* longer/rarer words are preferred; the bot then reads as a player rather than a trap-setter, and its final score becomes comparable to yours. (c) Adaptive difficulty that widens `branchCap`/`searchDepth` from the player's recent results — hides what the difficulty picker means and makes bug reports irreproducible.
**Rec.** (b), then (a) if the ladder still feels flat. Keep `Board` narrow — add `OutDegree(syllable)` only, and score inside the strategy; do not hand the strategy the Engine. **Size S-M.** **Depends on:** nothing.
### A2. Points are a number with no visible cause
**Problem.** `PlayedWord.points` (`game.proto:164-178`) is a single integer. Four terms went into it and the player sees none of them, so the one mechanic that rewards reaching for a three-syllable compound or a rare link teaches nothing. The client cannot re-derive it — it has no wordlist, by design.
**Options.** (a) Additive breakdown on `PlayedWord` (four small fields, or a `repeated PointPart{kind, value}`) plus a one-line "+15 hiếm" in the chain row. (b) A static rules/scoring screen in `vi.js` — no protocol change, but it is a table nobody reads mid-game. (c) Leave it.
**Rec.** (a) with a `repeated PointPart`, plus a short scoring paragraph on the help screen from C2. A repeated message means a future fifth term is not a wire break. **Size S.** **Depends on:** proto regen (routine here).
### A3. A dead end still costs the trapped player 30 seconds of nothing
**Problem.** `engine.go:250-258` deliberately leaves a dead end to the clock: the player handed an unanswerable syllable must time out. That is the right *rule* — but they sit through `NOITU_TURN_LIMIT` (30s) with nothing to do, and only afterwards learn from `PlayerEliminated.suggestions` that the position was empty. The engine knows instantly (`HasLegalMove`, `:340-347`); the bot gets to use that shortcut (`NoMove`, `:373-379`) and the human does not.
**Options.** (a) A `ClaimNoMove` client message: the server answers truthfully — if the position is dead, eliminate immediately with `EndNoLegalMove`; if not, the claim was wrong and costs the remaining speed bonus, or nothing at all. (b) Auto-resolve: the server settles a dead end without asking, which removes the "did I miss something?" moment that makes the game interesting. (c) Leave it; `Resign` exists, and `session.go:447-456` restricts it to the player to act, which is exactly the right shape to copy.
**Rec.** (a), modelled on `Resign`'s authorization. Converts up to 30s of dead air per elimination into one tap, and cannot be abused because the server verifies the claim. **Size S.** **Depends on:** proto addition.
### A4. One ruleset, one clock, no room options
**Problem.** `NOITU_TURN_LIMIT` is deployment-wide (`main.go:113`) and `RoomState` (`game.proto:343-366`) carries `max_players`/`min_players`/`grace_ms` but no game options. A room of four experienced players and a room of two beginners get the same 30s. `vietnamese.MinSyllables = 2` is a compile-time constant.
**Options.** (a) Nothing — one clock is one code path, and the README leans on that. (b) A `GameOptions` message on `CreateRoom`/`StartGame`, server-clamped, echoed in `RoomState` so the lobby draws what the server allows — the same trick `max_players` already uses. (c) Presets (chậm/thường/nhanh), which is (b) with the knobs hidden.
**Rec.** (c) — three presets, turn limit alone. Clamping an enum server-side is one switch; a free integer is a validation surface and an invitation to a 3-second room. **Size S-M.** **Depends on:** proto; keep the bot path on the same constant so "one implementation" survives.
---
## B. Product: progression, social, discovery
### B1. A room can only be reached by a code somebody already gave you
**Problem.** `hub.joinRoom` (`hub.go:126-137`) resolves a 6-character code and nothing else. There is no queue, no listing, no "play someone now". A first-time visitor with no friend online can only play the bot. That is the biggest ceiling on the online half of the product.
**Options.** (a) Quick-match: one in-memory FIFO in the hub keyed by nothing at all; two waiters become a room and the existing `createRoom`/`joinRoom` path does the rest. No moderation surface, no listing to abuse. (b) A public room list — needs a name, a listing message, and immediately raises "who can see my room". (c) Invite-by-link only, as today.
**Rec.** (a). It reuses the whole room machinery and adds one map plus a timeout; the client needs one button and one waiting state. **Size M.** **Depends on:** nothing. Second-order: a lone waiter must be offered the bot after N seconds, or quick-match becomes a dead end of its own.
### B2. Identity lasts exactly as long as a seat
**Problem.** `PlayerSlot.wins` resets when the seat is vacated (`game.proto:208-213`), nicknames live in `localStorage` (`settings.svelte.js:98,110-114`), and the solo record is a per-difficulty number in the same place (`:98,141-155`). Nothing survives a room, a device change or a restart — `docs/deployment.md:136-152` states this as a design position.
**Options.** (a) Keep it: honest, zero ops, and the game is fine as a drop-in. (b) An anonymous device token minted server-side plus a small read-write SQLite holding profile and aggregate stats, kept strictly separate from the read-only CC BY-SA dictionary file (licence hygiene matters — `dictionary.Open` opens `mode=ro` and closes the handle, `store.go:88-97`). (c) Real accounts/OAuth — a login wall on a casual word game.
**Rec.** (a) until D1 and D2 land. Progression without observability is a feature you cannot tell is working. When it comes, (b): two files, two licences, one binary. **Size L.** **Depends on:** D1.
### B3. A room turns away everyone once a game starts
**Problem.** `room.go:571` refuses a joiner with `game_in_progress` even when seats are free, and `README.md:83-86` defends it — correctly, for *players*. But an eliminated player already watches (`GameBoard.svelte:155`, `vi.js:140` `spectating`), so the rendering path for a non-acting participant exists. A friend who arrives two minutes late is simply told no.
**Options.** (a) Spectator seats: join as an observer, receive `TurnUpdate`/`ChatMessage`, seated in the lobby when the game ends. (b) A waiting list: the joiner is held and seated at the lobby return. (c) Leave it.
**Rec.** (b) first — no new per-recipient rendering, no spectator concept in `PlayerSlot`; the joiner waits in a pending list the room drains on lobby return. (a) is the better product and roughly triple the work. **Size M.** **Depends on:** proto (a pending/waiting flag in `RoomState`).
### B4. Solo has no reason to come back tomorrow
**Problem.** The only solo progression is `bestScores` per difficulty (`settings.svelte.js:141-155`), and every game opens on `RandomOpeningWord` (`store.go:409-418`), so no two runs are comparable and none is shareable.
**Options.** (a) A daily seeded solo challenge: the server derives the opening word and the bot RNG seed from the date, everyone gets the same board, and the transcript export (`lib/history-export.js`) becomes a shareable result. Fits the philosophy exactly — server-authoritative, nothing extra in the client. (b) Streaks/achievements — needs B2's persistence. (c) Nothing.
**Rec.** (a), without a leaderboard in v1. A shared board plus a shareable transcript is most of the loop; a leaderboard needs identity and anti-cheat the current model does not support. **Size M.** **Depends on:** a deterministic RNG path into `bot.New` (already injectable, `bot.go:70`) and a date-seeded opening pick.
---
## C. UX
### C1. What is left over from the 2026-09-10 reports
**Problem.** Two findings survive (verified above): the chat toggle and its unread badge sit below `ChainHistory`, which grows a row per turn (`online/+page.svelte:362`), so on a phone the badge is effectively unreachable mid-game; and `collapsible={playing && !wide}` means the lobby panel is never collapsible, so lobby chat has no badge at all while the log itself is below the fold.
**Options.** (a) Hoist a small unread pill into `GameBoard`'s `.top` row beside the connection badge, and drop `playing &&` from `collapsible`. (b) Reorder `.pane.talk` before `ChainHistory` in the stacked layout via `order` — puts the conversation above the game, which is wrong on a phone. (c) Leave it.
**Rec.** (a), both halves. **Size S.** **Depends on:** nothing.
### C2. Nothing anywhere states the rules
**Problem.** The landing page is a tagline, a nickname field, a difficulty picker and two buttons (`routes/+page.svelte:19-31`); `vi.js:10` is `'Trò chơi nối từ tiếng Việt'`. Grep finds no rules copy, no help screen, no scoring explanation anywhere in `web/src`. A player who does not already know *nối từ* learns the two-syllable minimum by having a word rejected, and learns scoring never.
**Options.** (a) A short `/help` route: the chain rule, the 2-syllable minimum, the clock, elimination, and the scoring terms from A2 — copy in `vi.js` plus one route. (b) A first-turn inline coach on `/play` — better conversion, more state, easy to get in the way on the second game. (c) A README link in the footer, which already carries the attribution (`AttributionFooter.svelte`).
**Rec.** (a), plus one line of rule text under the syllable on the *first* turn only. **Size S.** **Depends on:** A2 if scoring is to be explained truthfully.
### C3. "Không tìm thấy từ này trong từ điển" is a dead end on a near miss
**Problem.** `REJECT_REASON_NOT_IN_DICTIONARY` (`game.proto:37`) comes back whenever `dict.Resolve` fails (`engine.go:212-215`). In Vietnamese a miss is very often one tone mark. The builder already generates an alias table for known spelling variants (`build-dictionary/main.go:327-355`) and `Resolve` consults it (`store.go:345-353`), but anything outside that table is a flat refusal. The UI now shows the rejected word back (the 2026-09-10 P2-3 fix), which helps the eye and not the vocabulary.
**Options.** (a) Server-side diacritic-insensitive near-match: strip marks (`norm.NFD` then drop `Mn`, the technique already in `build-dictionary/filter.go:83-99`), and when exactly one real word matches the stripped form, say "ý bạn là …?" without accepting the move. (b) Full edit-distance suggestions — a hint engine that hands out words the player did not know, which changes the game. (c) Nothing.
**Rec.** (a), restricted to *diacritics only* and to a *unique* match. It corrects typing, not vocabulary. Needs a second index in `Store` populated at `Open` (a `map[stripped][]word`, a few MB), not a scan. **Size S-M.** **Depends on:** proto (a `suggestion` field on `MoveRejected`).
---
## D. Architecture and operability
### D1. Nothing persists, and the restart cost is a product decision, not just an ops one
**Problem.** Rooms are in-memory maps on the hub (`hub.go:45-50`); `docs/deployment.md:136-140` says a restart ends every live game and tells players so (`hub.shutdown`, `hub.go:213-231` — genuinely good behaviour). But it means you cannot deploy during the evening, and every feature in section B is blocked behind it.
**Options.** (a) Keep it and add a *drain* mode: stop accepting new rooms, let live ones finish (already bounded by the turn clock), then exit — turning "deploy when quiet" into "deploy whenever". (b) Snapshot rooms to disk and restore: engine state is fully copyable (`Snapshot`, `engine.go:521-546`) but sockets are not, so restore means every client re-handshaking into a restored room; large and subtle. (c) A read-write SQLite for durable data only (profiles, stats, disputes), live rooms staying in memory.
**Rec.** (a) now, (c) when B2/E3 need it, (b) probably never. Drain is one flag on the hub plus a `/readyz` that goes unhealthy while draining. **Size S** for (a), **M** for (c). **Depends on:** nothing.
### D2. You cannot answer a single question about how the game is actually played
**Problem.** Twelve `slog` call sites in the whole `wsapi` package (4 in `room.go`, 6 in `session.go`, 2 in `server.go`), no metrics, no `/metrics`, no expvar, no game-event stream. Unanswerable today: how many rooms are live, how often a word is rejected and for which reason, which syllables end games, how long games last, and — the important one — **which words players type that the dictionary does not have**. That last question is the input to every decision in section E.
**Options.** (a) `expvar` counters plus `GET /debug/vars`: stdlib, zero dependencies, matching the project's dependency discipline (`server/go.mod` is tiny). (b) A Prometheus client and `/metrics`: better tooling, one dependency and a scrape endpoint to protect. (c) A structured `slog` event per terminal game event and per rejection, read from container logs.
**Rec.** (a) for counters and (c) for the rejection stream — specifically one `slog.Info("word_rejected", "reason", …, "word", …)` line, which costs nothing and is the corpus feedback loop. Gate (b) on actually having a Prometheus. **Size S.** **Depends on:** nothing. Second-order: a typed word is user content — cap the length, log only the normalized form, and say so in the deployment doc.
### D3. Behind a proxy, one abuser can lock every player out of joining
**Problem.** `clientIP` uses `RemoteAddr` and deliberately ignores `X-Forwarded-For` (`server.go:150-162`), and `docs/deployment.md:113-119` states the consequence plainly: behind a reverse proxy — the *supported* deployment — every player shares one bucket at `joinsPerSecond = 1, joinBurst = 5` (`session.go:49-50`). One script drains the shared bucket and everybody else gets `too_many_attempts`. The reasoning for not trusting XFF is right; the resulting default is not safe in the supported topology.
**Options.** (a) `NOITU_TRUSTED_PROXY_CIDRS`: when `RemoteAddr` falls inside it, take the right-most XFF hop; otherwise `RemoteAddr` as today. Explicit, unset by default, no behaviour change for direct deploys. (b) A per-room failed-join counter — defends the room-code secret but not the shared bucket. (c) Leave it and raise the limit, trading a lockout for a weaker brute-force defence on a 6-character code.
**Rec.** (a) and (b) together; they defend different things, and (a) restores the limiter's stated purpose. **Size S.** **Depends on:** nothing.
### D4. Operational hygiene: no version stamp, no readiness, no capacity number
**Problem.** `/healthz` returns `"ok"` (`server.go:55-58`) and `docs/deployment.md:122-132` correctly calls it liveness-only; the binary carries no version stamp (no `-ldflags -X` in the `Dockerfile` or `Makefile`), so "is the new image live?" cannot be answered without playing a game; and nothing states how many concurrent rooms one binary handles — each room is a goroutine plus an engine plus timers (`room.go:326-486`) and the dictionary is ~7MB of shared maps (`store.go:8-14`), so the number is probably large, but it is unmeasured (speculation).
**Options.** (a) `/readyz` returning JSON with version, dictionary word count and live-room count, version injected at link time. (b) The same fields on `/healthz` — breaks the liveness contract by giving it a new way to fail. (c) A Go load harness (N concurrent bot games against the real corpus), run by hand, with the number recorded in `docs/deployment.md`.
**Rec.** (a) and (c). Both small; both turn deploy-day guesses into facts. **Size S.** **Depends on:** D1(a) for the draining state.
---
## E. Dictionary quality and word disputes
### E1. The corpus is thinner than it has been, and the journals said to watch for it
**Problem.** The shipped corpus is viwiktionary-only: `research-260908-1529` measured **36,200** accepted words from the dump against **61,026** for the undertheseanlp corpus and **64,110** for the union; the last kaikki-derived database was 34,813 words (`research-260910-0939`). `--min-words` defaults to 30,000 (`build-dictionary/main.go:68`). The journal for the switch ends with "Watch real play for thinness; the viwiktionary dump is the upgrade path, not GPL data". Nobody has watched, because of D2.
**Options.** (a) Union the dump with a second permissively-licensed list as an extra `--words` input, pinned by dated URL and checksum — the research report's own recommendation; licence compatibility must be re-checked per source. (b) A small curated additions/removals list kept in-repo (our own text, Apache-2.0, no upstream data) applied after the dump pass: cheap, targeted, and it keeps the two licence regimes separate as `NOTICE` requires. (c) Stay single-source.
**Rec.** (b) now — it is the mechanism that lets E3's disputes turn into fixes at all — and (a) once D2 says which words are actually missing. Doing (a) blind adds 28k words of unknown quality to a game whose problem may be the opposite (E2). **Size M.** **Depends on:** D2 for evidence, E3 for input.
### E2. Every Wiktionary entry is equally playable, including the ones nobody knows
**Problem.** `build-dictionary/filter.go:30-67` accepts any entry that is 2+ syllables, digit-free, punctuation-free and spelled from the Vietnamese alphabet. No frequency, no register, no proper-noun rule — and `research-260908-1529` showed the dump's proper-noun labels are *worse* than capitalisation as a filter. The bot searches exactly the graph the player is validated against (`bot.Board` over `WordsStartingWith`), so Hard can and does win with entries a native speaker has never met, which reads as cheating rather than as losing.
**Options.** (a) A `common` flag per word from a frequency list: validation accepts everything, the *bot* plays only common words, and `Suggestions` shows only common ones. Player-facing fairness without narrowing what a human may play. (b) Drop rare words entirely — narrows the game and throws away the corpus's long tail. (c) Nothing.
**Rec.** (a). Highest-value dictionary change available, and it takes no word away from a player. Needs a Vietnamese frequency source with a compatible licence — the open question for the research agent. **Size M-L.** **Depends on:** that list; a `words.common` column; one `Dictionary` method the bot's `Board` can see.
### E3. A player who is right has nowhere to say so
**Problem.** No dispute path exists anywhere — grep finds nothing in `server/` or `web/src`. A real word rejected as `NOT_IN_DICTIONARY` is simply lost, and the corpus journal records 1,138 words still carrying no meaning because the stripper does not know the `*form of` template family.
**Options.** (a) A `ReportWord` client message; the server records it (a log line via D2, or a row once D1(c) exists) and answers with a confirmation. (b) A prefilled GitHub issue link in the rejection line — zero server work, and it asks a casual player to have a GitHub account. (c) Nothing.
**Rec.** (a), logging only in v1, triaged by hand into E1(b)'s curated list. The full loop — report, triage, ship — is what makes the dictionary improve at all; every other option leaves it improving only when Wikimedia does. **Size S.** **Depends on:** D2 (where the report lands), E1(b) (where the fix lands), proto.
### E4. The dictionary cannot be rebuilt, only re-derived
**Problem.** `Makefile:10` and the `Dockerfile` both fetch `viwiktionary/latest/`, deliberately unpinned; `meta.source_sha256` identifies the bytes after the fact (`build-dictionary/main.go:508-519`) but there is no way to obtain those bytes again — Wikimedia repoints `latest/` monthly. So a regression traced to a corpus change cannot be reproduced or rolled back, and two journals name exactly this fallback ("keep the fetched file as a release asset") without having taken it.
**Options.** (a) Attach the derived `noitu.db` (a few MB) to each GitHub release and let a run consume it via `NOITU_DB_PATH`; rollback becomes redeploying an older image, which already carries its own copy. (b) Pin a dated dump URL and bump deliberately — contradicts the project's stated freshness preference and the house "moving tag over exact pin" rule. (c) Nothing.
**Rec.** (a). It keeps the moving-`latest` behaviour and makes the *output* recoverable, which is the property actually wanted. Ship `data/LICENSE` and `data/ATTRIBUTION.md` with the asset — CI already asserts this for the image (`ci.yml`, "The licence travels with the data"), and a release asset is distribution too. **Size S.** **Depends on:** a release workflow step.
---
## F. Developer experience
### F1. No linter on the JavaScript side
**Problem.** `web/package.json` has `svelte-check` (which does read the JSDoc types through `jsconfig.json`, so typing is genuinely covered) but no ESLint and no formatter check. House rules call for JavaScript + ESLint + JSDoc. Style consistency rests on discipline alone.
**Options.** (a) `eslint` with the Svelte and JSDoc plugins, wired into `npm run check` and CI. (b) A formatter check only. (c) Nothing.
**Rec.** (a) with a deliberately small rule set — the codebase is already clean, so a maximal config buys churn. **Size S.** **Depends on:** nothing.
### F2. The most valuable end-to-end coverage can only run in CI
**Problem.** `playwright.config.js` is Chromium-only, `workers: 1`, `fullyParallel: false`, 60s timeout; `e2e/pvp-game.spec.js` is 746 lines driving two-to-four browser contexts. Two journals record Chromium failing to install locally, and this workspace cannot run a browser at all (`CLAUDE.md`: none installed, no ARM64 builds, no root). Meanwhile `wsapi_test.go` is 2,292 lines of in-process multi-client coverage that runs everywhere in seconds.
**Options.** (a) Accept: e2e is a CI-only gate. (b) Move the multi-client *protocol* assertions (elimination order, grace expiry, owner handover, chat scoping) down into the Go suite where they already mostly live, and shrink Playwright to one smoke path per mode — faster signal, runnable locally, less flake (a four-player join flake is named in a journal). (c) Add a non-Chromium fallback — does not help; the constraint is the machine, not the engine.
**Rec.** (b). Keep e2e for what only a browser proves: IME composition, focus and keyboard behaviour, and the reconnect path through a real socket. **Size M.** **Depends on:** nothing.
### F3. Untrusted-input boundaries have no fuzz coverage; the proto workflow has one pin worth revisiting
**Problem.** Three functions parse hostile bytes — `Decode` (`codec.go`, capped at 4096 bytes), `sanitizeNickname` (`nickname.go`) and `vietnamese.Normalize` (`normalize.go:37-49`, shared by the corpus builder *and* the server, so a divergence silently unmatches the corpus). All are table-tested; none is fuzzed. Separately, `proto.yml` pins `buf` to exactly `1.69.0` against the house preference for a moving major tag, and there is no `buf format` check beside `buf lint`.
**Options.** (a) Three `FuzzXxx` functions seeded from the existing tables, run in CI for a bounded time. (b) A property test that `Normalize` is idempotent and agrees across both callers on the whole corpus. (c) Nothing.
**Rec.** (a) and (b) — both small, both guarding the one invariant the README stakes the game on. Move `buf` to a `1.x` tag and add `buf format --diff --exit-code`. **Size S.** **Depends on:** nothing.
### F4. Nothing watches the dependencies
**Problem.** `.github/` contains `workflows/` and nothing else — no Dependabot, no Renovate. `web/package.json` uses carets (fine) and `server/go.mod` pins exactly (as Go does), but nothing prompts an upgrade and nothing reports a CVE. The image is distroless and static, which limits the surface without removing it.
**Options.** (a) A `dependabot.yml` covering gomod, npm, GitHub Actions and Docker — four ecosystems, one file. (b) Renovate: better grouping, an app to install. (c) Manual.
**Rec.** (a). **Size S.** **Depends on:** nothing.
---
## G. Documentation
### G1. The README carries the architecture, the rules, the ops and the rationale
**Problem.** `README.md` is 316 lines and the best-written thing in the repo, but it now holds the rules, the online-play model, the frontend rationale, setup, configuration, every `make` target, testing and licensing. `docs/` holds one file. Meanwhile the *reasons* — why a dead end goes to the clock, why `is_me` is the only role encoding, why a departed player's name leaves the chat — live in code comments and in journals that explicitly say they are "not durable authority".
**Options.** (a) Split `docs/`: `rules.md`, `architecture.md`, `dictionary.md`, leaving `README.md` as philosophy plus quickstart plus licence. (b) A short `docs/decisions/` of ADRs extracted from the existing comments, each linking to the code that implements it rather than restating it. (c) Leave it; it is not yet painful.
**Rec.** (b) first, (a) when the README next grows. The decisions are the asset and are currently discoverable only by reading 20K LOC of (excellent) comments. Five or six ADRs — server authority, one rule implementation, proto as source of truth, two licences on two artifacts, no persistence — would cover most of it. **Size S.** **Depends on:** nothing.
### G2. The deployment doc is right about everything it covers and silent on capacity
**Problem.** `docs/deployment.md` covers configuration, proxy pitfalls, licence obligations, restart cost and dictionary updates, thoroughly. It has no capacity guidance, no log or metric reference (D2 would create one), and it documents the shared rate-limit bucket as a known limitation rather than as something configurable (D3).
**Rec.** Fold D2's metric names, D3's trusted-proxy variable and D4's measured room capacity into it as those land. Do not write the section ahead of the numbers. **Size S.** **Depends on:** D2, D3, D4.
---
## Top 10, ranked
Ranked on impact per unit of cost, weighted by fit with the stated philosophy: server-authoritative, one rule implementation, no wordlist in the client, proto as single source of truth, licences on separate artifacts, plain JS, Go preferred.
| # | Direction | Size | Why here |
|---|---|---|---|
| 1 | **D2** — expvar counters plus a rejection event line | S | Everything in section E is currently guesswork. Cheapest item on the list and it unblocks the most expensive ones. Pure Go, no dependency, no protocol change. |
| 2 | **D3** — trusted-proxy IP plus per-room join throttle | S | A documented, *known* denial of service in the supported topology. The deployment doc already describes the fix it declined to take. Security work with a bounded diff. |
| 3 | **A3** — claim a dead end instead of waiting out the clock | S | Removes up to 30s of dead air from every elimination. The server already computes the answer; the authorization shape is copied from `Resign`. Best felt-quality per line in the game itself. |
| 4 | **C2** — a rules/help surface | S | Nothing anywhere states the rules. Copy plus one route; the largest new-player gain available, and a precondition for B1 sending strangers into rooms. |
| 5 | **E3** — a word-report message | S | Turns "the dictionary is wrong" from a lost complaint into an input. Small alone, and the only thing that makes the dictionary improve on a cadence other than Wikimedia's. |
| 6 | **B1** — quick-match | M | The online half is unreachable without a friend. Reuses `createRoom`/`joinRoom` wholesale; one map and a timeout in the hub. Biggest product ceiling lifted for a mid-sized change. |
| 7 | **E2** — a `common` tier the bot is held to | M-L | Fixes the fairness complaint players will actually voice ("it won with a word nobody knows") without taking any word away from a human. Below its impact only because it waits on a licence-compatible frequency list. |
| 8 | **A2 + C3** — point breakdown, and a diacritic-only near-miss hint | S each | Both teach the player something the server already knows and the client cannot derive. Both additive proto changes, which this repo does routinely and safely (`buf breaking` in CI, committed generated code, cross-language binary fixtures). |
| 9 | **F2** — push multi-client coverage into Go, shrink Playwright to smoke | M | The protocol suite already exists and runs everywhere in seconds; the browser suite is CI-only, serial and has a known join flake. Better signal, less flake, and it makes the repo workable on the machine it is actually developed on. |
| 10 | **D1(a) + D4** — drain mode, `/readyz`, version stamp, measured capacity | S | Turns "deploy when the game is quiet" into a routine deploy, and three ops unknowns into printed numbers. Prerequisite for taking B2 and E1(b) seriously. |
Just below the line, and why. **F1/F3/F4** (lint, fuzz, dependency bot) are all S and all worth doing in whatever week has room; they are off the top 10 only because nothing currently hurts. **E1** (corpus breadth) is deliberately held behind D2: adding 28k words of unmeasured quality to a corpus whose real problem may be E2's long tail is hard to undo and harder to evaluate. **B2** (identity and progression) is the largest product prize here and is correctly gated behind persistence and measurement — starting it now means shipping progression with no way to tell whether anyone progresses. **A1** (bot objective) is genuinely interesting but changes a component that is tested, tuned and shipped; it should follow the measurement that says the ladder is wrong. **A4, B3, B4, C1, G1, G2** are all real and all fine as opportunistic work.
## Questions for the research agent
1. **A Vietnamese word-frequency list whose licence permits shipping inside a container image alongside CC BY-SA 4.0 data** — the single blocker on E2. Aggregate-only use is not enough: the ranking itself ships.
2. **Whether any open Vietnamese wordlist has meaningfully better compound coverage than viwiktionary** (E1) — the 2026-09-08 measurement covered undertheseanlp; is anything newer, and under what licence?
3. **How comparable word-chain games handle a no-legal-move claim** (A3) — does anyone let the trapped player assert it, and does it get abused?
4. **Realtime-server practice for single-binary room games:** typical goroutine-per-room budgets, and where people move to a shared registry (D4's capacity number; whether D1(b) snapshotting is ever worth it).
5. **Trusted-proxy header handling in Go** (D3) — is right-most-hop-with-CIDR-allowlist still the recommended shape, and does `coder/websocket` or any common middleware already provide it?
## Unresolved questions
1. **Is online play meant to grow, or is the bot the product?** B1/B3/B4 and B2 point in different directions; rank 6 assumes online growth matters. If it does not, the top 10 reorders around solo — A3, B4, C2, E2.
2. **Is a curated in-repo word list acceptable** (E1b), given the care taken to keep Apache-2.0 code and CC BY-SA data on separate artifacts? Our own additions are our own text, but they land in the same `.db` file; the licence statement needs a decision before that happens, not after.
3. **Does logging rejected words (D2) count as user content** the project wants to avoid holding? My assumption is that a normalized, length-capped word is fine and worth documenting; the alternative is counters only, which costs E1 and E3 most of their value.
4. **Is a fourth bot difficulty wanted at all,** or is the ladder deliberately three? A1 assumes the gap is the bot's *personality*, not its depth — unverified, and D2 would tell us.
5. **How long is a typical session, and how many rooms are live at peak?** Every capacity and persistence judgement in section D is speculation; no telemetry exists to check it, which is itself the argument for ranking D2 first.
6. **Report path:** the task brief names `…-260921-0016-improvement-directions.md` while the session hook advertises `…-260921-0017-{slug}.md`. Written to the brief's path; confirm which is canonical.
@@ -0,0 +1,328 @@
---
title: "Codebase health scan: noitu"
date: 2026-09-21
mode: codebase scan (not a PR review)
commit: dd3b463 (dev)
verdict: healthy; one silent feature regression, one availability gap, a handful of hygiene items
---
# Codebase health scan
Scope: whole repo at `dd3b463`. Generated trees (`server/gen`, `web/src/lib/proto`) not reviewed.
Prior findings from `code-review-260908-2302-codebase-cleanup.md` verified as fixed and not re-reported.
## Baseline (run here, ARM64 Linux)
| Check | Result |
|---|---|
| `go vet ./...` | clean |
| `go test ./... -race -cover` | all pass |
| coverage | vietnamese 100%, game 94.7%, bot 91.2%, wsapi 91.2%, build-dictionary 89.1%, dictionary 88.0%, `cmd/noitu-server` 0% |
| `npm run check` | 376 files, 0 errors, 0 warnings |
| `npm test` | 12 files, 190 tests, all pass |
| Playwright e2e | skipped (no browser on this host, per environment constraint) |
Code quality is above average for this size. Ownership between `game`, `wsapi` and the web store is
clean and deliberately documented; the engine is transport-free, the room owns the engine on one
goroutine, and the hub owns only registries. Comments explain *why* rather than restating code. The
findings below are gaps, not a pattern of carelessness.
---
# Confirmed findings (code path traced)
## C1 — `PlayedWord.player_id` is declared and consumed but never set. Chain attribution is dead in PvP
**Impact: high (silent feature loss in 3–4 player rooms). Size: S. Confidence: high.**
- `proto/noitu/v1/game.proto:170-173` declares `string player_id = 6` with the rationale "with four
people at the table the chain also has to say whose the other words were".
- `server/internal/wsapi/convert.go:107-116` — `PlayedWord()` sets `Word, Typed, ByMe, Points,
Syllables, Meanings`. **`PlayerId` is never assigned**, although `game.Move.Player` is right there
in the argument (`server/internal/game/engine.go:235`).
- `web/src/lib/stores/game.svelte.js:324` reads `playerId: played.playerId` → always `undefined`.
- `web/src/lib/components/ChainHistory.svelte:90-91` guards on `game.nameOf(entry.playerId)`, which
returns `''` for an absent id (`game.svelte.js:526-533`), so the byline never renders.
Net effect: in a 3- or 4-player room the chain shows *what* was played but never *who* played it —
exactly the gap `player_id` was added for in `4e2e9e1 feat(online)!: seat two to four players`.
Why both suites miss it — this is the AI-risk pattern worth naming:
- `server/internal/wsapi/convert_test.go:172-185` (`TestPlayedWordKeepsTypedInput`) asserts word,
typed, by_me, points, syllables. Not player_id.
- `server/internal/wsapi/wire_test.go:90` hand-writes `PlayerId: "p2"` into the cross-language
fixture, proving the *wire* can carry it.
- `web/tests/game-store.test.js:200` hand-writes `playerId: 'p1'` into the `played` fixture, proving
the *store* maps it.
Three tests touch the field; none exercises the producer. Each side is green against a fixture the
other side never produces.
Fix sketch:
```go
// convert.go
func PlayedWord(m game.Move, byMe bool, meanings []dictionary.Sense) *noituv1.PlayedWord {
return &noituv1.PlayedWord{
Word: m.Word, Typed: m.Typed, ByMe: byMe,
PlayerId: string(m.Player),
Points: uint32(m.Points), Syllables: uint32(m.Syllables),
Meanings: Senses(meanings),
}
}
```
Plus one assertion in `TestPlayedWordKeepsTypedInput`, and one end-to-end assertion in
`multiplayer_test.go` that a three-seat `turn_update` names the seat that played.
## C2 — No global room cap and no connection cap; the per-session limiter cannot bound either
**Impact: high (availability). Size: M. Confidence: high.**
- `server/internal/wsapi/server.go:73-91` — `handleWS` accepts every upgrade. Nothing counts
connections, per IP or in total.
- `server/internal/wsapi/session.go:54-55` — `roomsPerSecond = 0.2, roomBurst = 5`, and that bucket
is **per session** (`session.go:135`). It bounds one connection, not the process.
- `server/internal/wsapi/hub.go:140-154` — `newRegisteredRoom` has no ceiling on `len(h.rooms)`.
Each room is a goroutine, a 32-slot channel, an engine, up to three timers and a registry entry, held
for up to `defaultIdleWindow` (`room.go:67`, 10 minutes). N connections mint 5N rooms instantly and
0.2N/s thereafter. At 1 000 connections that is a 120 000-room steady state — each holding its
`chat` slice and engine — from a script, with no authentication anywhere in the protocol.
Fix sketch: a hub-level ceiling checked in `newRegisteredRoom` (return `errTooManyRooms` →
`room_start_failed`, or a new `server_busy` code), plus a per-IP concurrent-connection cap in
`handleWS` using the existing `keyedLimiter` shape. Both are cheap and neither changes the protocol.
## C3 — Behind the documented reverse proxy, every per-IP limiter is one global bucket
**Impact: high (availability), in the only supported deployment. Size: M. Confidence: high.**
- `server/internal/wsapi/server.go:150-162` — `clientIP` uses `RemoteAddr` only, deliberately
ignoring `X-Forwarded-For`. The reasoning is correct and I am not proposing reversing it.
- `docs/deployment.md` ("The client's own address") states the consequence plainly: *"Behind a proxy
every player therefore shares one bucket. If that becomes a problem, the fix is to make the proxy
the only source of the header and teach the server to trust it."*
- `docs/deployment.md` ("Behind a reverse proxy") names the container-behind-a-proxy shape as **the**
supported deployment.
So in the supported shape, `joinLimiter` (`session.go:49-50`, 1/s burst 5, keyed on the proxy's IP)
is a **single global limiter**. One client brute-forcing room codes spends the join budget for every
player on the server. That is a trivially reachable denial of service against a documented default.
The documented fix exists only as prose — there is no env var, no `Config` field, no code path for a
trusted-proxy mode. Fix sketch: add an opt-in `NOITU_TRUSTED_PROXY_HOPS` (default 0 = today's exact
behaviour) read in `loadConfig` (`cmd/noitu-server/main.go:109-118`) and threaded into `clientIP`, so
an operator who *has* made the proxy authoritative can say so. This adds the knob the doc already
names; it does not change the default or reverse the original decision.
## C4 — No per-connection message-rate ceiling; several dispatch arms are free
**Impact: medium (CPU exhaustion amplifier for C2). Size: S. Confidence: high.**
`server/internal/wsapi/session.go:399-493` rate-limits per message *type*, and three paths have no
budget at all:
- `session.go:489-490` — `ClientMessage_Ping` → `pongMsg` with no limiter. Self-limiting only because
a full outbox closes the session (`session.go:202-208`), so it costs the attacker their socket.
- A `ClientMessage` with **no payload set** matches no `case`, returns `nil`, and sends nothing. It
is completely free and can be replayed at line rate forever: one unmarshal + one dispatch per
frame, per connection, indefinitely.
- `session.go:409-418` — `StartBotGame` checks `Difficulty(...)` **before** `roomLimiter.allow`, so
an invalid difficulty is unlimited (it does cost an outbox slot, so it self-terminates).
`readLoop` (`session.go:289-305`) has no deadline and no frame counter. `maxFrameBytes` caps frame
*size* (`codec.go:195`) but not frame *rate*.
Fix sketch: one coarse `frameLimiter` bucket checked at the top of `dispatch`, generous enough that
no human hits it (say 30/s burst 60). Moving the `roomLimiter` check above the difficulty switch is a
one-line reorder.
## C5 — Raw player input crosses the trust boundary unsanitized as `PlayedWord.typed`
**Impact: medium (latent; not currently rendered to peers). Size: S. Confidence: high.**
- `server/internal/game/engine.go:236` — `Move{... Typed: raw ...}` stores the untrusted string.
- `server/internal/wsapi/convert.go:110` — `Typed: m.Typed`, and `room.go:943` builds this for
**every** recipient, not just the submitter.
- `server/internal/vietnamese/normalize.go:37-49` — `Normalize` does NFC, lowercase and
`strings.Fields`. It does **not** strip control or format characters, unlike `sanitizeText`
(`nickname.go:413-441`) which chat and nicknames go through.
A word is accepted whenever its *normalized* form resolves, so `raw` may legitimately contain any
`unicode.IsSpace` rune (U+000B, U+000C, U+0085, U+2028 LINE SEPARATOR, NBSP) in unlimited quantity up
to the 4 KiB frame cap, and those bytes are broadcast verbatim to every seat.
Today nothing renders it: `ChainHistory.svelte:102` gates the correction line on `entry.byMe`, so
only the author sees their own input. The defect is that the boundary is inconsistent — the server's
own rule (`nickname.go:406-412`: "text that is safe to render in a stranger's browser") is applied to
two of three player-authored strings — and the guarantee rests on a client-side `{#if}` rather than
on the server. Any future UI that shows who typed what turns this into a live rendering bug.
Fix sketch: `Typed: sanitizeText(m.Typed, maxNicknameRunes, maxNicknameMarks)` in
`PlayedWord`, or clear `Typed` for non-authors (`if !byMe { typed = "" }`), which is closer to what
the field is actually for.
## C6 — The dictionary builder opens SQLite without the path escaping the store fixed
**Impact: low. Size: S. Confidence: high.**
`server/internal/dictionary/store.go:142-148` documents and fixes a real trap: SQLite reads `#` in a
`file:` URI as a fragment delimiter, so a bare path silently opens a different file. The builder
builds the same URIs by hand and skips it:
- `server/cmd/build-dictionary/main.go:228` — `sql.Open("sqlite", "file:"+path+"?mode=ro")`
- `server/cmd/build-dictionary/main.go:393` — `sql.Open("sqlite", "file:"+path)`
Any output path containing `#` (or `?`) writes and verifies a different file than the one named, then
`write` renames the *intended* path over nothing. Fix: export `dictionary.DSN(path)` (or duplicate
the three-line helper) and use it in both call sites. DRY violation with a concrete failure mode.
## C7 — `write` deletes the good database before the rename
**Impact: low. Size: S. Confidence: high.**
`server/cmd/build-dictionary/main.go:357-390`. The doc comment promises the rename is what makes the
build safe — "a failure partway through … leaves an empty but syntactically valid database where a
good one used to be". Then `main.go:381-383` runs `os.Remove(path)` *before* `os.Rename`, for a
Windows constraint. On POSIX that reintroduces exactly the window the comment rules out: an interrupt
between the two calls leaves no dictionary at all.
Fix: guard the pre-remove with `if runtime.GOOS == "windows"`, so POSIX gets the atomic replace the
comment describes. Optionally `f.Sync()` the temp file before renaming.
## C8 — Room exit paths other than "everybody left" leave sessions attached
**Impact: low. Size: S. Confidence: medium-high.**
`server/internal/wsapi/room.go:456-461` (idle close) and `room.go:400-401` (context cancelled) return
without calling `vacate` on the remaining seats, so `session.room` keeps pointing at a dead room
(`session.go:175-182` is the only release path). Consequences:
- The room struct — engine, `chat` slice, `outWire` map — is retained for the life of the connection.
- Error copy degrades: `toRoom` correctly answers `not_in_a_room` (`session.go:512-514`), but
`handleSubmit` answers `busy` (`session.go:581-583`) for a room that is gone, not busy.
Fix: `for _, s := range r.seats { r.vacate(s) }` before those two returns.
---
# Plausible findings (worth a look, not traced to a failure)
## P1 — `Snapshot()` copies the whole history on every broadcast
`server/internal/game/engine.go:521-546` clones `history`, `scores`, `alive`, `outOrder` and rebuilds
`Standings()`; `room.go:910` calls it once per move, and `scoreRows` (`room.go:1098-1120`) allocates
per recipient. That is O(chain) per move, O(chain²) per game. At realistic chain lengths (tens of
words, ≤4 seats) this is noise — flagging it only because `broadcastTurn` is the hot path and the fix
is to pass the already-taken snapshot down rather than re-take it. **Size: S. Confidence: medium.**
## P2 — `readDump` aborts the whole build on a single malformed page
`server/cmd/build-dictionary/dump.go:128-130` returns a hard error when any page has no revision
text. Against a 61 MB monthly upstream that nobody pins, one bad page fails the entire dictionary
build (and therefore the release image job in `ci.yml:118-124`). A counted-and-skipped reject, with
the existing `minPages` floor (`main.go:106-109`) as the real guard, is more robust and loses
nothing. **Size: S. Confidence: medium** — this may be a deliberate fail-loud choice; the comment
does not say.
## P3 — One connection can hold two rooms for the grace window
`session.go:152-168` — `attach` releases the previous room with a `disconnectInput`, which starts a
grace window rather than vacating. A player who creates room A, starts a game, then joins room B
leaves A's opponents waiting out `NOITU_GRACE` for somebody who deliberately walked away. Probably
intended (it is the same code path as a refresh), and `TestOneConnectionCannotStrandRooms` proves
nothing leaks. Flagged as a product question, not a defect. **Size: S. Confidence: low.**
---
# Test-coverage gaps
| Behaviour | Where it lives | Covered? | Note |
|---|---|---|---|
| `PlayedWord.player_id` is populated by the server | `convert.go:107` | **No** | C1. Three fixtures hand-write the field; none produces it |
| Chain attribution reaches a 3rd/4th seat end-to-end | `room.go:929-946` | **No** | `multiplayer_test.go` checks turns and standings, not the chain byline |
| `cmd/noitu-server` config parsing | `main.go:109-157` | **No** | 0.0% coverage; `envDuration`/`envList` fallbacks are untested |
| Hub-level room ceiling | `hub.go:140` | n/a | No ceiling exists (C2) |
| Per-connection frame-rate ceiling | `session.go:399` | n/a | No ceiling exists (C4) |
| `PlayedWord.typed` sanitization | `convert.go:110` | **No** | `TestChatTextIsSanitizedAndCapped` and `TestOpponentNeverSeesAnUnsanitizedNickname` cover the other two strings |
| Fuzzing the untrusted-input boundary | `codec.go:215`, `nickname.go:413`, `normalize.go:37` | **No** | Zero `func Fuzz` in the repo. These three are ideal `testing.F` targets and would have surfaced C5 |
| SQLite DSN escaping in the builder | `build-dictionary/main.go:228,393` | **No** | `dictionary/store_test.go` covers the store's `dsn`; the builder's copies are untested |
| `write` interrupted between remove and rename | `build-dictionary/main.go:381` | **No** | C7 |
| Room close leaves no attached session | `room.go:456,400` | Partial | `TestIdleLobbyCloses` asserts the error and eviction, not seat release |
What *is* well covered and worth saying so, because it changes the risk calibration: 100 `wsapi`
tests including goroutine-baseline (`TestGoroutinesReturnToBaseline`), origin checking, oversize
frames, every rate limiter, seat authorization for strangers (`TestStrangerCannotSubmitForASeated
Player`), grace-window races, and exhaustive enum mapping tests. The error-code surface is closed:
all 33 `errorMsg` codes plus `room_idle_closed` have i18n entries (`web/src/lib/i18n/vi.js:204-239`).
---
# CI / build / dependency hygiene
| # | Item | Evidence | Size |
|---|---|---|---|
| H1 | CI does not run on the working branch | `.github/workflows/ci.yml:11-16` and `proto.yml:8-11` trigger on `push: branches: [main]` + `pull_request`. Active development is on `dev`, so pushes there are untested until a PR exists | S |
| H2 | No static analysis beyond `go vet` | `ci.yml:32-40`. The 260908 review ran `staticcheck` and `deadcode` by hand and they found real issues. Nothing keeps them green now | S |
| H3 | No `gofmt`/format gate | Absent from both workflows. Noted in the prior review as noisy under `core.autocrlf`; `gofmt -l` with `git config core.autocrlf input` in CI would still work | S |
| H4 | `buf` pinned to an exact version | `proto.yml:25-27` — `version: 1.69.0`. Repo convention prefers a moving major tag; `buf-setup-action` accepts a floating spec | S |
| H5 | No dependency-update automation | `.github/` contains only `workflows/`. No `dependabot.yml`, no renovate config. Given the "moving tags over pins" rule, a bot is the mechanism that rule assumes | S |
| H6 | Direct deps one minor behind | `go list -m -u all`: `golang.org/x/text v0.41.0 → v0.42.0`, `modernc.org/sqlite v1.58.0 → v1.59.0`. `coder/websocket v1.8.15` and `protobuf v1.36.12` are current. Not urgent | S |
| H7 | `allowScripts` not configured | `npm ci` warns: `esbuild@0.28.2 (postinstall) not yet covered by allowScripts`. Workspace convention (`/workspace/CLAUDE.md`, "Language Preferences") says install scripts are gated per package in `package.json#allowScripts`. `web/package.json` has no such block | S |
| H8 | No coverage floor | Coverage is measured only when asked for. A `-coverprofile` step with a floor would have made the `cmd/noitu-server` 0% visible | S |
Correctly done and worth not touching: the Dockerfile's three-stage split with the dump confined to
a builder stage, the CC BY-SA licence-travels-with-the-data assertion (`ci.yml:126-145`), the
`.dockerignore` dump exclusions, `buf breaking` against `origin/main`, and the `image` job depending
on `e2e` with a comment explaining why.
---
# Maintainability hot spots
| File | LOC | Note |
|---|---|---|
| `server/internal/wsapi/room.go` | 1609 | Room lifecycle, seating, lobby actions, chat store, engine bridging, per-recipient rendering and the frozen bot board in one file. All of it is genuinely room-goroutine state, so splitting by *concern* (`room_lobby.go`, `room_chat.go`, `room_broadcast.go`) preserves the ownership invariant while making the file navigable. Size: M |
| `server/internal/wsapi/wsapi_test.go` | 2292 | 68 tests in one file next to four focused test files. Same treatment: the names already cluster (chat, resume, lobby, limits) |
| `web/src/routes/online/+page.svelte` | 612 | Six interacting `$effect` blocks driving `resuming` / `needName` / `stalled` / `pending`, several using `untrack` to break cycles. This is the least inspectable code in the repo; the `untrack` calls are load-bearing, which is the signal. Extracting the join/resume state machine into a `.svelte.js` store — the way `bot-session.svelte.js` already does for the bot flow, for exactly the stated reason — would make it testable. Size: M |
No duplicated logic of consequence found. `vietnamese.Normalize` shared between builder and server is
the right call and its package doc says why. `bot.Board`/`frozenBoard` correctly copies engine state
rather than sharing it (`room.go:1570-1594`).
---
# Recommended actions, ranked
1. **C1** — set `PlayerId` in `PlayedWord`, assert it in `convert_test.go`, and add one multi-seat
end-to-end assertion. A shipped feature is silently absent. (S)
2. **C3** — add the opt-in trusted-proxy config the deployment doc already promises, defaulting to
today's behaviour. (M)
3. **C2** — hub-level room ceiling plus a per-IP concurrent-connection cap. (M)
4. **C4** — one coarse frame-rate bucket in `dispatch`; move the `roomLimiter` check above the
difficulty switch. (S)
5. **C5** — sanitize or drop `PlayedWord.typed` for non-authors. (S)
6. **H1** — add `dev` to the CI push triggers, or `branches-ignore: []`. (S)
7. **H2/H3** — `staticcheck` and a format gate in the Go job. (S)
8. **C6/C7** — share the SQLite DSN helper; make the pre-rename remove Windows-only. (S)
9. Add `testing.F` targets for `Decode`, `sanitizeText` and `vietnamese.Normalize`. (S)
10. **H5/H7** — dependabot config; `allowScripts` for `esbuild`. (S)
11. **C8**, then the three maintainability splits when the files are next touched. (S/M)
---
# Unresolved questions
1. **C1** — was `player_id` ever wired up and later lost, or never implemented? `git log -S` shows
the field arriving with `4e2e9e1` and no producer in any revision, which points at never. Worth
confirming before assuming a regression in a later refactor.
2. **C3** — is the container ever run without a reverse proxy in front of it? If the supported shape
is always proxied, the shared-bucket problem is not an edge case and should be ranked above C2.
3. **P2** — is the hard failure on a page with no revision text deliberate fail-loud, or an
unconsidered path? The comment does not say, and that decides whether it is a fix or a non-issue.
4. **C5** — is `typed` intended to be visible to anyone but its author? If not, clearing it for
non-authors is strictly better than sanitizing it, and also narrows the wire.
5. **H4** — is `buf 1.69.0` pinned because a newer release broke something? If so the reason belongs
in a comment in `proto.yml`, per the repo's own version-pinning rule.
@@ -0,0 +1,210 @@
# Research: noitu landscape scan — competitors, dictionary disputes, small-scale server practice
Prior reports (`research-260904-1058-noi-tu-game.md`, `research-2609{07,08,10}-*`) already cover
rules engine, bot AI, Vietnamese normalization, and an exhaustive dictionary-source comparison
(Wiktionary/undertheseanlp/Hoàng Phê). This report does not repeat those; it adds competitor
feature/complaint evidence, dispute-handling patterns, small-scale server practice, and ecosystem
notes.
## Summary
- Direct web competitors exist: `noitu.fun` and `wordfight.online` — both have ranked/Elo modes,
neither has noitu's spectator-on-elimination or bot-difficulty design.
- The #1 recurring complaint across mobile nối từ apps is dictionary coverage ("từ điển quá ít"),
not rules or UX — validates prior dictionary-focused work as the highest-leverage area.
- Every competitor with a fixed wordlist ships a manual "report word" channel; none of the ones
found do live player voting mid-game. One Discord bot layers a separate community-wordlist repo
on top of a base corpus — closest precedent to an allowlist overlay.
- Ranked/Elo matchmaking with wait-time-based window widening is the standard pattern for small
pools (chess.com/lichess-style); directly applicable to a 2-4 player room-code game with low CCU.
- SQLite-as-embedded-store (not just read-only dictionary) is a well-documented pattern for adding
room/stat persistence without a DB service — fits the existing `modernc.org/sqlite` dependency.
- `protovalidate` (Apache-2.0, Go+JS via `protovalidate-es`) is a concrete, license-compatible fit
for schema-level proto validation, complementing (not replacing) noitu's Vietnamese-specific
sanitization.
- Card-mechanic "Extended Mode" (skip/reverse/swap turn) and Elo ranked ladders are the clearest
"features players expect but noitu lacks" signal found.
- Casual house-rule turn timers run 5-10s vs noitu's 30s default — worth noting as a design choice,
not a defect.
- No changes found to Svelte 5/adapter-static or buf/protobuf-es that are breaking or urgent for
noitu's current setup.
## 1. Competing implementations — rules, features, complaints
### Direct web competitors
- **noitu.fun** — solo vs AI, "Ranked" mode with a leaderboard, and group/room-code mode (closest
analog to noitu's own room model). Also ships non-chain minigames under one brand: "Extended
Mode" adds card mechanics (reverse turn, assign-next-responder, swap opponent's word, skip turn),
plus a trivia-style "Character Arena" and a Wheel-of-Fortune-style "Letter Assembly" mode. Rules
text explicitly bans slang/shorthand and misspelled tone marks. [noitu.fun]
- **wordfight.online** — 3 modes: English word chain (shows IPA + Vietnamese gloss per word),
Vietnamese word chain (2-syllable, same rule as noitu), and a separate "King of Vietnamese"
word-puzzle mode. Turn timer 10-30s. **Elo-based leaderboard with skill-matched opponents**,
progressive difficulty across up to 999 levels, 3-star mastery rating per level, replay-resistant
randomized word lists, private shareable room links. [wordfight.online]
- Neither site documents spectator viewing of eliminated players, reconnect grace windows, or
per-seat colored room chat — these look like noitu differentiators, not table stakes to add.
### Discord bots (adjacent platform, same game)
- **minhqnd/Noi-Tu-Discord** ("Moi Nối Từ") — bot-vs-player and PvP-with-bot-as-referee modes, DM
play, `/leaderboard` and `/stats` (streak, personal record, wins), `/tratu` dictionary lookup
backed by 357k+ definitions via `dict.minhqnd.com`, emoji-reaction feedback per submission
(✅ correct / ❌ can't chain / 🔴 duplicate / ⚠️ format error). [github.com/minhqnd/Noi-Tu-Discord]
- **lvdat/bot-noi-tu** ("RaHub") — dictionary sourced from `undertheseanlp/dictionary` plus a
**separate community-contribution repo** (`phobo-contribute-words`) merged in at build/runtime.
This is the clearest real-world precedent for a base-corpus + community-overlay split (see §2).
[github.com/lvdat/bot-noi-tu]
### Mobile apps — features and complaints
- Galaxy Team "Nối từ - Word Chain": Survival / Time-Limited(3min) / Level-Challenge-vs-AI modes.
[play.google.com/.../com.galaxteam.wordchain]
- "Nối từ tiếng Việt" (iOS, MWM/id6449588406): 3 modes — Challenge (timed rounds + score
threshold), Duel (2p turn-based), Arena (4p elimination — same shape as noitu's 2-4 elimination
room), Game Center leaderboard integration, scoring by chain-word character count. **Rating 1.2/5
(19 reviews)**, dominant complaint is a too-small dictionary ("từ điển quá ít") rejecting valid
words and occasionally accepting non-Vietnamese junk; developer response cites Vietnamese's
richness as an excuse rather than fixing it. [apps.apple.com/vn/.../id6449588406]
- Cross-app pattern found via search: apps expose an in-app "Report Error" / "báo lỗi" flow for a
rejected word rather than any live dispute; devs publicly acknowledge coverage gaps as
unavoidable rather than committing to fixes. [search results, ktcc.blog context]
### Rule-variant survey (Vietnamese how-to-play content)
- Casual/offline house rules commonly run **5-10s per turn** (ktcc.blog), notably faster than
noitu's 30s default — a deliberate design choice for noitu given typed Vietnamese input
(Telex/VNI composition), not evidence of a gap.
- "Từ hiểm" (dead-end/trap words — syllables that start almost nothing) is a named, well-understood
strategic concept across the community, not something apps hide. This affirms noitu's design of
showing the leftover words in a dead-end position rather than treating it as a bug.
- Recurring dispute across communities: "từ ghép" (compound) vs "từ đơn" (single-syllable) and
which authority resolves disputes; common advice is to pre-agree a single reference dictionary
(often citing the Institute of Linguistics / Hoàng Phê) as sole arbiter — already covered in
`research-260910-0939-hoang-phe-in-noitu.md`.
### Implications for noitu
- Dictionary coverage is the make-or-break axis for player satisfaction industry-wide; any
roadmap item competing for effort against dictionary work should clear a high bar.
- A lightweight "report this rejection" action (word + turn context) is standard and cheap; noitu
has no such flow today per the README. Low-risk, high-precedent addition.
- Elo/ranked ladder and light "chaos" mechanics (skip/reverse/swap) are the two concrete feature
gaps versus the nearest web competitors, if PvP breadth is a goal.
## 2. Dictionary disputes — patterns beyond corpus choice
(Corpus/source comparison already exhaustively covered in prior reports; this is new: process
patterns for *handling disputes at runtime*, and licensing-separation precedent.)
- **Static base + community overlay repo** (lvdat/bot-noi-tu + phobo-contribute-words): base
dictionary from a versioned upstream corpus, disputed/missing words tracked in a *separate*
repository merged at build or load time. Mirrors an allowlist-overlay design: keeps the
overlay's provenance and license distinct from the base corpus, which matters for noitu's
Apache-2.0-code / CC BY-SA-4.0-data split — an overlay of player-submitted words would need its
own license decision (CC BY-SA if derived/mixed with Wiktionary text, or a fresh license if pure
word-list-no-definition additions, which are likely uncopyrightable facts). [github.com/lvdat]
- **Scrabble's live challenge**: any player may challenge a just-played word before the next turn;
in double-challenge scoring the *loser* of the challenge (wrong challenger or wrong player) loses
their turn — a real cost that discourages frivolous challenges. [en.wikipedia.org/Challenge_(Scrabble)]
- **Words With Friends' approach**: no live human challenge at all — the client just keeps
resubmitting until the server's own dictionary accepts something; the dictionary itself (Zynga's
ENABLE-derived ~173k word list, deliberately different from tournament Scrabble lists) is the
sole and silent arbiter. [wordfinder.yourdictionary.com; word.tips]
- No evidence found of any nối từ implementation running **live in-match player voting** on a
disputed word (i.e., other seated players vote accept/reject before the turn resolves). The
Scrabble-style post-hoc challenge with a real cost, or WWF's silent-server-arbiter, are the two
patterns actually used in the wild — an async "report → maintainer/community review → next
dictionary build" queue (closer to WWF, deferred) is simpler to implement correctly than a live
vote and has no real-time consistency/latency problem to solve.
### Implications for noitu
- If a self-serve "report word" pipeline is ever built, model it as WWF/lvdat's deferred pattern
(report now, reviewed and merged into a future `noitu.db` build) rather than live voting — no
new consensus/anti-brigading problem, and it fits the existing single-file, rebuilt-not-mutated
dictionary architecture.
- Any player-submitted word overlay must get its own explicit license decision before merging with
the CC BY-SA data tree; do not assume submitted words inherit Wiktionary's license by default —
bare word forms without wiktionary-derived definitions are likely factual/uncopyrightable, but a
submission that carries a definition text is a new derivative work needing its own care.
## 3. Small-scale realtime game server practice
- **Matchmaking for small pools**: standard pattern (chess.com/lichess-style, per Awesomenauts
postmortem) is `window = base_window + widen_rate × seconds_waited`, anchored on whichever queued
player has waited longest, capped at a max window — trades match quality for queue time as the
pool thins out. Directly portable to a "quick match" queue for noitu's current room-code-only
online mode. [joostdevblog.blogspot.com; medium.com/@deephavendatalabs]
- Batch pooling ("gather for ~1-2 min then match everyone at once") is the standard fallback for
genuinely tiny regional pools, cited as improving match quality up to ~300 concurrent queuers —
likely overkill for noitu's expected scale, worth knowing only if online play grows.
- **Persistence without a DB service**: SQLite is a well-documented fit for exactly this
("zero-configuration... ideal for small to medium services," pure-Go `modernc.org/sqlite` needs
no CGO, single static binary). Patterns seen: JSON-blob-per-row for full game/room state
(simplest, fine at noitu's scale), or row-per-update table if write volume grows. A SQLite
changelog-table + trigger + Go worker pattern exists for push-driven updates but is unnecessary
complexity for a single authoritative in-process game engine like noitu's. noitu already
depends on `modernc.org/sqlite` for the dictionary — reusing it (a second read-write DB file, or
a second schema in-process) for room/series-stat persistence across restarts costs no new
dependency. [oneuptime.com/.../sqlite-go; gist.github.com/rusco]
- **Rate limiting / abuse**: `golang.org/x/time/rate`'s token-bucket `Limiter` (Allow/Wait/Reserve)
is the standing idiomatic choice for per-connection message throttling in a Go WebSocket server;
described as covering "90% of cases" without needing Redis-backed distributed limiting, which is
irrelevant to a single-binary deployment. Combine with a concurrent-connection cap per IP.
[dev.to/lovestaco; oneuptime.com/.../websocket-rate-limiting]
- **Anonymous vs identity**: industry best practice (PlayFab, AWS Games Industry Lens) is
zero-friction anonymous login by default, with an *optional* upgrade path to a recoverable
identity later — matches noitu's current no-account, nickname-only model; nothing here argues
for forcing accounts. Reconnect best practice: identity (not the transient socket/session token)
is the thing that says two connections are the same player, and the server-issued session token
should stay stable across a reconnect within the grace window — consistent with noitu's existing
disconnect-grace-window design per its README. [learn.microsoft.com/playfab; docs.aws.amazon.com]
- **Observability**: not surfaced as a distinct concern in results beyond generic Prometheus/metrics
advice seen in one unrelated Go arena-server repo (`nguyenbatam/arena_game_server`, Redis-backed,
10k-CCU target) — that project's scale and dependency footprint (Redis, k8s) is not a fit for
noitu's single-binary constraint; flagging only because Prometheus text-format `/metrics` next to
the existing `/healthz` endpoint is a low-cost, dependency-light addition if observability becomes
a goal.
### Implications for noitu
- A quick-match queue (as opposed to room-code-only) is implementable with the widen-by-wait-time
formula and no new infrastructure; it's a pure in-memory queue in the existing Go process.
- Room/series persistence across restarts (currently implied in-memory, since README describes
rooms closing after 10 idle minutes with no persistence mention) could reuse the SQLite
dependency already in the binary rather than adding Redis/Postgres.
- `x/time/rate` per-connection + a connection-count cap per IP is the standard, low-effort answer
to WS abuse; no evidence any competitor does more than this at this scale.
## 4. Ecosystem notes (brief)
- **SvelteKit/Svelte 5 + adapter-static**: no breaking change found for 2026. One caution
surfaced repeatedly in 2026 guides: module-scope runes and `$derived` wrapping a store's `$`
auto-subscription can break specifically under prerendering/hydration; recommended fix is
keeping runes component-local and bridging global stores explicitly via `$effect`. Only relevant
if noitu's frontend uses module-level `$state`/`$derived` for anything beyond the documented
client-owned state (theme, personal best, input box) — worth a quick grep, not a redesign.
[svelte.dev/docs/kit/adapter-static; khromov.se]
- **buf/protobuf-es**: `protovalidate` (Apache-2.0) plus `protovalidate-es` gives schema-level
field constraints (e.g. string length/pattern) enforced identically from the same `.proto` file
on both the Go and JS generated types — same "one wire contract" philosophy noitu already uses
for message shape. Could formalize bounds already enforced by hand (e.g. nickname length) as
proto annotations instead of duplicated Go+JS logic — but noitu's nickname sanitization (control
chars, whitespace collapse, combining-mark cap) is Vietnamese-text-specific logic CEL/protovalidate
doesn't replace; treat it as a complement for simple bounds, not a replacement for sanitization.
[github.com/bufbuild/protovalidate]
- No 2026 buf CLI or `protoc-gen-es`/`@bufbuild/protobuf` breaking-change advisory found that
affects noitu's current generation setup.
## Unresolved questions
- Does noitu currently persist rooms/series scores across a server restart at all, or is
everything in-memory? README doesn't say; determines whether §3's SQLite-persistence point is
a new feature or filling a known gap.
- Would a "report rejected word" flow be worth the review-queue maintenance cost given the project
has one apparent maintainer, versus just improving corpus coverage directly (already the subject
of 5 prior reports)?
- Is competitive/ranked PvP actually a goal for noitu, or is vs-bot + casual room-code play the
intended scope? The Elo/quick-match findings only pay off if ranked play is in scope.
- Could not verify `phobo-contribute-words`' actual review/merge workflow or license (GitHub
fetch returned only nav chrome, no README text) — the "allowlist overlay" characterization is
inferred from the bot's own README description, not confirmed from that repo directly.
Status: DONE
Summary: Competitor scan of noitu.fun, wordfight.online, and 4 Discord/mobile nối từ implementations plus small-scale server-practice research (matchmaking, SQLite persistence, rate limiting) written to the report; dictionary coverage confirmed as the industry's #1 complaint, and a base+overlay community-wordlist precedent and protovalidate found as new, concrete inputs.
Concerns: none blocking; one item (phobo-contribute-words internals) unverified due to GitHub README not rendering through WebFetch — noted in Unresolved questions.
@@ -0,0 +1,71 @@
# noitu improvement brief — synthesis of three agent reports
Date: 2026-09-21. Branch: `dev` at `dd3b463`. Sources (same directory):
- `researcher-260921-0016-noitu-landscape-and-improvement-inputs.md` — competitors, dispute patterns, small-server practice
- `brainstormer-260921-0016-improvement-directions.md` — 25 codebase-grounded directions, ranked top 10
- `code-reviewer-260921-0016-codebase-health-scan.md` — health scan; all tests green, 8 confirmed + 3 plausible findings
## Headline
Codebase is healthy (Go coverage 88–100% per package, 190 JS tests green, prior review findings all fixed).
Three things stand out across all reports:
1. **One shipped feature is silently dead**: `PlayedWord.player_id` is never set by the server
(`server/internal/wsapi/convert.go:108`, verified), so 3–4 player rooms never show who played
each word. Three fixture-based tests mask it.
2. **The supported deployment has a known, documented, unfixed DoS**: behind the reverse proxy all
per-IP limiters collapse to one bucket (`server.go:150-162`, `docs/deployment.md:113-119`), and
there is no global room or connection cap.
3. **Zero telemetry** — no metrics, ~12 log sites in `wsapi`. Every dictionary/bot/capacity
decision is currently a guess. Research confirms dictionary coverage is the #1 complaint across
every competing nối từ app, so measuring rejections is the cheapest high-leverage step.
## Fix now (bugs and security, all S–M, no product decision needed)
| # | Item | Source | Size |
|---|---|---|---|
| 1 | Set `PlayerId` in `wsapi.PlayedWord`; add a producer-side test | reviewer C1 | S |
| 2 | Trusted-proxy mode (`NOITU_TRUSTED_PROXY` / `X-Forwarded-For`) + per-room join throttle | reviewer C3, brainstorm D3 | S–M |
| 3 | Global room cap + connection cap in hub/server | reviewer C2 | M |
| 4 | Per-connection frame-rate ceiling; charge `Ping` and empty `ClientMessage` | reviewer C4 | S |
| 5 | Sanitize `Move.Typed` server-side (reuse `sanitizeText`) instead of relying on client `byMe` gate | reviewer C5 | S |
| 6 | Builder: escape `#` in SQLite DSN like the store does; rename-over instead of remove-then-rename | reviewer C6, C7 | S |
| 7 | `vacate` sessions on idle/cancel room exits | reviewer C8 | S |
## Build next (ranked, brainstormer top 10 cross-checked with research)
| # | Direction | Why | Size | Depends on |
|---|---|---|---|---|
| 1 | **Observability**: expvar counters + one `word_rejected` slog line (normalized, length-capped) | Unblocks every dictionary/bot decision below | S | decision on logging rejected words |
| 2 | **Claim a dead end** (`ClaimDeadEnd` msg, shape copied from `Resign`) | Server already knows via `HasLegalMove`; removes 30s dead air per elimination | S | — |
| 3 | **Rules/help surface** | No rules copy anywhere in `web/src`; competitors all state rules; precondition for strangers meeting via quick-match | S | — |
| 4 | **`ReportWord` message** + simple review queue | Every competitor with a fixed wordlist has one; research found no live-vote precedent, so keep it async | S | #1 for triage |
| 5 | **Quick-match** with wait-time-widening window | Rooms reachable only by code today; lichess-style widening fits small pools | M | rules surface (#3) |
| 6 | **`common` word tier the bot is held to** | Fixes "Hard won with a word nobody knows" without removing words from humans | M–L | licence-compatible frequency list (not yet found) |
| 7 | **Point breakdown on `PlayedWord`** + diacritics-only near-miss hint on `MoveRejected` | Additive proto changes; teaches what the server already knows | S each | — |
| 8 | **Move multi-client assertions into Go suite; Playwright → smoke** | e2e is CI-only, serial, flaky, cannot run on the dev box | M | — |
| 9 | **Drain mode, `/readyz`, version stamp, measured capacity** | Makes deploys routine; prerequisite for persistence talk | S | — |
| 10 | **Corpus widening / community allowlist overlay** (Discord-bot precedent) | Held behind #1: unmeasured 28k-word addition is hard to evaluate or undo | L | #1, licence decision |
Below the line, opportunistic: JS lint (ESLint per house rule), fuzz targets for `Decode`/`sanitizeText`/`Normalize`,
dependabot/renovate, `allowScripts` block in `web/package.json`, CI on `dev` branch (`ci.yml` runs push only on `main`),
`buf` pinned to exact 1.69.0 against the moving-tag rule, split `room.go` (1609 LOC) and `online/+page.svelte` (612 LOC, six `$effect`s).
Two leftovers from the 2026-09-10 UX reports: chat toggle renders after `ChainHistory`; lobby chat has no unread badge.
Research-only signals, not recommended yet: Elo/ranked ladder (noitu.fun, wordfight.online both have one),
card-mechanic "chaos" mode (skip/reverse/swap), 5–10s casual timers vs noitu's 30s. All depend on the answer to Q1.
## Decisions needed from the maintainer
1. **Is online play meant to grow, or is vs-bot the product?** Decides whether quick-match/ranked (rows 5, research signals) rank above solo items (dead-end claim, daily mode).
2. **May the server log rejected words** (normalized, capped)? Counters-only kills most of the value of rows 1, 4, 10.
3. **Is a curated in-repo word overlay acceptable** under the Apache-2.0 / CC BY-SA split? Our additions land in the same `.db`; licence statement must be settled first.
4. **Is the container ever run unproxied?** Decides whether the proxy-bucket fix (C3) or the global caps (C2) ship first.
5. **Is a fourth bot difficulty wanted**, or is the ladder deliberately three?
## Unresolved
- No licence-compatible Vietnamese word-frequency list identified; gates row 6.
- `phobo-contribute-words` (Discord-bot community wordlist) internals unverified; overlay characterization is inferred.
- Peak rooms / session length unknown; all capacity claims are speculation until row 1 lands.