The browser downloaded 45 MB of gzipped SQLite before it could answer anything. Now sql.js-httpvfs asks for the pages a query touches and the databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip stream is not a byte range of a database. That only works if every query the site issues is index-driven, and measured against the real 2016 file, most were not: so_bao_danh = ? SEARCH via PK ~20 KB ho_ten_ascii LIKE '%x%' SCAN 127 MB ho_ten_ascii LIKE 'x%' SCAN 127 MB COUNT(*) covering index scan 20 MB ORDER BY toan DESC LIMIT 10 SCAN + temp b-tree 127 MB Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE index; a range comparison does use the index. So the schema changed to suit the access pattern rather than the search changing to suit the schema. name_word holds one row per word of each name, WITHOUT ROWID so the table is the index, carrying ho_ten_ascii so a multi-word query is resolved inside a single b-tree. name_word_freq says which word of a query is rarest — the vocabulary is 4,397 words across 2.87M entries, so "buu loc" seeks on 287 entries rather than walking the 300,000 that "thi" would. Searching by any word of a name survives, at a few hundred KB a query. idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either. Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL presets off a full scan. The footer's candidate count now comes from datasets.json instead of COUNT(*). 2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged. The SQL tab is the one place a user can still write a query that reads the whole table, so it asks before it opens, runs under a byte budget that stops a runaway query, and shows what each query actually fetched. Verified: row counts through the assembler guards, every app query index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206 with a correct Content-Range. Not verified in a browser — this machine has none — and the library refuses to open a file the host compresses, so the deployed response headers need a look.
3.7 KiB
thptqg
Exam-score lookup for Vietnam's national high school graduation exam. Four
stages, each independent: crawler/ (Go) fetches spreadsheets into data/<id>/,
parser/ (Go) turns them into <id>.db, assembler/ (Go) verifies, compresses
and builds _site/, and web/ (SvelteKit + Tailwind) serves every dataset from
one app.
datasets.json at the root is the registry the Go stages and the web app both
read. See docs/ for architecture, pipeline and deployment.
Verifying a change
Run parser and assembler:
go -C parser test ./...
go -C assembler test ./...
Add the web checks when frontend files changed:
(cd web && npm test && npm run lint && npm run build)
npm run lint is ESLint plus svelte-check, so it is the type check too.
Do not run the crawler suite as part of routine verification. The crawler is
not part of the build — data/<id>/ is committed and a crawl only refreshes it
by hand — so its tests do not gate ordinary changes. Run them only when the
change actually touches crawler/. CI still runs all three modules.
parser/internal/reader is the slow package (~9s) because the fidelity suite
hashes every real input file. That is the point of it; do not skip it.
Things that look wrong but are not
data/2016/filenames are content hashes and must stay verbatim. The parser sorts inputs bytewise and inserts last-wins, so filenames would decide which row survives a duplicate exam number. Renaming them also breaksparser/testdata/reader-fidelity-hashes.tsv, which is keyed by full path.parser/testdata/reader-fidelity-hashes.tsvis frozen and cannot be regenerated. A mismatch is a reader bug until proven otherwise, never a cue to refresh the file.parser/cmd/dumpcellsnarrows a failure to the cell.ToAsciifilters the literal range U+0300..U+036F, notunicode.Mn. It must matchtoAsciiinweb/src/lib/to-ascii.ts, or accent-insensitive search silently misses rows.to-ascii.test.tspins the pairs both sides must agree on.- The 2016 files use four different layouts, and detection is per sheet.
Two of them publish scores in one column per subject instead of a
DIEM_THIsentence, and one puts a three-row ministry title block above its header.parser/internal/ingest/detect2016.goholds the header tables; they are observations about 119 specific files, not a rule to generalise. paths.relativeis false inweb/svelte.config.js. Every route is prerendered to its own file and asset URLs stay absolute, so the copy ofindex.htmlserving as404.htmlworks at any depth. No SPA 404-fallback is used — a fallback would break the?q=deep links.dbSizeMbindatasets.jsonis a build guard, not just a label. The assembler refuses to publish an artifact that falls below a ratio of it.- The databases ship uncompressed, as
<id>.sqlite3. The browser reads byte ranges of them, and a range of a gzip stream is not a range of the database. The host must not applyContent-Encodingeither — check withcurl -sIafter a deploy. - Every query the site runs must be index-driven. Over range requests an
unindexed query fetches the whole table. Hence no index on
ho_ten(nothing can use one),name_wordfor name search, partial indexes for the score presets, and the footer count read fromdatasets.jsoninstead ofCOUNT(*).
Conventions
- Comments state current behavior and why. History belongs in git and
docs/, not in code comments. - Go files are CRLF in this checkout, so
gofmt -lflags every file. It is not usable as a formatting gate as things stand. - Conventional commits, no AI references.