4.7 KiB
thptqg
Exam-score lookup for Vietnam's national high school graduation exam. Four
stages, each independent: crawler/ (Go) fetches spreadsheets into data/<id>/,
parser/ (Go) turns them into <id>.sqlite3, assembler/ (Go) verifies them
and builds _site/, and web/ (SvelteKit + Tailwind) serves every dataset from
one app.
datasets.json at the root is the registry the Go stages and the web app both
read. See docs/ for architecture, pipeline and deployment.
Verifying a change
Run parser and assembler:
go -C parser test ./...
go -C assembler test ./...
Add the web checks when frontend files changed:
(cd web && npm test && npm run lint && npm run build)
npm run lint is ESLint. The web app is plain JavaScript — no TypeScript, no
type-check step.
Do not run the crawler suite as part of routine verification. The crawler is
not part of the build — data/<id>/ is committed and a crawl only refreshes it
by hand — so its tests do not gate ordinary changes. Run them only when the
change actually touches crawler/. CI still runs all three modules.
parser/internal/reader is the slow package (~9s) because the fidelity suite
hashes every real input file. That is the point of it; do not skip it.
Things that look wrong but are not
data/2016/filenames are content hashes and must stay verbatim. The parser sorts inputs bytewise and inserts last-wins, so filenames would decide which row survives a duplicate exam number. Renaming them also breaksparser/testdata/reader-fidelity-hashes.tsv, which is keyed by full path.parser/testdata/reader-fidelity-hashes.tsvis frozen and cannot be regenerated. A mismatch is a reader bug until proven otherwise, never a cue to refresh the file.parser/cmd/dumpcellsnarrows a failure to the cell.ToAsciifilters the literal range U+0300..U+036F, notunicode.Mn. It must matchtoAsciiinweb/src/lib/to-ascii.js, or accent-insensitive search silently misses rows.to-ascii.test.jspins the pairs both sides must agree on.- The 2016 files use four different layouts, and detection is per sheet.
Two of them publish scores in one column per subject instead of a
DIEM_THIsentence, and one puts a three-row ministry title block above its header.parser/internal/ingest/detect2016.goholds the header tables; they are observations about 119 specific files, not a rule to generalise. paths.relativeis false inweb/svelte.config.js. Every route is prerendered to its own file and asset URLs stay absolute, so the copy ofindex.htmlserving as404.htmlworks at any depth. No SPA 404-fallback is used — a fallback would break the?q=deep links.dbSizeMbindatasets.jsonis a build guard, not just a label. The assembler refuses to publish an artifact that falls below a ratio of it.- The databases ship uncompressed, as
<id>.sqlite30. The trailing 0 is a chunk index, not a typo — see the next point. The browser reads byte ranges of the file, and a range of a gzip stream is not a range of the database. - The site uses chunked mode over a single chunk, and that is deliberate.
GitHub Pages gzips
application/octet-stream, so the HEAD requestsql.js-httpvfssizes a file with reports the compressed length, and the library refuses to open the file. Chunked mode is the only mode whose config takes a length (databaseLengthBytes); in full mode the worker hardcodes it toundefined, so a length passed there is silently dropped. One chunk holds the database, so the index is always 0 and every request goes to<id>.sqlite3+0.web/src/lib/db-probe.jssupplies the length by reading the file header over a range request. - Ranged reads were never affected by the compression, because browsers
must send
Accept-Encoding: identitywhenever a request carries aRangeheader. Verify the way a browser asks —curl -s -r 0-14 -H 'Accept-Encoding: identity' …must printSQLite format 3— never a barecurl -sI, which advertises no encoding and so passes whatever the host does. - Every query the site runs must be index-driven. Over range requests an
unindexed query fetches the whole table. Hence no index on
ho_ten(nothing can use one),name_wordfor name search, partial indexes for the score presets, and the footer count read fromdatasets.jsoninstead ofCOUNT(*).
Conventions
- Comments state current behavior and why. History belongs in git and
docs/, not in code comments. - Go files are CRLF in this checkout, so
gofmt -lflags every file. It is not usable as a formatting gate as things stand. - Conventional commits, no AI references.