GitHub Pages serves .sqlite3 as application/octet-stream, which mime-db marks compressible, so an un-ranged response comes back gzipped with the compressed length. sql.js-httpvfs sizes a file with a HEAD request, sees that length is unusable and refuses to open the database: Length of the file not known. It must either be supplied in the config or given by the HTTP server. Page reads were never affected. The Fetch standard requires browsers to send Accept-Encoding: identity on any request carrying a Range header, and the live site returns 206 with raw bytes to one. So the length is probed the same way and passed as fileLength, which is the escape hatch the library's own message points at. The probe reads the first 100 bytes, so it also checks the file starts with the SQLite magic and that its page size matches the request size — a host that ever compresses a ranged response now fails with a clear message rather than feeding the library the wrong bytes. The post-deploy check the docs prescribed could not have caught this: a bare `curl -sI` advertises no encoding, so it reports success whatever the host does. It is replaced, in the docs and in CI, by a ranged read that verifies the bytes.
4.2 KiB
thptqg
Exam-score lookup for Vietnam's national high school graduation exam. Four
stages, each independent: crawler/ (Go) fetches spreadsheets into data/<id>/,
parser/ (Go) turns them into <id>.db, assembler/ (Go) verifies, compresses
and builds _site/, and web/ (SvelteKit + Tailwind) serves every dataset from
one app.
datasets.json at the root is the registry the Go stages and the web app both
read. See docs/ for architecture, pipeline and deployment.
Verifying a change
Run parser and assembler:
go -C parser test ./...
go -C assembler test ./...
Add the web checks when frontend files changed:
(cd web && npm test && npm run lint && npm run build)
npm run lint is ESLint. The web app is plain JavaScript — no TypeScript, no
type-check step.
Do not run the crawler suite as part of routine verification. The crawler is
not part of the build — data/<id>/ is committed and a crawl only refreshes it
by hand — so its tests do not gate ordinary changes. Run them only when the
change actually touches crawler/. CI still runs all three modules.
parser/internal/reader is the slow package (~9s) because the fidelity suite
hashes every real input file. That is the point of it; do not skip it.
Things that look wrong but are not
data/2016/filenames are content hashes and must stay verbatim. The parser sorts inputs bytewise and inserts last-wins, so filenames would decide which row survives a duplicate exam number. Renaming them also breaksparser/testdata/reader-fidelity-hashes.tsv, which is keyed by full path.parser/testdata/reader-fidelity-hashes.tsvis frozen and cannot be regenerated. A mismatch is a reader bug until proven otherwise, never a cue to refresh the file.parser/cmd/dumpcellsnarrows a failure to the cell.ToAsciifilters the literal range U+0300..U+036F, notunicode.Mn. It must matchtoAsciiinweb/src/lib/to-ascii.js, or accent-insensitive search silently misses rows.to-ascii.test.jspins the pairs both sides must agree on.- The 2016 files use four different layouts, and detection is per sheet.
Two of them publish scores in one column per subject instead of a
DIEM_THIsentence, and one puts a three-row ministry title block above its header.parser/internal/ingest/detect2016.goholds the header tables; they are observations about 119 specific files, not a rule to generalise. paths.relativeis false inweb/svelte.config.js. Every route is prerendered to its own file and asset URLs stay absolute, so the copy ofindex.htmlserving as404.htmlworks at any depth. No SPA 404-fallback is used — a fallback would break the?q=deep links.dbSizeMbindatasets.jsonis a build guard, not just a label. The assembler refuses to publish an artifact that falls below a ratio of it.- The databases ship uncompressed, as
<id>.sqlite3. The browser reads byte ranges of them, and a range of a gzip stream is not a range of the database. - The file length comes from a range request, not from the host's HEAD.
GitHub Pages gzips
application/octet-stream, so a HEAD reports the compressed size andsql.js-httpvfsrefuses to open the file. Ranged reads are unaffected — browsers must sendAccept-Encoding: identitywhenever a request carries aRangeheader — soweb/src/lib/db-probe.jsreads the header over a range and passesfileLength. Verify the way a browser asks:curl -sI -r 0-99 -H 'Accept-Encoding: identity' …, never a barecurl -sI, which advertises no encoding and hides the problem. - Every query the site runs must be index-driven. Over range requests an
unindexed query fetches the whole table. Hence no index on
ho_ten(nothing can use one),name_wordfor name search, partial indexes for the score presets, and the footer count read fromdatasets.jsoninstead ofCOUNT(*).
Conventions
- Comments state current behavior and why. History belongs in git and
docs/, not in code comments. - Go files are CRLF in this checkout, so
gofmt -lflags every file. It is not usable as a formatting gate as things stand. - Conventional commits, no AI references.