Files
thptqg/CLAUDE.md
T
tiennm99 9571fb7d55 fix(web): take the database length from a range probe, not the host
GitHub Pages serves .sqlite3 as application/octet-stream, which mime-db
marks compressible, so an un-ranged response comes back gzipped with the
compressed length. sql.js-httpvfs sizes a file with a HEAD request, sees
that length is unusable and refuses to open the database:

  Length of the file not known. It must either be supplied in the config
  or given by the HTTP server.

Page reads were never affected. The Fetch standard requires browsers to
send Accept-Encoding: identity on any request carrying a Range header,
and the live site returns 206 with raw bytes to one. So the length is
probed the same way and passed as fileLength, which is the escape hatch
the library's own message points at.

The probe reads the first 100 bytes, so it also checks the file starts
with the SQLite magic and that its page size matches the request size —
a host that ever compresses a ranged response now fails with a clear
message rather than feeding the library the wrong bytes.

The post-deploy check the docs prescribed could not have caught this: a
bare `curl -sI` advertises no encoding, so it reports success whatever
the host does. It is replaced, in the docs and in CI, by a ranged read
that verifies the bytes.
2026-08-14 15:23:20 +07:00

4.2 KiB

thptqg

Exam-score lookup for Vietnam's national high school graduation exam. Four stages, each independent: crawler/ (Go) fetches spreadsheets into data/<id>/, parser/ (Go) turns them into <id>.db, assembler/ (Go) verifies, compresses and builds _site/, and web/ (SvelteKit + Tailwind) serves every dataset from one app.

datasets.json at the root is the registry the Go stages and the web app both read. See docs/ for architecture, pipeline and deployment.

Verifying a change

Run parser and assembler:

go -C parser test ./...
go -C assembler test ./...

Add the web checks when frontend files changed:

(cd web && npm test && npm run lint && npm run build)

npm run lint is ESLint. The web app is plain JavaScript — no TypeScript, no type-check step.

Do not run the crawler suite as part of routine verification. The crawler is not part of the build — data/<id>/ is committed and a crawl only refreshes it by hand — so its tests do not gate ordinary changes. Run them only when the change actually touches crawler/. CI still runs all three modules.

parser/internal/reader is the slow package (~9s) because the fidelity suite hashes every real input file. That is the point of it; do not skip it.

Things that look wrong but are not

  • data/2016/ filenames are content hashes and must stay verbatim. The parser sorts inputs bytewise and inserts last-wins, so filenames would decide which row survives a duplicate exam number. Renaming them also breaks parser/testdata/reader-fidelity-hashes.tsv, which is keyed by full path.
  • parser/testdata/reader-fidelity-hashes.tsv is frozen and cannot be regenerated. A mismatch is a reader bug until proven otherwise, never a cue to refresh the file. parser/cmd/dumpcells narrows a failure to the cell.
  • ToAscii filters the literal range U+0300..U+036F, not unicode.Mn. It must match toAscii in web/src/lib/to-ascii.js, or accent-insensitive search silently misses rows. to-ascii.test.js pins the pairs both sides must agree on.
  • The 2016 files use four different layouts, and detection is per sheet. Two of them publish scores in one column per subject instead of a DIEM_THI sentence, and one puts a three-row ministry title block above its header. parser/internal/ingest/detect2016.go holds the header tables; they are observations about 119 specific files, not a rule to generalise.
  • paths.relative is false in web/svelte.config.js. Every route is prerendered to its own file and asset URLs stay absolute, so the copy of index.html serving as 404.html works at any depth. No SPA 404-fallback is used — a fallback would break the ?q= deep links.
  • dbSizeMb in datasets.json is a build guard, not just a label. The assembler refuses to publish an artifact that falls below a ratio of it.
  • The databases ship uncompressed, as <id>.sqlite3. The browser reads byte ranges of them, and a range of a gzip stream is not a range of the database.
  • The file length comes from a range request, not from the host's HEAD. GitHub Pages gzips application/octet-stream, so a HEAD reports the compressed size and sql.js-httpvfs refuses to open the file. Ranged reads are unaffected — browsers must send Accept-Encoding: identity whenever a request carries a Range header — so web/src/lib/db-probe.js reads the header over a range and passes fileLength. Verify the way a browser asks: curl -sI -r 0-99 -H 'Accept-Encoding: identity' …, never a bare curl -sI, which advertises no encoding and hides the problem.
  • Every query the site runs must be index-driven. Over range requests an unindexed query fetches the whole table. Hence no index on ho_ten (nothing can use one), name_word for name search, partial indexes for the score presets, and the footer count read from datasets.json instead of COUNT(*).

Conventions

  • Comments state current behavior and why. History belongs in git and docs/, not in code comments.
  • Go files are CRLF in this checkout, so gofmt -l flags every file. It is not usable as a formatting gate as things stand.
  • Conventional commits, no AI references.