Files
thptqg/CLAUDE.md
T
tiennm99 d4f106f488 Merge branch 'fix/httpvfs-chunked-length'
# Conflicts:
#	web/src/lib/db-probe.test.js
#	web/src/lib/sqlite.svelte.js
2026-10-03 11:59:22 +07:00

4.7 KiB

thptqg

Exam-score lookup for Vietnam's national high school graduation exam. Four stages, each independent: crawler/ (Go) fetches spreadsheets into data/<id>/, parser/ (Go) turns them into <id>.sqlite3, assembler/ (Go) verifies them and builds _site/, and web/ (SvelteKit + Tailwind) serves every dataset from one app.

datasets.json at the root is the registry the Go stages and the web app both read. See docs/ for architecture, pipeline and deployment.

Verifying a change

Run parser and assembler:

go -C parser test ./...
go -C assembler test ./...

Add the web checks when frontend files changed:

(cd web && npm test && npm run lint && npm run build)

npm run lint is ESLint. The web app is plain JavaScript — no TypeScript, no type-check step.

Do not run the crawler suite as part of routine verification. The crawler is not part of the build — data/<id>/ is committed and a crawl only refreshes it by hand — so its tests do not gate ordinary changes. Run them only when the change actually touches crawler/. CI still runs all three modules.

parser/internal/reader is the slow package (~9s) because the fidelity suite hashes every real input file. That is the point of it; do not skip it.

Things that look wrong but are not

  • data/2016/ filenames are content hashes and must stay verbatim. The parser sorts inputs bytewise and inserts last-wins, so filenames would decide which row survives a duplicate exam number. Renaming them also breaks parser/testdata/reader-fidelity-hashes.tsv, which is keyed by full path.
  • parser/testdata/reader-fidelity-hashes.tsv is frozen and cannot be regenerated. A mismatch is a reader bug until proven otherwise, never a cue to refresh the file. parser/cmd/dumpcells narrows a failure to the cell.
  • ToAscii filters the literal range U+0300..U+036F, not unicode.Mn. It must match toAscii in web/src/lib/to-ascii.js, or accent-insensitive search silently misses rows. to-ascii.test.js pins the pairs both sides must agree on.
  • The 2016 files use four different layouts, and detection is per sheet. Two of them publish scores in one column per subject instead of a DIEM_THI sentence, and one puts a three-row ministry title block above its header. parser/internal/ingest/detect2016.go holds the header tables; they are observations about 119 specific files, not a rule to generalise.
  • paths.relative is false in web/svelte.config.js. Every route is prerendered to its own file and asset URLs stay absolute, so the copy of index.html serving as 404.html works at any depth. No SPA 404-fallback is used — a fallback would break the ?q= deep links.
  • dbSizeMb in datasets.json is a build guard, not just a label. The assembler refuses to publish an artifact that falls below a ratio of it.
  • The databases ship uncompressed, as <id>.sqlite30. The trailing 0 is a chunk index, not a typo — see the next point. The browser reads byte ranges of the file, and a range of a gzip stream is not a range of the database.
  • The site uses chunked mode over a single chunk, and that is deliberate. GitHub Pages gzips application/octet-stream, so the HEAD request sql.js-httpvfs sizes a file with reports the compressed length, and the library refuses to open the file. Chunked mode is the only mode whose config takes a length (databaseLengthBytes); in full mode the worker hardcodes it to undefined, so a length passed there is silently dropped. One chunk holds the database, so the index is always 0 and every request goes to <id>.sqlite3 + 0. web/src/lib/db-probe.js supplies the length by reading the file header over a range request.
  • Ranged reads were never affected by the compression, because browsers must send Accept-Encoding: identity whenever a request carries a Range header. Verify the way a browser asks — curl -s -r 0-14 -H 'Accept-Encoding: identity' … must print SQLite format 3 — never a bare curl -sI, which advertises no encoding and so passes whatever the host does.
  • Every query the site runs must be index-driven. Over range requests an unindexed query fetches the whole table. Hence no index on ho_ten (nothing can use one), name_word for name search, partial indexes for the score presets, and the footer count read from datasets.json instead of COUNT(*).

Conventions

  • Comments state current behavior and why. History belongs in git and docs/, not in code comments.
  • Go files are CRLF in this checkout, so gofmt -l flags every file. It is not usable as a formatting gate as things stand.
  • Conventional commits, no AI references.