Files
thptqg/parser
tiennm99 dbc23c25c5 feat: read the databases over HTTP range requests
The browser downloaded 45 MB of gzipped SQLite before it could answer
anything. Now sql.js-httpvfs asks for the pages a query touches and the
databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip
stream is not a byte range of a database.

That only works if every query the site issues is index-driven, and
measured against the real 2016 file, most were not:

  so_bao_danh = ?              SEARCH via PK          ~20 KB
  ho_ten_ascii LIKE '%x%'      SCAN                   127 MB
  ho_ten_ascii LIKE 'x%'       SCAN                   127 MB
  COUNT(*)                     covering index scan     20 MB
  ORDER BY toan DESC LIMIT 10  SCAN + temp b-tree     127 MB

Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE
index; a range comparison does use the index. So the schema changed to
suit the access pattern rather than the search changing to suit the
schema.

name_word holds one row per word of each name, WITHOUT ROWID so the
table is the index, carrying ho_ten_ascii so a multi-word query is
resolved inside a single b-tree. name_word_freq says which word of a
query is rarest — the vocabulary is 4,397 words across 2.87M entries, so
"buu loc" seeks on 287 entries rather than walking the 300,000 that
"thi" would. Searching by any word of a name survives, at a few hundred
KB a query.

idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either.
Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL
presets off a full scan. The footer's candidate count now comes from
datasets.json instead of COUNT(*).

2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site
is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged.

The SQL tab is the one place a user can still write a query that reads
the whole table, so it asks before it opens, runs under a byte budget
that stops a runaway query, and shows what each query actually fetched.

Verified: row counts through the assembler guards, every app query
index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206
with a correct Content-Range. Not verified in a browser — this machine
has none — and the library refuses to open a file the host compresses,
so the deployed response headers need a look.
2026-08-14 12:42:48 +07:00
..

parser

Reads the .xls/.xlsx source spreadsheets in data/ and writes one SQLite database per dataset.

go -C parser build -o bin/xlsxread ./cmd/xlsxread   # compile
go -C parser test ./...                             # unit tests + the reader-fidelity suite
xlsxread build --schema parser/configs/<id>.yml --input data/<id> --output <db>
xlsxread audit --schema parser/configs/<id>.yml --input data/<id> --db <db>

This stage only produces a database. Verifying it against the expected row count, compressing it and publishing it belong to assembler/, which compiles this binary and drives it per dataset:

go -C assembler run ./cmd/assemble db

Layout

path role
internal/reader spreadsheet reading; the only place that knows about file formats
internal/ingest dataset policy — sheet selection, header skipping, blank rows, the build loop, and the 2016 per-sheet layout detection
internal/transform ToAscii, score-regex parsing, row validation
internal/schema the canonical 22-column table: DDL, INSERT, subject regexes
internal/config per-dataset YAML parse rules
internal/writer SQLite lifecycle and the stats block
internal/audit source-vs-database SBD comparison

The reader deliberately knows nothing about datasets: it reports every sheet and every row verbatim. All policy lives in ingest. That split is what made the reader independently verifiable against a hash oracle.

Behaviour worth knowing

  • ToAscii strips combining marks in the literal range U+0300–U+036F rather than by Unicode category. That covers every Vietnamese diacritic and must stay identical to toAscii in web/src/lib/to-ascii.ts, or accent-insensitive search misses rows.
  • Gender is normalised to Nam/Nữ; the Cần Thơ files write 0/1 instead and are translated. Anything else becomes NULL.
  • A score of 0 is a real score — the candidate sat the paper and scored nothing — and is stored, not dropped.
  • Birth dates are stored as dd/mm/yyyy. The Cần Thơ files' compact ddmmyy is expanded; the century is always 19xx, since a 2016 candidate born later would have sat the exam under age.

This code began as a port of a Rust crate that occupied the same path, and was gated on a field-by-field comparison against it. That comparison is over: correctness against the source spreadsheets decides behaviour now, not agreement with the old implementation. assemble verify is the comparator that gated it, and still compares any two sets of built databases.

Verification

testdata/reader-fidelity-hashes.tsv holds a SHA-256 per input file over a canonical dump of every cell of every sheet. It is frozen: the tool that produced it no longer exists, so it cannot be regenerated. It still fails if any single cell of any input file reads differently, which is what makes it useful — cell rendering and sheet geometry are settled, and a change there is a regression until proven otherwise. Dataset policy lives in ingest, so fixing a layout never touches it.

A mismatch names the file but not the cell. cmd/dumpcells prints the stream the hash is taken over, so two runs can be diffed:

go -C parser run ./cmd/dumpcells ../data/2017/an-giang.xls out.tsv

The assembler refuses to publish a database whose row count does not match the known figure, or whose artifact is under 90% of its usual size.