The whole design rests on a claim nobody has measured: that a lookup
fetches a few hundred KB rather than the file. Every query now reports
what it actually cost.
[httpvfs] name "nguyen buu loc" seeking "buu" — 14 request(s), 78 KB,
212 ms · session 31 request(s), 180 KB of 302.4 MB
The numbers come from the worker's getStats() rather than its bytesRead
counter: bytesRead is the budget accumulator and resets itself to zero
the moment a query trips the ceiling, so it would under-report exactly
when the number matters. getStats() reports per-file totals including
what the read heads prefetched, which is the honest figure — prefetch
overfetch is invisible to a page count.
Opening a database logs too, since the header and schema pages are read
before any query and would otherwise inflate the first search.
The SQL tab shows the same pair on screen: this query, then the session.
The browser fetches this file one page per HTTP request, so the page size
is the granularity of every read. At SQLite's 4 KiB default a row reached
by an index seek dragged 4 KB across the network; at 1 KiB it drags 1 KB.
A name search returns up to 100 scattered rows, so its row fetches fall
from about 400 KB to about 100 KB.
Measured on the rebuilt 2016 file: 6.3 rows share a page where 27 did.
The index walks are sequential and unaffected in bytes — the library's
read-ahead already collapses those into few requests.
Cost is 4% file size: 2016 288.6 -> 302.4 MB, 2017 237.7 -> 247.3 MB,
the site 528 -> 552 MB against the 1 GB GitHub Pages limit. Both
sql.js-httpvfs and sqlite-wasm-http recommend this page size.
The PRAGMA has to run before the DDL, since a page size is fixed once a
table exists, and requestChunkSize on the client has to match or every
page read spans two requests.
Row counts unchanged and through the assembler guards; query plans
re-checked and still index-driven on the rebuilt files.
The browser downloaded 45 MB of gzipped SQLite before it could answer
anything. Now sql.js-httpvfs asks for the pages a query touches and the
databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip
stream is not a byte range of a database.
That only works if every query the site issues is index-driven, and
measured against the real 2016 file, most were not:
so_bao_danh = ? SEARCH via PK ~20 KB
ho_ten_ascii LIKE '%x%' SCAN 127 MB
ho_ten_ascii LIKE 'x%' SCAN 127 MB
COUNT(*) covering index scan 20 MB
ORDER BY toan DESC LIMIT 10 SCAN + temp b-tree 127 MB
Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE
index; a range comparison does use the index. So the schema changed to
suit the access pattern rather than the search changing to suit the
schema.
name_word holds one row per word of each name, WITHOUT ROWID so the
table is the index, carrying ho_ten_ascii so a multi-word query is
resolved inside a single b-tree. name_word_freq says which word of a
query is rarest — the vocabulary is 4,397 words across 2.87M entries, so
"buu loc" seeks on 287 entries rather than walking the 300,000 that
"thi" would. Searching by any word of a name survives, at a few hundred
KB a query.
idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either.
Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL
presets off a full scan. The footer's candidate count now comes from
datasets.json instead of COUNT(*).
2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site
is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged.
The SQL tab is the one place a user can still write a query that reads
the whole table, so it asks before it opens, runs under a byte budget
that stops a runaway query, and shows what each query actually fetched.
Verified: row counts through the assembler guards, every app query
index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206
with a correct Content-Range. Not verified in a browser — this machine
has none — and the library refuses to open a file the host compresses,
so the deployed response headers need a look.
The app was a single-entry React SPA: one index.html that the assembler
copied to every dataset path, with a hand-rolled router resolving the
dataset from window.location and the title patched in at runtime because
one file had to serve every route.
SvelteKit prerenders a real page per route instead. The entry generator
in routes/[dataset]/+page.ts reads the same datasets.json the assembler
does, so the set of pages and the set of databases cannot drift apart,
and each dataset page ships its own <title> and description.
Everything framework-free moved across unchanged in behaviour and became
typed: the admission blocks, subject list, query classifier and SQL
presets. `Student` now mirrors the 22-column table, so a mistyped column
name fails the build rather than rendering blank.
Tailwind replaces the stylesheet. Theme tokens are CSS variables, which
keeps dark mode a single block of overrides rather than a `dark:` variant
on every class. The score tiers stay hand-written CSS: the class is
chosen at runtime from a score, and no utility generator can see that.
Adds the tests the frontend never had, over the three modules where a
silent wrong answer is possible — most importantly that toAscii here
folds exactly as ToAscii does in the parser, which is what makes
accent-insensitive search find anything.
The assembler stops copying index.html per dataset and checks that the
build prerendered each one instead. The database presence, raw-artifact
and idempotence guards are untouched.
Verified: 25 tests, ESLint and svelte-check clean, and a full
`assemble site` producing /thptqg/, /thptqg/2016/, /thptqg/2017/ and
404.html with absolute asset URLs. Not verified in a browser — this
machine has none.
The 2016 exam numbers are nine digits or a letter-prefixed cluster code,
so the eight-digit "17006021" matched nothing and query-mode.js read it
as a 2017 number. TKG002747 / Nguyễn Bửu Lộc is a real row.
Comments across the tree justified the code by pointing at a Rust
implementation that is no longer in the repository, citing files and line
numbers (config.rs:132, schema.rs:26-54, reader.rs:42) that cannot be
opened, plus crates and datasets that are equally gone. A reader could not
check any of it.
Every invariant those comments carried is kept and restated so it stands on
its own: the bytewise sort that decides which row survives a duplicate exam
number, the literal U+0300..U+036F range that must match the site's toAscii,
the trailing space in "SINH ", the BIFF and shared-string corrections, the
VACUUM-after-COMMIT rule, the deploy-from-main guard.
The reader's contract is now anchored to the frozen oracle in
parser/testdata, which still exists and is still checked, rather than to the
tool that originally produced it.
TestDDLMatchesRust becomes TestDDLIsFrozen: it compares against a copy of
the DDL inside the test and never read schema.rs, so both the name and the
failure message were misleading.
ToAscii no longer claims the d-replacement must precede lowercasing. Both
cases map to 'd' and ToLower runs last, so the order has no effect.
The site credited 2016 to Bộ GD&ĐT, which those files were never fetched
from. Both datasets come from published articles: 2016 from an aggregator
on dtnt.bacninh.edu.vn listing one spreadsheet per exam cluster, 2017 from
baotintuc.vn. The README, the architecture table and the web footer all
repeated the ministry claim.
The footer now shows each dataset's full article URL as a link rather than
a bare host, so the citation can be checked. That needs overflow-wrap on
the footer: the 2016 URL is 110 characters with no break opportunity and
would otherwise scroll the page sideways on a phone.
Also records that a full 2016 crawl has been run successfully. The host was
marked unconfirmed and data/2016/ described as the only recoverable copy;
both datasets are now rebuildable from source.
The pipeline is now Go outside web/. differential-parity.mjs becomes
assembler/internal/verify, reachable as `assemble verify A B`. The port fixed a
real weakness: the JavaScript hashed each row's fields joined bare, so a value
shifted across a column boundary produced the same digest. A test now pins that.
The hub still rendered "Phiên bản cũ của trang 2017" above a permanently empty
list — it split datasets on id.includes("old"), and both such datasets are gone.
The heading and the filter are removed. index.html titled every page "THPT QG
2017", including 2016 and the hub, because one file is copied to every route;
the static title is now neutral and the app sets the dataset's own.
Dead code removed: the isOld2/containsOld branches in the stats block, which
only 2017-old2 could ever reach; SUBJECT_LABELS, DATASET_IDS and the unread
`short` subject field; an unused vite.svg and a favicon link to a file that
never existed; two unused CSS rules and --shadow-sm; site.Paths.Root.
Corrected comments that were confidently wrong rather than merely stale: the
reader claimed to be row-streaming when both implementations decode the whole
workbook into memory first, and the fidelity oracle still spoke of 299 input
files when it covers 182. Candidate counts in the hub now derive from
datasets.json instead of being written a second time as prose.
plans/ is emptied. The parity report it held was cited by docs/data-pipeline.md,
so the evidence that the recovered foreign-language scores are real — not the
citation, the four arguments themselves — is now inline there.
Verified: 2017 rebuilt after the writer change hashes identically to the build
before it.
The repository now reads as the pipeline it is: crawler fetches, parser
converts, assembler verifies and publishes, with data/ and web/ as the stores
they hand work through. go-parser is renamed parser now that there is no other.
The assembler replaces build-db.js and assemble-site.js. It compiles the
parser, builds and verifies each database, compresses it, runs the Vite build
and assembles _site — one command, and the only place that knows the order.
It also closes a real hole: nothing previously asserted that a database reached
the site. An empty staging directory assembled happily, so every page rendered,
every query 404d and CI stayed green. The row-count and size guards could not
catch that, since they only run when a database was built at all.
Removing Node from the root forced the dataset list out of web/src/datasets.js,
which the assembler cannot import. datasets.json is now the registry both sides
read — JSON because Go and the browser both parse it without a dependency —
while presentation stays in the web app, keyed by id and cross-checked against
the registry so a half-added dataset fails instead of half-working.
Guards verified by making each one fail: a missing database, and an expected
row count one higher than the truth.
Move the frontend into web/, the repo's only npm workspace, and replace the
JS crawler with a Go module covering both remaining datasets. The crawler
writes to a .part file and renames on completion: writing straight to the
destination left truncated files that the skip-if-present check would then
skip forever.
Remove the 2017-old and 2017-old2 datasets. They were successive publications
of the same exam, kept side by side so the disagreement stayed inspectable;
the current 2017 supersedes them and they remain in git history.
Recover the 2016 crawler source from the Internet Archive's copy of the
aggregator article, whose original host no longer resolves. All 119 filenames
are verified against data/2016 in both directions, but no archive captured the
spreadsheets themselves, so the host still serving them is unconfirmed and
data/2016 remains the only confirmed copy.
Filenames are load-bearing throughout: go-parser sorts inputs bytewise and
inserts last-wins, so they decide which row survives a duplicate exam number.