Reading the file where it lay never worked well enough. Two costs were structural rather than bugs: reads are serial, because the worker uses synchronous XHR, so a name search touching 390 pages waited 17 seconds to move 608 KB — roughly one request per result row, which no page size removes — and the first visitor after each deploy waited ~26 seconds for the CDN to fill its cache with a 288 MB object. The browser now downloads the whole database once and queries it in memory with sql.js. A dataset page is gated behind that: the gate states what it will cost, in transfer and in memory, and offers only the download, because there is nothing to show without it. Dropping the structures that existed to make range-request queries index-driven halved the file. name_word carried one row per word of every name, about 3.5 million of them, and with the partial score indexes it was more than half of what every visitor would now download. Measured on rebuilt databases: 2016 went 288.6 -> 142.5 MB (31 MB gzipped on the wire), 2017 237.7 -> 119.3 MB, both with row counts and audits unchanged. Queries on the result: an exam number is immediate, a name scans all 877,460 rows in about 240 ms. Alternatives were measured before choosing this. sqlite-wasm-http sizes files from a HEAD Content-Length with no override, so on a host that gzips it silently uses the compressed size. DuckDB-WASM ships 32-37 MB of WebAssembly before its Parquet extension, more than this whole download. Static pre-generated shards are the most robust option but cannot answer arbitrary SQL, and cannot stop early the way LIMIT does. The published name loses its chunk index, the byte budgets and the SQL consent modal go with the range reads that made them necessary, and the docs no longer describe a design the site does not use.
3.9 KiB
thptqg
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
school graduation exam. Client-side SQL over a SQLite database the browser
downloads once and queries in memory, built
from the published .xls/.xlsx score files by the Go parser module. Where
those files come from: data pipeline.
Live at tiennm99.github.io/thptqg.
| Dataset | Exam | Candidates | Site |
|---|---|---|---|
2016 |
2016 | 877,460 | /2016/ |
2017 |
2017 | 861,068 | /2017/ |
Two earlier 2017 publications (2017-old, 2017-old2) were kept for a while
because they disagreed with the current one. They have been removed; they remain
in git history.
Layout
The repository is one directory per pipeline stage, plus the two stores they pass between them.
crawler/ Go — re-fetches the source spreadsheets → data/
parser/ Go — Excel to SQLite data/ → .db
assembler/ Go — verifies, compresses, builds, assembles .db + web/ → _site/
web/ npm — the frontend, one SvelteKit app for every dataset
data/<id>/ raw Excel files, one directory per dataset
datasets.json the registry: which datasets exist, and their expected size
docs/ architecture, data pipeline, deployment
Each stage runs on its own and hands its output to the next through the stores.
web/ is the only npm project; the three stages are independent Go modules.
datasets.json is the contract between them. It is JSON because Go and the web
app both read it and neither needs a dependency to do so; presentation stays in
web/src/lib/datasets.js, keyed by id, which fails loudly if the two disagree.
The dataset id is one identifier end to end:
data/2017/ → parser/configs/2017.yml → db/2017.sqlite3 → /thptqg/2017/
Build
(cd web && npm ci)
go -C assembler run ./cmd/assemble # databases, then the site, into _site/
npx serve _site
That one command compiles the parser, builds and verifies each database against
its registry row count, compresses it, builds the web app and assembles _site —
refusing to continue if a database is short, an artifact looks truncated, or one
is missing altogether. Sub-steps when iterating:
go -C assembler run ./cmd/assemble db 2017 # one database
go -C assembler run ./cmd/assemble site # web build and _site only
go -C assembler run ./cmd/assemble verify A B # compare two sets of databases
(cd web && npm run dev) # the app against staged databases
The source spreadsheets are committed, so a crawl is only needed to refresh them:
go -C crawler run ./cmd/crawl 2016
go -C crawler run ./cmd/crawl 2017
Each reads the download links out of the article that published the dataset, so no link list is kept in the repository. Crawling is idempotent — files already present are skipped — and is never part of the build.
Pushing to main runs the same steps in
.github/workflows/deploy-pages.yml and publishes to GitHub Pages.
Adding a dataset
- Put the Excel files in
data/<id>/ - Add
parser/configs/<id>.yml— sheet mode, column indices, validation guards. No SQL; the schema is canonical. - Add an entry to
datasets.jsonwith its expected row count and size - Add the matching presentation to
CONTENTinweb/src/lib/datasets.js
Everything else follows: the assembler, the router and the hub all read the registry, and the UI adapts to whichever columns the dataset fills. Steps 3 and 4 check each other, so forgetting either one fails rather than half-working.
Docs
See docs/ — overview,
architecture,
data pipeline,
deployment.