The docs still described two standalone projects with separate frontends,
separate parsers and three Vite variants. Rewritten around what the repo now
is, merging both projects' copies rather than keeping one and discarding the
other — deployment-guide.md and system-architecture.md existed in both and
documented different pipelines.
project-overview.md goal, scope, constraints, the four datasets, history
system-architecture.md data flow, canonical schema, routing, how one
frontend serves both exam years without branching
data-pipeline.md per-dataset Excel formats, the three 2016 layouts,
overflow-sheet gotcha, expected row counts
deployment-guide.md the single-build workflow, adding a dataset,
why no uncompressed database can ship
Records the release gate in plans/reports/parser-parity-result.md: row counts,
all 18 pre-existing per-column non-NULL counts and the deterministic student
samples are identical across all four datasets, against databases decompressed
from the exact bytes the pipeline publishes. The 1,691 recovered
foreign-language scores are documented with the evidence they are real.
Also drops a machine-specific absolute path from a comment in
format_detect_2016.rs. The build-database.js citations there are kept: that
file no longer exists in this repo, but the references explain why several
parsing rules look arbitrary.
thptqg
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
school graduation exam. Client-side SQL (sql.js) over a SQLite database built
from the ministry's raw .xls score files by the Rust xlsxread parser.
Live at tiennm99.github.io/thptqg.
| Dataset | Exam | Candidates | Site |
|---|---|---|---|
2016 |
2016 | 877,461 | /2016/ |
2017 |
2017 | 861,068 | /2017/ |
2017-old |
2017 | 847,348 | /2017-old/ |
2017-old2 |
2017 | 679,764 | /2017-old2/ |
The three 2017 datasets are successive publications of the same exam and they disagree; all three are kept so the differences stay inspectable.
Layout
index.html + src/ the frontend — one app serving all four datasets and the hub
datasets.js the four dataset ids and their per-dataset content
router.js pathname → dataset
data/<id>/ raw Excel files, one directory per dataset
parser/ the Rust parser
src/schema.rs canonical 22-column table: DDL, INSERT, subject regexes
configs/<id>.toml per-dataset parse rules only, no SQL
scripts/ database build, crawler, parity verification
scripts/ site assembly
docs/ architecture, data pipeline, deployment
The dataset id is one identifier end to end:
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
Build
npm ci
npm run build:rust # compile the parser
npm run build:db # build + gzip all four databases (add an id for just one)
npm run build:site # one Vite build, then assemble into _site/
npx serve _site
Pushing to main runs the same steps in
.github/workflows/deploy-pages.yml and publishes to GitHub Pages.
Adding a dataset
- Put the Excel files in
data/<id>/ - Add
parser/configs/<id>.toml— sheet mode, column indices, validation guards. No SQL; the schema is canonical. - Add an entry to
DATASETSinsrc/datasets.js
Everything else follows: the build script, the site assembly and the router all read that one list, and the UI adapts to whichever columns the dataset fills.
Docs
See docs/ — overview,
architecture,
data pipeline,
deployment.