Move the frontend into web/, the repo's only npm workspace, and replace the JS crawler with a Go module covering both remaining datasets. The crawler writes to a .part file and renames on completion: writing straight to the destination left truncated files that the skip-if-present check would then skip forever. Remove the 2017-old and 2017-old2 datasets. They were successive publications of the same exam, kept side by side so the disagreement stayed inspectable; the current 2017 supersedes them and they remain in git history. Recover the 2016 crawler source from the Internet Archive's copy of the aggregator article, whose original host no longer resolves. All 119 filenames are verified against data/2016 in both directions, but no archive captured the spreadsheets themselves, so the host still serving them is unconfirmed and data/2016 remains the only confirmed copy. Filenames are load-bearing throughout: go-parser sorts inputs bytewise and inserts last-wins, so they decide which row survives a duplicate exam number.
2.6 KiB
Project Overview
Goal
A public lookup tool for Vietnam's National High School Graduation Exam scores, running entirely in the browser and hosted for free on GitHub Pages. Covers the 2016 and 2017 exams — 1.7 million candidates across two datasets.
Scope
- Lookup by exam ID or full name, with Vietnamese diacritics handled
- Read-only SQL queries against a single
studenttable - Admission-block (khối thi) totals computed per candidate
- Static datasets — both exams are long over and the data is frozen
Target users
- Former candidates checking their scores
- Education researchers and data journalists running aggregate statistics
- Developers exploring SQL against a real-world dataset
Constraints
- Zero backend. The full database (44–48 MB gzipped per dataset) is downloaded to the browser and queried in-process.
- Read-only.
INSERT/UPDATE/DELETEare rejected, so nobody is misled into thinking edits persist.sql.jsis in-memory anyway. - Row caps. 100 rows for lookups, 1000 for custom SQL, to prevent browser hangs.
- Vietnamese-first UI. App labels and data are Vietnamese; documentation is English.
Datasets
| id | Exam | Candidates | Notes |
|---|---|---|---|
2016 |
2016 | 877,461 | 119 files, three column layouts |
2017 |
2017 | 861,068 | current generation, reproducible from source |
Two further 2017 datasets (2017-old, 2017-old2) were kept alongside these
because the three publications disagreed. They have been removed; git history
still has them.
Both datasets now have a crawler source. 2017 comes from the baotintuc.vn CDN,
which is still live. 2016 comes from an aggregator article whose original host
(dtntbacgiang.edu.vn) no longer resolves — its link list was recovered from
the Internet Archive and is pointed at a mirror that is still online. See
data-pipeline for what that does and does not
guarantee.
History
Each year began as a standalone repository (thptqg2016, thptqg2017), merged
here with full history. They initially kept separate frontends and separate
copies of the same Rust parser, synchronised by hand. That duplication was
removed: there is now one frontend, one parser, and one canonical schema, with
per-dataset differences confined to one small config file and one registry
entry each.
The unification also fixed a latent data-loss bug — neither year's parser
config listed the complete set of subjects, so 1,691 candidates were missing
their foreign-language score. See data-pipeline.md.
Status
Stable, data frozen. Work is limited to UX polish and keeping the pipeline maintainable.