Files
thptqg/docs/project-overview.md
T
tiennm99 ceb694a747 refactor: split into web/crawler/go-parser and drop the 2017 archives
Move the frontend into web/, the repo's only npm workspace, and replace the
JS crawler with a Go module covering both remaining datasets. The crawler
writes to a .part file and renames on completion: writing straight to the
destination left truncated files that the skip-if-present check would then
skip forever.

Remove the 2017-old and 2017-old2 datasets. They were successive publications
of the same exam, kept side by side so the disagreement stayed inspectable;
the current 2017 supersedes them and they remain in git history.

Recover the 2016 crawler source from the Internet Archive's copy of the
aggregator article, whose original host no longer resolves. All 119 filenames
are verified against data/2016 in both directions, but no archive captured the
spreadsheets themselves, so the host still serving them is unconfirmed and
data/2016 remains the only confirmed copy.

Filenames are load-bearing throughout: go-parser sorts inputs bytewise and
inserts last-wins, so they decide which row survives a duplicate exam number.
2026-08-13 21:45:05 +07:00

2.6 KiB
Raw Blame History

Project Overview

Goal

A public lookup tool for Vietnam's National High School Graduation Exam scores, running entirely in the browser and hosted for free on GitHub Pages. Covers the 2016 and 2017 exams — 1.7 million candidates across two datasets.

Scope

  • Lookup by exam ID or full name, with Vietnamese diacritics handled
  • Read-only SQL queries against a single student table
  • Admission-block (khối thi) totals computed per candidate
  • Static datasets — both exams are long over and the data is frozen

Target users

  • Former candidates checking their scores
  • Education researchers and data journalists running aggregate statistics
  • Developers exploring SQL against a real-world dataset

Constraints

  • Zero backend. The full database (44–48 MB gzipped per dataset) is downloaded to the browser and queried in-process.
  • Read-only. INSERT/UPDATE/DELETE are rejected, so nobody is misled into thinking edits persist. sql.js is in-memory anyway.
  • Row caps. 100 rows for lookups, 1000 for custom SQL, to prevent browser hangs.
  • Vietnamese-first UI. App labels and data are Vietnamese; documentation is English.

Datasets

id Exam Candidates Notes
2016 2016 877,461 119 files, three column layouts
2017 2017 861,068 current generation, reproducible from source

Two further 2017 datasets (2017-old, 2017-old2) were kept alongside these because the three publications disagreed. They have been removed; git history still has them.

Both datasets now have a crawler source. 2017 comes from the baotintuc.vn CDN, which is still live. 2016 comes from an aggregator article whose original host (dtntbacgiang.edu.vn) no longer resolves — its link list was recovered from the Internet Archive and is pointed at a mirror that is still online. See data-pipeline for what that does and does not guarantee.

History

Each year began as a standalone repository (thptqg2016, thptqg2017), merged here with full history. They initially kept separate frontends and separate copies of the same Rust parser, synchronised by hand. That duplication was removed: there is now one frontend, one parser, and one canonical schema, with per-dataset differences confined to one small config file and one registry entry each.

The unification also fixed a latent data-loss bug — neither year's parser config listed the complete set of subjects, so 1,691 candidates were missing their foreign-language score. See data-pipeline.md.

Status

Stable, data frozen. Work is limited to UX polish and keeping the pipeline maintainable.