Files
thptqg/docs/project-overview.md
T
tiennm99 c988bfafcf fix: correct the 2016 source attribution and link full article URLs
The site credited 2016 to Bộ GD&ĐT, which those files were never fetched
from. Both datasets come from published articles: 2016 from an aggregator
on dtnt.bacninh.edu.vn listing one spreadsheet per exam cluster, 2017 from
baotintuc.vn. The README, the architecture table and the web footer all
repeated the ministry claim.

The footer now shows each dataset's full article URL as a link rather than
a bare host, so the citation can be checked. That needs overflow-wrap on
the footer: the 2016 URL is 110 characters with no break opportunity and
would otherwise scroll the page sideways on a phone.

Also records that a full 2016 crawl has been run successfully. The host was
marked unconfirmed and data/2016/ described as the only recoverable copy;
both datasets are now rebuildable from source.
2026-08-14 09:26:27 +07:00

2.6 KiB
Raw Blame History

Project Overview

Goal

A public lookup tool for Vietnam's National High School Graduation Exam scores, running entirely in the browser and hosted for free on GitHub Pages. Covers the 2016 and 2017 exams — 1.7 million candidates across two datasets.

Scope

  • Lookup by exam ID or full name, with Vietnamese diacritics handled
  • Read-only SQL queries against a single student table
  • Admission-block (khối thi) totals computed per candidate
  • Static datasets — both exams are long over and the data is frozen

Target users

  • Former candidates checking their scores
  • Education researchers and data journalists running aggregate statistics
  • Developers exploring SQL against a real-world dataset

Constraints

  • Zero backend. The full database (44–48 MB gzipped per dataset) is downloaded to the browser and queried in-process.
  • Read-only. INSERT/UPDATE/DELETE are rejected, so nobody is misled into thinking edits persist. sql.js is in-memory anyway.
  • Row caps. 100 rows for lookups, 1000 for custom SQL, to prevent browser hangs.
  • Vietnamese-first UI. App labels and data are Vietnamese; documentation is English.

Datasets

id Exam Candidates Notes
2016 2016 877,461 119 files, three column layouts
2017 2017 861,068 current generation of three publications

Two further 2017 datasets (2017-old, 2017-old2) were kept alongside these because the three publications disagreed. They have been removed; git history still has them.

Both datasets now have a crawler source. 2017 comes from the baotintuc.vn CDN, which is still live. 2016 comes from an aggregator article on dtnt.bacninh.edu.vn, also still online. Both datasets have been crawled successfully, so either can be rebuilt from source — see data-pipeline for the full article URLs.

History

Each year began as a standalone repository (thptqg2016, thptqg2017), merged here with full history. They initially kept separate frontends and separate copies of the same Rust parser, synchronised by hand. That duplication was removed: there is now one frontend, one parser, and one canonical schema, with per-dataset differences confined to one small config file and one registry entry each.

The unification also fixed a latent data-loss bug — neither year's parser config listed the complete set of subjects, so 1,691 candidates were missing their foreign-language score. See data-pipeline.md.

Status

Stable, data frozen. Work is limited to UX polish and keeping the pipeline maintainable.