Files
thptqg/docs/project-overview.md
T
tiennm99 219c7c6a69 fix(parser): read the two 2016 layouts that were column-shifted
Four of the 119 files in data/2016 publish one score column per subject
instead of a DIEM_THI sentence, and none of them was being read correctly.

The ĐH Công nghiệp Thực phẩm file puts a three-row ministry title block
above its header, so no header was recognised and the positional fallback
shifted every column by one: the serial number became so_bao_danh, the
exam number became ho_ten, the name became ngay_sinh, and the national ID
became the score cell. All 7,833 rows were unusable. The three ĐH Cần Thơ
files name an SBD column but no DIEM_THI, so they fell to the same
fallback: surname into ngay_sinh, given name into ten_cum_thi, birth date
into the score cell, and 12,152 candidates with no scores at all.

Both are now read by FormatSubjectColumns, which resolves identity and one
column per subject from the header. The header is searched for in the
first five rows, so a title block no longer hides it.

The Cần Thơ score columns are numbered rather than named. They follow the
order the exam was sat — each morning an essay paper, each afternoon a
multiple-choice one — which is what identifies them: columns 1/3/5/7
quantise to 0.25 and 2/4/6/8 do not, and each column's mean lands within
0.5 of the same subject's mean across the rest of the dataset. The
foreign language is filed under the subject its N1..N6 code names.

Gender now accepts the 0/1 encoding those files use: of the rows marked
1, 53% carry "Thị" in the name against 1% of those marked 0. Birth dates
in the compact ddmmyy form are expanded so the column holds one format.

A score of 0 is stored rather than dropped, recovering 302 real scores
that a JavaScript falsy check had been turning into NULL.

Row count falls by one, to 877,460: the removed row is the title line
"ĐƠN VỊ: / TRƯỜNG ĐẠI HỌC CÔNG NGHIỆP THỰC PHẨM TP. HỒ CHÍ MINH", which
had been stored as a student. The dataset has no duplicate exam numbers;
the three rows previously described as collapsing were that same file's
title and header lines being counted and then rejected.

Also drops behaviour that existed only to match the parser this one
replaced: the inert "SINH " header token, the untrimmed diem_thi cell, an
unreachable blank-row branch, and a cross-check test against a database
that can no longer exist. None of them changes output.

Verified by rebuilding both datasets: 877,460 and 861,068 rows, both
artifacts through the assembler's row and size guards, and the reader
fidelity suite unchanged across all 182 files.
2026-08-14 10:38:59 +07:00

2.6 KiB
Raw Blame History

Project Overview

Goal

A public lookup tool for Vietnam's National High School Graduation Exam scores, running entirely in the browser and hosted for free on GitHub Pages. Covers the 2016 and 2017 exams — 1.7 million candidates across two datasets.

Scope

  • Lookup by exam ID or full name, with Vietnamese diacritics handled
  • Read-only SQL queries against a single student table
  • Admission-block (khối thi) totals computed per candidate
  • Static datasets — both exams are long over and the data is frozen

Target users

  • Former candidates checking their scores
  • Education researchers and data journalists running aggregate statistics
  • Developers exploring SQL against a real-world dataset

Constraints

  • Zero backend. The full database (44–48 MB gzipped per dataset) is downloaded to the browser and queried in-process.
  • Read-only. INSERT/UPDATE/DELETE are rejected, so nobody is misled into thinking edits persist. sql.js is in-memory anyway.
  • Row caps. 100 rows for lookups, 1000 for custom SQL, to prevent browser hangs.
  • Vietnamese-first UI. App labels and data are Vietnamese; documentation is English.

Datasets

id Exam Candidates Notes
2016 2016 877,460 119 files, four column layouts
2017 2017 861,068 current generation of three publications

Two further 2017 datasets (2017-old, 2017-old2) were kept alongside these because the three publications disagreed. They have been removed; git history still has them.

Both datasets now have a crawler source. 2017 comes from the baotintuc.vn CDN, which is still live. 2016 comes from an aggregator article on dtnt.bacninh.edu.vn, also still online. Both datasets have been crawled successfully, so either can be rebuilt from source — see data-pipeline for the full article URLs.

History

Each year began as a standalone repository (thptqg2016, thptqg2017), merged here with full history. They initially kept separate frontends and separate copies of the same Rust parser, synchronised by hand. That duplication was removed: there is now one frontend, one parser, and one canonical schema, with per-dataset differences confined to one small config file and one registry entry each.

The unification also fixed a latent data-loss bug — neither year's parser config listed the complete set of subjects, so 1,691 candidates were missing their foreign-language score. See data-pipeline.md.

Status

Stable, data frozen. Work is limited to UX polish and keeping the pipeline maintainable.