Files
thptqg/2016
tiennm99 cd4b07f9cb refactor(parser): define one canonical student schema for all datasets
The 2016 and 2017 parsers were the same crate with divergent SQL. DDL, INSERT
and the subject regex table lived in each dataset's TOML config, so four copies
had to be kept in step by hand — which is how the two schemas drifted apart.

Move all of it into src/schema.rs as a single 22-column definition: 6 identity
columns (adding ten_cum_thi and gioi_tinh, previously 2016-only) and 16 subject
columns (the union of both exam years). Columns a dataset carries no data for
bind NULL, costing ~1 byte per row.

This collapses writer.rs to one insert path: insert_row_2016, SCORE_FIELDS_2016,
SCORE_FIELDS_2017 and the SCORE_FIELDS alias all go away. Configs shrink to the
parse rules that genuinely vary per dataset — sheet mode, column indices, SBD
validation, header tokens, blank-row stripping — and carry no SQL at all.

Config parsing now denies unknown fields, so a leftover [schema] block fails
loudly instead of looking effective while schema.rs drives the build.

idx_ten_cum_thi is partial, so it holds zero entries on the three 2017 datasets
where the column is always NULL.

Adds db-stats.js and verify-parity.js to prove no data moved: row counts,
per-column non-NULL counts and a deterministic student sample are compared
against databases built from the previous code.

Verified across all four datasets — row counts, every pre-existing column count,
and all sampled students are identical. Database size grows 1.4-2.2%.
2026-08-13 11:16:24 +07:00
..
2026-04-14 19:02:17 +07:00

thptqg2016

Lookup tool for Vietnam's 2016 National High School Graduation Exam (THPT Quốc gia) scores — 877,461 candidates nationwide.

Fully static app running entirely in the browser (SQLite via sql.js). No backend, no query logging.

Features

  • Quick lookup by exam ID (số báo danh) or full name (diacritics-insensitive)
  • Custom read-only SQL (SELECT / PRAGMA / EXPLAIN / WITH) with 7 built-in preset queries
  • Complete data: 12 subjects (Math, Literature, Physics, Chemistry, Biology, History, Geography, English, French, German, Japanese, Chinese), exam cluster, date of birth, gender
  • Safety caps: 100 rows for lookup, 1000 rows for custom SQL
  • Dark mode, Ctrl+Enter shortcut to run queries

Demo

https://tiennm99.github.io/thptqg/2016/

Development

# Build the SQLite database from source Excel files (requires Rust stable)
pnpm run build:db

# Or build in two steps:
pnpm run build:rust          # compile the xlsxread binary
./tools/xlsxread/target/release/xlsxread build \
  --schema tools/xlsxread/configs/thptqg2016-data.toml \
  --input data \
  --output public/thptqg2016.db

pnpm run dev         # Vite dev server
pnpm run build       # Production bundle → dist/
pnpm run lint        # ESLint

The database is built by the xlsxread Rust binary (tools/xlsxread/), which reads the 119 mixed .xls/.xlsx files and auto-detects the column layout per file (separate-scores, mapped, or positional default). No Node.js Excel library is required at build time.

The GitHub Actions workflow (.github/workflows/deploy.yml) compiles xlsxread, builds the DB, gzips it, and deploys to GitHub Pages on every push to main.

Tech stack

React 19 · Vite · sql.js (WASM) · xlsxread (Rust, build-time) · GitHub Pages

Documentation

Source: Originally published at https://dtntbacgiang.edu.vn/tin-tuc/tin-tuc-su-kien/cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html (currently inaccessible). Data is for reference only.