tiennm99 0eb174721d refactor(parser): reimplement the parser in Go alongside the Rust crate
Adds go-parser/, a Go reimplementation of the xlsxread parser, verified
byte-for-byte against the Rust original before any cutover.

Reader fidelity is exact across all 299 input files: the canonical cell dump
of every sheet matches calamine's, locked in as a test against a committed
hash oracle. Reaching that required replacing extrame/xls, which corrupted
69% of cells and dropped a further 28% on the BIFF corpus, with pbnjay/grate;
correcting excelize's number-format application and trailing-cell trimming;
restoring carriage returns that XML line-ending normalisation strips from
2,233 ten_cum_thi values; and gating numeric re-rendering on cell type so
shared strings that merely look numeric keep their leading zeros.

The differential gate compares both parsers over all four datasets:
3,265,641 rows with identical full-table SHA-256, identical per-column
non-NULL counts, identical schema metadata and identical stdout.

Config moves from TOML to YAML for both parsers, so they keep reading the
same files and the gate stays meaningful. Verified by rebuilding 2016 and
2017-old2 with Rust under the new configs and matching the recorded counts.

build-db.js now refuses to publish a database whose row count does not match
the known figure, closing a path where an under-producing parser could ship a
truncated public dataset with green CI. The deploy workflow gains a
pull_request trigger and guards deploy to main, so branch verification can no
longer publish to production.
2026-08-13 20:27:44 +07:00

thptqg

Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high school graduation exam. Client-side SQL (sql.js) over a SQLite database built from the ministry's raw .xls score files by the Rust xlsxread parser.

Live at tiennm99.github.io/thptqg.

Dataset Exam Candidates Site
2016 2016 877,461 /2016/
2017 2017 861,068 /2017/
2017-old 2017 847,348 /2017-old/
2017-old2 2017 679,764 /2017-old2/

The three 2017 datasets are successive publications of the same exam and they disagree; all three are kept so the differences stay inspectable.

Layout

index.html + src/     the frontend — one app serving all four datasets and the hub
  datasets.js           the four dataset ids and their per-dataset content
  router.js             pathname → dataset
data/<id>/            raw Excel files, one directory per dataset
parser/               the Rust parser
  src/schema.rs         canonical 22-column table: DDL, INSERT, subject regexes
  configs/<id>.yml     per-dataset parse rules only, no SQL
  scripts/              database build, crawler, parity verification
scripts/              site assembly
docs/                 architecture, data pipeline, deployment

The dataset id is one identifier end to end:

data/2017-old/ → parser/configs/2017-old.yml → db/2017-old.db.gz → /thptqg/2017-old/

Build

npm ci
npm run build:go     # compile the parser
npm run build:db       # build + gzip all four databases (add an id for just one)
npm run build:site     # one Vite build, then assemble into _site/
npx serve _site

Pushing to main runs the same steps in .github/workflows/deploy-pages.yml and publishes to GitHub Pages.

Adding a dataset

  1. Put the Excel files in data/<id>/
  2. Add parser/configs/<id>.yml — sheet mode, column indices, validation guards. No SQL; the schema is canonical.
  3. Add an entry to DATASETS in src/datasets.js

Everything else follows: the build script, the site assembly and the router all read that one list, and the UI adapts to whichever columns the dataset fills.

Docs

See docs/ — overview, architecture, data pipeline, deployment.

S
Description
Tra cứu điểm thi THPT Quốc gia 2016–2017 — 1,7 triệu thí sinh · Truy vấn SQL client-side với sql.js
Readme Apache-2.0
165 MiB
0 Stars 1 Watchers 0 Forks
Languages
Go 67.1%
JavaScript 18.9%
Svelte 11.6%
CSS 2.2%
HTML 0.2%