docs: rewrite for the unified repo and record the parity gate

The docs still described two standalone projects with separate frontends,
separate parsers and three Vite variants. Rewritten around what the repo now
is, merging both projects' copies rather than keeping one and discarding the
other — deployment-guide.md and system-architecture.md existed in both and
documented different pipelines.

  project-overview.md    goal, scope, constraints, the four datasets, history
  system-architecture.md data flow, canonical schema, routing, how one
                         frontend serves both exam years without branching
  data-pipeline.md       per-dataset Excel formats, the three 2016 layouts,
                         overflow-sheet gotcha, expected row counts
  deployment-guide.md    the single-build workflow, adding a dataset,
                         why no uncompressed database can ship

Records the release gate in plans/reports/parser-parity-result.md: row counts,
all 18 pre-existing per-column non-NULL counts and the deterministic student
samples are identical across all four datasets, against databases decompressed
from the exact bytes the pipeline publishes. The 1,691 recovered
foreign-language scores are documented with the evidence they are real.

Also drops a machine-specific absolute path from a comment in
format_detect_2016.rs. The build-database.js citations there are kept: that
file no longer exists in this repo, but the references explain why several
parsing rules look arbitrary.
This commit is contained in:
tiennm99 committed 2026-08-13 13:25:46 +07:00
1 parent a5904a5d78
commit 3fcfe4862d
18 files changed
+591 -448

No files matched your search

+59 -10
View File
@@ -2,18 +2,67 @@
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
school graduation exam. Client-side SQL (sql.js) over a SQLite database built
from the ministry's raw `.xls` score files by the Rust `xlsxread` CLI.
from the ministry's raw `.xls` score files by the Rust `xlsxread` parser.
Served as one site at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
Live at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
| Directory | Year | Size | Site |
| Dataset | Exam | Candidates | Site |
| --- | --- | --- | --- |
| [`2016/`](./2016/) | THPT QG 2016 — 877,461 candidates | ~50 MB raw data | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
| [`2017/`](./2017/) | THPT QG 2017 — ~861,000 candidates | ~117 MB raw data | [/2017/](https://tiennm99.github.io/thptqg/2017/) (+ [old](https://tiennm99.github.io/thptqg/2017/old/), [old2](https://tiennm99.github.io/thptqg/2017/old2/) generations) |
| `2016` | 2016 | 877,461 | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
| `2017` | 2017 | 861,068 | [/2017/](https://tiennm99.github.io/thptqg/2017/) |
| `2017-old` | 2017 | 847,348 | [/2017-old/](https://tiennm99.github.io/thptqg/2017-old/) |
| `2017-old2` | 2017 | 679,764 | [/2017-old2/](https://tiennm99.github.io/thptqg/2017-old2/) |
Each year was a standalone repository (`thptqg2016`, `thptqg2017`), merged here
with full history. Each keeps its own `tools/xlsxread` copy because the raw
data formats differ per year.
The three 2017 datasets are successive publications of the same exam and they
disagree; all three are kept so the differences stay inspectable.
The `Deploy to GitHub Pages` workflow builds both years' databases and site
bundles and publishes them behind a root index.
## Layout
```
index.html + src/ the frontend — one app serving all four datasets and the hub
datasets.js the four dataset ids and their per-dataset content
router.js pathname → dataset
data/<id>/ raw Excel files, one directory per dataset
parser/ the Rust parser
src/schema.rs canonical 22-column table: DDL, INSERT, subject regexes
configs/<id>.toml per-dataset parse rules only, no SQL
scripts/ database build, crawler, parity verification
scripts/ site assembly
docs/ architecture, data pipeline, deployment
```
The dataset id is one identifier end to end:
```
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
```
## Build
```bash
npm ci
npm run build:rust # compile the parser
npm run build:db # build + gzip all four databases (add an id for just one)
npm run build:site # one Vite build, then assemble into _site/
npx serve _site
```
Pushing to `main` runs the same steps in
`.github/workflows/deploy-pages.yml` and publishes to GitHub Pages.
## Adding a dataset
1. Put the Excel files in `data/<id>/`
2. Add `parser/configs/<id>.toml` — sheet mode, column indices, validation
guards. No SQL; the schema is canonical.
3. Add an entry to `DATASETS` in `src/datasets.js`
Everything else follows: the build script, the site assembly and the router all
read that one list, and the UI adapts to whichever columns the dataset fills.
## Docs
See [`docs/`](./docs/) — [overview](./docs/project-overview.md),
[architecture](./docs/system-architecture.md),
[data pipeline](./docs/data-pipeline.md),
[deployment](./docs/deployment-guide.md).
+4 -3
View File
@@ -1,5 +1,6 @@
# Docs
- [`system-architecture.md`](./system-architecture.md) — data flow, schema, 3-variant deploy, score-tier model, admission-block catalog
- [`data-pipeline.md`](./data-pipeline.md) — Excel parse quirks, overflow-sheet gotcha, audit scripts, refresh flow
- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, adding a new variant, rollback
- [`project-overview.md`](./project-overview.md) — goal, scope, constraints, the four datasets, history
- [`system-architecture.md`](./system-architecture.md) — data flow, canonical schema, routing, how one frontend serves both exam years
- [`data-pipeline.md`](./data-pipeline.md) — Excel parse quirks, per-dataset formats, overflow-sheet gotcha, expected row counts, verifying a rebuild
- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, adding a dataset, rollback, troubleshooting
-63
View File
@@ -1,63 +0,0 @@
# Codebase Summary
## Directory layout
```
thptqg2016/
├── data/ # Source Excel files (~100, mixed formats)
├── scripts/
│ └── build-database.js # Parse Excel → SQLite (build-time, Node + better-sqlite3)
├── public/
│ └── thptqg2016.db # Generated DB, gzipped during CI
├── src/
│ ├── main.jsx # React entry
│ ├── App.jsx # Root: tabs, lookup logic, useSqlite wiring
│ ├── App.css / index.css # Design tokens, dark mode, a11y styles
│ ├── hooks/
│ │ └── use-sqlite.js # Fetch .db.gz + decompress + init sql.js
│ └── components/
│ ├── search-form.jsx # Input for exam ID / full name
│ ├── score-table.jsx # Result table for lookups
│ └── custom-query.jsx # SQL editor + presets + result grid
├── .github/workflows/deploy.yml # CI: build db → gzip → vite build → Pages
├── vite.config.js # base: "/thptqg2016/"
└── eslint.config.js
```
## Key modules
### `scripts/build-database.js`
Build-time only. Reads every `.xlsx/.xls` in `data/`, detects the header format (three variants), parses the `DIEM_THI` string via regex or separate score columns, normalizes gender, derives a diacritics-stripped `ho_ten_ascii` column for accent-insensitive search, and inserts into SQLite with three indexes (`ho_ten`, `ho_ten_ascii`, `ten_cum_thi`).
### `src/hooks/use-sqlite.js`
Streams `.db.gz` with download progress, decompresses via `DecompressionStream("gzip")`, loads `sql.js` (WASM served from the `sql.js.org` CDN), and returns `{ db, loading, error, progress }`.
### `src/App.jsx`
Two tabs: **Lookup** and **Custom SQL**. Lookup auto-detects exam IDs (regex `^[A-Z]{2,4}\d+$`) vs names and picks one of three query paths: exact exam ID / ASCII LIKE / original + ASCII LIKE. Capped at 100 rows.
### `src/components/custom-query.jsx`
Whitelists leading keywords (`SELECT`, `PRAGMA`, `EXPLAIN`, `WITH`), auto-appends `LIMIT 1000` when missing, measures `performance.now()` execution time, and ships 7 preset analytics queries.
## `student` table schema
```sql
so_bao_danh TEXT PRIMARY KEY -- exam ID
ho_ten TEXT NOT NULL -- full name
ho_ten_ascii TEXT NOT NULL -- diacritics stripped, lowercased
ngay_sinh TEXT -- date of birth
ten_cum_thi TEXT -- exam cluster name
gioi_tinh TEXT -- "Nam" | "Nữ" | NULL
toan, ngu_van, vat_ly, hoa_hoc, -- REAL (nullable) subject scores
sinh_hoc, lich_su, dia_ly,
tieng_anh, tieng_phap, tieng_duc,
tieng_nhat, tieng_trung
```
Indexes: `idx_ho_ten`, `idx_ho_ten_ascii`, `idx_ten_cum_thi`.
## Conventions
- JS/JSX filenames: **kebab-case** (e.g., `search-form.jsx`, `use-sqlite.js`)
- React components: `PascalCase` named exports
- UI strings: Vietnamese (target audience)
- Code comments: English; explain *why*, not *what*
+109 -70
View File
@@ -2,109 +2,148 @@
From raw Excel files to a compressed SQLite file the browser can load.
## Canonical source (data/)
One Rust binary (`parser/`) builds every dataset. What differs per dataset is
parse rules only — sheet strategy, column layout, validation guards — declared
in `parser/configs/<id>.toml`. The table shape, the INSERT and the subject
regexes are canonical and live in `parser/src/schema.rs`.
63 Excel files from baotintuc.vn CDN. Source article:
`https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm`
## Sources
Re-download anytime:
| id | Files | Reproducible | Origin |
| --- | --- | --- | --- |
| `2016` | 4 `.xls` + 115 `.xlsx` | no | Bộ GD&ĐT, collected 2016 |
| `2017` | 63 `.xls` | **yes** | baotintuc.vn CDN |
| `2017-old` | 63 `.xlsx` | no | pre-refresh archive |
| `2017-old2` | 54 `.xlsx` | no | corrected re-export |
Only `2017` can be re-fetched:
```bash
node scripts/crawl-baotintuc.js
node parser/scripts/crawl-baotintuc.js
```
Idempotent — skips files already present. Saves to `data/<ascii-kebab-province>.xls`.
Idempotent — skips files already present, saves to `data/2017/`. Source article:
`https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm`
## Reference sources (data-old/, data-old2/)
The other three were collected from Vietnamese news sites at the time and the
publisher URLs were not recorded. **The files committed in git are the only
copy.**
Both were collected from Vietnamese news sites around the time of the 2017 exam.
Specific publisher URLs were not recorded and cannot be recovered, so these
datasets are **not reproducible** — the files committed in git are the only
copy. They are kept for historical comparison and to expose the deploy matrix;
the main deployment always uses the reproducible `data/` source.
## Source Excel shapes
## Source Excel shape
Every file is a single sheet (for `data-old/`) or up to 3 sheets (for `data/` .xls) with columns:
The 2017 datasets share one layout:
| Col | Name | Content |
|---|---|---|
| 0 | HO_TEN | full name in Vietnamese |
| 1 | NGAY_SINH | `dd/mm/yyyy` |
| 2 | SOBAODANH | 8-digit string, first 2 digits = province |
| 3 | DIEM_THI | concatenated per-subject score string in the source language, e.g. `"Toan: 6.80 Van: 5.25 ... English: 5.80"` |
| --- | --- | --- |
| 0 | HO_TEN | full name in Vietnamese |
| 1 | NGAY_SINH | `dd/mm/yyyy` |
| 2 | SOBAODANH | 8-digit string, first 2 digits = province |
| 3 | DIEM_THI | concatenated per-subject scores, e.g. `"Toán: 6.80 Ngữ văn: 5.25 …"` |
Some files lack a header row. `isHeaderRow()` in `build-lib.js` detects and skips.
2016 has **three** layouts across its 119 files, chosen per file at runtime via
`format_detection = "thptqg2016"` in its config:
| Format | Detected by | Notes |
| --- | --- | --- |
| `separate-scores` | `row[0] == "SBD"` and `row[2] == "TOAN"` | one column per subject; scores read directly, no regex |
| `mapped` | header contains `SOBAODANH`/`SBD` plus `DIEM_THI` | column indices derived from the header |
| `default` | no recognised header | positional 6-column layout |
The `mapped` and `default` layouts also carry `TEN_CUMTHI` and `GIOI_TINH`,
which is why only 2016 populates those columns.
## Score text parsing
`SCORE_PATTERNS` in `build-lib.js` defines one regex per subject. The patterns must match the exact subject labels the source files use (Vietnamese).
`SCORE_PATTERNS` in `parser/src/schema.rs` defines one regex per subject, and
**all 16 run against every dataset**. A subject a given exam year did not offer
simply never matches and stays NULL.
Subjects supported (DB column / label in source text):
| Column | Source label | English | 2016 | 2017 |
| --- | --- | --- | --- | --- |
| toan | Toán | Math | ✓ | ✓ |
| ngu_van | Ngữ văn | Literature | ✓ | ✓ |
| vat_ly | Vật lí | Physics | ✓ | ✓ |
| hoa_hoc | Hóa học | Chemistry | ✓ | ✓ |
| sinh_hoc | Sinh học | Biology | ✓ | ✓ |
| khtn | KHTN | Natural Sciences combined | — | ✓ |
| lich_su | Lịch sử | History | ✓ | ✓ |
| dia_ly | Địa lí | Geography | ✓ | ✓ |
| gdcd | GDCD | Civic Education | — | ✓ |
| khxh | KHXH | Social Sciences combined | — | ✓ |
| tieng_anh | Tiếng Anh | English | ✓ | ✓ |
| tieng_phap | Tiếng Pháp | French | ✓ | ✓ |
| tieng_nga | Tiếng Nga | Russian | ✓ | ✓ |
| tieng_duc | Tiếng Đức | German | ✓ | ✓ |
| tieng_nhat | Tiếng Nhật | Japanese | ✓ | ✓ |
| tieng_trung | Tiếng Trung | Chinese | ✓ | ✓ |
| Column | Source label | English |
|---|---|---|
| toan | Toán | Math |
| ngu_van | Ngữ văn | Literature |
| vat_ly | Vật lí | Physics |
| hoa_hoc | Hóa học | Chemistry |
| sinh_hoc | Sinh học | Biology |
| khtn | KHTN | Natural Sciences combined |
| lich_su | Lịch sử | History |
| dia_ly | Địa lí | Geography |
| gdcd | GDCD | Civic Education |
| khxh | KHXH | Social Sciences combined |
| tieng_anh | Tiếng Anh | English |
| tieng_phap | Tiếng Pháp | French |
| tieng_nga | Tiếng Nga | Russian |
| tieng_trung | Tiếng Trung | Chinese |
### The foreign-language recovery
Missing-in-string → `NULL` in DB (student didn't take that subject).
Before the parser was unified, the 2016 config listed 12 subject patterns and
the 2017 configs listed 14. Neither list was complete: candidates could sit
German, Japanese and Russian in both years, so 1,691 students ended up with **no
foreign-language score at all**.
## Why three parsers
Running all 16 patterns everywhere recovered them — 182 Russian in 2016, and
German and Japanese across all three 2017 generations. See
`plans/reports/parser-parity-result.md` for the evidence that these are real
scores rather than false matches.
One shared lib + three thin parsers, not one parser with flags. Each dataset's quirks stay visible in its own file:
## Per-dataset quirks
| File | Source dir | Sheet strategy | Extra guard |
|---|---|---|---|
| `build-database.js` | `data/` | iterate ALL sheets (.xls row overflow) | — |
| `build-database-old.js` | `data-old/` | sheet 0 only | reject rows where `so_bao_danh` isn't digits-only (rejects an `HO_TEN` → `SOBAODANH` header leak present in this export) |
| `build-database-old2.js` | `data-old2/` | iterate ALL sheets (HCM overflow) | skip fully blank rows before counting |
| id | Sheet strategy | Guards |
| --- | --- | --- |
| `2016` | all sheets | per-file format detection; header-token rows rejected |
| `2017` | all sheets — Hà Nội and HCM overflow | none |
| `2017-old` | first sheet only | reject non-numeric SBD (rejects a header leak in this export) |
| `2017-old2` | all sheets — HCM overflows | reject non-numeric SBD; skip blank rows before counting |
## Overflow-sheet gotcha
### Overflow-sheet gotcha
The old `.xls` binary format caps at 65,536 rows per sheet. Hanoi and Ho Chi Minh City exceed that in the baotintuc export, so Excel auto-splits the overflow into `Sheet2`. **Only reading `Sheet1` silently drops 13,720 students** (Hanoi +7,275, HCM +6,445).
The old `.xls` format caps at 65,536 rows per sheet. Hanoi and Ho Chi Minh City
exceed that, so Excel splits the overflow into `Sheet2`. **Reading only `Sheet1`
silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what
`sheet_mode = "all"` exists for.
The audit `node scripts/audit-row-counts.js` catches this by comparing source-row-count (across all sheets) to DB-row-count.
## Expected row counts
## Audits
| id | Source rows | Skipped | DB rows |
| --- | --- | --- | --- |
| `2016` | 877,464 | 3 duplicate SBDs collapsed | **877,461** |
| `2017` | 861,131 | 63 empty | **861,068** |
| `2017-old` | 847,349 | 1 header leak | **847,348** |
| `2017-old2` | 679,764 | 0 | **679,764** |
Run any time after a rebuild:
## Verifying a rebuild
`parser/scripts/db-stats.js` dumps row counts, per-column non-NULL counts, file
size and a deterministic student sample as JSON.
`parser/scripts/verify-parity.js` diffs two such files and exits non-zero on any
mismatch.
```bash
node scripts/audit-row-counts.js # source rows vs DB rows (zero-loss check)
node scripts/check-duplicates.js # md5 + row-content duplicates in data/
node scripts/diff-datasets.js # compare two DBs (needs backup/ dir)
node parser/scripts/db-stats.js 2016=<path>.db … > current.json
node parser/scripts/verify-parity.js plans/reports/parser-parity-baseline.json current.json
```
Expected baseline:
The committed baseline was built with the pre-refactor code and cannot be
regenerated — the two old crates no longer exist. Both scripts use the built-in
`node:sqlite`, so they need no dependencies.
| DB | Source rows | Skipped | DB rows |
|---|---|---|---|
| main | 861,131 | 63 empty rows | **861,068** |
| old | 847,349 | 1 header-leak | **847,348** |
| old2 | 679,764 | 0 | **679,764** |
## Refreshing the data
## Refreshing the 2017 data
```bash
rm data/*.xls # clear current snapshot
node scripts/crawl-baotintuc.js # re-download from CDN
npm run build:db # rebuild
node scripts/audit-row-counts.js # verify
gzip -kf -9 public/thptqg2017.db # compress for shipping
rm data/2017/*.xls
node parser/scripts/crawl-baotintuc.js
node parser/scripts/build-db.js 2017
```
## Adding a new source dataset
Then re-run the parity check above and confirm the row count still matches.
See `docs/deployment-guide.md` → "Adding a new variant".
## Legacy scripts
`parser/scripts/check-duplicates.js` and `diff-datasets.js` are one-off audits
that were already broken before the repo was unified — a hardcoded Windows path
in one, an undeclared `better-sqlite3` dependency and stale paths in the other.
Each carries a comment saying so. For comparing two builds, use `db-stats.js`
plus `verify-parity.js` instead.
-50
View File
@@ -1,50 +0,0 @@
# Deployment Guide
## Automatic (recommended)
Push to `main` → GitHub Actions builds and deploys to GitHub Pages automatically.
Workflow: `.github/workflows/deploy.yml`
CI steps:
1. `npm ci`
2. `npm run build:db` — generate `public/thptqg2016.db` from `data/*.xlsx`
3. `gzip -k -9 public/thptqg2016.db` — max compression, keep original
4. `npm run build` — Vite bundles `dist/`
5. `rm -f dist/thptqg2016.db` — ship only the gzipped copy
6. `actions/upload-pages-artifact@v3` + `actions/deploy-pages@v4`
One-time setup: **Settings → Pages → Source: GitHub Actions**.
## Manual (local verification)
```bash
npm install
npm run build:db
gzip -k -9 public/thptqg2016.db # Linux/macOS; on Windows use 7zip or WSL
npm run build
rm dist/thptqg2016.db # optional, shrinks artifact
npm run preview # serve dist/ locally
```
Open <http://localhost:4173/thptqg2016/>.
## Base path
`vite.config.js` sets `base: "/thptqg2016/"`. If you fork under a different repo name, update this to match `<repo-name>` so assets resolve correctly on GitHub Pages.
## Updating data
1. Add the new Excel file to `data/`.
2. If its header is unfamiliar, open `scripts/build-database.js` and extend `detectFormat()` or the `processSeparateScoresRow` / `processMappedRow` helpers.
3. Run `npm run build:db` locally to check row counts and error skips.
4. Commit + push → CI redeploys.
## Troubleshooting
| Symptom | Typical cause |
|---------|---------------|
| Blank page, 404 on assets | `base` in `vite.config.js` doesn't match the repo name |
| `Failed to fetch database: 404` | gzip step skipped, or `.db.gz` removed from `dist/` |
| WASM fails to load | `sql.js.org` blocked / offline — self-host `sql-wasm.wasm` in `public/` and update `SQL_WASM_URL` in `use-sqlite.js` |
| Missing rows after build | Excel file has an unknown header — check console for `Failed to read` or `errorCount` |
+82 -29
View File
@@ -1,52 +1,105 @@
# Deployment Guide
Deploys to GitHub Pages via `.github/workflows/deploy.yml`. Every push to `main` rebuilds and redeploys all 3 variants.
Deploys to GitHub Pages via `.github/workflows/deploy-pages.yml`. Every push to
`main` rebuilds all four datasets and redeploys the whole site.
One-time setup: **Settings → Pages → Source: GitHub Actions**.
## What the workflow does
1. Checkout, setup Node 20, `npm ci`
2. `npm run build:db:all` — builds 3 SQLite DBs (main, old, old2)
3. Gzip each DB at level 9
4. `npm run build:all` — builds 3 Vite bundles into `dist/`, `dist/old/`, `dist/old2/`
5. Remove uncompressed `.db` files from `dist/` (only `.gz` ships)
6. `actions/upload-pages-artifact` + `actions/deploy-pages`
1. Checkout, Rust toolchain, Node 24, `npm ci`
2. `npm run build:rust` — one parser binary
3. `npm run build:db` — builds and gzips all four databases into
`.build/public/db/`
4. `npm run build:site` — one Vite build, then `scripts/assemble-site.js`
5. `actions/upload-pages-artifact` + `actions/deploy-pages`
Total CI time ≈ 4–6 min (DB build dominates).
The database build dominates the runtime: roughly 419 MB of Excel is parsed on
every deploy.
## Resulting URLs
- `https://<user>.github.io/thptqg2017/`
- `https://<user>.github.io/thptqg2017/old/`
- `https://<user>.github.io/thptqg2017/old2/`
```
https://<user>.github.io/thptqg/
https://<user>.github.io/thptqg/2016/
https://<user>.github.io/thptqg/2017/
https://<user>.github.io/thptqg/2017-old/
https://<user>.github.io/thptqg/2017-old2/
```
`/thptqg/2017/old/` and `/thptqg/2017/old2/` were the pre-flattening URLs. They
are still served, and the router rewrites them to the flat form with the query
string intact.
## Local reproduction
```bash
npm ci
npm run build:db:all
gzip -kf -9 public/thptqg2017.db public-old/thptqg2017.db public-old2/thptqg2017.db
npm run build:all
npx serve dist # or any static server
npm run build:rust
npm run build:db # all four; pass an id to build just one
npm run build:site # vite build + assemble into _site/
npx serve _site
```
## Adding a new variant
To rebuild a single dataset:
1. Drop source Excel files into a new `data-vN/` folder
2. Add `public-vN/` to `.gitignore` patterns for the `.db` + `.db.gz`
3. Copy `scripts/build-database-old.js` → `scripts/build-database-vN.js`; update `SRC_DIR` / `DB_PATH`; adjust parse logic for the new quirks (multi-sheet? blank rows? numeric guard?)
4. Add variant to `VARIANT_CONFIG` in `vite.config.js`:
```js
vN: { base: "/thptqg2017/vN/", publicDir: "public-vN", outDir: "dist/vN" }
```
5. Add `build:db:vN` and `build:vN` npm scripts; wire them into `build:db:all` and `build:all`
6. Update `.github/workflows/deploy.yml` — add gzip step + rm step for the new DB
```bash
node parser/scripts/build-db.js 2017-old
```
## Base path
`vite.config.js` sets `base: "/thptqg/"`. If you fork under a different repo
name, update it to match — assets are referenced absolutely, so a mismatch shows
up as a blank page with 404s on `/assets/...`.
## Adding a dataset
1. Put the Excel files in `data/<id>/`
2. Add `parser/configs/<id>.toml` with the parse rules — sheet mode, column
indices, SBD validation, header tokens, blank-row stripping. No SQL: the
schema is canonical and lives in `parser/src/schema.rs`
3. Add an entry to `DATASETS` in `src/datasets.js`
Nothing else. The build script, the site assembly and the router all read that
one list, and the frontend adapts to whichever columns the dataset populates.
## Why no uncompressed database can ship
`build-db.js` runs `gzip -9` **without** `-k`, so the raw file does not survive
the build. `assemble-site.js` then fails the job if any `.db`, `.db-journal`,
`.db-wal` or `.db-shm` reached the output.
Both guards exist because the previous pipeline wrote a 100+ MB uncompressed
database into the source tree and relied on an `rm` step to keep it out of the
artifact — one missing line away from publishing it.
## Notes
- **DB gzip is non-cacheable across deploys** — every rebuild produces a new `.db.gz` (not byte-identical due to SQLite page shuffling). First-visit users pay the 47 MB download; subsequent visits hit browser cache until the next deploy.
- **GitHub Pages file size limit**: individual files ≤ 100 MB. Main gzipped DB is ~47 MB — safe. Uncompressed 159 MB DB would not fit; that's why we ship `.gz` and the frontend decompresses via `DecompressionStream`.
- **No server-side compression assumption.** The app reads the `.gz` file directly (not `Content-Encoding: gzip`). GH Pages does not reliably gzip on-the-fly for arbitrary paths; shipping pre-gzipped bytes is deterministic.
- **The gzipped database is not cacheable across deploys.** Every rebuild
produces a different `.db.gz`, because SQLite does not lay pages out
deterministically. First-visit users pay the full download; later visits hit
browser cache until the next deploy.
- **GitHub Pages caps individual files at 100 MB.** The largest gzipped database
is about 48 MB. Uncompressed they run 135–234 MB and would not fit — which is
why the browser decompresses via `DecompressionStream`.
- **No server-side compression is assumed.** The app fetches the `.gz` bytes
directly rather than relying on `Content-Encoding: gzip`; Pages does not
reliably compress arbitrary paths on the fly.
- **Total artifact is about 177 MB**, well inside the 1 GB site limit.
## Rollback
Deploys are stateless snapshots. To rollback, revert the commit on `main` and push — the next workflow rebuilds the older state. There's no state to migrate.
Deploys are stateless snapshots. Revert the commit on `main` and push; the next
run rebuilds the older state. There is no data to migrate.
## Troubleshooting
| Symptom | Typical cause |
| --- | --- |
| Blank page, 404 on assets | `base` in `vite.config.js` does not match the repo name |
| `Failed to fetch database: 404` | Dataset id in `src/datasets.js` does not match the file in `db/` |
| A route 404s | `assemble-site.js` did not run, or the id is missing from `DATASETS` |
| WASM fails to load | `sql.js.org` unreachable — self-host `sql-wasm.wasm` and update `SQL_WASM_URL` in `use-sqlite.js` |
| Deploy fails on assembly | An uncompressed database artefact reached the output; the error names the files |
| Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints |
-34
View File
@@ -1,34 +0,0 @@
# Project Overview — thptqg2016
## Goal
Provide a public lookup tool for Vietnam's 2016 National High School Graduation Exam scores (877,461 candidates), running entirely on the client, hosted for free on GitHub Pages.
## Scope
- Lookup by exam ID or full name (with Vietnamese diacritics handling)
- Read-only SQL queries against a single `student` table
- Static dataset — no updates (the 2016 exam is long over)
## Target users
- Former 2016 candidates checking their scores
- Education researchers / data journalists running aggregate stats
- Developers exploring SQL on a real-world dataset
## Constraints
- **Zero backend**: the full DB (tens of MB gzipped) is downloaded to the browser
- **Read-only**: INSERT/UPDATE/DELETE rejected to avoid the illusion that user edits persist
- **Row caps**: 100 rows (lookup), 1000 rows (SQL) to prevent browser hangs
- **Vietnamese-first UI**: app labels and data are Vietnamese; documentation is English
## Data sources
Excel files (`.xlsx`/`.xls`) collected from newspapers and exam clusters in 2016, stored in `data/`. One file per cluster, with several different column layouts (see `scripts/build-database.js`).
Original aggregator link: <https://dtntbacgiang.edu.vn/tin-tuc/tin-tuc-su-kien/cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html> — no longer accessible, so the raw Excel files are mirrored in this repo under `data/`.
## Status
Stable. Data is frozen. Recent work focuses on UX polish (dark mode, accessibility, diacritics-insensitive search).
+66
View File
@@ -0,0 +1,66 @@
# Project Overview
## Goal
A public lookup tool for Vietnam's National High School Graduation Exam scores,
running entirely in the browser and hosted for free on GitHub Pages. Covers the
2016 and 2017 exams — 1.7 million candidates across four datasets.
## Scope
- Lookup by exam ID or full name, with Vietnamese diacritics handled
- Read-only SQL queries against a single `student` table
- Admission-block (khối thi) totals computed per candidate
- Static datasets — both exams are long over and the data is frozen
## Target users
- Former candidates checking their scores
- Education researchers and data journalists running aggregate statistics
- Developers exploring SQL against a real-world dataset
## Constraints
- **Zero backend.** The full database (38–48 MB gzipped per dataset) is
downloaded to the browser and queried in-process.
- **Read-only.** `INSERT`/`UPDATE`/`DELETE` are rejected, so nobody is misled
into thinking edits persist. `sql.js` is in-memory anyway.
- **Row caps.** 100 rows for lookups, 1000 for custom SQL, to prevent browser
hangs.
- **Vietnamese-first UI.** App labels and data are Vietnamese; documentation is
English.
## Datasets
| id | Exam | Candidates | Notes |
| --- | --- | --- | --- |
| `2016` | 2016 | 877,461 | 119 files, three column layouts |
| `2017` | 2017 | 861,068 | current generation, reproducible from source |
| `2017-old` | 2017 | 847,348 | pre-refresh archive |
| `2017-old2` | 2017 | 679,764 | corrected re-export |
The three 2017 datasets are kept side by side because they disagree, and the
disagreement is itself informative. Only `2017` is re-fetchable; the rest exist
solely as the copies committed here.
Original 2016 aggregator link
(<https://dtntbacgiang.edu.vn/tin-tuc/tin-tuc-su-kien/cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html>)
is no longer accessible, which is why the raw files are mirrored in `data/2016/`.
## History
Each year began as a standalone repository (`thptqg2016`, `thptqg2017`), merged
here with full history. They initially kept separate frontends and separate
copies of the same Rust parser, synchronised by hand. That duplication was
removed: there is now one frontend, one parser, and one canonical schema, with
per-dataset differences confined to four small config files and one registry
entry each.
The unification also fixed a latent data-loss bug — neither year's parser
config listed the complete set of subjects, so 1,691 candidates were missing
their foreign-language score. See `data-pipeline.md`.
## Status
Stable, data frozen. Work is limited to UX polish and keeping the pipeline
maintainable.
-83
View File
@@ -1,83 +0,0 @@
# System Architecture
## Overview
A **static serverless** design: the entire dataset is packaged into a single SQLite file, gzip-compressed, and served as a static asset via GitHub Pages. The browser downloads it, decompresses it, and queries it in-process using `sql.js` (SQLite compiled to WebAssembly).
```
┌─────────────────┐ build ┌──────────────────────┐
│ data/*.xlsx │ ───────────▶ │ scripts/ │
│ (mixed formats)│ │ build-database.js │
└─────────────────┘ │ (Node + xlsx + │
│ better-sqlite3) │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ public/thptqg2016.db │
└──────────┬───────────┘
│ gzip -9 (CI)
▼
┌──────────────────────┐
│ dist/thptqg2016.db.gz│
│ dist/assets/* │ ◀── Vite build
└──────────┬───────────┘
│ upload-pages-artifact
▼
GitHub Pages CDN
│
▼
┌──────────────────────────────────┐
│ Browser │
│ ┌────────────────────────────┐ │
│ │ useSqlite hook │ │
│ │ fetch(.db.gz) + stream │ │
│ │ DecompressionStream gzip │ │
│ │ sql.js WASM (from CDN) │ │
│ └─────────────┬──────────────┘ │
│ ▼ │
│ ┌────────────────────────────┐ │
│ │ React UI │ │
│ │ - SearchForm / ScoreTable │ │
│ │ - CustomQuery (SQL editor)│ │
│ └────────────────────────────┘ │
└──────────────────────────────────┘
```
## Build-time data flow
1. Developer drops Excel files into `data/`.
2. `npm run build:db` reads every file; for each one:
- Detects the header row against a `KNOWN_HEADERS` set.
- Picks a format: `separate-scores` (one column per subject) vs. `mapped` (single `DIEM_THI` string).
- Parses each row into a canonical 18-column record.
- `INSERT OR REPLACE` into SQLite (primary key = `so_bao_danh` handles duplicates).
3. `VACUUM` shrinks the file.
4. CI runs `gzip -k -9` → `.db.gz`.
## Runtime flow
1. Page loads → React mounts → `useSqlite("thptqg2016.db.gz")`.
2. Streaming fetch with a progress bar (driven by `Content-Length`).
3. `DecompressionStream("gzip")` decompresses on the fly.
4. `sql.js` loads its WASM from `https://sql.js.org/dist/sql-wasm.wasm`.
5. `new SQL.Database(Uint8Array)` — DB is now in RAM.
6. Each search / query → `db.prepare()` + `stmt.step()` loop → render.
## Design decisions
| Concern | Choice | Rationale |
|---------|--------|-----------|
| Storage | Static SQLite file | No backend needed; dataset is frozen |
| Compression | gzip in CI, `DecompressionStream` in browser | Native browser API; no extra library |
| WASM hosting | `sql.js.org` CDN | Smaller self-hosted artifact |
| Diacritics search | Pre-computed `ho_ten_ascii` column | `LOWER(REPLACE(...))` at query time defeats the index |
| SQL safety | Leading-keyword allowlist | `sql.js` is in-memory so writes don't persist, but the allowlist prevents user confusion |
| Row caps | 100 (lookup), 1000 (SQL) | Keep DOM render sizes reasonable |
## Risks and limitations
- **DB size**: tens of MB gzipped — slow links have a visible wait; mitigated by the progress bar.
- **Browser memory**: the full DB lives in RAM; older mobile devices may OOM.
- **Dependency on `sql.js.org`**: if that CDN is unreachable, WASM fails to load.
- **Excel format drift**: a new source file with an unseen header layout needs a new branch in `detectFormat()`.
+143 -60
View File
@@ -1,97 +1,180 @@
# System Architecture
Static site. No backend. Browser downloads a compressed SQLite file at boot, then every query runs locally via `sql.js` (WASM).
Static site, no backend. The browser downloads a compressed SQLite file at boot
and every query runs locally via `sql.js` (SQLite compiled to WebAssembly).
One frontend, one parser, one schema, four datasets.
## Data flow
```
Excel files (.xls / .xlsx)
data/<id>/*.xls(x)
│
▼ build-database*.js (Node)
SQLite DB (public*/thptqg2017.db)
▼ parser/ (Rust, one binary, one config per dataset)
.build/public/db/<id>.db
│
▼ gzip -9
thptqg2017.db.gz (~47 MB for main variant)
▼ gzip -9 (no -k: the raw file does not survive)
.build/public/db/<id>.db.gz
│
▼ Vite build copies public*/ into dist/
Static site on GitHub Pages
▼ vite build (publicDir = .build/public)
dist/
│
▼ browser loads
sql.js (WASM) opens the .db → queries run client-side
▼ scripts/assemble-site.js
_site/ → GitHub Pages
│
▼ browser
sql.js (WASM) opens the .db → queries run client-side
```
## Three deployment variants
## The dataset id
One repo → three independent sites, same frontend, different dataset.
One identifier ties the whole pipeline together:
| Variant | Route | Source dir | Public dir | Build cmd |
|---|---|---|---|---|
| main | `/thptqg2017/` | `data/` | `public/` | `npm run build` |
| old | `/thptqg2017/old/` | `data-old/` | `public-old/` | `npm run build:old` |
| old2 | `/thptqg2017/old2/` | `data-old2/` | `public-old2/` | `npm run build:old2` |
```
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
```
`vite.config.js` reads `process.env.VARIANT` and switches `base`, `publicDir`, `outDir` accordingly. `emptyOutDir` is on only for the main build so the variant builds merge into `dist/old/`, `dist/old2/` cleanly.
`src/datasets.js` declares the four ids once. The frontend, the database build
(`parser/scripts/build-db.js`) and the site assembly all import that list, so
adding a dataset means adding one entry and one config file.
## Schema (all 3 DBs share this)
| id | Exam | Rows | Source |
| --- | --- | --- | --- |
| `2016` | 2016 | 877,461 | Bộ GD&ĐT |
| `2017` | 2017 | 861,068 | baotintuc.vn |
| `2017-old` | 2017 | 847,348 | pre-refresh archive |
| `2017-old2` | 2017 | 679,764 | corrected re-export |
## Canonical schema
Defined once in `parser/src/schema.rs` — DDL, INSERT, column order and the 16
subject regexes. The four TOML configs carry no SQL at all, only per-dataset
parse rules. Config parsing uses `deny_unknown_fields`, so a leftover `[schema]`
block fails loudly instead of looking effective while `schema.rs` drives the
build.
```sql
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY, -- exam ID (8 digits, first 2 = province)
so_bao_danh TEXT PRIMARY KEY, -- 2017: 8 digits; 2016: 9 digits or a
-- 2-4 letter cluster code then digits
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL, -- NFD-stripped lowercase for accent-insensitive search
ngay_sinh TEXT, -- dd/mm/yyyy
ho_ten_ascii TEXT NOT NULL, -- NFD-stripped lowercase, for accent-insensitive search
ngay_sinh TEXT, -- dd/mm/yyyy
ten_cum_thi TEXT, -- 2016 only
gioi_tinh TEXT, -- 2016 only
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn,
lich_su, dia_ly, gdcd, khxh,
tieng_anh, tieng_phap, tieng_nga, tieng_trung REAL
tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
```
No German / Japanese language columns — neither appears in any source file.
Every dataset gets all 22 columns; ones it has no data for are NULL, costing
about a byte per row. `khtn`, `khxh` and `gdcd` are empty on 2016;
`ten_cum_thi` and `gioi_tinh` are empty on the 2017 datasets.
Each score column is `NULL` when the student didn't take that subject. A student's "admission block" total is only computed when all 3 required subjects are non-null.
`idx_ten_cum_thi` is partial, so it holds zero entries where the column is
always NULL.
## Parse quirks (why 3 builder scripts)
## Routing
Each dataset had a different quirk. Rather than one mega-parser with mode flags, each builder script is ~70 lines and isolates its own workarounds.
URLs are flat, one segment per dataset, and the segment is the id:
| Dataset | Quirk | Mitigation |
|---|---|---|
| data/ (baotintuc .xls) | Hanoi + HCM overflow past the 65,536 row-per-sheet .xls limit | Iterate all `wb.SheetNames` |
| data-old/ (xlsx) | Single-sheet always; one bogus header row (`SOBAODANH`/`HO_TEN`) leaked into an earlier DB | Strict numeric-SBD guard rejects the leak |
| data-old2/ (xlsx) | HCM overflow + many blank trailing rows | Multi-sheet walk + blank-row pre-filter |
```
/thptqg/ hub
/thptqg/2016/
/thptqg/2017/
/thptqg/2017-old/
/thptqg/2017-old2/
```
Shared concerns (regex, schema, `toAscii`, header detection) live in `scripts/build-lib.js`.
`src/router.js` is an exact match on that segment. The nested form used before
(`/thptqg/2017/old/`) would have needed longest-prefix matching, since it also
starts with `/thptqg/2017/`. Those two legacy URLs still resolve: the router
rewrites them to the flat equivalent with `history.replaceState`, preserving the
query string so `?q=` deep links survive.
A single Vite build emits one `index.html`. Because `base` is absolute
(`/thptqg/`), that file references `/thptqg/assets/...` regardless of the
directory it is served from, so the assemble step copies it to every route and
each URL is a real static file. **No SPA 404-fallback redirect is used** — the
usual hack rewrites URLs and would interfere with the deep links.
## Serving both exam years without branching
No component contains a per-dataset conditional. Two mechanisms do the work:
- **All-NULL columns are hidden.** `score-table.jsx` drops any column where
every row in the result set is NULL, so 2016 rows surface Cụm thi / GT / Đức /
Nhật and 2017 rows surface KHTN / KHXH / GDCD / Nga.
- **Incomplete admission blocks are skipped.** `computeBlocks()` only returns a
block when the student has all three subjects, so one block list covers both
years: GDCD blocks self-exclude on 2016, German and Japanese blocks
self-exclude wherever those languages were not sat.
Anything genuinely per-dataset — title, source, database size, search examples,
SQL presets — lives in `src/datasets.js`.
## Exam ID formats
`src/lib/query-mode.js` decides whether a query is an exam ID or a name, and is
shared by `App.jsx` and `search-form.jsx` (they previously held separate copies
and had drifted apart on exactly this rule).
| Form | Example | Where |
| --- | --- | --- |
| 8 digits | `49008235` | 2017 — first two digits are the province |
| 9 digits with leading zero | `017006021` | 2016 |
| 2-4 letters then digits | `BAL000001` | 2016 — exam cluster code |
The letter-prefixed form is the majority case for 2016: 616,593 of 877,461
candidates (70.3%). Letter prefixes are upper-cased before lookup, so
`bal000001` resolves.
## Score tiers
Six-level ladder in `scoreTier()` (`src/lib/admission-blocks.js`), paired with a
symbol so meaning is never colour-only.
| Tier | Range | Vietnamese |
| --- | --- | --- |
| common | ≤ 1 | Điểm liệt |
| uncommon | < 5 | Chưa đạt |
| rare | 5–6.5 | Trung bình |
| epic | 6.5–8 | Khá |
| legendary | 8–9 | Giỏi |
| prismatic | 9–10 | Xuất sắc |
## Admission blocks
Vietnamese universities admit students based on 3-subject combinations called "admission blocks" (khối thi). `src/lib/admission-blocks.js` lists all 49 official 2017 blocks computable from our schema (A00–A11, B00–B08, C00–C20, D01–D15). Each entry is `{ code, subjects: [3 keys], label }`.
Vietnamese universities admit on three-subject combinations (khối thi).
`src/lib/admission-blocks.js` lists the blocks computable from this schema
(A00–A11, B00–B08, C00–C20, D01–D15, plus D05/D06 for German and Japanese).
`computeBlocks(student)` returns those where all three scores exist, sorted by
total descending.
`computeBlocks(student)` returns only blocks where the student has all 3 subject scores, sorted by total desc. Used by the student detail card.
## Design decisions
The SQL preset "Top 10 best-block in Long An" materialises all 49 blocks as a `UNION ALL` CTE, picks each student's best block via `ROW_NUMBER() OVER (PARTITION BY so_bao_danh ORDER BY s DESC, k)`, then ranks across students.
| Concern | Choice | Rationale |
| --- | --- | --- |
| Storage | Static SQLite file | No backend; the datasets are frozen |
| Compression | gzip in CI, `DecompressionStream` in the browser | Native API, no extra library |
| WASM hosting | `sql.js.org` CDN | Smaller self-hosted artifact |
| Diacritics search | Pre-computed `ho_ten_ascii` | `LOWER(REPLACE(...))` at query time defeats the index |
| SQL safety | Leading-keyword allowlist | `sql.js` is in-memory so writes cannot persist; the allowlist prevents confusion |
| Row caps | 100 (lookup), 1000 (SQL) | Keeps DOM render sizes reasonable |
| Routing | Hand-rolled, ~20 lines | Five static routes do not justify a router dependency |
## Score tiers (UI coding)
## Risks and limitations
6-level TFT rarity ladder (white → green → blue → purple → gold → prismatic). Defined in `scoreTier()` in `src/lib/admission-blocks.js`.
| Tier | Range | UI label (Vietnamese) | English |
|---|---|---|---|
| common | ≤ 1 | Điểm liệt | Disqualifying |
| uncommon | < 5 | Chưa đạt | Below passing |
| rare | 5–6.5 | Trung bình | Average |
| epic | 6.5–8 | Khá | Good |
| legendary | 8–9 | Giỏi | Very good |
| prismatic | 9–10 | Xuất sắc | Excellent |
Color + icon + text label — never color-only.
## Frontend key behaviors
- Query state lives in `App.jsx` and is synced to URL `?q=...` via `history.replaceState` — bookmarkable, shareable.
- `SearchForm` is a controlled component that 300ms-debounces input changes and auto-triggers search once the query is long enough (3+ digits for SBD, 2+ chars for name).
- Detection: all-digits → exact `so_bao_danh` match; ASCII-only → folded `ho_ten_ascii LIKE`; has Vietnamese diacritics → search both `ho_ten` and folded column.
- Exactly 1 result → `StudentDetail` card; 0 or >1 → `ScoreTable`.
- Global `/` key focuses the search input when not already typing.
- Share button on the detail card uses Web Share API when present, else clipboard with a pre-formatted summary + deep-link URL.
- **Database size.** 38–48 MB gzipped per dataset; slow links wait, mitigated by
a progress bar.
- **Browser memory.** The full database lives in RAM; older mobile devices may
run out.
- **`sql.js.org` dependency.** If that CDN is unreachable, the WASM fails to
load. Self-hosting `sql-wasm.wasm` and updating `SQL_WASM_URL` in
`use-sqlite.js` is the fix.
- **Excel format drift.** A new source file with an unseen header layout needs a
new branch in `format_detect_2016.rs` or a new config.
+4 -1
View File
@@ -19,7 +19,10 @@
/// SBD(0) HO_TEN(1) NGAY_SINH(2) TEN_CUMTHI(3) GIOI_TINH(4) DIEM_THI(5)
/// → JS: build-database.js:149–151 DEFAULT_MAP + processMappedRow
///
/// JS citations are line numbers in /config/workspace/tiennm99/thptqg2016/scripts/build-database.js.
/// The `build-database.js` citations throughout this file refer to the Node
/// script this parser replaced, in the original standalone thptqg2016 repo.
/// That file no longer exists here; the references are kept because they
/// explain why several of the rules below look arbitrary.
use calamine::Data;
use crate::transform::{parse_scores, to_ascii, CompiledPatterns, ParsedRow};
@@ -1,7 +1,7 @@
---
phase: 1
title: "Standard schema and unified parser"
status: pending
status: completed
priority: P1
dependencies: []
effort: ""
@@ -1,7 +1,7 @@
---
phase: 2
title: "Repo restructure"
status: pending
status: completed
priority: P1
dependencies: [1]
effort: ""
@@ -131,12 +131,12 @@ component.
## Success Criteria
- [ ] Target layout matches `plan.md` exactly
- [ ] `2016/` and `2017/` no longer exist
- [ ] Old hub page content captured before deletion
- [ ] Git reports renames, repository size unchanged
- [ ] `package-lock.json` committed; `npm ci` succeeds; zero pnpm files remain
- [ ] `cargo test` green from the new location
- [x] Target layout matches `plan.md` exactly
- [x] `2016/` and `2017/` no longer exist
- [x] Old hub page content captured before deletion
- [x] Git reports renames, repository size unchanged
- [x] `package-lock.json` committed; `npm ci` succeeds; zero pnpm files remain
- [x] `cargo test` green from the new location
## Risk Assessment
@@ -1,7 +1,7 @@
---
phase: 3
title: "Unified frontend and dataset registry"
status: pending
status: completed
priority: P1
dependencies: [2]
effort: ""
@@ -225,16 +225,16 @@ this repo today):
## Success Criteria
- [ ] One `src/` serving four datasets **and** the hub, zero `if (dataset === ...)` in components
- [ ] Subject list defined once in `src/lib/subjects.js`
- [ ] Hub route fetches no database
- [ ] Dataset path and DB URL are derived from `id`, not stored per entry
- [ ] Legacy `/2017/old/` and `/2017/old2/` redirect to flat URLs, `?q=` preserved
- [ ] 2016 route renders cluster + gender; 2017 routes do not
- [ ] 2016 route has live search, deep links, student detail, score tiers
- [ ] Admission blocks correct for both exam years
- [ ] No routing library added
- [ ] `npm run lint` green
- [x] One `src/` serving four datasets **and** the hub, zero `if (dataset === ...)` in components
- [x] Subject list defined once in `src/lib/subjects.js`
- [x] Hub route fetches no database
- [x] Dataset path and DB URL are derived from `id`, not stored per entry
- [x] Legacy `/2017/old/` and `/2017/old2/` redirect to flat URLs, `?q=` preserved
- [x] 2016 route renders cluster + gender; 2017 routes do not
- [x] 2016 route has live search, deep links, student detail, score tiers
- [x] Admission blocks correct for both exam years
- [x] No routing library added
- [x] `npm run lint` green
## Risk Assessment
@@ -1,7 +1,7 @@
---
phase: 4
title: "Build and deploy pipeline"
status: pending
status: completed
priority: P1
dependencies: [3]
effort: ""
@@ -1,7 +1,7 @@
---
phase: 5
title: "Parity verification and docs"
status: pending
status: completed
priority: P1
dependencies: [4]
effort: ""
@@ -110,13 +110,13 @@ so reruns compare the same students.
## Success Criteria
- [ ] `verify-parity.js` committed and exiting 0 on all four datasets
- [ ] Parity result report committed under `plans/reports/`
- [ ] Zero unexpected non-NULL columns
- [ ] Manual checklist passes on the hub and all four dataset routes
- [ ] `verify-parity.js` has zero npm dependencies (`node:sqlite` only)
- [ ] `README.md` and `docs/` describe the actual repo, with npm commands
- [ ] No stale path, pnpm, or build-variant references anywhere outside `plans/`
- [x] `verify-parity.js` committed and exiting 0 on all four datasets
- [x] Parity result report committed under `plans/reports/`
- [x] Zero unexpected non-NULL columns
- [x] Manual checklist passes on the hub and all four dataset routes
- [x] `verify-parity.js` has zero npm dependencies (`node:sqlite` only)
- [x] `README.md` and `docs/` describe the actual repo, with npm commands
- [x] No stale path, pnpm, or build-variant references anywhere outside `plans/`
## Risk Assessment
@@ -1,7 +1,7 @@
---
title: "Unify frontend, standardize SQL schema, restructure repo"
description: "One 2017-based frontend, one canonical 22-column student schema, one parser crate, four datasets under data/"
status: pending
status: completed
priority: P2
branch: "main"
tags: [refactor, schema, frontend, parser]
@@ -134,11 +134,11 @@ CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT N
| Phase | Name | Status |
|-------|------|--------|
| 1 | [Standard schema and unified parser](./phase-01-standard-schema-and-unified-parser.md) | Pending |
| 2 | [Repo restructure](./phase-02-repo-restructure.md) | Pending |
| 3 | [Unified frontend and dataset registry](./phase-03-unified-frontend-and-dataset-registry.md) | Pending |
| 4 | [Build and deploy pipeline](./phase-04-build-and-deploy-pipeline.md) | Pending |
| 5 | [Parity verification and docs](./phase-05-parity-verification-and-docs.md) | Pending |
| 1 | [Standard schema and unified parser](./phase-01-standard-schema-and-unified-parser.md) | Completed |
| 2 | [Repo restructure](./phase-02-repo-restructure.md) | Completed |
| 3 | [Unified frontend and dataset registry](./phase-03-unified-frontend-and-dataset-registry.md) | Completed |
| 4 | [Build and deploy pipeline](./phase-04-build-and-deploy-pipeline.md) | Completed |
| 5 | [Parity verification and docs](./phase-05-parity-verification-and-docs.md) | Completed |
## Dependencies
@@ -151,17 +151,17 @@ No cross-plan dependencies (`plans/` was empty before this plan).
## Acceptance Criteria
- [ ] One Rust crate; `2016/tools/` and `2017/tools/` gone
- [ ] One `src/`; `2016/src/` and `2017/src/` gone
- [ ] One Vite build producing all five pages
- [ ] All four DBs built from the same DDL, same INSERT, same 16 regexes
- [ ] Per-dataset row count and per-column non-NULL count identical to pre-refactor baseline
- [ ] All five published URLs functional under the flat scheme, deep links included
- [ ] Legacy `/2017/old/` and `/2017/old2/` redirect to their flat equivalents, preserving `?q=`
- [ ] 2016 site still shows `ten_cum_thi` + `gioi_tinh`; 2017 sites do not
- [ ] 2016 site gains 2017's features (deep links, student detail, tiers, live search)
- [ ] npm only: `package-lock.json` committed, no pnpm files or CI steps remain
- [ ] `cargo test` green; `npm run lint` green
- [x] One Rust crate; `2016/tools/` and `2017/tools/` gone
- [x] One `src/`; `2016/src/` and `2017/src/` gone
- [x] One Vite build producing all five pages
- [x] All four DBs built from the same DDL, same INSERT, same 16 regexes
- [x] Per-dataset row count and per-column non-NULL count identical to pre-refactor baseline
- [x] All five published URLs functional under the flat scheme, deep links included
- [x] Legacy `/2017/old/` and `/2017/old2/` redirect to their flat equivalents, preserving `?q=`
- [x] 2016 site still shows `ten_cum_thi` + `gioi_tinh`; 2017 sites do not
- [x] 2016 site gains 2017's features (deep links, student detail, tiers, live search)
- [x] npm only: `package-lock.json` committed, no pnpm files or CI steps remain
- [x] `cargo test` green; `npm run lint` green
## Rollback
+79
View File
@@ -0,0 +1,79 @@
# Parser parity result
Comparison of the databases the unified pipeline ships against a baseline built
from the pre-refactor code. Run on 2026-08-13 against the final tree
(`npm run build:rust && npm run build:db`), gate exit code 0.
Inputs:
- `plans/reports/parser-parity-baseline.json` — built with the two separate
crates, before any change
- `plans/reports/parser-parity-current.json` — decompressed from
`.build/public/db/*.db.gz`, i.e. the exact bytes published
## Result
| Dataset | Rows | Columns | Size |
| --- | --- | --- | --- |
| 2016 | 877,461 | 18 → 22 | +1.4% |
| 2017 | 861,068 | 18 → 22 | +2.2% |
| 2017-old | 847,348 | 18 → 22 | +2.2% |
| 2017-old2 | 679,764 | 18 → 22 | +2.2% |
Unchanged, for every dataset:
- row count
- non-NULL count for all 18 pre-existing columns
- every field of a deterministic student sample (SBDs ending `0000`)
## Approved recoveries
Unifying the subject regexes recovered 1,691 foreign-language scores that the
old per-year configs discarded. The 2016 config listed 12 subject patterns and
the 2017 configs 14; neither list was complete, so candidates who sat German,
Japanese or Russian ended up with no foreign-language score at all.
| Dataset | Recovered |
| --- | --- |
| 2016 | 182 × `tieng_nga` |
| 2017 | 93 × `tieng_duc`, 512 × `tieng_nhat` |
| 2017-old | 85 × `tieng_duc`, 484 × `tieng_nhat` |
| 2017-old2 | 22 × `tieng_duc`, 313 × `tieng_nhat` |
Confirmed real rather than spurious matches:
- Every student in all four datasets holds zero or exactly one foreign
language — never two — so nothing is duplicated.
- Each affected student had all language columns NULL beforehand. SBD
`01003198` went from all-NULL to `tieng_duc = 8`.
- Counts track dataset size across the three 2017 generations (93/85/22
German, 512/484/313 Japanese).
- Values are ordinary exam scores in the 0–10 range.
These exact counts are pinned as `APPROVED_RECOVERY` in
`parser/scripts/verify-parity.js`. Any other newly-populated column, or drift
in these numbers, still fails the gate.
## Columns that must stay empty
Verified 0 non-NULL:
- 2016: `khtn`, `khxh`, `gdcd`
- 2017, 2017-old, 2017-old2: `ten_cum_thi`, `gioi_tinh`
## Reproducing
```sh
npm run build:rust && npm run build:db
node parser/scripts/db-stats.js \
2016=<db> 2017=<db> 2017-old=<db> 2017-old2=<db> > current.json
node parser/scripts/verify-parity.js \
plans/reports/parser-parity-baseline.json current.json
```
The baseline cannot be regenerated — it required the two pre-refactor crates,
which no longer exist. It is committed for that reason.
## Unresolved questions
None.