diff --git a/README.md b/README.md index a50ca4a..da6a3df 100644 --- a/README.md +++ b/README.md @@ -2,18 +2,67 @@ Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high school graduation exam. Client-side SQL (sql.js) over a SQLite database built -from the ministry's raw `.xls` score files by the Rust `xlsxread` CLI. +from the ministry's raw `.xls` score files by the Rust `xlsxread` parser. -Served as one site at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**. +Live at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**. -| Directory | Year | Size | Site | +| Dataset | Exam | Candidates | Site | | --- | --- | --- | --- | -| [`2016/`](./2016/) | THPT QG 2016 — 877,461 candidates | ~50 MB raw data | [/2016/](https://tiennm99.github.io/thptqg/2016/) | -| [`2017/`](./2017/) | THPT QG 2017 — ~861,000 candidates | ~117 MB raw data | [/2017/](https://tiennm99.github.io/thptqg/2017/) (+ [old](https://tiennm99.github.io/thptqg/2017/old/), [old2](https://tiennm99.github.io/thptqg/2017/old2/) generations) | +| `2016` | 2016 | 877,461 | [/2016/](https://tiennm99.github.io/thptqg/2016/) | +| `2017` | 2017 | 861,068 | [/2017/](https://tiennm99.github.io/thptqg/2017/) | +| `2017-old` | 2017 | 847,348 | [/2017-old/](https://tiennm99.github.io/thptqg/2017-old/) | +| `2017-old2` | 2017 | 679,764 | [/2017-old2/](https://tiennm99.github.io/thptqg/2017-old2/) | -Each year was a standalone repository (`thptqg2016`, `thptqg2017`), merged here -with full history. Each keeps its own `tools/xlsxread` copy because the raw -data formats differ per year. +The three 2017 datasets are successive publications of the same exam and they +disagree; all three are kept so the differences stay inspectable. -The `Deploy to GitHub Pages` workflow builds both years' databases and site -bundles and publishes them behind a root index. +## Layout + +``` +index.html + src/ the frontend — one app serving all four datasets and the hub + datasets.js the four dataset ids and their per-dataset content + router.js pathname → dataset +data// raw Excel files, one directory per dataset +parser/ the Rust parser + src/schema.rs canonical 22-column table: DDL, INSERT, subject regexes + configs/.toml per-dataset parse rules only, no SQL + scripts/ database build, crawler, parity verification +scripts/ site assembly +docs/ architecture, data pipeline, deployment +``` + +The dataset id is one identifier end to end: + +``` +data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/ +``` + +## Build + +```bash +npm ci +npm run build:rust # compile the parser +npm run build:db # build + gzip all four databases (add an id for just one) +npm run build:site # one Vite build, then assemble into _site/ +npx serve _site +``` + +Pushing to `main` runs the same steps in +`.github/workflows/deploy-pages.yml` and publishes to GitHub Pages. + +## Adding a dataset + +1. Put the Excel files in `data//` +2. Add `parser/configs/.toml` — sheet mode, column indices, validation + guards. No SQL; the schema is canonical. +3. Add an entry to `DATASETS` in `src/datasets.js` + +Everything else follows: the build script, the site assembly and the router all +read that one list, and the UI adapts to whichever columns the dataset fills. + +## Docs + +See [`docs/`](./docs/) — [overview](./docs/project-overview.md), +[architecture](./docs/system-architecture.md), +[data pipeline](./docs/data-pipeline.md), +[deployment](./docs/deployment-guide.md). diff --git a/docs/README.md b/docs/README.md index a16c8b6..b2ad89a 100644 --- a/docs/README.md +++ b/docs/README.md @@ -1,5 +1,6 @@ # Docs -- [`system-architecture.md`](./system-architecture.md) — data flow, schema, 3-variant deploy, score-tier model, admission-block catalog -- [`data-pipeline.md`](./data-pipeline.md) — Excel parse quirks, overflow-sheet gotcha, audit scripts, refresh flow -- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, adding a new variant, rollback +- [`project-overview.md`](./project-overview.md) — goal, scope, constraints, the four datasets, history +- [`system-architecture.md`](./system-architecture.md) — data flow, canonical schema, routing, how one frontend serves both exam years +- [`data-pipeline.md`](./data-pipeline.md) — Excel parse quirks, per-dataset formats, overflow-sheet gotcha, expected row counts, verifying a rebuild +- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, adding a dataset, rollback, troubleshooting diff --git a/docs/codebase-summary-2016-legacy.md b/docs/codebase-summary-2016-legacy.md deleted file mode 100644 index 29626e7..0000000 --- a/docs/codebase-summary-2016-legacy.md +++ /dev/null @@ -1,63 +0,0 @@ -# Codebase Summary - -## Directory layout - -``` -thptqg2016/ -├── data/ # Source Excel files (~100, mixed formats) -├── scripts/ -│ └── build-database.js # Parse Excel → SQLite (build-time, Node + better-sqlite3) -├── public/ -│ └── thptqg2016.db # Generated DB, gzipped during CI -├── src/ -│ ├── main.jsx # React entry -│ ├── App.jsx # Root: tabs, lookup logic, useSqlite wiring -│ ├── App.css / index.css # Design tokens, dark mode, a11y styles -│ ├── hooks/ -│ │ └── use-sqlite.js # Fetch .db.gz + decompress + init sql.js -│ └── components/ -│ ├── search-form.jsx # Input for exam ID / full name -│ ├── score-table.jsx # Result table for lookups -│ └── custom-query.jsx # SQL editor + presets + result grid -├── .github/workflows/deploy.yml # CI: build db → gzip → vite build → Pages -├── vite.config.js # base: "/thptqg2016/" -└── eslint.config.js -``` - -## Key modules - -### `scripts/build-database.js` -Build-time only. Reads every `.xlsx/.xls` in `data/`, detects the header format (three variants), parses the `DIEM_THI` string via regex or separate score columns, normalizes gender, derives a diacritics-stripped `ho_ten_ascii` column for accent-insensitive search, and inserts into SQLite with three indexes (`ho_ten`, `ho_ten_ascii`, `ten_cum_thi`). - -### `src/hooks/use-sqlite.js` -Streams `.db.gz` with download progress, decompresses via `DecompressionStream("gzip")`, loads `sql.js` (WASM served from the `sql.js.org` CDN), and returns `{ db, loading, error, progress }`. - -### `src/App.jsx` -Two tabs: **Lookup** and **Custom SQL**. Lookup auto-detects exam IDs (regex `^[A-Z]{2,4}\d+$`) vs names and picks one of three query paths: exact exam ID / ASCII LIKE / original + ASCII LIKE. Capped at 100 rows. - -### `src/components/custom-query.jsx` -Whitelists leading keywords (`SELECT`, `PRAGMA`, `EXPLAIN`, `WITH`), auto-appends `LIMIT 1000` when missing, measures `performance.now()` execution time, and ships 7 preset analytics queries. - -## `student` table schema - -```sql -so_bao_danh TEXT PRIMARY KEY -- exam ID -ho_ten TEXT NOT NULL -- full name -ho_ten_ascii TEXT NOT NULL -- diacritics stripped, lowercased -ngay_sinh TEXT -- date of birth -ten_cum_thi TEXT -- exam cluster name -gioi_tinh TEXT -- "Nam" | "Nữ" | NULL -toan, ngu_van, vat_ly, hoa_hoc, -- REAL (nullable) subject scores -sinh_hoc, lich_su, dia_ly, -tieng_anh, tieng_phap, tieng_duc, -tieng_nhat, tieng_trung -``` - -Indexes: `idx_ho_ten`, `idx_ho_ten_ascii`, `idx_ten_cum_thi`. - -## Conventions - -- JS/JSX filenames: **kebab-case** (e.g., `search-form.jsx`, `use-sqlite.js`) -- React components: `PascalCase` named exports -- UI strings: Vietnamese (target audience) -- Code comments: English; explain *why*, not *what* diff --git a/docs/data-pipeline.md b/docs/data-pipeline.md index 589f7d4..aa8d477 100644 --- a/docs/data-pipeline.md +++ b/docs/data-pipeline.md @@ -2,109 +2,148 @@ From raw Excel files to a compressed SQLite file the browser can load. -## Canonical source (data/) +One Rust binary (`parser/`) builds every dataset. What differs per dataset is +parse rules only — sheet strategy, column layout, validation guards — declared +in `parser/configs/.toml`. The table shape, the INSERT and the subject +regexes are canonical and live in `parser/src/schema.rs`. -63 Excel files from baotintuc.vn CDN. Source article: -`https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm` +## Sources -Re-download anytime: +| id | Files | Reproducible | Origin | +| --- | --- | --- | --- | +| `2016` | 4 `.xls` + 115 `.xlsx` | no | Bộ GD&ĐT, collected 2016 | +| `2017` | 63 `.xls` | **yes** | baotintuc.vn CDN | +| `2017-old` | 63 `.xlsx` | no | pre-refresh archive | +| `2017-old2` | 54 `.xlsx` | no | corrected re-export | + +Only `2017` can be re-fetched: ```bash -node scripts/crawl-baotintuc.js +node parser/scripts/crawl-baotintuc.js ``` -Idempotent — skips files already present. Saves to `data/.xls`. +Idempotent — skips files already present, saves to `data/2017/`. Source article: +`https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm` -## Reference sources (data-old/, data-old2/) +The other three were collected from Vietnamese news sites at the time and the +publisher URLs were not recorded. **The files committed in git are the only +copy.** -Both were collected from Vietnamese news sites around the time of the 2017 exam. -Specific publisher URLs were not recorded and cannot be recovered, so these -datasets are **not reproducible** — the files committed in git are the only -copy. They are kept for historical comparison and to expose the deploy matrix; -the main deployment always uses the reproducible `data/` source. +## Source Excel shapes -## Source Excel shape - -Every file is a single sheet (for `data-old/`) or up to 3 sheets (for `data/` .xls) with columns: +The 2017 datasets share one layout: | Col | Name | Content | -|---|---|---| -| 0 | HO_TEN | full name in Vietnamese | -| 1 | NGAY_SINH | `dd/mm/yyyy` | -| 2 | SOBAODANH | 8-digit string, first 2 digits = province | -| 3 | DIEM_THI | concatenated per-subject score string in the source language, e.g. `"Toan: 6.80 Van: 5.25 ... English: 5.80"` | +| --- | --- | --- | +| 0 | HO_TEN | full name in Vietnamese | +| 1 | NGAY_SINH | `dd/mm/yyyy` | +| 2 | SOBAODANH | 8-digit string, first 2 digits = province | +| 3 | DIEM_THI | concatenated per-subject scores, e.g. `"Toán: 6.80 Ngữ văn: 5.25 …"` | -Some files lack a header row. `isHeaderRow()` in `build-lib.js` detects and skips. +2016 has **three** layouts across its 119 files, chosen per file at runtime via +`format_detection = "thptqg2016"` in its config: + +| Format | Detected by | Notes | +| --- | --- | --- | +| `separate-scores` | `row[0] == "SBD"` and `row[2] == "TOAN"` | one column per subject; scores read directly, no regex | +| `mapped` | header contains `SOBAODANH`/`SBD` plus `DIEM_THI` | column indices derived from the header | +| `default` | no recognised header | positional 6-column layout | + +The `mapped` and `default` layouts also carry `TEN_CUMTHI` and `GIOI_TINH`, +which is why only 2016 populates those columns. ## Score text parsing -`SCORE_PATTERNS` in `build-lib.js` defines one regex per subject. The patterns must match the exact subject labels the source files use (Vietnamese). +`SCORE_PATTERNS` in `parser/src/schema.rs` defines one regex per subject, and +**all 16 run against every dataset**. A subject a given exam year did not offer +simply never matches and stays NULL. -Subjects supported (DB column / label in source text): +| Column | Source label | English | 2016 | 2017 | +| --- | --- | --- | --- | --- | +| toan | Toán | Math | ✓ | ✓ | +| ngu_van | Ngữ văn | Literature | ✓ | ✓ | +| vat_ly | Vật lí | Physics | ✓ | ✓ | +| hoa_hoc | Hóa học | Chemistry | ✓ | ✓ | +| sinh_hoc | Sinh học | Biology | ✓ | ✓ | +| khtn | KHTN | Natural Sciences combined | — | ✓ | +| lich_su | Lịch sử | History | ✓ | ✓ | +| dia_ly | Địa lí | Geography | ✓ | ✓ | +| gdcd | GDCD | Civic Education | — | ✓ | +| khxh | KHXH | Social Sciences combined | — | ✓ | +| tieng_anh | Tiếng Anh | English | ✓ | ✓ | +| tieng_phap | Tiếng Pháp | French | ✓ | ✓ | +| tieng_nga | Tiếng Nga | Russian | ✓ | ✓ | +| tieng_duc | Tiếng Đức | German | ✓ | ✓ | +| tieng_nhat | Tiếng Nhật | Japanese | ✓ | ✓ | +| tieng_trung | Tiếng Trung | Chinese | ✓ | ✓ | -| Column | Source label | English | -|---|---|---| -| toan | Toán | Math | -| ngu_van | Ngữ văn | Literature | -| vat_ly | Vật lí | Physics | -| hoa_hoc | Hóa học | Chemistry | -| sinh_hoc | Sinh học | Biology | -| khtn | KHTN | Natural Sciences combined | -| lich_su | Lịch sử | History | -| dia_ly | Địa lí | Geography | -| gdcd | GDCD | Civic Education | -| khxh | KHXH | Social Sciences combined | -| tieng_anh | Tiếng Anh | English | -| tieng_phap | Tiếng Pháp | French | -| tieng_nga | Tiếng Nga | Russian | -| tieng_trung | Tiếng Trung | Chinese | +### The foreign-language recovery -Missing-in-string → `NULL` in DB (student didn't take that subject). +Before the parser was unified, the 2016 config listed 12 subject patterns and +the 2017 configs listed 14. Neither list was complete: candidates could sit +German, Japanese and Russian in both years, so 1,691 students ended up with **no +foreign-language score at all**. -## Why three parsers +Running all 16 patterns everywhere recovered them — 182 Russian in 2016, and +German and Japanese across all three 2017 generations. See +`plans/reports/parser-parity-result.md` for the evidence that these are real +scores rather than false matches. -One shared lib + three thin parsers, not one parser with flags. Each dataset's quirks stay visible in its own file: +## Per-dataset quirks -| File | Source dir | Sheet strategy | Extra guard | -|---|---|---|---| -| `build-database.js` | `data/` | iterate ALL sheets (.xls row overflow) | — | -| `build-database-old.js` | `data-old/` | sheet 0 only | reject rows where `so_bao_danh` isn't digits-only (rejects an `HO_TEN` → `SOBAODANH` header leak present in this export) | -| `build-database-old2.js` | `data-old2/` | iterate ALL sheets (HCM overflow) | skip fully blank rows before counting | +| id | Sheet strategy | Guards | +| --- | --- | --- | +| `2016` | all sheets | per-file format detection; header-token rows rejected | +| `2017` | all sheets — Hà Nội and HCM overflow | none | +| `2017-old` | first sheet only | reject non-numeric SBD (rejects a header leak in this export) | +| `2017-old2` | all sheets — HCM overflows | reject non-numeric SBD; skip blank rows before counting | -## Overflow-sheet gotcha +### Overflow-sheet gotcha -The old `.xls` binary format caps at 65,536 rows per sheet. Hanoi and Ho Chi Minh City exceed that in the baotintuc export, so Excel auto-splits the overflow into `Sheet2`. **Only reading `Sheet1` silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). +The old `.xls` format caps at 65,536 rows per sheet. Hanoi and Ho Chi Minh City +exceed that, so Excel splits the overflow into `Sheet2`. **Reading only `Sheet1` +silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what +`sheet_mode = "all"` exists for. -The audit `node scripts/audit-row-counts.js` catches this by comparing source-row-count (across all sheets) to DB-row-count. +## Expected row counts -## Audits +| id | Source rows | Skipped | DB rows | +| --- | --- | --- | --- | +| `2016` | 877,464 | 3 duplicate SBDs collapsed | **877,461** | +| `2017` | 861,131 | 63 empty | **861,068** | +| `2017-old` | 847,349 | 1 header leak | **847,348** | +| `2017-old2` | 679,764 | 0 | **679,764** | -Run any time after a rebuild: +## Verifying a rebuild + +`parser/scripts/db-stats.js` dumps row counts, per-column non-NULL counts, file +size and a deterministic student sample as JSON. +`parser/scripts/verify-parity.js` diffs two such files and exits non-zero on any +mismatch. ```bash -node scripts/audit-row-counts.js # source rows vs DB rows (zero-loss check) -node scripts/check-duplicates.js # md5 + row-content duplicates in data/ -node scripts/diff-datasets.js # compare two DBs (needs backup/ dir) +node parser/scripts/db-stats.js 2016=.db … > current.json +node parser/scripts/verify-parity.js plans/reports/parser-parity-baseline.json current.json ``` -Expected baseline: +The committed baseline was built with the pre-refactor code and cannot be +regenerated — the two old crates no longer exist. Both scripts use the built-in +`node:sqlite`, so they need no dependencies. -| DB | Source rows | Skipped | DB rows | -|---|---|---|---| -| main | 861,131 | 63 empty rows | **861,068** | -| old | 847,349 | 1 header-leak | **847,348** | -| old2 | 679,764 | 0 | **679,764** | - -## Refreshing the data +## Refreshing the 2017 data ```bash -rm data/*.xls # clear current snapshot -node scripts/crawl-baotintuc.js # re-download from CDN -npm run build:db # rebuild -node scripts/audit-row-counts.js # verify -gzip -kf -9 public/thptqg2017.db # compress for shipping +rm data/2017/*.xls +node parser/scripts/crawl-baotintuc.js +node parser/scripts/build-db.js 2017 ``` -## Adding a new source dataset +Then re-run the parity check above and confirm the row count still matches. -See `docs/deployment-guide.md` → "Adding a new variant". +## Legacy scripts + +`parser/scripts/check-duplicates.js` and `diff-datasets.js` are one-off audits +that were already broken before the repo was unified — a hardcoded Windows path +in one, an undeclared `better-sqlite3` dependency and stale paths in the other. +Each carries a comment saying so. For comparing two builds, use `db-stats.js` +plus `verify-parity.js` instead. diff --git a/docs/deployment-guide-2016-legacy.md b/docs/deployment-guide-2016-legacy.md deleted file mode 100644 index 87e4a03..0000000 --- a/docs/deployment-guide-2016-legacy.md +++ /dev/null @@ -1,50 +0,0 @@ -# Deployment Guide - -## Automatic (recommended) - -Push to `main` → GitHub Actions builds and deploys to GitHub Pages automatically. - -Workflow: `.github/workflows/deploy.yml` - -CI steps: -1. `npm ci` -2. `npm run build:db` — generate `public/thptqg2016.db` from `data/*.xlsx` -3. `gzip -k -9 public/thptqg2016.db` — max compression, keep original -4. `npm run build` — Vite bundles `dist/` -5. `rm -f dist/thptqg2016.db` — ship only the gzipped copy -6. `actions/upload-pages-artifact@v3` + `actions/deploy-pages@v4` - -One-time setup: **Settings → Pages → Source: GitHub Actions**. - -## Manual (local verification) - -```bash -npm install -npm run build:db -gzip -k -9 public/thptqg2016.db # Linux/macOS; on Windows use 7zip or WSL -npm run build -rm dist/thptqg2016.db # optional, shrinks artifact -npm run preview # serve dist/ locally -``` - -Open . - -## Base path - -`vite.config.js` sets `base: "/thptqg2016/"`. If you fork under a different repo name, update this to match `` so assets resolve correctly on GitHub Pages. - -## Updating data - -1. Add the new Excel file to `data/`. -2. If its header is unfamiliar, open `scripts/build-database.js` and extend `detectFormat()` or the `processSeparateScoresRow` / `processMappedRow` helpers. -3. Run `npm run build:db` locally to check row counts and error skips. -4. Commit + push → CI redeploys. - -## Troubleshooting - -| Symptom | Typical cause | -|---------|---------------| -| Blank page, 404 on assets | `base` in `vite.config.js` doesn't match the repo name | -| `Failed to fetch database: 404` | gzip step skipped, or `.db.gz` removed from `dist/` | -| WASM fails to load | `sql.js.org` blocked / offline — self-host `sql-wasm.wasm` in `public/` and update `SQL_WASM_URL` in `use-sqlite.js` | -| Missing rows after build | Excel file has an unknown header — check console for `Failed to read` or `errorCount` | diff --git a/docs/deployment-guide.md b/docs/deployment-guide.md index 4410294..f5b29b8 100644 --- a/docs/deployment-guide.md +++ b/docs/deployment-guide.md @@ -1,52 +1,105 @@ # Deployment Guide -Deploys to GitHub Pages via `.github/workflows/deploy.yml`. Every push to `main` rebuilds and redeploys all 3 variants. +Deploys to GitHub Pages via `.github/workflows/deploy-pages.yml`. Every push to +`main` rebuilds all four datasets and redeploys the whole site. + +One-time setup: **Settings → Pages → Source: GitHub Actions**. ## What the workflow does -1. Checkout, setup Node 20, `npm ci` -2. `npm run build:db:all` — builds 3 SQLite DBs (main, old, old2) -3. Gzip each DB at level 9 -4. `npm run build:all` — builds 3 Vite bundles into `dist/`, `dist/old/`, `dist/old2/` -5. Remove uncompressed `.db` files from `dist/` (only `.gz` ships) -6. `actions/upload-pages-artifact` + `actions/deploy-pages` +1. Checkout, Rust toolchain, Node 24, `npm ci` +2. `npm run build:rust` — one parser binary +3. `npm run build:db` — builds and gzips all four databases into + `.build/public/db/` +4. `npm run build:site` — one Vite build, then `scripts/assemble-site.js` +5. `actions/upload-pages-artifact` + `actions/deploy-pages` -Total CI time ≈ 4–6 min (DB build dominates). +The database build dominates the runtime: roughly 419 MB of Excel is parsed on +every deploy. ## Resulting URLs -- `https://.github.io/thptqg2017/` -- `https://.github.io/thptqg2017/old/` -- `https://.github.io/thptqg2017/old2/` +``` +https://.github.io/thptqg/ +https://.github.io/thptqg/2016/ +https://.github.io/thptqg/2017/ +https://.github.io/thptqg/2017-old/ +https://.github.io/thptqg/2017-old2/ +``` + +`/thptqg/2017/old/` and `/thptqg/2017/old2/` were the pre-flattening URLs. They +are still served, and the router rewrites them to the flat form with the query +string intact. ## Local reproduction ```bash npm ci -npm run build:db:all -gzip -kf -9 public/thptqg2017.db public-old/thptqg2017.db public-old2/thptqg2017.db -npm run build:all -npx serve dist # or any static server +npm run build:rust +npm run build:db # all four; pass an id to build just one +npm run build:site # vite build + assemble into _site/ +npx serve _site ``` -## Adding a new variant +To rebuild a single dataset: -1. Drop source Excel files into a new `data-vN/` folder -2. Add `public-vN/` to `.gitignore` patterns for the `.db` + `.db.gz` -3. Copy `scripts/build-database-old.js` → `scripts/build-database-vN.js`; update `SRC_DIR` / `DB_PATH`; adjust parse logic for the new quirks (multi-sheet? blank rows? numeric guard?) -4. Add variant to `VARIANT_CONFIG` in `vite.config.js`: - ```js - vN: { base: "/thptqg2017/vN/", publicDir: "public-vN", outDir: "dist/vN" } - ``` -5. Add `build:db:vN` and `build:vN` npm scripts; wire them into `build:db:all` and `build:all` -6. Update `.github/workflows/deploy.yml` — add gzip step + rm step for the new DB +```bash +node parser/scripts/build-db.js 2017-old +``` + +## Base path + +`vite.config.js` sets `base: "/thptqg/"`. If you fork under a different repo +name, update it to match — assets are referenced absolutely, so a mismatch shows +up as a blank page with 404s on `/assets/...`. + +## Adding a dataset + +1. Put the Excel files in `data//` +2. Add `parser/configs/.toml` with the parse rules — sheet mode, column + indices, SBD validation, header tokens, blank-row stripping. No SQL: the + schema is canonical and lives in `parser/src/schema.rs` +3. Add an entry to `DATASETS` in `src/datasets.js` + +Nothing else. The build script, the site assembly and the router all read that +one list, and the frontend adapts to whichever columns the dataset populates. + +## Why no uncompressed database can ship + +`build-db.js` runs `gzip -9` **without** `-k`, so the raw file does not survive +the build. `assemble-site.js` then fails the job if any `.db`, `.db-journal`, +`.db-wal` or `.db-shm` reached the output. + +Both guards exist because the previous pipeline wrote a 100+ MB uncompressed +database into the source tree and relied on an `rm` step to keep it out of the +artifact — one missing line away from publishing it. ## Notes -- **DB gzip is non-cacheable across deploys** — every rebuild produces a new `.db.gz` (not byte-identical due to SQLite page shuffling). First-visit users pay the 47 MB download; subsequent visits hit browser cache until the next deploy. -- **GitHub Pages file size limit**: individual files ≤ 100 MB. Main gzipped DB is ~47 MB — safe. Uncompressed 159 MB DB would not fit; that's why we ship `.gz` and the frontend decompresses via `DecompressionStream`. -- **No server-side compression assumption.** The app reads the `.gz` file directly (not `Content-Encoding: gzip`). GH Pages does not reliably gzip on-the-fly for arbitrary paths; shipping pre-gzipped bytes is deterministic. +- **The gzipped database is not cacheable across deploys.** Every rebuild + produces a different `.db.gz`, because SQLite does not lay pages out + deterministically. First-visit users pay the full download; later visits hit + browser cache until the next deploy. +- **GitHub Pages caps individual files at 100 MB.** The largest gzipped database + is about 48 MB. Uncompressed they run 135–234 MB and would not fit — which is + why the browser decompresses via `DecompressionStream`. +- **No server-side compression is assumed.** The app fetches the `.gz` bytes + directly rather than relying on `Content-Encoding: gzip`; Pages does not + reliably compress arbitrary paths on the fly. +- **Total artifact is about 177 MB**, well inside the 1 GB site limit. ## Rollback -Deploys are stateless snapshots. To rollback, revert the commit on `main` and push — the next workflow rebuilds the older state. There's no state to migrate. +Deploys are stateless snapshots. Revert the commit on `main` and push; the next +run rebuilds the older state. There is no data to migrate. + +## Troubleshooting + +| Symptom | Typical cause | +| --- | --- | +| Blank page, 404 on assets | `base` in `vite.config.js` does not match the repo name | +| `Failed to fetch database: 404` | Dataset id in `src/datasets.js` does not match the file in `db/` | +| A route 404s | `assemble-site.js` did not run, or the id is missing from `DATASETS` | +| WASM fails to load | `sql.js.org` unreachable — self-host `sql-wasm.wasm` and update `SQL_WASM_URL` in `use-sqlite.js` | +| Deploy fails on assembly | An uncompressed database artefact reached the output; the error names the files | +| Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints | diff --git a/docs/project-overview-pdr-2016-legacy.md b/docs/project-overview-pdr-2016-legacy.md deleted file mode 100644 index 9249e0a..0000000 --- a/docs/project-overview-pdr-2016-legacy.md +++ /dev/null @@ -1,34 +0,0 @@ -# Project Overview — thptqg2016 - -## Goal - -Provide a public lookup tool for Vietnam's 2016 National High School Graduation Exam scores (877,461 candidates), running entirely on the client, hosted for free on GitHub Pages. - -## Scope - -- Lookup by exam ID or full name (with Vietnamese diacritics handling) -- Read-only SQL queries against a single `student` table -- Static dataset — no updates (the 2016 exam is long over) - -## Target users - -- Former 2016 candidates checking their scores -- Education researchers / data journalists running aggregate stats -- Developers exploring SQL on a real-world dataset - -## Constraints - -- **Zero backend**: the full DB (tens of MB gzipped) is downloaded to the browser -- **Read-only**: INSERT/UPDATE/DELETE rejected to avoid the illusion that user edits persist -- **Row caps**: 100 rows (lookup), 1000 rows (SQL) to prevent browser hangs -- **Vietnamese-first UI**: app labels and data are Vietnamese; documentation is English - -## Data sources - -Excel files (`.xlsx`/`.xls`) collected from newspapers and exam clusters in 2016, stored in `data/`. One file per cluster, with several different column layouts (see `scripts/build-database.js`). - -Original aggregator link: — no longer accessible, so the raw Excel files are mirrored in this repo under `data/`. - -## Status - -Stable. Data is frozen. Recent work focuses on UX polish (dark mode, accessibility, diacritics-insensitive search). diff --git a/docs/project-overview.md b/docs/project-overview.md new file mode 100644 index 0000000..af8dfdd --- /dev/null +++ b/docs/project-overview.md @@ -0,0 +1,66 @@ +# Project Overview + +## Goal + +A public lookup tool for Vietnam's National High School Graduation Exam scores, +running entirely in the browser and hosted for free on GitHub Pages. Covers the +2016 and 2017 exams — 1.7 million candidates across four datasets. + +## Scope + +- Lookup by exam ID or full name, with Vietnamese diacritics handled +- Read-only SQL queries against a single `student` table +- Admission-block (khối thi) totals computed per candidate +- Static datasets — both exams are long over and the data is frozen + +## Target users + +- Former candidates checking their scores +- Education researchers and data journalists running aggregate statistics +- Developers exploring SQL against a real-world dataset + +## Constraints + +- **Zero backend.** The full database (38–48 MB gzipped per dataset) is + downloaded to the browser and queried in-process. +- **Read-only.** `INSERT`/`UPDATE`/`DELETE` are rejected, so nobody is misled + into thinking edits persist. `sql.js` is in-memory anyway. +- **Row caps.** 100 rows for lookups, 1000 for custom SQL, to prevent browser + hangs. +- **Vietnamese-first UI.** App labels and data are Vietnamese; documentation is + English. + +## Datasets + +| id | Exam | Candidates | Notes | +| --- | --- | --- | --- | +| `2016` | 2016 | 877,461 | 119 files, three column layouts | +| `2017` | 2017 | 861,068 | current generation, reproducible from source | +| `2017-old` | 2017 | 847,348 | pre-refresh archive | +| `2017-old2` | 2017 | 679,764 | corrected re-export | + +The three 2017 datasets are kept side by side because they disagree, and the +disagreement is itself informative. Only `2017` is re-fetchable; the rest exist +solely as the copies committed here. + +Original 2016 aggregator link +() +is no longer accessible, which is why the raw files are mirrored in `data/2016/`. + +## History + +Each year began as a standalone repository (`thptqg2016`, `thptqg2017`), merged +here with full history. They initially kept separate frontends and separate +copies of the same Rust parser, synchronised by hand. That duplication was +removed: there is now one frontend, one parser, and one canonical schema, with +per-dataset differences confined to four small config files and one registry +entry each. + +The unification also fixed a latent data-loss bug — neither year's parser +config listed the complete set of subjects, so 1,691 candidates were missing +their foreign-language score. See `data-pipeline.md`. + +## Status + +Stable, data frozen. Work is limited to UX polish and keeping the pipeline +maintainable. diff --git a/docs/system-architecture-2016-legacy.md b/docs/system-architecture-2016-legacy.md deleted file mode 100644 index 4c98b2d..0000000 --- a/docs/system-architecture-2016-legacy.md +++ /dev/null @@ -1,83 +0,0 @@ -# System Architecture - -## Overview - -A **static serverless** design: the entire dataset is packaged into a single SQLite file, gzip-compressed, and served as a static asset via GitHub Pages. The browser downloads it, decompresses it, and queries it in-process using `sql.js` (SQLite compiled to WebAssembly). - -``` -┌─────────────────┐ build ┌──────────────────────┐ -│ data/*.xlsx │ ───────────▶ │ scripts/ │ -│ (mixed formats)│ │ build-database.js │ -└─────────────────┘ │ (Node + xlsx + │ - │ better-sqlite3) │ - └──────────┬───────────┘ - │ - ▼ - ┌──────────────────────┐ - │ public/thptqg2016.db │ - └──────────┬───────────┘ - │ gzip -9 (CI) - ▼ - ┌──────────────────────┐ - │ dist/thptqg2016.db.gz│ - │ dist/assets/* │ ◀── Vite build - └──────────┬───────────┘ - │ upload-pages-artifact - ▼ - GitHub Pages CDN - │ - ▼ - ┌──────────────────────────────────┐ - │ Browser │ - │ ┌────────────────────────────┐ │ - │ │ useSqlite hook │ │ - │ │ fetch(.db.gz) + stream │ │ - │ │ DecompressionStream gzip │ │ - │ │ sql.js WASM (from CDN) │ │ - │ └─────────────┬──────────────┘ │ - │ ▼ │ - │ ┌────────────────────────────┐ │ - │ │ React UI │ │ - │ │ - SearchForm / ScoreTable │ │ - │ │ - CustomQuery (SQL editor)│ │ - │ └────────────────────────────┘ │ - └──────────────────────────────────┘ -``` - -## Build-time data flow - -1. Developer drops Excel files into `data/`. -2. `npm run build:db` reads every file; for each one: - - Detects the header row against a `KNOWN_HEADERS` set. - - Picks a format: `separate-scores` (one column per subject) vs. `mapped` (single `DIEM_THI` string). - - Parses each row into a canonical 18-column record. - - `INSERT OR REPLACE` into SQLite (primary key = `so_bao_danh` handles duplicates). -3. `VACUUM` shrinks the file. -4. CI runs `gzip -k -9` → `.db.gz`. - -## Runtime flow - -1. Page loads → React mounts → `useSqlite("thptqg2016.db.gz")`. -2. Streaming fetch with a progress bar (driven by `Content-Length`). -3. `DecompressionStream("gzip")` decompresses on the fly. -4. `sql.js` loads its WASM from `https://sql.js.org/dist/sql-wasm.wasm`. -5. `new SQL.Database(Uint8Array)` — DB is now in RAM. -6. Each search / query → `db.prepare()` + `stmt.step()` loop → render. - -## Design decisions - -| Concern | Choice | Rationale | -|---------|--------|-----------| -| Storage | Static SQLite file | No backend needed; dataset is frozen | -| Compression | gzip in CI, `DecompressionStream` in browser | Native browser API; no extra library | -| WASM hosting | `sql.js.org` CDN | Smaller self-hosted artifact | -| Diacritics search | Pre-computed `ho_ten_ascii` column | `LOWER(REPLACE(...))` at query time defeats the index | -| SQL safety | Leading-keyword allowlist | `sql.js` is in-memory so writes don't persist, but the allowlist prevents user confusion | -| Row caps | 100 (lookup), 1000 (SQL) | Keep DOM render sizes reasonable | - -## Risks and limitations - -- **DB size**: tens of MB gzipped — slow links have a visible wait; mitigated by the progress bar. -- **Browser memory**: the full DB lives in RAM; older mobile devices may OOM. -- **Dependency on `sql.js.org`**: if that CDN is unreachable, WASM fails to load. -- **Excel format drift**: a new source file with an unseen header layout needs a new branch in `detectFormat()`. diff --git a/docs/system-architecture.md b/docs/system-architecture.md index aa290cb..7eb5cce 100644 --- a/docs/system-architecture.md +++ b/docs/system-architecture.md @@ -1,97 +1,180 @@ # System Architecture -Static site. No backend. Browser downloads a compressed SQLite file at boot, then every query runs locally via `sql.js` (WASM). +Static site, no backend. The browser downloads a compressed SQLite file at boot +and every query runs locally via `sql.js` (SQLite compiled to WebAssembly). + +One frontend, one parser, one schema, four datasets. ## Data flow ``` -Excel files (.xls / .xlsx) +data//*.xls(x) │ - ▼ build-database*.js (Node) -SQLite DB (public*/thptqg2017.db) + ▼ parser/ (Rust, one binary, one config per dataset) + .build/public/db/.db │ - ▼ gzip -9 -thptqg2017.db.gz (~47 MB for main variant) + ▼ gzip -9 (no -k: the raw file does not survive) + .build/public/db/.db.gz │ - ▼ Vite build copies public*/ into dist/ -Static site on GitHub Pages + ▼ vite build (publicDir = .build/public) + dist/ │ - ▼ browser loads -sql.js (WASM) opens the .db → queries run client-side + ▼ scripts/assemble-site.js + _site/ → GitHub Pages + │ + ▼ browser + sql.js (WASM) opens the .db → queries run client-side ``` -## Three deployment variants +## The dataset id -One repo → three independent sites, same frontend, different dataset. +One identifier ties the whole pipeline together: -| Variant | Route | Source dir | Public dir | Build cmd | -|---|---|---|---|---| -| main | `/thptqg2017/` | `data/` | `public/` | `npm run build` | -| old | `/thptqg2017/old/` | `data-old/` | `public-old/` | `npm run build:old` | -| old2 | `/thptqg2017/old2/` | `data-old2/` | `public-old2/` | `npm run build:old2` | +``` +data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/ +``` -`vite.config.js` reads `process.env.VARIANT` and switches `base`, `publicDir`, `outDir` accordingly. `emptyOutDir` is on only for the main build so the variant builds merge into `dist/old/`, `dist/old2/` cleanly. +`src/datasets.js` declares the four ids once. The frontend, the database build +(`parser/scripts/build-db.js`) and the site assembly all import that list, so +adding a dataset means adding one entry and one config file. -## Schema (all 3 DBs share this) +| id | Exam | Rows | Source | +| --- | --- | --- | --- | +| `2016` | 2016 | 877,461 | Bộ GD&ĐT | +| `2017` | 2017 | 861,068 | baotintuc.vn | +| `2017-old` | 2017 | 847,348 | pre-refresh archive | +| `2017-old2` | 2017 | 679,764 | corrected re-export | + +## Canonical schema + +Defined once in `parser/src/schema.rs` — DDL, INSERT, column order and the 16 +subject regexes. The four TOML configs carry no SQL at all, only per-dataset +parse rules. Config parsing uses `deny_unknown_fields`, so a leftover `[schema]` +block fails loudly instead of looking effective while `schema.rs` drives the +build. ```sql CREATE TABLE student ( - so_bao_danh TEXT PRIMARY KEY, -- exam ID (8 digits, first 2 = province) + so_bao_danh TEXT PRIMARY KEY, -- 2017: 8 digits; 2016: 9 digits or a + -- 2-4 letter cluster code then digits ho_ten TEXT NOT NULL, - ho_ten_ascii TEXT NOT NULL, -- NFD-stripped lowercase for accent-insensitive search - ngay_sinh TEXT, -- dd/mm/yyyy + ho_ten_ascii TEXT NOT NULL, -- NFD-stripped lowercase, for accent-insensitive search + ngay_sinh TEXT, -- dd/mm/yyyy + ten_cum_thi TEXT, -- 2016 only + gioi_tinh TEXT, -- 2016 only toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn, lich_su, dia_ly, gdcd, khxh, - tieng_anh, tieng_phap, tieng_nga, tieng_trung REAL + tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung REAL ); -CREATE INDEX idx_ho_ten ON student(ho_ten); -CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii); +CREATE INDEX idx_ho_ten ON student(ho_ten); +CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii); +CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL; ``` -No German / Japanese language columns — neither appears in any source file. +Every dataset gets all 22 columns; ones it has no data for are NULL, costing +about a byte per row. `khtn`, `khxh` and `gdcd` are empty on 2016; +`ten_cum_thi` and `gioi_tinh` are empty on the 2017 datasets. -Each score column is `NULL` when the student didn't take that subject. A student's "admission block" total is only computed when all 3 required subjects are non-null. +`idx_ten_cum_thi` is partial, so it holds zero entries where the column is +always NULL. -## Parse quirks (why 3 builder scripts) +## Routing -Each dataset had a different quirk. Rather than one mega-parser with mode flags, each builder script is ~70 lines and isolates its own workarounds. +URLs are flat, one segment per dataset, and the segment is the id: -| Dataset | Quirk | Mitigation | -|---|---|---| -| data/ (baotintuc .xls) | Hanoi + HCM overflow past the 65,536 row-per-sheet .xls limit | Iterate all `wb.SheetNames` | -| data-old/ (xlsx) | Single-sheet always; one bogus header row (`SOBAODANH`/`HO_TEN`) leaked into an earlier DB | Strict numeric-SBD guard rejects the leak | -| data-old2/ (xlsx) | HCM overflow + many blank trailing rows | Multi-sheet walk + blank-row pre-filter | +``` +/thptqg/ hub +/thptqg/2016/ +/thptqg/2017/ +/thptqg/2017-old/ +/thptqg/2017-old2/ +``` -Shared concerns (regex, schema, `toAscii`, header detection) live in `scripts/build-lib.js`. +`src/router.js` is an exact match on that segment. The nested form used before +(`/thptqg/2017/old/`) would have needed longest-prefix matching, since it also +starts with `/thptqg/2017/`. Those two legacy URLs still resolve: the router +rewrites them to the flat equivalent with `history.replaceState`, preserving the +query string so `?q=` deep links survive. + +A single Vite build emits one `index.html`. Because `base` is absolute +(`/thptqg/`), that file references `/thptqg/assets/...` regardless of the +directory it is served from, so the assemble step copies it to every route and +each URL is a real static file. **No SPA 404-fallback redirect is used** — the +usual hack rewrites URLs and would interfere with the deep links. + +## Serving both exam years without branching + +No component contains a per-dataset conditional. Two mechanisms do the work: + +- **All-NULL columns are hidden.** `score-table.jsx` drops any column where + every row in the result set is NULL, so 2016 rows surface Cụm thi / GT / Đức / + Nhật and 2017 rows surface KHTN / KHXH / GDCD / Nga. +- **Incomplete admission blocks are skipped.** `computeBlocks()` only returns a + block when the student has all three subjects, so one block list covers both + years: GDCD blocks self-exclude on 2016, German and Japanese blocks + self-exclude wherever those languages were not sat. + +Anything genuinely per-dataset — title, source, database size, search examples, +SQL presets — lives in `src/datasets.js`. + +## Exam ID formats + +`src/lib/query-mode.js` decides whether a query is an exam ID or a name, and is +shared by `App.jsx` and `search-form.jsx` (they previously held separate copies +and had drifted apart on exactly this rule). + +| Form | Example | Where | +| --- | --- | --- | +| 8 digits | `49008235` | 2017 — first two digits are the province | +| 9 digits with leading zero | `017006021` | 2016 | +| 2-4 letters then digits | `BAL000001` | 2016 — exam cluster code | + +The letter-prefixed form is the majority case for 2016: 616,593 of 877,461 +candidates (70.3%). Letter prefixes are upper-cased before lookup, so +`bal000001` resolves. + +## Score tiers + +Six-level ladder in `scoreTier()` (`src/lib/admission-blocks.js`), paired with a +symbol so meaning is never colour-only. + +| Tier | Range | Vietnamese | +| --- | --- | --- | +| common | ≤ 1 | Điểm liệt | +| uncommon | < 5 | Chưa đạt | +| rare | 5–6.5 | Trung bình | +| epic | 6.5–8 | Khá | +| legendary | 8–9 | Giỏi | +| prismatic | 9–10 | Xuất sắc | ## Admission blocks -Vietnamese universities admit students based on 3-subject combinations called "admission blocks" (khối thi). `src/lib/admission-blocks.js` lists all 49 official 2017 blocks computable from our schema (A00–A11, B00–B08, C00–C20, D01–D15). Each entry is `{ code, subjects: [3 keys], label }`. +Vietnamese universities admit on three-subject combinations (khối thi). +`src/lib/admission-blocks.js` lists the blocks computable from this schema +(A00–A11, B00–B08, C00–C20, D01–D15, plus D05/D06 for German and Japanese). +`computeBlocks(student)` returns those where all three scores exist, sorted by +total descending. -`computeBlocks(student)` returns only blocks where the student has all 3 subject scores, sorted by total desc. Used by the student detail card. +## Design decisions -The SQL preset "Top 10 best-block in Long An" materialises all 49 blocks as a `UNION ALL` CTE, picks each student's best block via `ROW_NUMBER() OVER (PARTITION BY so_bao_danh ORDER BY s DESC, k)`, then ranks across students. +| Concern | Choice | Rationale | +| --- | --- | --- | +| Storage | Static SQLite file | No backend; the datasets are frozen | +| Compression | gzip in CI, `DecompressionStream` in the browser | Native API, no extra library | +| WASM hosting | `sql.js.org` CDN | Smaller self-hosted artifact | +| Diacritics search | Pre-computed `ho_ten_ascii` | `LOWER(REPLACE(...))` at query time defeats the index | +| SQL safety | Leading-keyword allowlist | `sql.js` is in-memory so writes cannot persist; the allowlist prevents confusion | +| Row caps | 100 (lookup), 1000 (SQL) | Keeps DOM render sizes reasonable | +| Routing | Hand-rolled, ~20 lines | Five static routes do not justify a router dependency | -## Score tiers (UI coding) +## Risks and limitations -6-level TFT rarity ladder (white → green → blue → purple → gold → prismatic). Defined in `scoreTier()` in `src/lib/admission-blocks.js`. - -| Tier | Range | UI label (Vietnamese) | English | -|---|---|---|---| -| common | ≤ 1 | Điểm liệt | Disqualifying | -| uncommon | < 5 | Chưa đạt | Below passing | -| rare | 5–6.5 | Trung bình | Average | -| epic | 6.5–8 | Khá | Good | -| legendary | 8–9 | Giỏi | Very good | -| prismatic | 9–10 | Xuất sắc | Excellent | - -Color + icon + text label — never color-only. - -## Frontend key behaviors - -- Query state lives in `App.jsx` and is synced to URL `?q=...` via `history.replaceState` — bookmarkable, shareable. -- `SearchForm` is a controlled component that 300ms-debounces input changes and auto-triggers search once the query is long enough (3+ digits for SBD, 2+ chars for name). -- Detection: all-digits → exact `so_bao_danh` match; ASCII-only → folded `ho_ten_ascii LIKE`; has Vietnamese diacritics → search both `ho_ten` and folded column. -- Exactly 1 result → `StudentDetail` card; 0 or >1 → `ScoreTable`. -- Global `/` key focuses the search input when not already typing. -- Share button on the detail card uses Web Share API when present, else clipboard with a pre-formatted summary + deep-link URL. +- **Database size.** 38–48 MB gzipped per dataset; slow links wait, mitigated by + a progress bar. +- **Browser memory.** The full database lives in RAM; older mobile devices may + run out. +- **`sql.js.org` dependency.** If that CDN is unreachable, the WASM fails to + load. Self-hosting `sql-wasm.wasm` and updating `SQL_WASM_URL` in + `use-sqlite.js` is the fix. +- **Excel format drift.** A new source file with an unseen header layout needs a + new branch in `format_detect_2016.rs` or a new config. diff --git a/parser/src/format_detect_2016.rs b/parser/src/format_detect_2016.rs index 8103a4e..fd62f8c 100644 --- a/parser/src/format_detect_2016.rs +++ b/parser/src/format_detect_2016.rs @@ -19,7 +19,10 @@ /// SBD(0) HO_TEN(1) NGAY_SINH(2) TEN_CUMTHI(3) GIOI_TINH(4) DIEM_THI(5) /// → JS: build-database.js:149–151 DEFAULT_MAP + processMappedRow /// -/// JS citations are line numbers in /config/workspace/tiennm99/thptqg2016/scripts/build-database.js. +/// The `build-database.js` citations throughout this file refer to the Node +/// script this parser replaced, in the original standalone thptqg2016 repo. +/// That file no longer exists here; the references are kept because they +/// explain why several of the rules below look arbitrary. use calamine::Data; use crate::transform::{parse_scores, to_ascii, CompiledPatterns, ParsedRow}; diff --git a/plans/260813-0956-unify-frontend-standard-schema/phase-01-standard-schema-and-unified-parser.md b/plans/260813-0956-unify-frontend-standard-schema/phase-01-standard-schema-and-unified-parser.md index dc984a5..761f9f5 100644 --- a/plans/260813-0956-unify-frontend-standard-schema/phase-01-standard-schema-and-unified-parser.md +++ b/plans/260813-0956-unify-frontend-standard-schema/phase-01-standard-schema-and-unified-parser.md @@ -1,7 +1,7 @@ --- phase: 1 title: "Standard schema and unified parser" -status: pending +status: completed priority: P1 dependencies: [] effort: "" diff --git a/plans/260813-0956-unify-frontend-standard-schema/phase-02-repo-restructure.md b/plans/260813-0956-unify-frontend-standard-schema/phase-02-repo-restructure.md index 03f8b16..1665226 100644 --- a/plans/260813-0956-unify-frontend-standard-schema/phase-02-repo-restructure.md +++ b/plans/260813-0956-unify-frontend-standard-schema/phase-02-repo-restructure.md @@ -1,7 +1,7 @@ --- phase: 2 title: "Repo restructure" -status: pending +status: completed priority: P1 dependencies: [1] effort: "" @@ -131,12 +131,12 @@ component. ## Success Criteria -- [ ] Target layout matches `plan.md` exactly -- [ ] `2016/` and `2017/` no longer exist -- [ ] Old hub page content captured before deletion -- [ ] Git reports renames, repository size unchanged -- [ ] `package-lock.json` committed; `npm ci` succeeds; zero pnpm files remain -- [ ] `cargo test` green from the new location +- [x] Target layout matches `plan.md` exactly +- [x] `2016/` and `2017/` no longer exist +- [x] Old hub page content captured before deletion +- [x] Git reports renames, repository size unchanged +- [x] `package-lock.json` committed; `npm ci` succeeds; zero pnpm files remain +- [x] `cargo test` green from the new location ## Risk Assessment diff --git a/plans/260813-0956-unify-frontend-standard-schema/phase-03-unified-frontend-and-dataset-registry.md b/plans/260813-0956-unify-frontend-standard-schema/phase-03-unified-frontend-and-dataset-registry.md index 1d6eea9..e36e839 100644 --- a/plans/260813-0956-unify-frontend-standard-schema/phase-03-unified-frontend-and-dataset-registry.md +++ b/plans/260813-0956-unify-frontend-standard-schema/phase-03-unified-frontend-and-dataset-registry.md @@ -1,7 +1,7 @@ --- phase: 3 title: "Unified frontend and dataset registry" -status: pending +status: completed priority: P1 dependencies: [2] effort: "" @@ -225,16 +225,16 @@ this repo today): ## Success Criteria -- [ ] One `src/` serving four datasets **and** the hub, zero `if (dataset === ...)` in components -- [ ] Subject list defined once in `src/lib/subjects.js` -- [ ] Hub route fetches no database -- [ ] Dataset path and DB URL are derived from `id`, not stored per entry -- [ ] Legacy `/2017/old/` and `/2017/old2/` redirect to flat URLs, `?q=` preserved -- [ ] 2016 route renders cluster + gender; 2017 routes do not -- [ ] 2016 route has live search, deep links, student detail, score tiers -- [ ] Admission blocks correct for both exam years -- [ ] No routing library added -- [ ] `npm run lint` green +- [x] One `src/` serving four datasets **and** the hub, zero `if (dataset === ...)` in components +- [x] Subject list defined once in `src/lib/subjects.js` +- [x] Hub route fetches no database +- [x] Dataset path and DB URL are derived from `id`, not stored per entry +- [x] Legacy `/2017/old/` and `/2017/old2/` redirect to flat URLs, `?q=` preserved +- [x] 2016 route renders cluster + gender; 2017 routes do not +- [x] 2016 route has live search, deep links, student detail, score tiers +- [x] Admission blocks correct for both exam years +- [x] No routing library added +- [x] `npm run lint` green ## Risk Assessment diff --git a/plans/260813-0956-unify-frontend-standard-schema/phase-04-build-and-deploy-pipeline.md b/plans/260813-0956-unify-frontend-standard-schema/phase-04-build-and-deploy-pipeline.md index c94c595..9c308d8 100644 --- a/plans/260813-0956-unify-frontend-standard-schema/phase-04-build-and-deploy-pipeline.md +++ b/plans/260813-0956-unify-frontend-standard-schema/phase-04-build-and-deploy-pipeline.md @@ -1,7 +1,7 @@ --- phase: 4 title: "Build and deploy pipeline" -status: pending +status: completed priority: P1 dependencies: [3] effort: "" diff --git a/plans/260813-0956-unify-frontend-standard-schema/phase-05-parity-verification-and-docs.md b/plans/260813-0956-unify-frontend-standard-schema/phase-05-parity-verification-and-docs.md index a075a51..848b3db 100644 --- a/plans/260813-0956-unify-frontend-standard-schema/phase-05-parity-verification-and-docs.md +++ b/plans/260813-0956-unify-frontend-standard-schema/phase-05-parity-verification-and-docs.md @@ -1,7 +1,7 @@ --- phase: 5 title: "Parity verification and docs" -status: pending +status: completed priority: P1 dependencies: [4] effort: "" @@ -110,13 +110,13 @@ so reruns compare the same students. ## Success Criteria -- [ ] `verify-parity.js` committed and exiting 0 on all four datasets -- [ ] Parity result report committed under `plans/reports/` -- [ ] Zero unexpected non-NULL columns -- [ ] Manual checklist passes on the hub and all four dataset routes -- [ ] `verify-parity.js` has zero npm dependencies (`node:sqlite` only) -- [ ] `README.md` and `docs/` describe the actual repo, with npm commands -- [ ] No stale path, pnpm, or build-variant references anywhere outside `plans/` +- [x] `verify-parity.js` committed and exiting 0 on all four datasets +- [x] Parity result report committed under `plans/reports/` +- [x] Zero unexpected non-NULL columns +- [x] Manual checklist passes on the hub and all four dataset routes +- [x] `verify-parity.js` has zero npm dependencies (`node:sqlite` only) +- [x] `README.md` and `docs/` describe the actual repo, with npm commands +- [x] No stale path, pnpm, or build-variant references anywhere outside `plans/` ## Risk Assessment diff --git a/plans/260813-0956-unify-frontend-standard-schema/plan.md b/plans/260813-0956-unify-frontend-standard-schema/plan.md index 98f6ac4..dbedb6b 100644 --- a/plans/260813-0956-unify-frontend-standard-schema/plan.md +++ b/plans/260813-0956-unify-frontend-standard-schema/plan.md @@ -1,7 +1,7 @@ --- title: "Unify frontend, standardize SQL schema, restructure repo" description: "One 2017-based frontend, one canonical 22-column student schema, one parser crate, four datasets under data/" -status: pending +status: completed priority: P2 branch: "main" tags: [refactor, schema, frontend, parser] @@ -134,11 +134,11 @@ CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT N | Phase | Name | Status | |-------|------|--------| -| 1 | [Standard schema and unified parser](./phase-01-standard-schema-and-unified-parser.md) | Pending | -| 2 | [Repo restructure](./phase-02-repo-restructure.md) | Pending | -| 3 | [Unified frontend and dataset registry](./phase-03-unified-frontend-and-dataset-registry.md) | Pending | -| 4 | [Build and deploy pipeline](./phase-04-build-and-deploy-pipeline.md) | Pending | -| 5 | [Parity verification and docs](./phase-05-parity-verification-and-docs.md) | Pending | +| 1 | [Standard schema and unified parser](./phase-01-standard-schema-and-unified-parser.md) | Completed | +| 2 | [Repo restructure](./phase-02-repo-restructure.md) | Completed | +| 3 | [Unified frontend and dataset registry](./phase-03-unified-frontend-and-dataset-registry.md) | Completed | +| 4 | [Build and deploy pipeline](./phase-04-build-and-deploy-pipeline.md) | Completed | +| 5 | [Parity verification and docs](./phase-05-parity-verification-and-docs.md) | Completed | ## Dependencies @@ -151,17 +151,17 @@ No cross-plan dependencies (`plans/` was empty before this plan). ## Acceptance Criteria -- [ ] One Rust crate; `2016/tools/` and `2017/tools/` gone -- [ ] One `src/`; `2016/src/` and `2017/src/` gone -- [ ] One Vite build producing all five pages -- [ ] All four DBs built from the same DDL, same INSERT, same 16 regexes -- [ ] Per-dataset row count and per-column non-NULL count identical to pre-refactor baseline -- [ ] All five published URLs functional under the flat scheme, deep links included -- [ ] Legacy `/2017/old/` and `/2017/old2/` redirect to their flat equivalents, preserving `?q=` -- [ ] 2016 site still shows `ten_cum_thi` + `gioi_tinh`; 2017 sites do not -- [ ] 2016 site gains 2017's features (deep links, student detail, tiers, live search) -- [ ] npm only: `package-lock.json` committed, no pnpm files or CI steps remain -- [ ] `cargo test` green; `npm run lint` green +- [x] One Rust crate; `2016/tools/` and `2017/tools/` gone +- [x] One `src/`; `2016/src/` and `2017/src/` gone +- [x] One Vite build producing all five pages +- [x] All four DBs built from the same DDL, same INSERT, same 16 regexes +- [x] Per-dataset row count and per-column non-NULL count identical to pre-refactor baseline +- [x] All five published URLs functional under the flat scheme, deep links included +- [x] Legacy `/2017/old/` and `/2017/old2/` redirect to their flat equivalents, preserving `?q=` +- [x] 2016 site still shows `ten_cum_thi` + `gioi_tinh`; 2017 sites do not +- [x] 2016 site gains 2017's features (deep links, student detail, tiers, live search) +- [x] npm only: `package-lock.json` committed, no pnpm files or CI steps remain +- [x] `cargo test` green; `npm run lint` green ## Rollback diff --git a/plans/reports/parser-parity-result.md b/plans/reports/parser-parity-result.md new file mode 100644 index 0000000..39d62a8 --- /dev/null +++ b/plans/reports/parser-parity-result.md @@ -0,0 +1,79 @@ +# Parser parity result + +Comparison of the databases the unified pipeline ships against a baseline built +from the pre-refactor code. Run on 2026-08-13 against the final tree +(`npm run build:rust && npm run build:db`), gate exit code 0. + +Inputs: + +- `plans/reports/parser-parity-baseline.json` — built with the two separate + crates, before any change +- `plans/reports/parser-parity-current.json` — decompressed from + `.build/public/db/*.db.gz`, i.e. the exact bytes published + +## Result + +| Dataset | Rows | Columns | Size | +| --- | --- | --- | --- | +| 2016 | 877,461 | 18 → 22 | +1.4% | +| 2017 | 861,068 | 18 → 22 | +2.2% | +| 2017-old | 847,348 | 18 → 22 | +2.2% | +| 2017-old2 | 679,764 | 18 → 22 | +2.2% | + +Unchanged, for every dataset: + +- row count +- non-NULL count for all 18 pre-existing columns +- every field of a deterministic student sample (SBDs ending `0000`) + +## Approved recoveries + +Unifying the subject regexes recovered 1,691 foreign-language scores that the +old per-year configs discarded. The 2016 config listed 12 subject patterns and +the 2017 configs 14; neither list was complete, so candidates who sat German, +Japanese or Russian ended up with no foreign-language score at all. + +| Dataset | Recovered | +| --- | --- | +| 2016 | 182 × `tieng_nga` | +| 2017 | 93 × `tieng_duc`, 512 × `tieng_nhat` | +| 2017-old | 85 × `tieng_duc`, 484 × `tieng_nhat` | +| 2017-old2 | 22 × `tieng_duc`, 313 × `tieng_nhat` | + +Confirmed real rather than spurious matches: + +- Every student in all four datasets holds zero or exactly one foreign + language — never two — so nothing is duplicated. +- Each affected student had all language columns NULL beforehand. SBD + `01003198` went from all-NULL to `tieng_duc = 8`. +- Counts track dataset size across the three 2017 generations (93/85/22 + German, 512/484/313 Japanese). +- Values are ordinary exam scores in the 0–10 range. + +These exact counts are pinned as `APPROVED_RECOVERY` in +`parser/scripts/verify-parity.js`. Any other newly-populated column, or drift +in these numbers, still fails the gate. + +## Columns that must stay empty + +Verified 0 non-NULL: + +- 2016: `khtn`, `khxh`, `gdcd` +- 2017, 2017-old, 2017-old2: `ten_cum_thi`, `gioi_tinh` + +## Reproducing + +```sh +npm run build:rust && npm run build:db +node parser/scripts/db-stats.js \ + 2016= 2017= 2017-old= 2017-old2= > current.json +node parser/scripts/verify-parity.js \ + plans/reports/parser-parity-baseline.json current.json +``` + +The baseline cannot be regenerated — it required the two pre-refactor crates, +which no longer exist. It is committed for that reason. + +## Unresolved questions + +None.