mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
docs: rewrite for the unified repo and record the parity gate
The docs still described two standalone projects with separate frontends,
separate parsers and three Vite variants. Rewritten around what the repo now
is, merging both projects' copies rather than keeping one and discarding the
other — deployment-guide.md and system-architecture.md existed in both and
documented different pipelines.
project-overview.md goal, scope, constraints, the four datasets, history
system-architecture.md data flow, canonical schema, routing, how one
frontend serves both exam years without branching
data-pipeline.md per-dataset Excel formats, the three 2016 layouts,
overflow-sheet gotcha, expected row counts
deployment-guide.md the single-build workflow, adding a dataset,
why no uncompressed database can ship
Records the release gate in plans/reports/parser-parity-result.md: row counts,
all 18 pre-existing per-column non-NULL counts and the deterministic student
samples are identical across all four datasets, against databases decompressed
from the exact bytes the pipeline publishes. The 1,691 recovered
foreign-language scores are documented with the evidence they are real.
Also drops a machine-specific absolute path from a comment in
format_detect_2016.rs. The build-database.js citations there are kept: that
file no longer exists in this repo, but the references explain why several
parsing rules look arbitrary.
This commit is contained in:
1 parent
a5904a5d78
commit
3fcfe4862d
18 files changed
+591
-448
No files matched your search
@@ -2,18 +2,67 @@
|
||||
|
||||
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
|
||||
school graduation exam. Client-side SQL (sql.js) over a SQLite database built
|
||||
from the ministry's raw `.xls` score files by the Rust `xlsxread` CLI.
|
||||
from the ministry's raw `.xls` score files by the Rust `xlsxread` parser.
|
||||
|
||||
Served as one site at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
|
||||
Live at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
|
||||
|
||||
| Directory | Year | Size | Site |
|
||||
| Dataset | Exam | Candidates | Site |
|
||||
| --- | --- | --- | --- |
|
||||
| [`2016/`](./2016/) | THPT QG 2016 — 877,461 candidates | ~50 MB raw data | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
|
||||
| [`2017/`](./2017/) | THPT QG 2017 — ~861,000 candidates | ~117 MB raw data | [/2017/](https://tiennm99.github.io/thptqg/2017/) (+ [old](https://tiennm99.github.io/thptqg/2017/old/), [old2](https://tiennm99.github.io/thptqg/2017/old2/) generations) |
|
||||
| `2016` | 2016 | 877,461 | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
|
||||
| `2017` | 2017 | 861,068 | [/2017/](https://tiennm99.github.io/thptqg/2017/) |
|
||||
| `2017-old` | 2017 | 847,348 | [/2017-old/](https://tiennm99.github.io/thptqg/2017-old/) |
|
||||
| `2017-old2` | 2017 | 679,764 | [/2017-old2/](https://tiennm99.github.io/thptqg/2017-old2/) |
|
||||
|
||||
Each year was a standalone repository (`thptqg2016`, `thptqg2017`), merged here
|
||||
with full history. Each keeps its own `tools/xlsxread` copy because the raw
|
||||
data formats differ per year.
|
||||
The three 2017 datasets are successive publications of the same exam and they
|
||||
disagree; all three are kept so the differences stay inspectable.
|
||||
|
||||
The `Deploy to GitHub Pages` workflow builds both years' databases and site
|
||||
bundles and publishes them behind a root index.
|
||||
## Layout
|
||||
|
||||
```
|
||||
index.html + src/ the frontend — one app serving all four datasets and the hub
|
||||
datasets.js the four dataset ids and their per-dataset content
|
||||
router.js pathname → dataset
|
||||
data/<id>/ raw Excel files, one directory per dataset
|
||||
parser/ the Rust parser
|
||||
src/schema.rs canonical 22-column table: DDL, INSERT, subject regexes
|
||||
configs/<id>.toml per-dataset parse rules only, no SQL
|
||||
scripts/ database build, crawler, parity verification
|
||||
scripts/ site assembly
|
||||
docs/ architecture, data pipeline, deployment
|
||||
```
|
||||
|
||||
The dataset id is one identifier end to end:
|
||||
|
||||
```
|
||||
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
|
||||
```
|
||||
|
||||
## Build
|
||||
|
||||
```bash
|
||||
npm ci
|
||||
npm run build:rust # compile the parser
|
||||
npm run build:db # build + gzip all four databases (add an id for just one)
|
||||
npm run build:site # one Vite build, then assemble into _site/
|
||||
npx serve _site
|
||||
```
|
||||
|
||||
Pushing to `main` runs the same steps in
|
||||
`.github/workflows/deploy-pages.yml` and publishes to GitHub Pages.
|
||||
|
||||
## Adding a dataset
|
||||
|
||||
1. Put the Excel files in `data/<id>/`
|
||||
2. Add `parser/configs/<id>.toml` — sheet mode, column indices, validation
|
||||
guards. No SQL; the schema is canonical.
|
||||
3. Add an entry to `DATASETS` in `src/datasets.js`
|
||||
|
||||
Everything else follows: the build script, the site assembly and the router all
|
||||
read that one list, and the UI adapts to whichever columns the dataset fills.
|
||||
|
||||
## Docs
|
||||
|
||||
See [`docs/`](./docs/) — [overview](./docs/project-overview.md),
|
||||
[architecture](./docs/system-architecture.md),
|
||||
[data pipeline](./docs/data-pipeline.md),
|
||||
[deployment](./docs/deployment-guide.md).
|
||||
+4
-3
@@ -1,5 +1,6 @@
|
||||
# Docs
|
||||
|
||||
- [`system-architecture.md`](./system-architecture.md) — data flow, schema, 3-variant deploy, score-tier model, admission-block catalog
|
||||
- [`data-pipeline.md`](./data-pipeline.md) — Excel parse quirks, overflow-sheet gotcha, audit scripts, refresh flow
|
||||
- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, adding a new variant, rollback
|
||||
- [`project-overview.md`](./project-overview.md) — goal, scope, constraints, the four datasets, history
|
||||
- [`system-architecture.md`](./system-architecture.md) — data flow, canonical schema, routing, how one frontend serves both exam years
|
||||
- [`data-pipeline.md`](./data-pipeline.md) — Excel parse quirks, per-dataset formats, overflow-sheet gotcha, expected row counts, verifying a rebuild
|
||||
- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, adding a dataset, rollback, troubleshooting
|
||||
@@ -1,63 +0,0 @@
|
||||
# Codebase Summary
|
||||
|
||||
## Directory layout
|
||||
|
||||
```
|
||||
thptqg2016/
|
||||
├── data/ # Source Excel files (~100, mixed formats)
|
||||
├── scripts/
|
||||
│ └── build-database.js # Parse Excel → SQLite (build-time, Node + better-sqlite3)
|
||||
├── public/
|
||||
│ └── thptqg2016.db # Generated DB, gzipped during CI
|
||||
├── src/
|
||||
│ ├── main.jsx # React entry
|
||||
│ ├── App.jsx # Root: tabs, lookup logic, useSqlite wiring
|
||||
│ ├── App.css / index.css # Design tokens, dark mode, a11y styles
|
||||
│ ├── hooks/
|
||||
│ │ └── use-sqlite.js # Fetch .db.gz + decompress + init sql.js
|
||||
│ └── components/
|
||||
│ ├── search-form.jsx # Input for exam ID / full name
|
||||
│ ├── score-table.jsx # Result table for lookups
|
||||
│ └── custom-query.jsx # SQL editor + presets + result grid
|
||||
├── .github/workflows/deploy.yml # CI: build db → gzip → vite build → Pages
|
||||
├── vite.config.js # base: "/thptqg2016/"
|
||||
└── eslint.config.js
|
||||
```
|
||||
|
||||
## Key modules
|
||||
|
||||
### `scripts/build-database.js`
|
||||
Build-time only. Reads every `.xlsx/.xls` in `data/`, detects the header format (three variants), parses the `DIEM_THI` string via regex or separate score columns, normalizes gender, derives a diacritics-stripped `ho_ten_ascii` column for accent-insensitive search, and inserts into SQLite with three indexes (`ho_ten`, `ho_ten_ascii`, `ten_cum_thi`).
|
||||
|
||||
### `src/hooks/use-sqlite.js`
|
||||
Streams `.db.gz` with download progress, decompresses via `DecompressionStream("gzip")`, loads `sql.js` (WASM served from the `sql.js.org` CDN), and returns `{ db, loading, error, progress }`.
|
||||
|
||||
### `src/App.jsx`
|
||||
Two tabs: **Lookup** and **Custom SQL**. Lookup auto-detects exam IDs (regex `^[A-Z]{2,4}\d+$`) vs names and picks one of three query paths: exact exam ID / ASCII LIKE / original + ASCII LIKE. Capped at 100 rows.
|
||||
|
||||
### `src/components/custom-query.jsx`
|
||||
Whitelists leading keywords (`SELECT`, `PRAGMA`, `EXPLAIN`, `WITH`), auto-appends `LIMIT 1000` when missing, measures `performance.now()` execution time, and ships 7 preset analytics queries.
|
||||
|
||||
## `student` table schema
|
||||
|
||||
```sql
|
||||
so_bao_danh TEXT PRIMARY KEY -- exam ID
|
||||
ho_ten TEXT NOT NULL -- full name
|
||||
ho_ten_ascii TEXT NOT NULL -- diacritics stripped, lowercased
|
||||
ngay_sinh TEXT -- date of birth
|
||||
ten_cum_thi TEXT -- exam cluster name
|
||||
gioi_tinh TEXT -- "Nam" | "Nữ" | NULL
|
||||
toan, ngu_van, vat_ly, hoa_hoc, -- REAL (nullable) subject scores
|
||||
sinh_hoc, lich_su, dia_ly,
|
||||
tieng_anh, tieng_phap, tieng_duc,
|
||||
tieng_nhat, tieng_trung
|
||||
```
|
||||
|
||||
Indexes: `idx_ho_ten`, `idx_ho_ten_ascii`, `idx_ten_cum_thi`.
|
||||
|
||||
## Conventions
|
||||
|
||||
- JS/JSX filenames: **kebab-case** (e.g., `search-form.jsx`, `use-sqlite.js`)
|
||||
- React components: `PascalCase` named exports
|
||||
- UI strings: Vietnamese (target audience)
|
||||
- Code comments: English; explain *why*, not *what*
|
||||
+109
-70
@@ -2,109 +2,148 @@
|
||||
|
||||
From raw Excel files to a compressed SQLite file the browser can load.
|
||||
|
||||
## Canonical source (data/)
|
||||
One Rust binary (`parser/`) builds every dataset. What differs per dataset is
|
||||
parse rules only — sheet strategy, column layout, validation guards — declared
|
||||
in `parser/configs/<id>.toml`. The table shape, the INSERT and the subject
|
||||
regexes are canonical and live in `parser/src/schema.rs`.
|
||||
|
||||
63 Excel files from baotintuc.vn CDN. Source article:
|
||||
`https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm`
|
||||
## Sources
|
||||
|
||||
Re-download anytime:
|
||||
| id | Files | Reproducible | Origin |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 4 `.xls` + 115 `.xlsx` | no | Bộ GD&ĐT, collected 2016 |
|
||||
| `2017` | 63 `.xls` | **yes** | baotintuc.vn CDN |
|
||||
| `2017-old` | 63 `.xlsx` | no | pre-refresh archive |
|
||||
| `2017-old2` | 54 `.xlsx` | no | corrected re-export |
|
||||
|
||||
Only `2017` can be re-fetched:
|
||||
|
||||
```bash
|
||||
node scripts/crawl-baotintuc.js
|
||||
node parser/scripts/crawl-baotintuc.js
|
||||
```
|
||||
|
||||
Idempotent — skips files already present. Saves to `data/<ascii-kebab-province>.xls`.
|
||||
Idempotent — skips files already present, saves to `data/2017/`. Source article:
|
||||
`https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm`
|
||||
|
||||
## Reference sources (data-old/, data-old2/)
|
||||
The other three were collected from Vietnamese news sites at the time and the
|
||||
publisher URLs were not recorded. **The files committed in git are the only
|
||||
copy.**
|
||||
|
||||
Both were collected from Vietnamese news sites around the time of the 2017 exam.
|
||||
Specific publisher URLs were not recorded and cannot be recovered, so these
|
||||
datasets are **not reproducible** — the files committed in git are the only
|
||||
copy. They are kept for historical comparison and to expose the deploy matrix;
|
||||
the main deployment always uses the reproducible `data/` source.
|
||||
## Source Excel shapes
|
||||
|
||||
## Source Excel shape
|
||||
|
||||
Every file is a single sheet (for `data-old/`) or up to 3 sheets (for `data/` .xls) with columns:
|
||||
The 2017 datasets share one layout:
|
||||
|
||||
| Col | Name | Content |
|
||||
|---|---|---|
|
||||
| 0 | HO_TEN | full name in Vietnamese |
|
||||
| 1 | NGAY_SINH | `dd/mm/yyyy` |
|
||||
| 2 | SOBAODANH | 8-digit string, first 2 digits = province |
|
||||
| 3 | DIEM_THI | concatenated per-subject score string in the source language, e.g. `"Toan: 6.80 Van: 5.25 ... English: 5.80"` |
|
||||
| --- | --- | --- |
|
||||
| 0 | HO_TEN | full name in Vietnamese |
|
||||
| 1 | NGAY_SINH | `dd/mm/yyyy` |
|
||||
| 2 | SOBAODANH | 8-digit string, first 2 digits = province |
|
||||
| 3 | DIEM_THI | concatenated per-subject scores, e.g. `"Toán: 6.80 Ngữ văn: 5.25 …"` |
|
||||
|
||||
Some files lack a header row. `isHeaderRow()` in `build-lib.js` detects and skips.
|
||||
2016 has **three** layouts across its 119 files, chosen per file at runtime via
|
||||
`format_detection = "thptqg2016"` in its config:
|
||||
|
||||
| Format | Detected by | Notes |
|
||||
| --- | --- | --- |
|
||||
| `separate-scores` | `row[0] == "SBD"` and `row[2] == "TOAN"` | one column per subject; scores read directly, no regex |
|
||||
| `mapped` | header contains `SOBAODANH`/`SBD` plus `DIEM_THI` | column indices derived from the header |
|
||||
| `default` | no recognised header | positional 6-column layout |
|
||||
|
||||
The `mapped` and `default` layouts also carry `TEN_CUMTHI` and `GIOI_TINH`,
|
||||
which is why only 2016 populates those columns.
|
||||
|
||||
## Score text parsing
|
||||
|
||||
`SCORE_PATTERNS` in `build-lib.js` defines one regex per subject. The patterns must match the exact subject labels the source files use (Vietnamese).
|
||||
`SCORE_PATTERNS` in `parser/src/schema.rs` defines one regex per subject, and
|
||||
**all 16 run against every dataset**. A subject a given exam year did not offer
|
||||
simply never matches and stays NULL.
|
||||
|
||||
Subjects supported (DB column / label in source text):
|
||||
| Column | Source label | English | 2016 | 2017 |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| toan | Toán | Math | ✓ | ✓ |
|
||||
| ngu_van | Ngữ văn | Literature | ✓ | ✓ |
|
||||
| vat_ly | Vật lí | Physics | ✓ | ✓ |
|
||||
| hoa_hoc | Hóa học | Chemistry | ✓ | ✓ |
|
||||
| sinh_hoc | Sinh học | Biology | ✓ | ✓ |
|
||||
| khtn | KHTN | Natural Sciences combined | — | ✓ |
|
||||
| lich_su | Lịch sử | History | ✓ | ✓ |
|
||||
| dia_ly | Địa lí | Geography | ✓ | ✓ |
|
||||
| gdcd | GDCD | Civic Education | — | ✓ |
|
||||
| khxh | KHXH | Social Sciences combined | — | ✓ |
|
||||
| tieng_anh | Tiếng Anh | English | ✓ | ✓ |
|
||||
| tieng_phap | Tiếng Pháp | French | ✓ | ✓ |
|
||||
| tieng_nga | Tiếng Nga | Russian | ✓ | ✓ |
|
||||
| tieng_duc | Tiếng Đức | German | ✓ | ✓ |
|
||||
| tieng_nhat | Tiếng Nhật | Japanese | ✓ | ✓ |
|
||||
| tieng_trung | Tiếng Trung | Chinese | ✓ | ✓ |
|
||||
|
||||
| Column | Source label | English |
|
||||
|---|---|---|
|
||||
| toan | Toán | Math |
|
||||
| ngu_van | Ngữ văn | Literature |
|
||||
| vat_ly | Vật lí | Physics |
|
||||
| hoa_hoc | Hóa học | Chemistry |
|
||||
| sinh_hoc | Sinh học | Biology |
|
||||
| khtn | KHTN | Natural Sciences combined |
|
||||
| lich_su | Lịch sử | History |
|
||||
| dia_ly | Địa lí | Geography |
|
||||
| gdcd | GDCD | Civic Education |
|
||||
| khxh | KHXH | Social Sciences combined |
|
||||
| tieng_anh | Tiếng Anh | English |
|
||||
| tieng_phap | Tiếng Pháp | French |
|
||||
| tieng_nga | Tiếng Nga | Russian |
|
||||
| tieng_trung | Tiếng Trung | Chinese |
|
||||
### The foreign-language recovery
|
||||
|
||||
Missing-in-string → `NULL` in DB (student didn't take that subject).
|
||||
Before the parser was unified, the 2016 config listed 12 subject patterns and
|
||||
the 2017 configs listed 14. Neither list was complete: candidates could sit
|
||||
German, Japanese and Russian in both years, so 1,691 students ended up with **no
|
||||
foreign-language score at all**.
|
||||
|
||||
## Why three parsers
|
||||
Running all 16 patterns everywhere recovered them — 182 Russian in 2016, and
|
||||
German and Japanese across all three 2017 generations. See
|
||||
`plans/reports/parser-parity-result.md` for the evidence that these are real
|
||||
scores rather than false matches.
|
||||
|
||||
One shared lib + three thin parsers, not one parser with flags. Each dataset's quirks stay visible in its own file:
|
||||
## Per-dataset quirks
|
||||
|
||||
| File | Source dir | Sheet strategy | Extra guard |
|
||||
|---|---|---|---|
|
||||
| `build-database.js` | `data/` | iterate ALL sheets (.xls row overflow) | — |
|
||||
| `build-database-old.js` | `data-old/` | sheet 0 only | reject rows where `so_bao_danh` isn't digits-only (rejects an `HO_TEN` → `SOBAODANH` header leak present in this export) |
|
||||
| `build-database-old2.js` | `data-old2/` | iterate ALL sheets (HCM overflow) | skip fully blank rows before counting |
|
||||
| id | Sheet strategy | Guards |
|
||||
| --- | --- | --- |
|
||||
| `2016` | all sheets | per-file format detection; header-token rows rejected |
|
||||
| `2017` | all sheets — Hà Nội and HCM overflow | none |
|
||||
| `2017-old` | first sheet only | reject non-numeric SBD (rejects a header leak in this export) |
|
||||
| `2017-old2` | all sheets — HCM overflows | reject non-numeric SBD; skip blank rows before counting |
|
||||
|
||||
## Overflow-sheet gotcha
|
||||
### Overflow-sheet gotcha
|
||||
|
||||
The old `.xls` binary format caps at 65,536 rows per sheet. Hanoi and Ho Chi Minh City exceed that in the baotintuc export, so Excel auto-splits the overflow into `Sheet2`. **Only reading `Sheet1` silently drops 13,720 students** (Hanoi +7,275, HCM +6,445).
|
||||
The old `.xls` format caps at 65,536 rows per sheet. Hanoi and Ho Chi Minh City
|
||||
exceed that, so Excel splits the overflow into `Sheet2`. **Reading only `Sheet1`
|
||||
silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what
|
||||
`sheet_mode = "all"` exists for.
|
||||
|
||||
The audit `node scripts/audit-row-counts.js` catches this by comparing source-row-count (across all sheets) to DB-row-count.
|
||||
## Expected row counts
|
||||
|
||||
## Audits
|
||||
| id | Source rows | Skipped | DB rows |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 877,464 | 3 duplicate SBDs collapsed | **877,461** |
|
||||
| `2017` | 861,131 | 63 empty | **861,068** |
|
||||
| `2017-old` | 847,349 | 1 header leak | **847,348** |
|
||||
| `2017-old2` | 679,764 | 0 | **679,764** |
|
||||
|
||||
Run any time after a rebuild:
|
||||
## Verifying a rebuild
|
||||
|
||||
`parser/scripts/db-stats.js` dumps row counts, per-column non-NULL counts, file
|
||||
size and a deterministic student sample as JSON.
|
||||
`parser/scripts/verify-parity.js` diffs two such files and exits non-zero on any
|
||||
mismatch.
|
||||
|
||||
```bash
|
||||
node scripts/audit-row-counts.js # source rows vs DB rows (zero-loss check)
|
||||
node scripts/check-duplicates.js # md5 + row-content duplicates in data/
|
||||
node scripts/diff-datasets.js # compare two DBs (needs backup/ dir)
|
||||
node parser/scripts/db-stats.js 2016=<path>.db … > current.json
|
||||
node parser/scripts/verify-parity.js plans/reports/parser-parity-baseline.json current.json
|
||||
```
|
||||
|
||||
Expected baseline:
|
||||
The committed baseline was built with the pre-refactor code and cannot be
|
||||
regenerated — the two old crates no longer exist. Both scripts use the built-in
|
||||
`node:sqlite`, so they need no dependencies.
|
||||
|
||||
| DB | Source rows | Skipped | DB rows |
|
||||
|---|---|---|---|
|
||||
| main | 861,131 | 63 empty rows | **861,068** |
|
||||
| old | 847,349 | 1 header-leak | **847,348** |
|
||||
| old2 | 679,764 | 0 | **679,764** |
|
||||
|
||||
## Refreshing the data
|
||||
## Refreshing the 2017 data
|
||||
|
||||
```bash
|
||||
rm data/*.xls # clear current snapshot
|
||||
node scripts/crawl-baotintuc.js # re-download from CDN
|
||||
npm run build:db # rebuild
|
||||
node scripts/audit-row-counts.js # verify
|
||||
gzip -kf -9 public/thptqg2017.db # compress for shipping
|
||||
rm data/2017/*.xls
|
||||
node parser/scripts/crawl-baotintuc.js
|
||||
node parser/scripts/build-db.js 2017
|
||||
```
|
||||
|
||||
## Adding a new source dataset
|
||||
Then re-run the parity check above and confirm the row count still matches.
|
||||
|
||||
See `docs/deployment-guide.md` → "Adding a new variant".
|
||||
## Legacy scripts
|
||||
|
||||
`parser/scripts/check-duplicates.js` and `diff-datasets.js` are one-off audits
|
||||
that were already broken before the repo was unified — a hardcoded Windows path
|
||||
in one, an undeclared `better-sqlite3` dependency and stale paths in the other.
|
||||
Each carries a comment saying so. For comparing two builds, use `db-stats.js`
|
||||
plus `verify-parity.js` instead.
|
||||
@@ -1,50 +0,0 @@
|
||||
# Deployment Guide
|
||||
|
||||
## Automatic (recommended)
|
||||
|
||||
Push to `main` → GitHub Actions builds and deploys to GitHub Pages automatically.
|
||||
|
||||
Workflow: `.github/workflows/deploy.yml`
|
||||
|
||||
CI steps:
|
||||
1. `npm ci`
|
||||
2. `npm run build:db` — generate `public/thptqg2016.db` from `data/*.xlsx`
|
||||
3. `gzip -k -9 public/thptqg2016.db` — max compression, keep original
|
||||
4. `npm run build` — Vite bundles `dist/`
|
||||
5. `rm -f dist/thptqg2016.db` — ship only the gzipped copy
|
||||
6. `actions/upload-pages-artifact@v3` + `actions/deploy-pages@v4`
|
||||
|
||||
One-time setup: **Settings → Pages → Source: GitHub Actions**.
|
||||
|
||||
## Manual (local verification)
|
||||
|
||||
```bash
|
||||
npm install
|
||||
npm run build:db
|
||||
gzip -k -9 public/thptqg2016.db # Linux/macOS; on Windows use 7zip or WSL
|
||||
npm run build
|
||||
rm dist/thptqg2016.db # optional, shrinks artifact
|
||||
npm run preview # serve dist/ locally
|
||||
```
|
||||
|
||||
Open <http://localhost:4173/thptqg2016/>.
|
||||
|
||||
## Base path
|
||||
|
||||
`vite.config.js` sets `base: "/thptqg2016/"`. If you fork under a different repo name, update this to match `<repo-name>` so assets resolve correctly on GitHub Pages.
|
||||
|
||||
## Updating data
|
||||
|
||||
1. Add the new Excel file to `data/`.
|
||||
2. If its header is unfamiliar, open `scripts/build-database.js` and extend `detectFormat()` or the `processSeparateScoresRow` / `processMappedRow` helpers.
|
||||
3. Run `npm run build:db` locally to check row counts and error skips.
|
||||
4. Commit + push → CI redeploys.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Typical cause |
|
||||
|---------|---------------|
|
||||
| Blank page, 404 on assets | `base` in `vite.config.js` doesn't match the repo name |
|
||||
| `Failed to fetch database: 404` | gzip step skipped, or `.db.gz` removed from `dist/` |
|
||||
| WASM fails to load | `sql.js.org` blocked / offline — self-host `sql-wasm.wasm` in `public/` and update `SQL_WASM_URL` in `use-sqlite.js` |
|
||||
| Missing rows after build | Excel file has an unknown header — check console for `Failed to read` or `errorCount` |
|
||||
+82
-29
@@ -1,52 +1,105 @@
|
||||
# Deployment Guide
|
||||
|
||||
Deploys to GitHub Pages via `.github/workflows/deploy.yml`. Every push to `main` rebuilds and redeploys all 3 variants.
|
||||
Deploys to GitHub Pages via `.github/workflows/deploy-pages.yml`. Every push to
|
||||
`main` rebuilds all four datasets and redeploys the whole site.
|
||||
|
||||
One-time setup: **Settings → Pages → Source: GitHub Actions**.
|
||||
|
||||
## What the workflow does
|
||||
|
||||
1. Checkout, setup Node 20, `npm ci`
|
||||
2. `npm run build:db:all` — builds 3 SQLite DBs (main, old, old2)
|
||||
3. Gzip each DB at level 9
|
||||
4. `npm run build:all` — builds 3 Vite bundles into `dist/`, `dist/old/`, `dist/old2/`
|
||||
5. Remove uncompressed `.db` files from `dist/` (only `.gz` ships)
|
||||
6. `actions/upload-pages-artifact` + `actions/deploy-pages`
|
||||
1. Checkout, Rust toolchain, Node 24, `npm ci`
|
||||
2. `npm run build:rust` — one parser binary
|
||||
3. `npm run build:db` — builds and gzips all four databases into
|
||||
`.build/public/db/`
|
||||
4. `npm run build:site` — one Vite build, then `scripts/assemble-site.js`
|
||||
5. `actions/upload-pages-artifact` + `actions/deploy-pages`
|
||||
|
||||
Total CI time ≈ 4–6 min (DB build dominates).
|
||||
The database build dominates the runtime: roughly 419 MB of Excel is parsed on
|
||||
every deploy.
|
||||
|
||||
## Resulting URLs
|
||||
|
||||
- `https://<user>.github.io/thptqg2017/`
|
||||
- `https://<user>.github.io/thptqg2017/old/`
|
||||
- `https://<user>.github.io/thptqg2017/old2/`
|
||||
```
|
||||
https://<user>.github.io/thptqg/
|
||||
https://<user>.github.io/thptqg/2016/
|
||||
https://<user>.github.io/thptqg/2017/
|
||||
https://<user>.github.io/thptqg/2017-old/
|
||||
https://<user>.github.io/thptqg/2017-old2/
|
||||
```
|
||||
|
||||
`/thptqg/2017/old/` and `/thptqg/2017/old2/` were the pre-flattening URLs. They
|
||||
are still served, and the router rewrites them to the flat form with the query
|
||||
string intact.
|
||||
|
||||
## Local reproduction
|
||||
|
||||
```bash
|
||||
npm ci
|
||||
npm run build:db:all
|
||||
gzip -kf -9 public/thptqg2017.db public-old/thptqg2017.db public-old2/thptqg2017.db
|
||||
npm run build:all
|
||||
npx serve dist # or any static server
|
||||
npm run build:rust
|
||||
npm run build:db # all four; pass an id to build just one
|
||||
npm run build:site # vite build + assemble into _site/
|
||||
npx serve _site
|
||||
```
|
||||
|
||||
## Adding a new variant
|
||||
To rebuild a single dataset:
|
||||
|
||||
1. Drop source Excel files into a new `data-vN/` folder
|
||||
2. Add `public-vN/` to `.gitignore` patterns for the `.db` + `.db.gz`
|
||||
3. Copy `scripts/build-database-old.js` → `scripts/build-database-vN.js`; update `SRC_DIR` / `DB_PATH`; adjust parse logic for the new quirks (multi-sheet? blank rows? numeric guard?)
|
||||
4. Add variant to `VARIANT_CONFIG` in `vite.config.js`:
|
||||
```js
|
||||
vN: { base: "/thptqg2017/vN/", publicDir: "public-vN", outDir: "dist/vN" }
|
||||
```
|
||||
5. Add `build:db:vN` and `build:vN` npm scripts; wire them into `build:db:all` and `build:all`
|
||||
6. Update `.github/workflows/deploy.yml` — add gzip step + rm step for the new DB
|
||||
```bash
|
||||
node parser/scripts/build-db.js 2017-old
|
||||
```
|
||||
|
||||
## Base path
|
||||
|
||||
`vite.config.js` sets `base: "/thptqg/"`. If you fork under a different repo
|
||||
name, update it to match — assets are referenced absolutely, so a mismatch shows
|
||||
up as a blank page with 404s on `/assets/...`.
|
||||
|
||||
## Adding a dataset
|
||||
|
||||
1. Put the Excel files in `data/<id>/`
|
||||
2. Add `parser/configs/<id>.toml` with the parse rules — sheet mode, column
|
||||
indices, SBD validation, header tokens, blank-row stripping. No SQL: the
|
||||
schema is canonical and lives in `parser/src/schema.rs`
|
||||
3. Add an entry to `DATASETS` in `src/datasets.js`
|
||||
|
||||
Nothing else. The build script, the site assembly and the router all read that
|
||||
one list, and the frontend adapts to whichever columns the dataset populates.
|
||||
|
||||
## Why no uncompressed database can ship
|
||||
|
||||
`build-db.js` runs `gzip -9` **without** `-k`, so the raw file does not survive
|
||||
the build. `assemble-site.js` then fails the job if any `.db`, `.db-journal`,
|
||||
`.db-wal` or `.db-shm` reached the output.
|
||||
|
||||
Both guards exist because the previous pipeline wrote a 100+ MB uncompressed
|
||||
database into the source tree and relied on an `rm` step to keep it out of the
|
||||
artifact — one missing line away from publishing it.
|
||||
|
||||
## Notes
|
||||
|
||||
- **DB gzip is non-cacheable across deploys** — every rebuild produces a new `.db.gz` (not byte-identical due to SQLite page shuffling). First-visit users pay the 47 MB download; subsequent visits hit browser cache until the next deploy.
|
||||
- **GitHub Pages file size limit**: individual files ≤ 100 MB. Main gzipped DB is ~47 MB — safe. Uncompressed 159 MB DB would not fit; that's why we ship `.gz` and the frontend decompresses via `DecompressionStream`.
|
||||
- **No server-side compression assumption.** The app reads the `.gz` file directly (not `Content-Encoding: gzip`). GH Pages does not reliably gzip on-the-fly for arbitrary paths; shipping pre-gzipped bytes is deterministic.
|
||||
- **The gzipped database is not cacheable across deploys.** Every rebuild
|
||||
produces a different `.db.gz`, because SQLite does not lay pages out
|
||||
deterministically. First-visit users pay the full download; later visits hit
|
||||
browser cache until the next deploy.
|
||||
- **GitHub Pages caps individual files at 100 MB.** The largest gzipped database
|
||||
is about 48 MB. Uncompressed they run 135–234 MB and would not fit — which is
|
||||
why the browser decompresses via `DecompressionStream`.
|
||||
- **No server-side compression is assumed.** The app fetches the `.gz` bytes
|
||||
directly rather than relying on `Content-Encoding: gzip`; Pages does not
|
||||
reliably compress arbitrary paths on the fly.
|
||||
- **Total artifact is about 177 MB**, well inside the 1 GB site limit.
|
||||
|
||||
## Rollback
|
||||
|
||||
Deploys are stateless snapshots. To rollback, revert the commit on `main` and push — the next workflow rebuilds the older state. There's no state to migrate.
|
||||
Deploys are stateless snapshots. Revert the commit on `main` and push; the next
|
||||
run rebuilds the older state. There is no data to migrate.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Typical cause |
|
||||
| --- | --- |
|
||||
| Blank page, 404 on assets | `base` in `vite.config.js` does not match the repo name |
|
||||
| `Failed to fetch database: 404` | Dataset id in `src/datasets.js` does not match the file in `db/` |
|
||||
| A route 404s | `assemble-site.js` did not run, or the id is missing from `DATASETS` |
|
||||
| WASM fails to load | `sql.js.org` unreachable — self-host `sql-wasm.wasm` and update `SQL_WASM_URL` in `use-sqlite.js` |
|
||||
| Deploy fails on assembly | An uncompressed database artefact reached the output; the error names the files |
|
||||
| Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints |
|
||||
@@ -1,34 +0,0 @@
|
||||
# Project Overview — thptqg2016
|
||||
|
||||
## Goal
|
||||
|
||||
Provide a public lookup tool for Vietnam's 2016 National High School Graduation Exam scores (877,461 candidates), running entirely on the client, hosted for free on GitHub Pages.
|
||||
|
||||
## Scope
|
||||
|
||||
- Lookup by exam ID or full name (with Vietnamese diacritics handling)
|
||||
- Read-only SQL queries against a single `student` table
|
||||
- Static dataset — no updates (the 2016 exam is long over)
|
||||
|
||||
## Target users
|
||||
|
||||
- Former 2016 candidates checking their scores
|
||||
- Education researchers / data journalists running aggregate stats
|
||||
- Developers exploring SQL on a real-world dataset
|
||||
|
||||
## Constraints
|
||||
|
||||
- **Zero backend**: the full DB (tens of MB gzipped) is downloaded to the browser
|
||||
- **Read-only**: INSERT/UPDATE/DELETE rejected to avoid the illusion that user edits persist
|
||||
- **Row caps**: 100 rows (lookup), 1000 rows (SQL) to prevent browser hangs
|
||||
- **Vietnamese-first UI**: app labels and data are Vietnamese; documentation is English
|
||||
|
||||
## Data sources
|
||||
|
||||
Excel files (`.xlsx`/`.xls`) collected from newspapers and exam clusters in 2016, stored in `data/`. One file per cluster, with several different column layouts (see `scripts/build-database.js`).
|
||||
|
||||
Original aggregator link: <https://dtntbacgiang.edu.vn/tin-tuc/tin-tuc-su-kien/cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html> — no longer accessible, so the raw Excel files are mirrored in this repo under `data/`.
|
||||
|
||||
## Status
|
||||
|
||||
Stable. Data is frozen. Recent work focuses on UX polish (dark mode, accessibility, diacritics-insensitive search).
|
||||
@@ -0,0 +1,66 @@
|
||||
# Project Overview
|
||||
|
||||
## Goal
|
||||
|
||||
A public lookup tool for Vietnam's National High School Graduation Exam scores,
|
||||
running entirely in the browser and hosted for free on GitHub Pages. Covers the
|
||||
2016 and 2017 exams — 1.7 million candidates across four datasets.
|
||||
|
||||
## Scope
|
||||
|
||||
- Lookup by exam ID or full name, with Vietnamese diacritics handled
|
||||
- Read-only SQL queries against a single `student` table
|
||||
- Admission-block (khối thi) totals computed per candidate
|
||||
- Static datasets — both exams are long over and the data is frozen
|
||||
|
||||
## Target users
|
||||
|
||||
- Former candidates checking their scores
|
||||
- Education researchers and data journalists running aggregate statistics
|
||||
- Developers exploring SQL against a real-world dataset
|
||||
|
||||
## Constraints
|
||||
|
||||
- **Zero backend.** The full database (38–48 MB gzipped per dataset) is
|
||||
downloaded to the browser and queried in-process.
|
||||
- **Read-only.** `INSERT`/`UPDATE`/`DELETE` are rejected, so nobody is misled
|
||||
into thinking edits persist. `sql.js` is in-memory anyway.
|
||||
- **Row caps.** 100 rows for lookups, 1000 for custom SQL, to prevent browser
|
||||
hangs.
|
||||
- **Vietnamese-first UI.** App labels and data are Vietnamese; documentation is
|
||||
English.
|
||||
|
||||
## Datasets
|
||||
|
||||
| id | Exam | Candidates | Notes |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 2016 | 877,461 | 119 files, three column layouts |
|
||||
| `2017` | 2017 | 861,068 | current generation, reproducible from source |
|
||||
| `2017-old` | 2017 | 847,348 | pre-refresh archive |
|
||||
| `2017-old2` | 2017 | 679,764 | corrected re-export |
|
||||
|
||||
The three 2017 datasets are kept side by side because they disagree, and the
|
||||
disagreement is itself informative. Only `2017` is re-fetchable; the rest exist
|
||||
solely as the copies committed here.
|
||||
|
||||
Original 2016 aggregator link
|
||||
(<https://dtntbacgiang.edu.vn/tin-tuc/tin-tuc-su-kien/cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html>)
|
||||
is no longer accessible, which is why the raw files are mirrored in `data/2016/`.
|
||||
|
||||
## History
|
||||
|
||||
Each year began as a standalone repository (`thptqg2016`, `thptqg2017`), merged
|
||||
here with full history. They initially kept separate frontends and separate
|
||||
copies of the same Rust parser, synchronised by hand. That duplication was
|
||||
removed: there is now one frontend, one parser, and one canonical schema, with
|
||||
per-dataset differences confined to four small config files and one registry
|
||||
entry each.
|
||||
|
||||
The unification also fixed a latent data-loss bug — neither year's parser
|
||||
config listed the complete set of subjects, so 1,691 candidates were missing
|
||||
their foreign-language score. See `data-pipeline.md`.
|
||||
|
||||
## Status
|
||||
|
||||
Stable, data frozen. Work is limited to UX polish and keeping the pipeline
|
||||
maintainable.
|
||||
@@ -1,83 +0,0 @@
|
||||
# System Architecture
|
||||
|
||||
## Overview
|
||||
|
||||
A **static serverless** design: the entire dataset is packaged into a single SQLite file, gzip-compressed, and served as a static asset via GitHub Pages. The browser downloads it, decompresses it, and queries it in-process using `sql.js` (SQLite compiled to WebAssembly).
|
||||
|
||||
```
|
||||
┌─────────────────┐ build ┌──────────────────────┐
|
||||
│ data/*.xlsx │ ───────────▶ │ scripts/ │
|
||||
│ (mixed formats)│ │ build-database.js │
|
||||
└─────────────────┘ │ (Node + xlsx + │
|
||||
│ better-sqlite3) │
|
||||
└──────────┬───────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────┐
|
||||
│ public/thptqg2016.db │
|
||||
└──────────┬───────────┘
|
||||
│ gzip -9 (CI)
|
||||
▼
|
||||
┌──────────────────────┐
|
||||
│ dist/thptqg2016.db.gz│
|
||||
│ dist/assets/* │ ◀── Vite build
|
||||
└──────────┬───────────┘
|
||||
│ upload-pages-artifact
|
||||
▼
|
||||
GitHub Pages CDN
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────────┐
|
||||
│ Browser │
|
||||
│ ┌────────────────────────────┐ │
|
||||
│ │ useSqlite hook │ │
|
||||
│ │ fetch(.db.gz) + stream │ │
|
||||
│ │ DecompressionStream gzip │ │
|
||||
│ │ sql.js WASM (from CDN) │ │
|
||||
│ └─────────────┬──────────────┘ │
|
||||
│ ▼ │
|
||||
│ ┌────────────────────────────┐ │
|
||||
│ │ React UI │ │
|
||||
│ │ - SearchForm / ScoreTable │ │
|
||||
│ │ - CustomQuery (SQL editor)│ │
|
||||
│ └────────────────────────────┘ │
|
||||
└──────────────────────────────────┘
|
||||
```
|
||||
|
||||
## Build-time data flow
|
||||
|
||||
1. Developer drops Excel files into `data/`.
|
||||
2. `npm run build:db` reads every file; for each one:
|
||||
- Detects the header row against a `KNOWN_HEADERS` set.
|
||||
- Picks a format: `separate-scores` (one column per subject) vs. `mapped` (single `DIEM_THI` string).
|
||||
- Parses each row into a canonical 18-column record.
|
||||
- `INSERT OR REPLACE` into SQLite (primary key = `so_bao_danh` handles duplicates).
|
||||
3. `VACUUM` shrinks the file.
|
||||
4. CI runs `gzip -k -9` → `.db.gz`.
|
||||
|
||||
## Runtime flow
|
||||
|
||||
1. Page loads → React mounts → `useSqlite("thptqg2016.db.gz")`.
|
||||
2. Streaming fetch with a progress bar (driven by `Content-Length`).
|
||||
3. `DecompressionStream("gzip")` decompresses on the fly.
|
||||
4. `sql.js` loads its WASM from `https://sql.js.org/dist/sql-wasm.wasm`.
|
||||
5. `new SQL.Database(Uint8Array)` — DB is now in RAM.
|
||||
6. Each search / query → `db.prepare()` + `stmt.step()` loop → render.
|
||||
|
||||
## Design decisions
|
||||
|
||||
| Concern | Choice | Rationale |
|
||||
|---------|--------|-----------|
|
||||
| Storage | Static SQLite file | No backend needed; dataset is frozen |
|
||||
| Compression | gzip in CI, `DecompressionStream` in browser | Native browser API; no extra library |
|
||||
| WASM hosting | `sql.js.org` CDN | Smaller self-hosted artifact |
|
||||
| Diacritics search | Pre-computed `ho_ten_ascii` column | `LOWER(REPLACE(...))` at query time defeats the index |
|
||||
| SQL safety | Leading-keyword allowlist | `sql.js` is in-memory so writes don't persist, but the allowlist prevents user confusion |
|
||||
| Row caps | 100 (lookup), 1000 (SQL) | Keep DOM render sizes reasonable |
|
||||
|
||||
## Risks and limitations
|
||||
|
||||
- **DB size**: tens of MB gzipped — slow links have a visible wait; mitigated by the progress bar.
|
||||
- **Browser memory**: the full DB lives in RAM; older mobile devices may OOM.
|
||||
- **Dependency on `sql.js.org`**: if that CDN is unreachable, WASM fails to load.
|
||||
- **Excel format drift**: a new source file with an unseen header layout needs a new branch in `detectFormat()`.
|
||||
+143
-60
@@ -1,97 +1,180 @@
|
||||
# System Architecture
|
||||
|
||||
Static site. No backend. Browser downloads a compressed SQLite file at boot, then every query runs locally via `sql.js` (WASM).
|
||||
Static site, no backend. The browser downloads a compressed SQLite file at boot
|
||||
and every query runs locally via `sql.js` (SQLite compiled to WebAssembly).
|
||||
|
||||
One frontend, one parser, one schema, four datasets.
|
||||
|
||||
## Data flow
|
||||
|
||||
```
|
||||
Excel files (.xls / .xlsx)
|
||||
data/<id>/*.xls(x)
|
||||
│
|
||||
▼ build-database*.js (Node)
|
||||
SQLite DB (public*/thptqg2017.db)
|
||||
▼ parser/ (Rust, one binary, one config per dataset)
|
||||
.build/public/db/<id>.db
|
||||
│
|
||||
▼ gzip -9
|
||||
thptqg2017.db.gz (~47 MB for main variant)
|
||||
▼ gzip -9 (no -k: the raw file does not survive)
|
||||
.build/public/db/<id>.db.gz
|
||||
│
|
||||
▼ Vite build copies public*/ into dist/
|
||||
Static site on GitHub Pages
|
||||
▼ vite build (publicDir = .build/public)
|
||||
dist/
|
||||
│
|
||||
▼ browser loads
|
||||
sql.js (WASM) opens the .db → queries run client-side
|
||||
▼ scripts/assemble-site.js
|
||||
_site/ → GitHub Pages
|
||||
│
|
||||
▼ browser
|
||||
sql.js (WASM) opens the .db → queries run client-side
|
||||
```
|
||||
|
||||
## Three deployment variants
|
||||
## The dataset id
|
||||
|
||||
One repo → three independent sites, same frontend, different dataset.
|
||||
One identifier ties the whole pipeline together:
|
||||
|
||||
| Variant | Route | Source dir | Public dir | Build cmd |
|
||||
|---|---|---|---|---|
|
||||
| main | `/thptqg2017/` | `data/` | `public/` | `npm run build` |
|
||||
| old | `/thptqg2017/old/` | `data-old/` | `public-old/` | `npm run build:old` |
|
||||
| old2 | `/thptqg2017/old2/` | `data-old2/` | `public-old2/` | `npm run build:old2` |
|
||||
```
|
||||
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
|
||||
```
|
||||
|
||||
`vite.config.js` reads `process.env.VARIANT` and switches `base`, `publicDir`, `outDir` accordingly. `emptyOutDir` is on only for the main build so the variant builds merge into `dist/old/`, `dist/old2/` cleanly.
|
||||
`src/datasets.js` declares the four ids once. The frontend, the database build
|
||||
(`parser/scripts/build-db.js`) and the site assembly all import that list, so
|
||||
adding a dataset means adding one entry and one config file.
|
||||
|
||||
## Schema (all 3 DBs share this)
|
||||
| id | Exam | Rows | Source |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 2016 | 877,461 | Bộ GD&ĐT |
|
||||
| `2017` | 2017 | 861,068 | baotintuc.vn |
|
||||
| `2017-old` | 2017 | 847,348 | pre-refresh archive |
|
||||
| `2017-old2` | 2017 | 679,764 | corrected re-export |
|
||||
|
||||
## Canonical schema
|
||||
|
||||
Defined once in `parser/src/schema.rs` — DDL, INSERT, column order and the 16
|
||||
subject regexes. The four TOML configs carry no SQL at all, only per-dataset
|
||||
parse rules. Config parsing uses `deny_unknown_fields`, so a leftover `[schema]`
|
||||
block fails loudly instead of looking effective while `schema.rs` drives the
|
||||
build.
|
||||
|
||||
```sql
|
||||
CREATE TABLE student (
|
||||
so_bao_danh TEXT PRIMARY KEY, -- exam ID (8 digits, first 2 = province)
|
||||
so_bao_danh TEXT PRIMARY KEY, -- 2017: 8 digits; 2016: 9 digits or a
|
||||
-- 2-4 letter cluster code then digits
|
||||
ho_ten TEXT NOT NULL,
|
||||
ho_ten_ascii TEXT NOT NULL, -- NFD-stripped lowercase for accent-insensitive search
|
||||
ngay_sinh TEXT, -- dd/mm/yyyy
|
||||
ho_ten_ascii TEXT NOT NULL, -- NFD-stripped lowercase, for accent-insensitive search
|
||||
ngay_sinh TEXT, -- dd/mm/yyyy
|
||||
ten_cum_thi TEXT, -- 2016 only
|
||||
gioi_tinh TEXT, -- 2016 only
|
||||
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn,
|
||||
lich_su, dia_ly, gdcd, khxh,
|
||||
tieng_anh, tieng_phap, tieng_nga, tieng_trung REAL
|
||||
tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung REAL
|
||||
);
|
||||
CREATE INDEX idx_ho_ten ON student(ho_ten);
|
||||
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
|
||||
CREATE INDEX idx_ho_ten ON student(ho_ten);
|
||||
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
|
||||
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
|
||||
```
|
||||
|
||||
No German / Japanese language columns — neither appears in any source file.
|
||||
Every dataset gets all 22 columns; ones it has no data for are NULL, costing
|
||||
about a byte per row. `khtn`, `khxh` and `gdcd` are empty on 2016;
|
||||
`ten_cum_thi` and `gioi_tinh` are empty on the 2017 datasets.
|
||||
|
||||
Each score column is `NULL` when the student didn't take that subject. A student's "admission block" total is only computed when all 3 required subjects are non-null.
|
||||
`idx_ten_cum_thi` is partial, so it holds zero entries where the column is
|
||||
always NULL.
|
||||
|
||||
## Parse quirks (why 3 builder scripts)
|
||||
## Routing
|
||||
|
||||
Each dataset had a different quirk. Rather than one mega-parser with mode flags, each builder script is ~70 lines and isolates its own workarounds.
|
||||
URLs are flat, one segment per dataset, and the segment is the id:
|
||||
|
||||
| Dataset | Quirk | Mitigation |
|
||||
|---|---|---|
|
||||
| data/ (baotintuc .xls) | Hanoi + HCM overflow past the 65,536 row-per-sheet .xls limit | Iterate all `wb.SheetNames` |
|
||||
| data-old/ (xlsx) | Single-sheet always; one bogus header row (`SOBAODANH`/`HO_TEN`) leaked into an earlier DB | Strict numeric-SBD guard rejects the leak |
|
||||
| data-old2/ (xlsx) | HCM overflow + many blank trailing rows | Multi-sheet walk + blank-row pre-filter |
|
||||
```
|
||||
/thptqg/ hub
|
||||
/thptqg/2016/
|
||||
/thptqg/2017/
|
||||
/thptqg/2017-old/
|
||||
/thptqg/2017-old2/
|
||||
```
|
||||
|
||||
Shared concerns (regex, schema, `toAscii`, header detection) live in `scripts/build-lib.js`.
|
||||
`src/router.js` is an exact match on that segment. The nested form used before
|
||||
(`/thptqg/2017/old/`) would have needed longest-prefix matching, since it also
|
||||
starts with `/thptqg/2017/`. Those two legacy URLs still resolve: the router
|
||||
rewrites them to the flat equivalent with `history.replaceState`, preserving the
|
||||
query string so `?q=` deep links survive.
|
||||
|
||||
A single Vite build emits one `index.html`. Because `base` is absolute
|
||||
(`/thptqg/`), that file references `/thptqg/assets/...` regardless of the
|
||||
directory it is served from, so the assemble step copies it to every route and
|
||||
each URL is a real static file. **No SPA 404-fallback redirect is used** — the
|
||||
usual hack rewrites URLs and would interfere with the deep links.
|
||||
|
||||
## Serving both exam years without branching
|
||||
|
||||
No component contains a per-dataset conditional. Two mechanisms do the work:
|
||||
|
||||
- **All-NULL columns are hidden.** `score-table.jsx` drops any column where
|
||||
every row in the result set is NULL, so 2016 rows surface Cụm thi / GT / Đức /
|
||||
Nhật and 2017 rows surface KHTN / KHXH / GDCD / Nga.
|
||||
- **Incomplete admission blocks are skipped.** `computeBlocks()` only returns a
|
||||
block when the student has all three subjects, so one block list covers both
|
||||
years: GDCD blocks self-exclude on 2016, German and Japanese blocks
|
||||
self-exclude wherever those languages were not sat.
|
||||
|
||||
Anything genuinely per-dataset — title, source, database size, search examples,
|
||||
SQL presets — lives in `src/datasets.js`.
|
||||
|
||||
## Exam ID formats
|
||||
|
||||
`src/lib/query-mode.js` decides whether a query is an exam ID or a name, and is
|
||||
shared by `App.jsx` and `search-form.jsx` (they previously held separate copies
|
||||
and had drifted apart on exactly this rule).
|
||||
|
||||
| Form | Example | Where |
|
||||
| --- | --- | --- |
|
||||
| 8 digits | `49008235` | 2017 — first two digits are the province |
|
||||
| 9 digits with leading zero | `017006021` | 2016 |
|
||||
| 2-4 letters then digits | `BAL000001` | 2016 — exam cluster code |
|
||||
|
||||
The letter-prefixed form is the majority case for 2016: 616,593 of 877,461
|
||||
candidates (70.3%). Letter prefixes are upper-cased before lookup, so
|
||||
`bal000001` resolves.
|
||||
|
||||
## Score tiers
|
||||
|
||||
Six-level ladder in `scoreTier()` (`src/lib/admission-blocks.js`), paired with a
|
||||
symbol so meaning is never colour-only.
|
||||
|
||||
| Tier | Range | Vietnamese |
|
||||
| --- | --- | --- |
|
||||
| common | ≤ 1 | Điểm liệt |
|
||||
| uncommon | < 5 | Chưa đạt |
|
||||
| rare | 5–6.5 | Trung bình |
|
||||
| epic | 6.5–8 | Khá |
|
||||
| legendary | 8–9 | Giỏi |
|
||||
| prismatic | 9–10 | Xuất sắc |
|
||||
|
||||
## Admission blocks
|
||||
|
||||
Vietnamese universities admit students based on 3-subject combinations called "admission blocks" (khối thi). `src/lib/admission-blocks.js` lists all 49 official 2017 blocks computable from our schema (A00–A11, B00–B08, C00–C20, D01–D15). Each entry is `{ code, subjects: [3 keys], label }`.
|
||||
Vietnamese universities admit on three-subject combinations (khối thi).
|
||||
`src/lib/admission-blocks.js` lists the blocks computable from this schema
|
||||
(A00–A11, B00–B08, C00–C20, D01–D15, plus D05/D06 for German and Japanese).
|
||||
`computeBlocks(student)` returns those where all three scores exist, sorted by
|
||||
total descending.
|
||||
|
||||
`computeBlocks(student)` returns only blocks where the student has all 3 subject scores, sorted by total desc. Used by the student detail card.
|
||||
## Design decisions
|
||||
|
||||
The SQL preset "Top 10 best-block in Long An" materialises all 49 blocks as a `UNION ALL` CTE, picks each student's best block via `ROW_NUMBER() OVER (PARTITION BY so_bao_danh ORDER BY s DESC, k)`, then ranks across students.
|
||||
| Concern | Choice | Rationale |
|
||||
| --- | --- | --- |
|
||||
| Storage | Static SQLite file | No backend; the datasets are frozen |
|
||||
| Compression | gzip in CI, `DecompressionStream` in the browser | Native API, no extra library |
|
||||
| WASM hosting | `sql.js.org` CDN | Smaller self-hosted artifact |
|
||||
| Diacritics search | Pre-computed `ho_ten_ascii` | `LOWER(REPLACE(...))` at query time defeats the index |
|
||||
| SQL safety | Leading-keyword allowlist | `sql.js` is in-memory so writes cannot persist; the allowlist prevents confusion |
|
||||
| Row caps | 100 (lookup), 1000 (SQL) | Keeps DOM render sizes reasonable |
|
||||
| Routing | Hand-rolled, ~20 lines | Five static routes do not justify a router dependency |
|
||||
|
||||
## Score tiers (UI coding)
|
||||
## Risks and limitations
|
||||
|
||||
6-level TFT rarity ladder (white → green → blue → purple → gold → prismatic). Defined in `scoreTier()` in `src/lib/admission-blocks.js`.
|
||||
|
||||
| Tier | Range | UI label (Vietnamese) | English |
|
||||
|---|---|---|---|
|
||||
| common | ≤ 1 | Điểm liệt | Disqualifying |
|
||||
| uncommon | < 5 | Chưa đạt | Below passing |
|
||||
| rare | 5–6.5 | Trung bình | Average |
|
||||
| epic | 6.5–8 | Khá | Good |
|
||||
| legendary | 8–9 | Giỏi | Very good |
|
||||
| prismatic | 9–10 | Xuất sắc | Excellent |
|
||||
|
||||
Color + icon + text label — never color-only.
|
||||
|
||||
## Frontend key behaviors
|
||||
|
||||
- Query state lives in `App.jsx` and is synced to URL `?q=...` via `history.replaceState` — bookmarkable, shareable.
|
||||
- `SearchForm` is a controlled component that 300ms-debounces input changes and auto-triggers search once the query is long enough (3+ digits for SBD, 2+ chars for name).
|
||||
- Detection: all-digits → exact `so_bao_danh` match; ASCII-only → folded `ho_ten_ascii LIKE`; has Vietnamese diacritics → search both `ho_ten` and folded column.
|
||||
- Exactly 1 result → `StudentDetail` card; 0 or >1 → `ScoreTable`.
|
||||
- Global `/` key focuses the search input when not already typing.
|
||||
- Share button on the detail card uses Web Share API when present, else clipboard with a pre-formatted summary + deep-link URL.
|
||||
- **Database size.** 38–48 MB gzipped per dataset; slow links wait, mitigated by
|
||||
a progress bar.
|
||||
- **Browser memory.** The full database lives in RAM; older mobile devices may
|
||||
run out.
|
||||
- **`sql.js.org` dependency.** If that CDN is unreachable, the WASM fails to
|
||||
load. Self-hosting `sql-wasm.wasm` and updating `SQL_WASM_URL` in
|
||||
`use-sqlite.js` is the fix.
|
||||
- **Excel format drift.** A new source file with an unseen header layout needs a
|
||||
new branch in `format_detect_2016.rs` or a new config.
|
||||
@@ -19,7 +19,10 @@
|
||||
/// SBD(0) HO_TEN(1) NGAY_SINH(2) TEN_CUMTHI(3) GIOI_TINH(4) DIEM_THI(5)
|
||||
/// → JS: build-database.js:149–151 DEFAULT_MAP + processMappedRow
|
||||
///
|
||||
/// JS citations are line numbers in /config/workspace/tiennm99/thptqg2016/scripts/build-database.js.
|
||||
/// The `build-database.js` citations throughout this file refer to the Node
|
||||
/// script this parser replaced, in the original standalone thptqg2016 repo.
|
||||
/// That file no longer exists here; the references are kept because they
|
||||
/// explain why several of the rules below look arbitrary.
|
||||
use calamine::Data;
|
||||
|
||||
use crate::transform::{parse_scores, to_ascii, CompiledPatterns, ParsedRow};
|
||||
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 1
|
||||
title: "Standard schema and unified parser"
|
||||
status: pending
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies: []
|
||||
effort: ""
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 2
|
||||
title: "Repo restructure"
|
||||
status: pending
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies: [1]
|
||||
effort: ""
|
||||
@@ -131,12 +131,12 @@ component.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] Target layout matches `plan.md` exactly
|
||||
- [ ] `2016/` and `2017/` no longer exist
|
||||
- [ ] Old hub page content captured before deletion
|
||||
- [ ] Git reports renames, repository size unchanged
|
||||
- [ ] `package-lock.json` committed; `npm ci` succeeds; zero pnpm files remain
|
||||
- [ ] `cargo test` green from the new location
|
||||
- [x] Target layout matches `plan.md` exactly
|
||||
- [x] `2016/` and `2017/` no longer exist
|
||||
- [x] Old hub page content captured before deletion
|
||||
- [x] Git reports renames, repository size unchanged
|
||||
- [x] `package-lock.json` committed; `npm ci` succeeds; zero pnpm files remain
|
||||
- [x] `cargo test` green from the new location
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
|
||||
+11
-11
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 3
|
||||
title: "Unified frontend and dataset registry"
|
||||
status: pending
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies: [2]
|
||||
effort: ""
|
||||
@@ -225,16 +225,16 @@ this repo today):
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] One `src/` serving four datasets **and** the hub, zero `if (dataset === ...)` in components
|
||||
- [ ] Subject list defined once in `src/lib/subjects.js`
|
||||
- [ ] Hub route fetches no database
|
||||
- [ ] Dataset path and DB URL are derived from `id`, not stored per entry
|
||||
- [ ] Legacy `/2017/old/` and `/2017/old2/` redirect to flat URLs, `?q=` preserved
|
||||
- [ ] 2016 route renders cluster + gender; 2017 routes do not
|
||||
- [ ] 2016 route has live search, deep links, student detail, score tiers
|
||||
- [ ] Admission blocks correct for both exam years
|
||||
- [ ] No routing library added
|
||||
- [ ] `npm run lint` green
|
||||
- [x] One `src/` serving four datasets **and** the hub, zero `if (dataset === ...)` in components
|
||||
- [x] Subject list defined once in `src/lib/subjects.js`
|
||||
- [x] Hub route fetches no database
|
||||
- [x] Dataset path and DB URL are derived from `id`, not stored per entry
|
||||
- [x] Legacy `/2017/old/` and `/2017/old2/` redirect to flat URLs, `?q=` preserved
|
||||
- [x] 2016 route renders cluster + gender; 2017 routes do not
|
||||
- [x] 2016 route has live search, deep links, student detail, score tiers
|
||||
- [x] Admission blocks correct for both exam years
|
||||
- [x] No routing library added
|
||||
- [x] `npm run lint` green
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 4
|
||||
title: "Build and deploy pipeline"
|
||||
status: pending
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies: [3]
|
||||
effort: ""
|
||||
|
||||
+8
-8
@@ -1,7 +1,7 @@
|
||||
---
|
||||
phase: 5
|
||||
title: "Parity verification and docs"
|
||||
status: pending
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies: [4]
|
||||
effort: ""
|
||||
@@ -110,13 +110,13 @@ so reruns compare the same students.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] `verify-parity.js` committed and exiting 0 on all four datasets
|
||||
- [ ] Parity result report committed under `plans/reports/`
|
||||
- [ ] Zero unexpected non-NULL columns
|
||||
- [ ] Manual checklist passes on the hub and all four dataset routes
|
||||
- [ ] `verify-parity.js` has zero npm dependencies (`node:sqlite` only)
|
||||
- [ ] `README.md` and `docs/` describe the actual repo, with npm commands
|
||||
- [ ] No stale path, pnpm, or build-variant references anywhere outside `plans/`
|
||||
- [x] `verify-parity.js` committed and exiting 0 on all four datasets
|
||||
- [x] Parity result report committed under `plans/reports/`
|
||||
- [x] Zero unexpected non-NULL columns
|
||||
- [x] Manual checklist passes on the hub and all four dataset routes
|
||||
- [x] `verify-parity.js` has zero npm dependencies (`node:sqlite` only)
|
||||
- [x] `README.md` and `docs/` describe the actual repo, with npm commands
|
||||
- [x] No stale path, pnpm, or build-variant references anywhere outside `plans/`
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
title: "Unify frontend, standardize SQL schema, restructure repo"
|
||||
description: "One 2017-based frontend, one canonical 22-column student schema, one parser crate, four datasets under data/"
|
||||
status: pending
|
||||
status: completed
|
||||
priority: P2
|
||||
branch: "main"
|
||||
tags: [refactor, schema, frontend, parser]
|
||||
@@ -134,11 +134,11 @@ CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT N
|
||||
|
||||
| Phase | Name | Status |
|
||||
|-------|------|--------|
|
||||
| 1 | [Standard schema and unified parser](./phase-01-standard-schema-and-unified-parser.md) | Pending |
|
||||
| 2 | [Repo restructure](./phase-02-repo-restructure.md) | Pending |
|
||||
| 3 | [Unified frontend and dataset registry](./phase-03-unified-frontend-and-dataset-registry.md) | Pending |
|
||||
| 4 | [Build and deploy pipeline](./phase-04-build-and-deploy-pipeline.md) | Pending |
|
||||
| 5 | [Parity verification and docs](./phase-05-parity-verification-and-docs.md) | Pending |
|
||||
| 1 | [Standard schema and unified parser](./phase-01-standard-schema-and-unified-parser.md) | Completed |
|
||||
| 2 | [Repo restructure](./phase-02-repo-restructure.md) | Completed |
|
||||
| 3 | [Unified frontend and dataset registry](./phase-03-unified-frontend-and-dataset-registry.md) | Completed |
|
||||
| 4 | [Build and deploy pipeline](./phase-04-build-and-deploy-pipeline.md) | Completed |
|
||||
| 5 | [Parity verification and docs](./phase-05-parity-verification-and-docs.md) | Completed |
|
||||
|
||||
## Dependencies
|
||||
|
||||
@@ -151,17 +151,17 @@ No cross-plan dependencies (`plans/` was empty before this plan).
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] One Rust crate; `2016/tools/` and `2017/tools/` gone
|
||||
- [ ] One `src/`; `2016/src/` and `2017/src/` gone
|
||||
- [ ] One Vite build producing all five pages
|
||||
- [ ] All four DBs built from the same DDL, same INSERT, same 16 regexes
|
||||
- [ ] Per-dataset row count and per-column non-NULL count identical to pre-refactor baseline
|
||||
- [ ] All five published URLs functional under the flat scheme, deep links included
|
||||
- [ ] Legacy `/2017/old/` and `/2017/old2/` redirect to their flat equivalents, preserving `?q=`
|
||||
- [ ] 2016 site still shows `ten_cum_thi` + `gioi_tinh`; 2017 sites do not
|
||||
- [ ] 2016 site gains 2017's features (deep links, student detail, tiers, live search)
|
||||
- [ ] npm only: `package-lock.json` committed, no pnpm files or CI steps remain
|
||||
- [ ] `cargo test` green; `npm run lint` green
|
||||
- [x] One Rust crate; `2016/tools/` and `2017/tools/` gone
|
||||
- [x] One `src/`; `2016/src/` and `2017/src/` gone
|
||||
- [x] One Vite build producing all five pages
|
||||
- [x] All four DBs built from the same DDL, same INSERT, same 16 regexes
|
||||
- [x] Per-dataset row count and per-column non-NULL count identical to pre-refactor baseline
|
||||
- [x] All five published URLs functional under the flat scheme, deep links included
|
||||
- [x] Legacy `/2017/old/` and `/2017/old2/` redirect to their flat equivalents, preserving `?q=`
|
||||
- [x] 2016 site still shows `ten_cum_thi` + `gioi_tinh`; 2017 sites do not
|
||||
- [x] 2016 site gains 2017's features (deep links, student detail, tiers, live search)
|
||||
- [x] npm only: `package-lock.json` committed, no pnpm files or CI steps remain
|
||||
- [x] `cargo test` green; `npm run lint` green
|
||||
|
||||
## Rollback
|
||||
|
||||
|
||||
@@ -0,0 +1,79 @@
|
||||
# Parser parity result
|
||||
|
||||
Comparison of the databases the unified pipeline ships against a baseline built
|
||||
from the pre-refactor code. Run on 2026-08-13 against the final tree
|
||||
(`npm run build:rust && npm run build:db`), gate exit code 0.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `plans/reports/parser-parity-baseline.json` — built with the two separate
|
||||
crates, before any change
|
||||
- `plans/reports/parser-parity-current.json` — decompressed from
|
||||
`.build/public/db/*.db.gz`, i.e. the exact bytes published
|
||||
|
||||
## Result
|
||||
|
||||
| Dataset | Rows | Columns | Size |
|
||||
| --- | --- | --- | --- |
|
||||
| 2016 | 877,461 | 18 → 22 | +1.4% |
|
||||
| 2017 | 861,068 | 18 → 22 | +2.2% |
|
||||
| 2017-old | 847,348 | 18 → 22 | +2.2% |
|
||||
| 2017-old2 | 679,764 | 18 → 22 | +2.2% |
|
||||
|
||||
Unchanged, for every dataset:
|
||||
|
||||
- row count
|
||||
- non-NULL count for all 18 pre-existing columns
|
||||
- every field of a deterministic student sample (SBDs ending `0000`)
|
||||
|
||||
## Approved recoveries
|
||||
|
||||
Unifying the subject regexes recovered 1,691 foreign-language scores that the
|
||||
old per-year configs discarded. The 2016 config listed 12 subject patterns and
|
||||
the 2017 configs 14; neither list was complete, so candidates who sat German,
|
||||
Japanese or Russian ended up with no foreign-language score at all.
|
||||
|
||||
| Dataset | Recovered |
|
||||
| --- | --- |
|
||||
| 2016 | 182 × `tieng_nga` |
|
||||
| 2017 | 93 × `tieng_duc`, 512 × `tieng_nhat` |
|
||||
| 2017-old | 85 × `tieng_duc`, 484 × `tieng_nhat` |
|
||||
| 2017-old2 | 22 × `tieng_duc`, 313 × `tieng_nhat` |
|
||||
|
||||
Confirmed real rather than spurious matches:
|
||||
|
||||
- Every student in all four datasets holds zero or exactly one foreign
|
||||
language — never two — so nothing is duplicated.
|
||||
- Each affected student had all language columns NULL beforehand. SBD
|
||||
`01003198` went from all-NULL to `tieng_duc = 8`.
|
||||
- Counts track dataset size across the three 2017 generations (93/85/22
|
||||
German, 512/484/313 Japanese).
|
||||
- Values are ordinary exam scores in the 0–10 range.
|
||||
|
||||
These exact counts are pinned as `APPROVED_RECOVERY` in
|
||||
`parser/scripts/verify-parity.js`. Any other newly-populated column, or drift
|
||||
in these numbers, still fails the gate.
|
||||
|
||||
## Columns that must stay empty
|
||||
|
||||
Verified 0 non-NULL:
|
||||
|
||||
- 2016: `khtn`, `khxh`, `gdcd`
|
||||
- 2017, 2017-old, 2017-old2: `ten_cum_thi`, `gioi_tinh`
|
||||
|
||||
## Reproducing
|
||||
|
||||
```sh
|
||||
npm run build:rust && npm run build:db
|
||||
node parser/scripts/db-stats.js \
|
||||
2016=<db> 2017=<db> 2017-old=<db> 2017-old2=<db> > current.json
|
||||
node parser/scripts/verify-parity.js \
|
||||
plans/reports/parser-parity-baseline.json current.json
|
||||
```
|
||||
|
||||
The baseline cannot be regenerated — it required the two pre-refactor crates,
|
||||
which no longer exist. It is committed for that reason.
|
||||
|
||||
## Unresolved questions
|
||||
|
||||
None.
|
||||
Reference in new issue
Block a user