The repo held two near-duplicate projects. 2016/ and 2017/ each carried their
own React frontend, their own copy of the same Rust crate, and their own
package manager setup. 2016/tools/sync-from-thptqg2017.sh existed purely to
copy the parser source between them.
New layout:
index.html + src/ the 2017 frontend, now the only one
data/<id>/ 2016, 2017, 2017-old, 2017-old2
parser/ the single Rust crate, configs renamed to <id>.toml
docs/ both projects' docs, 2016 copies suffixed -2016-legacy
pending the merge pass
<id> is now one identifier end to end: data/<id>/ feeds parser/configs/<id>.toml
and produces db/<id>.db.gz.
pnpm gives way to npm. pnpm-workspace.yaml existed only to whitelist
better-sqlite3's native build, which npm permits by default, so it has no
equivalent and is simply gone. Lockfiles cannot be converted; package-lock.json
is generated fresh. The migration direction is safe — pnpm's strict layout
forbids phantom dependencies, so anything that resolved under pnpm resolves
under npm's flat tree.
Adds parser/scripts/build-db.js and src/datasets.js: the four dataset IDs are
declared once and read by both the build tooling and (from the next phase) the
frontend.
Follow-on fixes the move made necessary:
- eslint's Node-globals override pointed at scripts/, now parser/scripts/
- crawl-baotintuc.js wrote to <root>/data, now data/2017
- golden tests loaded configs by their old thptqg*-data.toml names
Drops the #[ignore]d Rust-vs-Node golden test. It shelled out to pnpm to run
scripts/build-database.js, a file removed when the parser was ported to Rust,
so it could never pass. check-duplicates.js and diff-datasets.js were already
broken before this change and are annotated as such rather than half-fixed.
63 Rust tests pass and clippy is clean from the new location.
cargo clippy --all-targets -- -D warnings failed with 8 errors before the schema
refactor and was never part of CI. Clearing them so the gate is meaningful from
here on.
All behaviour-preserving: slice contains() over iter().any(), is_multiple_of()
over a modulo, a needless borrow in a test helper, and a doc-list indent.
process_mapped_row keeps its 8 parameters behind an allow attribute — the
signature mirrors the JS column map one-for-one, and bundling the indices into
a struct would hide that correspondence.
Kept separate from the schema change so that diff stays readable.
The 2016 and 2017 parsers were the same crate with divergent SQL. DDL, INSERT
and the subject regex table lived in each dataset's TOML config, so four copies
had to be kept in step by hand — which is how the two schemas drifted apart.
Move all of it into src/schema.rs as a single 22-column definition: 6 identity
columns (adding ten_cum_thi and gioi_tinh, previously 2016-only) and 16 subject
columns (the union of both exam years). Columns a dataset carries no data for
bind NULL, costing ~1 byte per row.
This collapses writer.rs to one insert path: insert_row_2016, SCORE_FIELDS_2016,
SCORE_FIELDS_2017 and the SCORE_FIELDS alias all go away. Configs shrink to the
parse rules that genuinely vary per dataset — sheet mode, column indices, SBD
validation, header tokens, blank-row stripping — and carry no SQL at all.
Config parsing now denies unknown fields, so a leftover [schema] block fails
loudly instead of looking effective while schema.rs drives the build.
idx_ten_cum_thi is partial, so it holds zero entries on the three 2017 datasets
where the column is always NULL.
Adds db-stats.js and verify-parity.js to prove no data moved: row counts,
per-column non-NULL counts and a deterministic student sample are compared
against databases built from the previous code.
Verified across all four datasets — row counts, every pre-existing column count,
and all sampled students are identical. Database size grows 1.4-2.2%.
Each year keeps its full pipeline under its directory; vite bases move to
/thptqg/<year>/ (2017 keeps its old and old2 generations), one deploy
workflow builds both years' databases and bundles behind a root index.
* feat(xlsxread): vendor Rust binary cloned from thptqg2017
Copies the xlsxread Rust crate from thptqg2017@8b4a755 (chore/xlsxread-rust).
Adds format_detect_2016 module with per-file column-layout auto-detection
mirroring detectFormat() in scripts/build-database.js (lines 63-87):
- separate-scores: SBD/HOTEN/TOAN... fixed columns (dhhanghai files)
- mapped: header-derived SOBAODANH|SBD + DIEM_THI dynamic indices
- default: positional 6-col layout (no header)
Extends ParsedRow with ten_cum_thi and gioi_tinh fields.
Adds SCORE_FIELDS_2016 (12 cols: tieng_duc/tieng_nhat; no khtn/khxh/tieng_nga).
Adds thptqg2016-data.toml config with 18-column schema and format_detection flag.
58 tests pass (50 unit + 8 integration), 0 failures.
* feat(build): wire build:db to xlsxread CLI
Replaces the Node.js build:db script with the xlsxread Rust binary.
Adds build:rust script for the cargo compile step in isolation.
* chore: remove deprecated build-database.js
Superseded by the xlsxread Rust binary. All 119 source files (4 .xls +
115 .xlsx) are now processed by xlsxread with per-file format detection.
* ci: build xlsxread before running database build job
Adds dtolnay/rust-toolchain@stable and Swatinem/rust-cache@v2 steps
before the xlsxread build and database generation steps. Node/pnpm
steps now follow the Rust build rather than preceding it.
* chore: add Rust build artifacts to .gitignore
* docs: update README build instructions for xlsxread pipeline
* chore(deps): drop xlsx and better-sqlite3 from package.json and lockfile
* feat(xlsxread): stage 0 scaffold with pinned deps and clap CLI skeleton
* feat(xlsxread): stage 1 reader — calamine sheet enumeration and header skip
* feat(xlsxread): stage 2 transform — to_ascii, score regex, validation with 38 unit tests
* feat(xlsxread): stage 3 writer — SQLite DDL, INSERT OR REPLACE, VACUUM, stats output
* feat(xlsxread): stage 4 audit — distinct SBD scan vs DB count, mirrors audit-row-counts.js output
* feat(xlsxread): stage 5 golden tests — in-process xlsx fixtures, 8 integration tests pass
* chore(xlsxread): commit Cargo.lock for reproducible Rust builds
* feat(build): wire root build:db scripts to xlsxread CLI
Replace node scripts/build-database*.js invocations with the Rust
xlsxread binary. Each build:db* script now calls `pnpm build:rust`
(cargo build --release) before invoking the xlsxread build subcommand
with the matching per-dataset config.
Drop xlsx and better-sqlite3 from devDependencies — no Node script
consumes them anymore. sql.js (runtime DB reader in the SPA) is
unaffected and remains in dependencies.
* ci: build xlsxread before running database build jobs
Add dtolnay/rust-toolchain@stable and Swatinem/rust-cache@v2
(workspaces: tools/xlsxread) for warm incremental Rust builds.
Replace the single `pnpm build:db:all` step with explicit xlsxread
invocations so CI doesn't call pnpm build:rust redundantly three times.
The binary is built once, then each of the three datasets is processed
in sequence.
* chore: remove deprecated xlsx-based build scripts
Delete scripts/build-database.js, build-database-old.js,
build-database-old2.js, build-lib.js, and audit-row-counts.js.
Functionality replaced by the xlsxread Rust CLI configured via
tools/xlsxread/configs/*.toml. History preserved in git; one-click
revert available via the chore/migration-backup-260519 branch.
* docs: update README build instructions for xlsxread pipeline
Replace Node.js + xlsx references with Rust + xlsxread workflow.
Update requirements (Node 24+, pnpm, Rust stable), quickstart, scripts
table, and project layout tree to reflect the current state after the
xlsx-based build scripts were removed.
* chore(deps): drop xlsx and better-sqlite3 from package.json and lockfile
Remove xlsx (SheetJS, vulnerable: GHSA-4r6h-8v6p-xvw6, GHSA-5pgg-2g8v-p4x9)
and better-sqlite3 from devDependencies. Both were only used by the now-deleted
Node build scripts. The Rust xlsxread CLI vendors SQLite via rusqlite-bundled;
no Node-side SQLite dependency is needed. `pnpm audit` returns clean.
Replace package-lock.json with pnpm-lock.yaml, add packageManager field,
and update GitHub Actions workflow to use pnpm/action-setup@v4 with
frozen-lockfile installs.
Replace package-lock.json with pnpm-lock.yaml, add packageManager field,
update package.json scripts from npm run to pnpm, and update GitHub Actions
workflow to use pnpm/action-setup@v4 with frozen-lockfile installs.
- admission-blocks.js: expand from 8 core blocks to all 49 official 2017
blocks (A00-A11, B00-B08, C00-C20, D01-D15) the DB schema can support.
Detail card now shows every qualifying block, sorted best-first
- custom-query: rewrite "Top 100 Long An" preset using UNION ALL +
ROW_NUMBER() OVER PARTITION BY so we can label each row with the
winning block code, not just the score
Computes each student's max admission-block sum across 49 official 2017
blocks (A00-A11, B00-B08, C00-C20, D01-D15, restricted to subjects
present in our DB — no Đức/Nhật languages). One row per student, sorted
by their personal best, top 100.
SQL note: SQLite's MAX(x,y,...) scalar returns NULL if ANY arg is NULL.
Each block expression is wrapped in COALESCE(..., -1); NULLIF(..., -1)
restores NULL only for students with zero computable blocks.
Old palette reused gold for both "Trung bình" and "Xuất sắc" — confusing.
Replaces with League of Legends / TFT rarity ladder where each rank gets
a distinct hue that also communicates relative rarity.
Tiers:
≤ 1 Điểm liệt (common) white / gray
< 5 Chưa đạt (uncommon) green
5-6.5 Trung bình (rare) blue
6.5-8 Khá (epic) purple
8-9 Giỏi (legendary) gold
9-10 Xuất sắc (prismatic) multi-color gradient
- admission-blocks.js: scoreTier now returns 6 keys (common..prismatic)
plus the "điểm liệt" tier (≤1) that didn't exist before
- student-detail.jsx: TIER_LEGEND updated to 6 entries with ranges
- App.css: --tier-{common,uncommon,rare,epic,legendary,prismatic}-*
tokens, light + dark variants each AA-contrast verified; applied to
.score-cell, .score-tile, .tier-legend-item
- App.jsx owns query state, syncs to URL via ?q= (works for SBD or name);
hydrates initial search from URL on DB ready
- Loading panel now shows DB size (~47 MB) and one-time download note
- Footer reports total student count from the live DB
- Keyboard shortcut '/' focuses the search input when not already typing
- StudentDetail: "Chia sẻ" button uses Web Share API or clipboard with a
formatted multi-line summary including a deep-link URL
- StudentDetail: inline tier legend (▽ ○ ◆ ★ ✦) with score ranges
- ScoreTable cells now use background tint matching detail tiles
- CustomQuery: presets grouped into 4 categories; auto-runs PRAGMA
table_info(student) on first DB ready so the tab opens with the schema
visible instead of a blank textarea
- SearchForm accepts controlled value prop for URL hydration
- Add Be Vietnam Pro web font for consistent Vietnamese diacritic rendering
- custom-query: two new presets filtered by Long An (SBD prefix 49):
Top 10 điểm Toán, Top 100 khối A (Toán+Lý+Hóa)
- search-form: show clickable example pills ("49008235" /
"Nguyễn Minh Tiến") next to the empty-state hint; clicking fills
the input and triggers the debounced live search
- App.css: .example-btn pill styling matching preset buttons
Restore the 54 corrected-export files that previously lived in
data/raw/update/ (removed in 37cf9df). Kept alongside data-old/ for
historical reference; not consumed by build pipeline.
Dataset update:
- Crawl all 63 .xls province files from baotintuc.vn CDN (original source)
- Old xlsx dataset moved to data-old/ for reference
- Net: +13,719 students (Hà Nội +7,275, HCM +6,445) — the old .xls → xlsx
conversion silently dropped rows beyond the 65,536 per-sheet cap
- Also removes 1 bogus header row that had leaked into the old DB
- 100% identical scores on the 847,348 SBDs present in both datasets
Build pipeline:
- build-database.js: iterate ALL sheets per workbook (fixes the overflow
loss) and accept .xls in addition to .xlsx
Audit tooling:
- scripts/crawl-baotintuc.js: idempotent 63-province downloader
- scripts/diff-datasets.js: compares two DBs by SBD set and per-column
score deltas
- Move 63 Excel files from data/raw/ to data/ (single flat dir)
- Remove all 53 files in data/raw/update/: verified identical SBD
coverage to raw/ (847349 rows either way), so they added no new
students — only potential score corrections that can be reintroduced
later if source is recovered
- Update build-database.js to read data/ directly
- Add scripts/audit-row-counts.js: compares source row count to DB row
count to verify zero-loss parsing
- Point check-duplicates.js at new data/ location
- Drop 10_LamDong_GNFT (1) and 2.BacKan_YQNX(1): identical row content to
siblings (Excel metadata differs but file size & sheet rows match)
- Add scripts/check-duplicates.js to detect byte-identical and row-identical
files across data/raw and data/raw/update
Add ho_ten_ascii column with normalized names (no diacritics, lowercase)
so users can search "nguyen van a" to find "NGUYỄN VĂN A".
- ASCII input searches against ho_ten_ascii column
- Vietnamese input searches both ho_ten and ho_ten_ascii
- Indexed for fast lookups
- Remove Gradle build, Java sources, Hibernate config, old database.sqlite
- Move Excel data files from src/main/resources/raw/ to data/raw/
- Move Vite+React app from web/ to project root
- Merge package.json into single root-level config
- Update build script paths and CI workflow accordingly