The range-request work is merged, so its plan and the two reports behind
it describe a decision already taken. What they held that the code does
not is why the alternatives were turned down, so that moves into
system-architecture under "Considered and not taken": chunked serverMode
(GitHub Pages' ten-minute cache TTL cancels the caching it buys),
sqlite-wasm-http (its shared cache needs headers Pages cannot send, and
swapping now would leave two unverified variables), and substring search
(no index can serve it).
Everything else those documents recorded — the measurements, the query
plans, the sizes — is already in docs/ and in the commits that made the
changes. Git keeps the originals either way.
setup-go v7 exports GOTOOLCHAIN=local, so Go no longer downloads the
toolchain a module asks for — whatever setup-go installed has to satisfy
it. `go-version: '1.26'` resolves to the runner manifest's patch, which
is 1.26.5, against modules requiring 1.26.6 since this morning's CVE
bump. The parser tests stopped on that.
go-version-file installs exactly what the modules declare, and removes
the fourth place a Go version was written down.
The runner warns that the node20 action runtime is deprecated. Every
action in the workflow was on it:
checkout v4 -> v7
setup-go v5 -> v7
setup-node v4 -> v7
upload-pages-artifact v3 -> v5
deploy-pages v4 -> v5
The two Pages actions move together because the artifact format is
shared between them.
None of the inputs this workflow passes changed across those majors —
the releases are ESM migrations and the runtime bump. `node-version: 24`
was already the Node the build runs on; this is about the runtime the
actions themselves execute in.
The footer printed the full article URL, which runs to 120 characters
and wrapped across three lines on a phone — the reason that footer
carried an overflow-wrap rule at all. It now shows who published the
article: "Trường Phổ thông DTNT tỉnh Bắc Ninh" for 2016 and "Báo Tin tức
và Dân tộc - TTXVN" for 2017.
The link still points at the article and its title attribute still
carries the URL, so nothing is hidden from anyone who wants to know
where it goes. The wrap rule is gone with the reason for it.
sourceName joins the fields every dataset must define, checked at module
load alongside the registry cross-check — with the types gone, a missing
one would otherwise render as an empty link.
Every .ts becomes .js, every lang="ts" becomes lang-less, and the type
declarations go with them: types.ts held nothing but types, so it is
deleted outright.
Tooling follows. typescript, svelte-check, typescript-eslint and
@types/sql.js are uninstalled; tsconfig.json becomes jsconfig.json, which
still extends the generated SvelteKit config so $lib and $app resolve in
an editor; `npm run lint` is now ESLint alone, and CI's comment about it
covering the type check goes too.
What this gives up, stated plainly: a mistyped column name like
row.nguvan used to fail the build and now renders blank, and the
datasets.json-to-CONTENT cross-check is back to throwing at module load
rather than at compile time. The runtime guard for the latter is still
there and still throws loudly.
Two mechanical notes. The svelte/no-navigation-without-resolve rule
started flagging the footer's source link, which points at an off-site
article — without type information the rule can no longer tell an
external URL from a route, so that one line carries a disable comment.
And Vitest's include pattern had to follow the tests to .js.
Lint, 25 tests and the build all pass.
The whole design rests on a claim nobody has measured: that a lookup
fetches a few hundred KB rather than the file. Every query now reports
what it actually cost.
[httpvfs] name "nguyen buu loc" seeking "buu" — 14 request(s), 78 KB,
212 ms · session 31 request(s), 180 KB of 302.4 MB
The numbers come from the worker's getStats() rather than its bytesRead
counter: bytesRead is the budget accumulator and resets itself to zero
the moment a query trips the ceiling, so it would under-report exactly
when the number matters. getStats() reports per-file totals including
what the read heads prefetched, which is the honest figure — prefetch
overfetch is invisible to a page count.
Opening a database logs too, since the header and schema pages are read
before any query and would otherwise inflate the first search.
The SQL tab shows the same pair on screen: this query, then the session.
The browser fetches this file one page per HTTP request, so the page size
is the granularity of every read. At SQLite's 4 KiB default a row reached
by an index seek dragged 4 KB across the network; at 1 KiB it drags 1 KB.
A name search returns up to 100 scattered rows, so its row fetches fall
from about 400 KB to about 100 KB.
Measured on the rebuilt 2016 file: 6.3 rows share a page where 27 did.
The index walks are sequential and unaffected in bytes — the library's
read-ahead already collapses those into few requests.
Cost is 4% file size: 2016 288.6 -> 302.4 MB, 2017 237.7 -> 247.3 MB,
the site 528 -> 552 MB against the 1 GB GitHub Pages limit. Both
sql.js-httpvfs and sqlite-wasm-http recommend this page size.
The PRAGMA has to run before the DDL, since a page size is fixed once a
table exists, and requestChunkSize on the client has to match or every
page read spans two requests.
Row counts unchanged and through the assembler guards; query plans
re-checked and still index-driven on the rebuilt files.
The browser downloaded 45 MB of gzipped SQLite before it could answer
anything. Now sql.js-httpvfs asks for the pages a query touches and the
databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip
stream is not a byte range of a database.
That only works if every query the site issues is index-driven, and
measured against the real 2016 file, most were not:
so_bao_danh = ? SEARCH via PK ~20 KB
ho_ten_ascii LIKE '%x%' SCAN 127 MB
ho_ten_ascii LIKE 'x%' SCAN 127 MB
COUNT(*) covering index scan 20 MB
ORDER BY toan DESC LIMIT 10 SCAN + temp b-tree 127 MB
Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE
index; a range comparison does use the index. So the schema changed to
suit the access pattern rather than the search changing to suit the
schema.
name_word holds one row per word of each name, WITHOUT ROWID so the
table is the index, carrying ho_ten_ascii so a multi-word query is
resolved inside a single b-tree. name_word_freq says which word of a
query is rarest — the vocabulary is 4,397 words across 2.87M entries, so
"buu loc" seeks on 287 entries rather than walking the 300,000 that
"thi" would. Searching by any word of a name survives, at a few hundred
KB a query.
idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either.
Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL
presets off a full scan. The footer's candidate count now comes from
datasets.json instead of COUNT(*).
2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site
is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged.
The SQL tab is the one place a user can still write a query that reads
the whole table, so it asks before it opens, runs under a byte budget
that stops a runaway query, and shows what each query actually fetched.
Verified: row counts through the assembler guards, every app query
index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206
with a correct Content-Range. Not verified in a browser — this machine
has none — and the library refuses to open a file the host compresses,
so the deployed response headers need a look.
The app was a single-entry React SPA: one index.html that the assembler
copied to every dataset path, with a hand-rolled router resolving the
dataset from window.location and the title patched in at runtime because
one file had to serve every route.
SvelteKit prerenders a real page per route instead. The entry generator
in routes/[dataset]/+page.ts reads the same datasets.json the assembler
does, so the set of pages and the set of databases cannot drift apart,
and each dataset page ships its own <title> and description.
Everything framework-free moved across unchanged in behaviour and became
typed: the admission blocks, subject list, query classifier and SQL
presets. `Student` now mirrors the 22-column table, so a mistyped column
name fails the build rather than rendering blank.
Tailwind replaces the stylesheet. Theme tokens are CSS variables, which
keeps dark mode a single block of overrides rather than a `dark:` variant
on every class. The score tiers stay hand-written CSS: the class is
chosen at runtime from a score, and no utility generator can see that.
Adds the tests the frontend never had, over the three modules where a
silent wrong answer is possible — most importantly that toAscii here
folds exactly as ToAscii does in the parser, which is what makes
accent-insensitive search find anything.
The assembler stops copying index.html per dataset and checks that the
build prerendered each one instead. The database presence, raw-artifact
and idempotence guards are untouched.
Verified: 25 tests, ESLint and svelte-check clean, and a full
`assemble site` producing /thptqg/, /thptqg/2016/, /thptqg/2017/ and
404.html with absolute asset URLs. Not verified in a browser — this
machine has none.
govulncheck fails the pipeline on GO-2026-6088: encoding/xml decodes
without a recursion depth guard, reachable from excelize's OpenFile,
GetRows and GetSheetList and from buildCRFixups directly. The parser is
fed spreadsheets downloaded over the network by the crawler, so the path
is real.
Raising the go directive to 1.26.6 in all three modules puts the fix
below every build rather than leaving it to whichever patch release the
runner happens to install.
govulncheck is clean on all three modules, and every suite passes on the
new toolchain — including the reader fidelity sweep, which matters here
because buildCRFixups depends on how encoding/xml normalises line endings.
The 2016 exam numbers are nine digits or a letter-prefixed cluster code,
so the eight-digit "17006021" matched nothing and query-mode.js read it
as a 2017 number. TKG002747 / Nguyễn Bửu Lộc is a real row.
Four of the 119 files in data/2016 publish one score column per subject
instead of a DIEM_THI sentence, and none of them was being read correctly.
The ĐH Công nghiệp Thực phẩm file puts a three-row ministry title block
above its header, so no header was recognised and the positional fallback
shifted every column by one: the serial number became so_bao_danh, the
exam number became ho_ten, the name became ngay_sinh, and the national ID
became the score cell. All 7,833 rows were unusable. The three ĐH Cần Thơ
files name an SBD column but no DIEM_THI, so they fell to the same
fallback: surname into ngay_sinh, given name into ten_cum_thi, birth date
into the score cell, and 12,152 candidates with no scores at all.
Both are now read by FormatSubjectColumns, which resolves identity and one
column per subject from the header. The header is searched for in the
first five rows, so a title block no longer hides it.
The Cần Thơ score columns are numbered rather than named. They follow the
order the exam was sat — each morning an essay paper, each afternoon a
multiple-choice one — which is what identifies them: columns 1/3/5/7
quantise to 0.25 and 2/4/6/8 do not, and each column's mean lands within
0.5 of the same subject's mean across the rest of the dataset. The
foreign language is filed under the subject its N1..N6 code names.
Gender now accepts the 0/1 encoding those files use: of the rows marked
1, 53% carry "Thị" in the name against 1% of those marked 0. Birth dates
in the compact ddmmyy form are expanded so the column holds one format.
A score of 0 is stored rather than dropped, recovering 302 real scores
that a JavaScript falsy check had been turning into NULL.
Row count falls by one, to 877,460: the removed row is the title line
"ĐƠN VỊ: / TRƯỜNG ĐẠI HỌC CÔNG NGHIỆP THỰC PHẨM TP. HỒ CHÍ MINH", which
had been stored as a student. The dataset has no duplicate exam numbers;
the three rows previously described as collapsing were that same file's
title and header lines being counted and then rejected.
Also drops behaviour that existed only to match the parser this one
replaced: the inert "SINH " header token, the untrimmed diem_thi cell, an
unreachable blank-row branch, and a cross-check test against a database
that can no longer exist. None of them changes output.
Verified by rebuilding both datasets: 877,460 and 861,068 rows, both
artifacts through the assembler's row and size guards, and the reader
fidelity suite unchanged across all 182 files.
Routine verification is parser and assembler. The crawler is not part of the
build — data/<id>/ is committed and a crawl only refreshes it by hand — so its
tests do not gate ordinary changes, though CI still runs all three.
Also records the constraints that read as mistakes and get "fixed": the
verbatim data/2016/ filenames that decide duplicate-SBD resolution, the frozen
reader oracle, the literal combining-mark range ToAscii must share with the
site, the absolute Vite base, and dbSizeMb doubling as a truncation guard.
Comments across the tree justified the code by pointing at a Rust
implementation that is no longer in the repository, citing files and line
numbers (config.rs:132, schema.rs:26-54, reader.rs:42) that cannot be
opened, plus crates and datasets that are equally gone. A reader could not
check any of it.
Every invariant those comments carried is kept and restated so it stands on
its own: the bytewise sort that decides which row survives a duplicate exam
number, the literal U+0300..U+036F range that must match the site's toAscii,
the trailing space in "SINH ", the BIFF and shared-string corrections, the
VACUUM-after-COMMIT rule, the deploy-from-main guard.
The reader's contract is now anchored to the frozen oracle in
parser/testdata, which still exists and is still checked, rather than to the
tool that originally produced it.
TestDDLMatchesRust becomes TestDDLIsFrozen: it compares against a copy of
the DDL inside the test and never read schema.rs, so both the name and the
failure message were misleading.
ToAscii no longer claims the d-replacement must precede lowercasing. Both
cases map to 'd' and ToLower runs last, so the order has no effect.
The site credited 2016 to Bộ GD&ĐT, which those files were never fetched
from. Both datasets come from published articles: 2016 from an aggregator
on dtnt.bacninh.edu.vn listing one spreadsheet per exam cluster, 2017 from
baotintuc.vn. The README, the architecture table and the web footer all
repeated the ministry claim.
The footer now shows each dataset's full article URL as a link rather than
a bare host, so the citation can be checked. That needs overflow-wrap on
the footer: the 2016 URL is 110 characters with no break opportunity and
would otherwise scroll the page sideways on a phone.
Also records that a full 2016 crawl has been run successfully. The host was
marked unconfirmed and data/2016/ described as the only recoverable copy;
both datasets are now rebuildable from source.
The pipeline is now Go outside web/. differential-parity.mjs becomes
assembler/internal/verify, reachable as `assemble verify A B`. The port fixed a
real weakness: the JavaScript hashed each row's fields joined bare, so a value
shifted across a column boundary produced the same digest. A test now pins that.
The hub still rendered "Phiên bản cũ của trang 2017" above a permanently empty
list — it split datasets on id.includes("old"), and both such datasets are gone.
The heading and the filter are removed. index.html titled every page "THPT QG
2017", including 2016 and the hub, because one file is copied to every route;
the static title is now neutral and the app sets the dataset's own.
Dead code removed: the isOld2/containsOld branches in the stats block, which
only 2017-old2 could ever reach; SUBJECT_LABELS, DATASET_IDS and the unread
`short` subject field; an unused vite.svg and a favicon link to a file that
never existed; two unused CSS rules and --shadow-sm; site.Paths.Root.
Corrected comments that were confidently wrong rather than merely stale: the
reader claimed to be row-streaming when both implementations decode the whole
workbook into memory first, and the fidelity oracle still spoke of 299 input
files when it covers 182. Candidate counts in the hub now derive from
datasets.json instead of being written a second time as prose.
plans/ is emptied. The parity report it held was cited by docs/data-pipeline.md,
so the evidence that the recovered foreign-language scores are real — not the
citation, the four arguments themselves — is now inline there.
Verified: 2017 rebuilt after the writer change hashes identically to the build
before it.
Both plans are done and merged, and their contents now describe a tree that no
longer exists — go-parser/, four datasets, npm entry points, a Rust crate. Git
history keeps them.
The parity evidence stays: docs/data-pipeline.md cites parser-parity-result.md
for the recovered foreign-language scores, so that report and the two JSON
snapshots behind it remain, now marked as an archived record.
The repository now reads as the pipeline it is: crawler fetches, parser
converts, assembler verifies and publishes, with data/ and web/ as the stores
they hand work through. go-parser is renamed parser now that there is no other.
The assembler replaces build-db.js and assemble-site.js. It compiles the
parser, builds and verifies each database, compresses it, runs the Vite build
and assembles _site — one command, and the only place that knows the order.
It also closes a real hole: nothing previously asserted that a database reached
the site. An empty staging directory assembled happily, so every page rendered,
every query 404d and CI stayed green. The row-count and size guards could not
catch that, since they only run when a database was built at all.
Removing Node from the root forced the dataset list out of web/src/datasets.js,
which the assembler cannot import. datasets.json is now the registry both sides
read — JSON because Go and the browser both parse it without a dependency —
while presentation stays in the web app, keyed by id and cross-checked against
the registry so a half-added dataset fails instead of half-working.
Guards verified by making each one fail: a missing database, and an expected
row count one higher than the truth.
Both sources carried their download links as a hardcoded array, which is not a
crawl: the lists could drift from what the articles actually published, and
nothing would say so. A source now names the article and how to name what it
finds there, and internal/article reads the links out of that page at run time.
2016 takes its filenames straight from the URL. 2017 cannot — the CDN names are
inconsistent (Angiang.xls, 1BaRiaVungTau.xls, 23HaiPhong.xls) — so it derives
them from the province in the link text, transliterated to ASCII the same way
go-parser builds ho_ten_ascii.
Filenames stay load-bearing: go-parser sorts inputs bytewise and inserts
last-wins, so they decide which row survives a duplicate exam number. Saved
copies of both articles are committed as fixtures, and a test asserts that
reading them and applying each naming rule reproduces data/<id> exactly, in both
directions. Resolve also rejects a page that yields the wrong number of links or
two links that would write the same file, since either silently costs the
dataset files that only the row-count guard would notice afterwards.
Verified against the live 2017 article: a from-scratch crawl of all 63 files
leaves the committed data unchanged.
Move the frontend into web/, the repo's only npm workspace, and replace the
JS crawler with a Go module covering both remaining datasets. The crawler
writes to a .part file and renames on completion: writing straight to the
destination left truncated files that the skip-if-present check would then
skip forever.
Remove the 2017-old and 2017-old2 datasets. They were successive publications
of the same exam, kept side by side so the disagreement stayed inspectable;
the current 2017 supersedes them and they remain in git history.
Recover the 2016 crawler source from the Internet Archive's copy of the
aggregator article, whose original host no longer resolves. All 119 filenames
are verified against data/2016 in both directions, but no archive captured the
spreadsheets themselves, so the host still serving them is unconfirmed and
data/2016 remains the only confirmed copy.
Filenames are load-bearing throughout: go-parser sorts inputs bytewise and
inserts last-wins, so they decide which row survives a duplicate exam number.
The Go parser has matched the Rust one field-by-field across all four datasets,
so the Rust crate is retired and CI builds the Go binary instead.
crawl-baotintuc.js moves to go-parser/scripts/ — it is the only mechanism for
refreshing data/2017 and has a documented runbook. check-duplicates.js and
diff-datasets.js are dropped: both had been broken since before the repo was
unified, and neither had a caller. db-stats.js and verify-parity.js are dropped
as superseded by differential-parity.mjs, which compares more and cannot
silently skip a dataset.
The reader-fidelity oracle is kept and marked frozen. It was produced by the
Rust reader so it can no longer be regenerated, but it still fails if any single
cell of any of the 299 input files reads differently.
Source comments cite the original Rust by file and line; those paths resolve at
tag pre-go-parser-removal, recorded in go-parser/README.md.
Adds go-parser/, a Go reimplementation of the xlsxread parser, verified
byte-for-byte against the Rust original before any cutover.
Reader fidelity is exact across all 299 input files: the canonical cell dump
of every sheet matches calamine's, locked in as a test against a committed
hash oracle. Reaching that required replacing extrame/xls, which corrupted
69% of cells and dropped a further 28% on the BIFF corpus, with pbnjay/grate;
correcting excelize's number-format application and trailing-cell trimming;
restoring carriage returns that XML line-ending normalisation strips from
2,233 ten_cum_thi values; and gating numeric re-rendering on cell type so
shared strings that merely look numeric keep their leading zeros.
The differential gate compares both parsers over all four datasets:
3,265,641 rows with identical full-table SHA-256, identical per-column
non-NULL counts, identical schema metadata and identical stdout.
Config moves from TOML to YAML for both parsers, so they keep reading the
same files and the gate stays meaningful. Verified by rebuilding 2016 and
2017-old2 with Rust under the new configs and matching the recorded counts.
build-db.js now refuses to publish a database whose row count does not match
the known figure, closing a path where an under-producing parser could ship a
truncated public dataset with green CI. The deploy workflow gains a
pull_request trigger and guards deploy to main, so branch verification can no
longer publish to production.
The docs still described two standalone projects with separate frontends,
separate parsers and three Vite variants. Rewritten around what the repo now
is, merging both projects' copies rather than keeping one and discarding the
other — deployment-guide.md and system-architecture.md existed in both and
documented different pipelines.
project-overview.md goal, scope, constraints, the four datasets, history
system-architecture.md data flow, canonical schema, routing, how one
frontend serves both exam years without branching
data-pipeline.md per-dataset Excel formats, the three 2016 layouts,
overflow-sheet gotcha, expected row counts
deployment-guide.md the single-build workflow, adding a dataset,
why no uncompressed database can ship
Records the release gate in plans/reports/parser-parity-result.md: row counts,
all 18 pre-existing per-column non-NULL counts and the deterministic student
samples are identical across all four datasets, against databases decompressed
from the exact bytes the pipeline publishes. The 1,691 recovered
foreign-language scores are documented with the evidence they are real.
Also drops a machine-specific absolute path from a comment in
format_detect_2016.rs. The build-database.js citations there are kept: that
file no longer exists in this repo, but the references explain why several
parsing rules look arbitrary.
The four build variants existed only because the app could not resolve its own
dataset. Now that it can, vite.config.js is a single build with an absolute
base and no VARIANT switching, and the workflow compiles Rust once, installs
Node dependencies once, and builds the site once — it previously built the same
Rust crate twice and ran two separate pnpm installs.
Because base is absolute, the emitted index.html references /thptqg/assets/...
regardless of where it is served from, so the same file works as an entry point
at any depth. scripts/assemble-site.js copies it to each dataset path and to
the two legacy nested URLs, giving a real static file at every published route.
That is what removes the need for an SPA 404-fallback redirect, which would
otherwise have rewritten URLs and interfered with the ?q= deep links.
Generated databases move to a gitignored .build/public, which Vite consumes as
its publicDir. build-db.js gzips without -k, and the assemble step then refuses
to finish if any uncompressed database artefact reached the output — .db,
.db-journal, .db-wal or .db-shm. Previously a raw 100+ MB database was written
into the source tree and deleted afterwards by an rm in the workflow, so
shipping one was a missing cleanup step away.
Both build scripts import DATASET_IDS from src/datasets.js rather than
repeating the dataset list in workflow shell, so the four ids are declared in
exactly one place across the frontend, the database build and the assembly.
Verified by running the full pipeline locally and serving the artifact over
HTTP: all nine routes return 200, every entry point is byte-identical, asset
references are absolute, and the databases are fetchable. The guard was tested
by injecting a raw .db and .db-journal into the build, which failed the
assembly as intended.
eslint-plugin-react-hooks v7 flags setState inside an effect body. All four
occurrences predate this refactor and are one-shot initialisation or external
sync, not the cascading-render pattern the rule targets:
- hydrating a ?q= deep link once the database is ready
- reading the candidate count for the footer
- auto-running the schema preset when the SQL tab first opens
- mirroring the parent-owned query into SearchForm for deep links and clear
Each gets a scoped disable with a comment explaining why it is correct here,
rather than a blanket rule change. Restructuring these properly would mean
reworking deep-link hydration and auto-search with no test harness to catch a
regression; that is worth doing separately, not inside this refactor.
npm run lint is now clean.
The 2017 frontend becomes the only frontend. It now resolves which dataset to
show from the URL, so the four dataset views and the /thptqg/ landing page all
come from a single build instead of four builds plus a hand-written static
index.html.
Routing is an exact match on one path segment, because the segment is the
dataset id: /thptqg/2017-old/ -> "2017-old". The nested form this replaces
(/thptqg/2017/old/) would have needed longest-prefix matching, since it also
starts with /thptqg/2017/. Previously published nested URLs are rewritten to
their flat equivalent with the query string intact, so ?q= deep links keep
working.
Restores what the 2016 site had and the 2017 site lacked:
- Cụm thi and Giới tính columns, in the table and on the detail card
- exam-ID lookup for letter-prefixed số báo danh
That second one would have been a severe regression. 616,593 of 877,461
candidates in the 2016 dataset — 70.3% — have an SBD like BAL000001, and the
2017 app matched /^\d+$/ only, so exam-ID search would have silently failed for
most of that year. Detection now lives in lib/query-mode.js, shared by App and
SearchForm, which had already drifted apart on exactly this rule. Letter
prefixes are upper-cased before lookup so "bal000001" resolves.
Neither dataset needs a conditional in a component. Columns where every row is
NULL are hidden, so 2016 rows surface Cụm thi / Đức / Nhật while 2017 rows
surface KHTN / KHXH / GDCD / Nga. computeBlocks() likewise skips blocks with a
missing subject, so one admission-block list covers both years — D05 and D06
are added now that German and Japanese scores exist.
Per-dataset content moves into the registry: title, source, database size,
search examples and SQL presets. The 2016 presets are lifted from the deleted
2016 component; the Long An (49xxx) queries stay on 2017 only, where that SBD
prefix means something.
The subject list, previously maintained separately in score-table.jsx and
student-detail.jsx, is now declared once in lib/subjects.js mirroring
SCORE_FIELDS in the parser.
The repo held two near-duplicate projects. 2016/ and 2017/ each carried their
own React frontend, their own copy of the same Rust crate, and their own
package manager setup. 2016/tools/sync-from-thptqg2017.sh existed purely to
copy the parser source between them.
New layout:
index.html + src/ the 2017 frontend, now the only one
data/<id>/ 2016, 2017, 2017-old, 2017-old2
parser/ the single Rust crate, configs renamed to <id>.toml
docs/ both projects' docs, 2016 copies suffixed -2016-legacy
pending the merge pass
<id> is now one identifier end to end: data/<id>/ feeds parser/configs/<id>.toml
and produces db/<id>.db.gz.
pnpm gives way to npm. pnpm-workspace.yaml existed only to whitelist
better-sqlite3's native build, which npm permits by default, so it has no
equivalent and is simply gone. Lockfiles cannot be converted; package-lock.json
is generated fresh. The migration direction is safe — pnpm's strict layout
forbids phantom dependencies, so anything that resolved under pnpm resolves
under npm's flat tree.
Adds parser/scripts/build-db.js and src/datasets.js: the four dataset IDs are
declared once and read by both the build tooling and (from the next phase) the
frontend.
Follow-on fixes the move made necessary:
- eslint's Node-globals override pointed at scripts/, now parser/scripts/
- crawl-baotintuc.js wrote to <root>/data, now data/2017
- golden tests loaded configs by their old thptqg*-data.toml names
Drops the #[ignore]d Rust-vs-Node golden test. It shelled out to pnpm to run
scripts/build-database.js, a file removed when the parser was ported to Rust,
so it could never pass. check-duplicates.js and diff-datasets.js were already
broken before this change and are annotated as such rather than half-fixed.
63 Rust tests pass and clippy is clean from the new location.
cargo clippy --all-targets -- -D warnings failed with 8 errors before the schema
refactor and was never part of CI. Clearing them so the gate is meaningful from
here on.
All behaviour-preserving: slice contains() over iter().any(), is_multiple_of()
over a modulo, a needless borrow in a test helper, and a doc-list indent.
process_mapped_row keeps its 8 parameters behind an allow attribute — the
signature mirrors the JS column map one-for-one, and bundling the indices into
a struct would hide that correspondence.
Kept separate from the schema change so that diff stays readable.
The 2016 and 2017 parsers were the same crate with divergent SQL. DDL, INSERT
and the subject regex table lived in each dataset's TOML config, so four copies
had to be kept in step by hand — which is how the two schemas drifted apart.
Move all of it into src/schema.rs as a single 22-column definition: 6 identity
columns (adding ten_cum_thi and gioi_tinh, previously 2016-only) and 16 subject
columns (the union of both exam years). Columns a dataset carries no data for
bind NULL, costing ~1 byte per row.
This collapses writer.rs to one insert path: insert_row_2016, SCORE_FIELDS_2016,
SCORE_FIELDS_2017 and the SCORE_FIELDS alias all go away. Configs shrink to the
parse rules that genuinely vary per dataset — sheet mode, column indices, SBD
validation, header tokens, blank-row stripping — and carry no SQL at all.
Config parsing now denies unknown fields, so a leftover [schema] block fails
loudly instead of looking effective while schema.rs drives the build.
idx_ten_cum_thi is partial, so it holds zero entries on the three 2017 datasets
where the column is always NULL.
Adds db-stats.js and verify-parity.js to prove no data moved: row counts,
per-column non-NULL counts and a deterministic student sample are compared
against databases built from the previous code.
Verified across all four datasets — row counts, every pre-existing column count,
and all sampled students are identical. Database size grows 1.4-2.2%.
Each year keeps its full pipeline under its directory; vite bases move to
/thptqg/<year>/ (2017 keeps its old and old2 generations), one deploy
workflow builds both years' databases and bundles behind a root index.
* feat(xlsxread): vendor Rust binary cloned from thptqg2017
Copies the xlsxread Rust crate from thptqg2017@8b4a755 (chore/xlsxread-rust).
Adds format_detect_2016 module with per-file column-layout auto-detection
mirroring detectFormat() in scripts/build-database.js (lines 63-87):
- separate-scores: SBD/HOTEN/TOAN... fixed columns (dhhanghai files)
- mapped: header-derived SOBAODANH|SBD + DIEM_THI dynamic indices
- default: positional 6-col layout (no header)
Extends ParsedRow with ten_cum_thi and gioi_tinh fields.
Adds SCORE_FIELDS_2016 (12 cols: tieng_duc/tieng_nhat; no khtn/khxh/tieng_nga).
Adds thptqg2016-data.toml config with 18-column schema and format_detection flag.
58 tests pass (50 unit + 8 integration), 0 failures.
* feat(build): wire build:db to xlsxread CLI
Replaces the Node.js build:db script with the xlsxread Rust binary.
Adds build:rust script for the cargo compile step in isolation.
* chore: remove deprecated build-database.js
Superseded by the xlsxread Rust binary. All 119 source files (4 .xls +
115 .xlsx) are now processed by xlsxread with per-file format detection.
* ci: build xlsxread before running database build job
Adds dtolnay/rust-toolchain@stable and Swatinem/rust-cache@v2 steps
before the xlsxread build and database generation steps. Node/pnpm
steps now follow the Rust build rather than preceding it.
* chore: add Rust build artifacts to .gitignore
* docs: update README build instructions for xlsxread pipeline
* chore(deps): drop xlsx and better-sqlite3 from package.json and lockfile
* feat(xlsxread): stage 0 scaffold with pinned deps and clap CLI skeleton
* feat(xlsxread): stage 1 reader — calamine sheet enumeration and header skip
* feat(xlsxread): stage 2 transform — to_ascii, score regex, validation with 38 unit tests
* feat(xlsxread): stage 3 writer — SQLite DDL, INSERT OR REPLACE, VACUUM, stats output
* feat(xlsxread): stage 4 audit — distinct SBD scan vs DB count, mirrors audit-row-counts.js output
* feat(xlsxread): stage 5 golden tests — in-process xlsx fixtures, 8 integration tests pass
* chore(xlsxread): commit Cargo.lock for reproducible Rust builds
* feat(build): wire root build:db scripts to xlsxread CLI
Replace node scripts/build-database*.js invocations with the Rust
xlsxread binary. Each build:db* script now calls `pnpm build:rust`
(cargo build --release) before invoking the xlsxread build subcommand
with the matching per-dataset config.
Drop xlsx and better-sqlite3 from devDependencies — no Node script
consumes them anymore. sql.js (runtime DB reader in the SPA) is
unaffected and remains in dependencies.
* ci: build xlsxread before running database build jobs
Add dtolnay/rust-toolchain@stable and Swatinem/rust-cache@v2
(workspaces: tools/xlsxread) for warm incremental Rust builds.
Replace the single `pnpm build:db:all` step with explicit xlsxread
invocations so CI doesn't call pnpm build:rust redundantly three times.
The binary is built once, then each of the three datasets is processed
in sequence.
* chore: remove deprecated xlsx-based build scripts
Delete scripts/build-database.js, build-database-old.js,
build-database-old2.js, build-lib.js, and audit-row-counts.js.
Functionality replaced by the xlsxread Rust CLI configured via
tools/xlsxread/configs/*.toml. History preserved in git; one-click
revert available via the chore/migration-backup-260519 branch.
* docs: update README build instructions for xlsxread pipeline
Replace Node.js + xlsx references with Rust + xlsxread workflow.
Update requirements (Node 24+, pnpm, Rust stable), quickstart, scripts
table, and project layout tree to reflect the current state after the
xlsx-based build scripts were removed.
* chore(deps): drop xlsx and better-sqlite3 from package.json and lockfile
Remove xlsx (SheetJS, vulnerable: GHSA-4r6h-8v6p-xvw6, GHSA-5pgg-2g8v-p4x9)
and better-sqlite3 from devDependencies. Both were only used by the now-deleted
Node build scripts. The Rust xlsxread CLI vendors SQLite via rusqlite-bundled;
no Node-side SQLite dependency is needed. `pnpm audit` returns clean.
Replace package-lock.json with pnpm-lock.yaml, add packageManager field,
and update GitHub Actions workflow to use pnpm/action-setup@v4 with
frozen-lockfile installs.
Replace package-lock.json with pnpm-lock.yaml, add packageManager field,
update package.json scripts from npm run to pnpm, and update GitHub Actions
workflow to use pnpm/action-setup@v4 with frozen-lockfile installs.
- admission-blocks.js: expand from 8 core blocks to all 49 official 2017
blocks (A00-A11, B00-B08, C00-C20, D01-D15) the DB schema can support.
Detail card now shows every qualifying block, sorted best-first
- custom-query: rewrite "Top 100 Long An" preset using UNION ALL +
ROW_NUMBER() OVER PARTITION BY so we can label each row with the
winning block code, not just the score
Computes each student's max admission-block sum across 49 official 2017
blocks (A00-A11, B00-B08, C00-C20, D01-D15, restricted to subjects
present in our DB — no Đức/Nhật languages). One row per student, sorted
by their personal best, top 100.
SQL note: SQLite's MAX(x,y,...) scalar returns NULL if ANY arg is NULL.
Each block expression is wrapped in COALESCE(..., -1); NULLIF(..., -1)
restores NULL only for students with zero computable blocks.
Old palette reused gold for both "Trung bình" and "Xuất sắc" — confusing.
Replaces with League of Legends / TFT rarity ladder where each rank gets
a distinct hue that also communicates relative rarity.
Tiers:
≤ 1 Điểm liệt (common) white / gray
< 5 Chưa đạt (uncommon) green
5-6.5 Trung bình (rare) blue
6.5-8 Khá (epic) purple
8-9 Giỏi (legendary) gold
9-10 Xuất sắc (prismatic) multi-color gradient
- admission-blocks.js: scoreTier now returns 6 keys (common..prismatic)
plus the "điểm liệt" tier (≤1) that didn't exist before
- student-detail.jsx: TIER_LEGEND updated to 6 entries with ranges
- App.css: --tier-{common,uncommon,rare,epic,legendary,prismatic}-*
tokens, light + dark variants each AA-contrast verified; applied to
.score-cell, .score-tile, .tier-legend-item
- App.jsx owns query state, syncs to URL via ?q= (works for SBD or name);
hydrates initial search from URL on DB ready
- Loading panel now shows DB size (~47 MB) and one-time download note
- Footer reports total student count from the live DB
- Keyboard shortcut '/' focuses the search input when not already typing
- StudentDetail: "Chia sẻ" button uses Web Share API or clipboard with a
formatted multi-line summary including a deep-link URL
- StudentDetail: inline tier legend (▽ ○ ◆ ★ ✦) with score ranges
- ScoreTable cells now use background tint matching detail tiles
- CustomQuery: presets grouped into 4 categories; auto-runs PRAGMA
table_info(student) on first DB ready so the tab opens with the schema
visible instead of a blank textarea
- SearchForm accepts controlled value prop for URL hydration
- Add Be Vietnam Pro web font for consistent Vietnamese diacritic rendering
- custom-query: two new presets filtered by Long An (SBD prefix 49):
Top 10 điểm Toán, Top 100 khối A (Toán+Lý+Hóa)
- search-form: show clickable example pills ("49008235" /
"Nguyễn Minh Tiến") next to the empty-state hint; clicking fills
the input and triggers the debounced live search
- App.css: .example-btn pill styling matching preset buttons
Restore the 54 corrected-export files that previously lived in
data/raw/update/ (removed in 37cf9df). Kept alongside data-old/ for
historical reference; not consumed by build pipeline.
Dataset update:
- Crawl all 63 .xls province files from baotintuc.vn CDN (original source)
- Old xlsx dataset moved to data-old/ for reference
- Net: +13,719 students (Hà Nội +7,275, HCM +6,445) — the old .xls → xlsx
conversion silently dropped rows beyond the 65,536 per-sheet cap
- Also removes 1 bogus header row that had leaked into the old DB
- 100% identical scores on the 847,348 SBDs present in both datasets
Build pipeline:
- build-database.js: iterate ALL sheets per workbook (fixes the overflow
loss) and accept .xls in addition to .xlsx
Audit tooling:
- scripts/crawl-baotintuc.js: idempotent 63-province downloader
- scripts/diff-datasets.js: compares two DBs by SBD set and per-column
score deltas
- Move 63 Excel files from data/raw/ to data/ (single flat dir)
- Remove all 53 files in data/raw/update/: verified identical SBD
coverage to raw/ (847349 rows either way), so they added no new
students — only potential score corrections that can be reintroduced
later if source is recovered
- Update build-database.js to read data/ directly
- Add scripts/audit-row-counts.js: compares source row count to DB row
count to verify zero-loss parsing
- Point check-duplicates.js at new data/ location
- Drop 10_LamDong_GNFT (1) and 2.BacKan_YQNX(1): identical row content to
siblings (Excel metadata differs but file size & sheet rows match)
- Add scripts/check-duplicates.js to detect byte-identical and row-identical
files across data/raw and data/raw/update
Add ho_ten_ascii column with normalized names (no diacritics, lowercase)
so users can search "nguyen van a" to find "NGUYỄN VĂN A".
- ASCII input searches against ho_ten_ascii column
- Vietnamese input searches both ho_ten and ho_ten_ascii
- Indexed for fast lookups
- Remove Gradle build, Java sources, Hibernate config, old database.sqlite
- Move Excel data files from src/main/resources/raw/ to data/raw/
- Move Vite+React app from web/ to project root
- Merge package.json into single root-level config
- Update build script paths and CI workflow accordingly
- Node script parses 119 Excel files into SQLite (847K students)
- Vite + React frontend with sql.js for client-side querying
- Search by exam ID (số báo danh) or student name
- Gzipped DB (36MB) with download progress bar
- GitHub Actions workflow for GitHub Pages deployment