The browser fetches this file one page per HTTP request, so the page size
is the granularity of every read. At SQLite's 4 KiB default a row reached
by an index seek dragged 4 KB across the network; at 1 KiB it drags 1 KB.
A name search returns up to 100 scattered rows, so its row fetches fall
from about 400 KB to about 100 KB.
Measured on the rebuilt 2016 file: 6.3 rows share a page where 27 did.
The index walks are sequential and unaffected in bytes — the library's
read-ahead already collapses those into few requests.
Cost is 4% file size: 2016 288.6 -> 302.4 MB, 2017 237.7 -> 247.3 MB,
the site 528 -> 552 MB against the 1 GB GitHub Pages limit. Both
sql.js-httpvfs and sqlite-wasm-http recommend this page size.
The PRAGMA has to run before the DDL, since a page size is fixed once a
table exists, and requestChunkSize on the client has to match or every
page read spans two requests.
Row counts unchanged and through the assembler guards; query plans
re-checked and still index-driven on the rebuilt files.
The browser downloaded 45 MB of gzipped SQLite before it could answer
anything. Now sql.js-httpvfs asks for the pages a query touches and the
databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip
stream is not a byte range of a database.
That only works if every query the site issues is index-driven, and
measured against the real 2016 file, most were not:
so_bao_danh = ? SEARCH via PK ~20 KB
ho_ten_ascii LIKE '%x%' SCAN 127 MB
ho_ten_ascii LIKE 'x%' SCAN 127 MB
COUNT(*) covering index scan 20 MB
ORDER BY toan DESC LIMIT 10 SCAN + temp b-tree 127 MB
Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE
index; a range comparison does use the index. So the schema changed to
suit the access pattern rather than the search changing to suit the
schema.
name_word holds one row per word of each name, WITHOUT ROWID so the
table is the index, carrying ho_ten_ascii so a multi-word query is
resolved inside a single b-tree. name_word_freq says which word of a
query is rarest — the vocabulary is 4,397 words across 2.87M entries, so
"buu loc" seeks on 287 entries rather than walking the 300,000 that
"thi" would. Searching by any word of a name survives, at a few hundred
KB a query.
idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either.
Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL
presets off a full scan. The footer's candidate count now comes from
datasets.json instead of COUNT(*).
2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site
is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged.
The SQL tab is the one place a user can still write a query that reads
the whole table, so it asks before it opens, runs under a byte budget
that stops a runaway query, and shows what each query actually fetched.
Verified: row counts through the assembler guards, every app query
index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206
with a correct Content-Range. Not verified in a browser — this machine
has none — and the library refuses to open a file the host compresses,
so the deployed response headers need a look.
The pipeline is now Go outside web/. differential-parity.mjs becomes
assembler/internal/verify, reachable as `assemble verify A B`. The port fixed a
real weakness: the JavaScript hashed each row's fields joined bare, so a value
shifted across a column boundary produced the same digest. A test now pins that.
The hub still rendered "Phiên bản cũ của trang 2017" above a permanently empty
list — it split datasets on id.includes("old"), and both such datasets are gone.
The heading and the filter are removed. index.html titled every page "THPT QG
2017", including 2016 and the hub, because one file is copied to every route;
the static title is now neutral and the app sets the dataset's own.
Dead code removed: the isOld2/containsOld branches in the stats block, which
only 2017-old2 could ever reach; SUBJECT_LABELS, DATASET_IDS and the unread
`short` subject field; an unused vite.svg and a favicon link to a file that
never existed; two unused CSS rules and --shadow-sm; site.Paths.Root.
Corrected comments that were confidently wrong rather than merely stale: the
reader claimed to be row-streaming when both implementations decode the whole
workbook into memory first, and the fidelity oracle still spoke of 299 input
files when it covers 182. Candidate counts in the hub now derive from
datasets.json instead of being written a second time as prose.
plans/ is emptied. The parity report it held was cited by docs/data-pipeline.md,
so the evidence that the recovered foreign-language scores are real — not the
citation, the four arguments themselves — is now inline there.
Verified: 2017 rebuilt after the writer change hashes identically to the build
before it.
Both plans are done and merged, and their contents now describe a tree that no
longer exists — go-parser/, four datasets, npm entry points, a Rust crate. Git
history keeps them.
The parity evidence stays: docs/data-pipeline.md cites parser-parity-result.md
for the recovered foreign-language scores, so that report and the two JSON
snapshots behind it remain, now marked as an archived record.
Adds go-parser/, a Go reimplementation of the xlsxread parser, verified
byte-for-byte against the Rust original before any cutover.
Reader fidelity is exact across all 299 input files: the canonical cell dump
of every sheet matches calamine's, locked in as a test against a committed
hash oracle. Reaching that required replacing extrame/xls, which corrupted
69% of cells and dropped a further 28% on the BIFF corpus, with pbnjay/grate;
correcting excelize's number-format application and trailing-cell trimming;
restoring carriage returns that XML line-ending normalisation strips from
2,233 ten_cum_thi values; and gating numeric re-rendering on cell type so
shared strings that merely look numeric keep their leading zeros.
The differential gate compares both parsers over all four datasets:
3,265,641 rows with identical full-table SHA-256, identical per-column
non-NULL counts, identical schema metadata and identical stdout.
Config moves from TOML to YAML for both parsers, so they keep reading the
same files and the gate stays meaningful. Verified by rebuilding 2016 and
2017-old2 with Rust under the new configs and matching the recorded counts.
build-db.js now refuses to publish a database whose row count does not match
the known figure, closing a path where an under-producing parser could ship a
truncated public dataset with green CI. The deploy workflow gains a
pull_request trigger and guards deploy to main, so branch verification can no
longer publish to production.
The docs still described two standalone projects with separate frontends,
separate parsers and three Vite variants. Rewritten around what the repo now
is, merging both projects' copies rather than keeping one and discarding the
other — deployment-guide.md and system-architecture.md existed in both and
documented different pipelines.
project-overview.md goal, scope, constraints, the four datasets, history
system-architecture.md data flow, canonical schema, routing, how one
frontend serves both exam years without branching
data-pipeline.md per-dataset Excel formats, the three 2016 layouts,
overflow-sheet gotcha, expected row counts
deployment-guide.md the single-build workflow, adding a dataset,
why no uncompressed database can ship
Records the release gate in plans/reports/parser-parity-result.md: row counts,
all 18 pre-existing per-column non-NULL counts and the deterministic student
samples are identical across all four datasets, against databases decompressed
from the exact bytes the pipeline publishes. The 1,691 recovered
foreign-language scores are documented with the evidence they are real.
Also drops a machine-specific absolute path from a comment in
format_detect_2016.rs. The build-database.js citations there are kept: that
file no longer exists in this repo, but the references explain why several
parsing rules look arbitrary.
The four build variants existed only because the app could not resolve its own
dataset. Now that it can, vite.config.js is a single build with an absolute
base and no VARIANT switching, and the workflow compiles Rust once, installs
Node dependencies once, and builds the site once — it previously built the same
Rust crate twice and ran two separate pnpm installs.
Because base is absolute, the emitted index.html references /thptqg/assets/...
regardless of where it is served from, so the same file works as an entry point
at any depth. scripts/assemble-site.js copies it to each dataset path and to
the two legacy nested URLs, giving a real static file at every published route.
That is what removes the need for an SPA 404-fallback redirect, which would
otherwise have rewritten URLs and interfered with the ?q= deep links.
Generated databases move to a gitignored .build/public, which Vite consumes as
its publicDir. build-db.js gzips without -k, and the assemble step then refuses
to finish if any uncompressed database artefact reached the output — .db,
.db-journal, .db-wal or .db-shm. Previously a raw 100+ MB database was written
into the source tree and deleted afterwards by an rm in the workflow, so
shipping one was a missing cleanup step away.
Both build scripts import DATASET_IDS from src/datasets.js rather than
repeating the dataset list in workflow shell, so the four ids are declared in
exactly one place across the frontend, the database build and the assembly.
Verified by running the full pipeline locally and serving the artifact over
HTTP: all nine routes return 200, every entry point is byte-identical, asset
references are absolute, and the databases are fetchable. The guard was tested
by injecting a raw .db and .db-journal into the build, which failed the
assembly as intended.
The 2016 and 2017 parsers were the same crate with divergent SQL. DDL, INSERT
and the subject regex table lived in each dataset's TOML config, so four copies
had to be kept in step by hand — which is how the two schemas drifted apart.
Move all of it into src/schema.rs as a single 22-column definition: 6 identity
columns (adding ten_cum_thi and gioi_tinh, previously 2016-only) and 16 subject
columns (the union of both exam years). Columns a dataset carries no data for
bind NULL, costing ~1 byte per row.
This collapses writer.rs to one insert path: insert_row_2016, SCORE_FIELDS_2016,
SCORE_FIELDS_2017 and the SCORE_FIELDS alias all go away. Configs shrink to the
parse rules that genuinely vary per dataset — sheet mode, column indices, SBD
validation, header tokens, blank-row stripping — and carry no SQL at all.
Config parsing now denies unknown fields, so a leftover [schema] block fails
loudly instead of looking effective while schema.rs drives the build.
idx_ten_cum_thi is partial, so it holds zero entries on the three 2017 datasets
where the column is always NULL.
Adds db-stats.js and verify-parity.js to prove no data moved: row counts,
per-column non-NULL counts and a deterministic student sample are compared
against databases built from the previous code.
Verified across all four datasets — row counts, every pre-existing column count,
and all sampled students are identical. Database size grows 1.4-2.2%.