Files
thptqg/datasets.json
T
tiennm99 933adf70c9 refactor: comments state current behavior, not project history
Comments across the tree justified the code by pointing at a Rust
implementation that is no longer in the repository, citing files and line
numbers (config.rs:132, schema.rs:26-54, reader.rs:42) that cannot be
opened, plus crates and datasets that are equally gone. A reader could not
check any of it.

Every invariant those comments carried is kept and restated so it stands on
its own: the bytewise sort that decides which row survives a duplicate exam
number, the literal U+0300..U+036F range that must match the site's toAscii,
the trailing space in "SINH ", the BIFF and shared-string corrections, the
VACUUM-after-COMMIT rule, the deploy-from-main guard.

The reader's contract is now anchored to the frozen oracle in
parser/testdata, which still exists and is still checked, rather than to the
tool that originally produced it.

TestDDLMatchesRust becomes TestDDLIsFrozen: it compares against a copy of
the DDL inside the test and never read schema.rs, so both the name and the
failure message were misleading.

ToAscii no longer claims the d-replacement must precede lowercasing. Both
cases map to 'd' and ToLower runs last, so the order has no effect.
2026-08-14 09:40:41 +07:00

31 lines
1.1 KiB
JSON

{
"_comment": [
"The dataset registry: the one place every pipeline stage agrees on what",
"exists. JSON because the Go stages and the Vite app both read it.",
"",
" id the identifier used end to end, from data/<id>/ to /thptqg/<id>/",
" expectedRows exact row count of the built database, not an estimate. The",
" inputs are frozen exam results, so a deviation of even one row",
" means something changed unintentionally and the assembler",
" refuses to publish.",
" dbSizeMb usual size of the gzipped database. Required and non-zero:",
" the assembler rejects a build that comes out far smaller, and",
" the web app shows it while the download runs.",
"",
"Presentation (titles, labels, SQL presets) lives in web/src/datasets.js keyed",
"by id; that file throws at load if the two lists disagree."
],
"datasets": [
{
"id": "2016",
"expectedRows": 877461,
"dbSizeMb": 44
},
{
"id": "2017",
"expectedRows": 861068,
"dbSizeMb": 48
}
]
}