mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
The browser downloaded 45 MB of gzipped SQLite before it could answer anything. Now sql.js-httpvfs asks for the pages a query touches and the databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip stream is not a byte range of a database. That only works if every query the site issues is index-driven, and measured against the real 2016 file, most were not: so_bao_danh = ? SEARCH via PK ~20 KB ho_ten_ascii LIKE '%x%' SCAN 127 MB ho_ten_ascii LIKE 'x%' SCAN 127 MB COUNT(*) covering index scan 20 MB ORDER BY toan DESC LIMIT 10 SCAN + temp b-tree 127 MB Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE index; a range comparison does use the index. So the schema changed to suit the access pattern rather than the search changing to suit the schema. name_word holds one row per word of each name, WITHOUT ROWID so the table is the index, carrying ho_ten_ascii so a multi-word query is resolved inside a single b-tree. name_word_freq says which word of a query is rarest — the vocabulary is 4,397 words across 2.87M entries, so "buu loc" seeks on 287 entries rather than walking the 300,000 that "thi" would. Searching by any word of a name survives, at a few hundred KB a query. idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either. Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL presets off a full scan. The footer's candidate count now comes from datasets.json instead of COUNT(*). 2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged. The SQL tab is the one place a user can still write a query that reads the whole table, so it asks before it opens, runs under a byte budget that stops a runaway query, and shows what each query actually fetched. Verified: row counts through the assembler guards, every app query index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206 with a correct Content-Range. Not verified in a browser — this machine has none — and the library refuses to open a file the host compresses, so the deployed response headers need a look.
32 lines
1.2 KiB
JSON
32 lines
1.2 KiB
JSON
{
|
|
"_comment": [
|
|
"The dataset registry: the one place every pipeline stage agrees on what",
|
|
"exists. JSON because the Go stages and the Vite app both read it.",
|
|
"",
|
|
" id the identifier used end to end, from data/<id>/ to /thptqg/<id>/",
|
|
" expectedRows exact row count of the built database, not an estimate. The",
|
|
" inputs are frozen exam results, so a deviation of even one row",
|
|
" means something changed unintentionally and the assembler",
|
|
" refuses to publish.",
|
|
" dbSizeMb usual size of the published database. Required and non-zero:",
|
|
" the assembler rejects a build that comes out far smaller. The",
|
|
" file is served uncompressed and read a page at a time over",
|
|
" HTTP range requests, so nothing downloads it whole.",
|
|
"",
|
|
"Presentation (titles, labels, SQL presets) lives in web/src/datasets.js keyed",
|
|
"by id; that file throws at load if the two lists disagree."
|
|
],
|
|
"datasets": [
|
|
{
|
|
"id": "2016",
|
|
"expectedRows": 877460,
|
|
"dbSizeMb": 289
|
|
},
|
|
{
|
|
"id": "2017",
|
|
"expectedRows": 861068,
|
|
"dbSizeMb": 238
|
|
}
|
|
]
|
|
}
|