mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 12:28:58 +00:00
The 1 KiB page size optimised for the fewest bytes per seeked row. The first real measurement says that is the wrong quantity: a name search made 390 requests for 608 KB and took 16.9 seconds. The worker reads with synchronous XHR (lazyFile.ts opens every GET with async=false), so requests are strictly serial at roughly 40 ms each, and 390 x 40 ms is the whole of that time. The bytes were never the problem. Four times the page size means a quarter of the pages for the same scan, and a shallower b-tree for each of the 100 result rows the search seeks individually. It also costs about 5% less file.
Docs
project-overview.md— goal, scope, constraints, the datasets, historysystem-architecture.md— data flow, canonical schema, routing, how one frontend serves both exam yearsdata-pipeline.md— Excel parse quirks, per-dataset formats, overflow-sheet gotcha, expected row counts, verifying a rebuilddeployment-guide.md— GitHub Pages workflow, adding a dataset, rollback, troubleshooting