Files
thptqg/.github/workflows
tiennm99 d8b5b66fcf feat: keep the downloaded database, and stop rebuilding it every deploy
Two caches, one on each side.

In the browser, the response is stored in Cache Storage keyed by the
server's ETag, so the transfer is paid once per device rather than once
per visit. A stored copy opens without the gate: consent was given the
first time and reuse costs no network. A redeploy changes the ETag, so
the new version replaces the old instead of answering with last week's
data, and every older version of that database is dropped so one dataset
never occupies the disk twice. When the server cannot be reached at all,
any stored version is used, which incidentally makes the site work
offline. Storing is best effort — a full disk means downloading again
next time, which is no reason to fail the page.

The response is cached from a clone while the original is read for the
progress bar, rather than buffering the file a second time in the one
place where memory is already the binding constraint.

In CI, the built databases are cached on the inputs that determine them:
data/**, parser/** and datasets.json. Parsing 348 MB of spreadsheets is
the slow part of the job, and a web or docs change cannot alter a
database, so those pushes restore instead of rebuilding. Exact matches
only — no restore-keys, since a near-miss would publish databases built
from inputs the commit does not describe.
2026-08-14 17:52:09 +07:00
..