Files
thptqg/docs/deployment-guide.md
T
tiennm99 dbc23c25c5 feat: read the databases over HTTP range requests
The browser downloaded 45 MB of gzipped SQLite before it could answer
anything. Now sql.js-httpvfs asks for the pages a query touches and the
databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip
stream is not a byte range of a database.

That only works if every query the site issues is index-driven, and
measured against the real 2016 file, most were not:

  so_bao_danh = ?              SEARCH via PK          ~20 KB
  ho_ten_ascii LIKE '%x%'      SCAN                   127 MB
  ho_ten_ascii LIKE 'x%'       SCAN                   127 MB
  COUNT(*)                     covering index scan     20 MB
  ORDER BY toan DESC LIMIT 10  SCAN + temp b-tree     127 MB

Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE
index; a range comparison does use the index. So the schema changed to
suit the access pattern rather than the search changing to suit the
schema.

name_word holds one row per word of each name, WITHOUT ROWID so the
table is the index, carrying ho_ten_ascii so a multi-word query is
resolved inside a single b-tree. name_word_freq says which word of a
query is rarest — the vocabulary is 4,397 words across 2.87M entries, so
"buu loc" seeks on 287 entries rather than walking the 300,000 that
"thi" would. Searching by any word of a name survives, at a few hundred
KB a query.

idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either.
Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL
presets off a full scan. The footer's candidate count now comes from
datasets.json instead of COUNT(*).

2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site
is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged.

The SQL tab is the one place a user can still write a query that reads
the whole table, so it asks before it opens, runs under a byte budget
that stops a runaway query, and shows what each query actually fetched.

Verified: row counts through the assembler guards, every app query
index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206
with a correct Content-Range. Not verified in a browser — this machine
has none — and the library refuses to open a file the host compresses,
so the deployed response headers need a look.
2026-08-14 12:42:48 +07:00

111 lines
4.8 KiB
Markdown

# Deployment Guide
Deploys to GitHub Pages via `.github/workflows/deploy-pages.yml`. Every push to
`main` rebuilds both datasets and redeploys the whole site.
One-time setup: **Settings → Pages → Source: GitHub Actions**.
## What the workflow does
1. Checkout, Go toolchain, Node 24, `npm ci` in `web/`
2. Parser, crawler and assembler test suites, web lint, `govulncheck` over all three modules
3. `go -C assembler run ./cmd/assemble` — the whole pipeline: compile the
parser, build and verify each database, compress it into `.build/public/db/`,
run the web build, assemble `_site/`
4. `actions/upload-pages-artifact` + `actions/deploy-pages`
The database build dominates the runtime: roughly 348 MB of Excel is parsed on
every deploy.
## Resulting URLs
```
https://<user>.github.io/thptqg/
https://<user>.github.io/thptqg/2016/
https://<user>.github.io/thptqg/2017/
```
`/thptqg/2017/old/` and `/thptqg/2017/old2/` were the pre-flattening URLs for
the two removed 2017 archives. They are no longer served; like any unknown path
they now render the hub via `404.html`.
## Local reproduction
```bash
(cd web && npm ci)
go -C assembler run ./cmd/assemble
npx serve _site
```
To rebuild a single dataset, or only the site:
```bash
go -C assembler run ./cmd/assemble db 2017
go -C assembler run ./cmd/assemble site
```
## Base path
`svelte.config.js` sets `paths.base: "/thptqg"`. If you fork under a different
repo name, update it to match — assets are referenced absolutely, so a mismatch
shows up as a blank page with 404s on `/_app/...`.
## Adding a dataset
1. Put the Excel files in `data/<id>/`
2. Add `parser/configs/<id>.yml` with the parse rules — sheet mode, column
indices, SBD validation, header tokens, blank-row stripping. No SQL: the
schema is canonical and lives in `parser/internal/schema/schema.go`
3. Add an entry to `datasets.json` — id, `expectedRows`, `dbSizeMb`
4. Add its presentation to `CONTENT` in `web/src/datasets.js`
Nothing else. The assembler and the router both read the registry, and the
frontend adapts to whichever columns the dataset populates. The last two steps
check each other, so forgetting either fails rather than half-working.
## Why no uncompressed database can ship
The assembler deletes the source once compression succeeds, so the raw file
does not survive the build, and it then fails the job if any `.db`,
`.db-journal`, `.db-wal` or `.db-shm` reached the output.
Both guards exist because the previous pipeline wrote a 100+ MB uncompressed
database into the source tree and relied on an `rm` step to keep it out of the
artifact — one missing line away from publishing it.
## Notes
- **The database is not cacheable across deploys.** Every rebuild lays SQLite
pages out differently, so the file changes even when the data does not. Only
the pages a query touches are fetched, so this costs far less than it used
to, but a deploy does invalidate what a returning visitor had cached.
- **The 100 MB file limit is a Git limit, not a Pages one.** It applies to
files committed to a repository; the databases are built in CI and uploaded
as a Pages artifact, and the documented Pages limits are a 1 GB published
site and 100 GB/month of bandwidth, with no per-file figure. The two
databases are 289 MB and 238 MB.
- **Total artifact is about 528 MB**, inside the 1 GB site limit but with less
headroom than before: a third dataset of this size would not fit. The fallback
is `sql.js-httpvfs`'s chunked mode, which splits a database into parts.
- **The server must not compress the databases.** Ranges of a compressed body
address the wrong bytes, and the library refuses to open a file whose HEAD
carries a `Content-Encoding`. `.sqlite3` is an unknown type to Pages, so it is
served as `application/octet-stream` and left alone — verify after a deploy.
## Rollback
Deploys are stateless snapshots. Revert the commit on `main` and push; the next
run rebuilds the older state. There is no data to migrate.
## Troubleshooting
| Symptom | Typical cause |
| --- | --- |
| Blank page, 404 on assets | `paths.base` in `svelte.config.js` does not match the repo name |
| `Failed to fetch database: 404` | Dataset id in `datasets.json` does not match the file in `db/` |
| A route 404s | The site step did not run, or the id is missing from `datasets.json` |
| Database fails to open | The host compressed it. `curl -sI …/db/<id>.sqlite3` must show no `content-encoding`; ranges of a compressed body are unusable |
| Every query is slow or huge | It is not using an index. `EXPLAIN QUERY PLAN` it: a `SCAN` means the browser is fetching the whole table |
| Deploy fails on assembly | An uncompressed database artefact reached the output; the error names the files |
| Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints |