docs: describe the design the code has, not the two before it

Three reversals had landed without the documentation following them, so the
docs described a pipeline that compresses its output, a schema with three
secondary indexes, and a browser that re-downloads the file on every visit.
None of those are true any more.

- Compression: the assembler stopped producing .gz when the databases began
  shipping as .sqlite3. The deployment guide's "why no uncompressed database
  can ship" section explained a guard that now exists for the opposite reason
  — to keep .db, .gz and journals out, so .sqlite3 stays the only name.
- Indexes: the architecture printed a DDL with three CREATE INDEX statements
  and a paragraph on the partial one. schema.go carries none.
- Persistence: "the download is repeated every visit ... has not been done"
  was listed as an open risk after db-cache.js closed it. Replaced with the
  ETag flow, the offline fallback, and the risks that did replace it.

Measured both transfers rather than scaling one from the other, which would
have been wrong: 2016 is 142 MB stored and 31 MB delivered, 2017 is 119 MB and
36 MB. The smaller database is the larger download, so neither figure follows
from the stored size.

Also corrects a CHUNK_BYTES reference to a module that no longer exists, the
238-289 MB per-dataset figure, two paths to web/src/lib/datasets.js, and the
CI step list, which omitted npm test and the post-deploy header check.

The two code comments that said the same outdated things go with them.
This commit is contained in:
tiennm99 committed 2026-08-15 17:35:09 +07:00
1 parent d8b5b66fcf
commit 2c24943cc6
11 files changed
+112 -52

No files matched your search

+1 -1
View File
@@ -2,7 +2,7 @@
Exam-score lookup for Vietnam's national high school graduation exam. Four
stages, each independent: `crawler/` (Go) fetches spreadsheets into `data/<id>/`,
`parser/` (Go) turns them into `<id>.db`, `assembler/` (Go) verifies, compresses
`parser/` (Go) turns them into `<id>.sqlite3`, `assembler/` (Go) verifies them
and builds `_site/`, and `web/` (SvelteKit + Tailwind) serves every dataset from
one app.
+4 -4
View File
@@ -2,7 +2,7 @@
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
school graduation exam. Client-side SQL over a SQLite database the browser
downloads once and queries in memory, built
downloads once, keeps, and queries in memory, built
from the published `.xls`/`.xlsx` score files by the Go `parser` module. Where
those files come from: [data pipeline](./docs/data-pipeline.md#sources).
@@ -24,8 +24,8 @@ pass between them.
```
crawler/ Go — re-fetches the source spreadsheets → data/
parser/ Go — Excel to SQLite data/ → .db
assembler/ Go — verifies, compresses, builds, assembles .db + web/ → _site/
parser/ Go — Excel to SQLite data/ → .sqlite3
assembler/ Go — verifies, builds and assembles .sqlite3 + web/ → _site/
web/ npm — the frontend, one SvelteKit app for every dataset
data/<id>/ raw Excel files, one directory per dataset
datasets.json the registry: which datasets exist, and their expected size
@@ -54,7 +54,7 @@ npx serve _site
```
That one command compiles the parser, builds and verifies each database against
its registry row count, compresses it, builds the web app and assembles `_site` —
its registry row count and size, builds the web app and assembles `_site` —
refusing to continue if a database is short, an artifact looks truncated, or one
is missing altogether. Sub-steps when iterating:
+5 -3
View File
@@ -73,10 +73,12 @@ func BuildParser(p Paths) (string, error) {
return bin, nil
}
// Build runs the parser for one dataset, verifies the result and compresses it.
// Build runs the parser for one dataset and verifies the result.
//
// Only the .gz survives: shipping a 100+ MB uncompressed database is made
// structurally impossible rather than left to a cleanup step.
// Two guards, both refusing to publish rather than warning: the row count must
// equal the registry's exactly, and the file must be at least minSizeRatio of
// the size the registry records. A truncated or short database is the failure
// that would otherwise reach the site with a green pipeline.
func Build(p Paths, bin string, d registry.Dataset) error {
if err := os.MkdirAll(p.OutDir, 0o755); err != nil {
return err
+2 -2
View File
@@ -13,8 +13,8 @@
" is also the memory figure the download gate shows, since the",
" browser holds the whole file while the tab is open.",
"",
"Presentation (titles, labels, SQL presets) lives in web/src/datasets.js keyed",
"by id; that file throws at load if the two lists disagree."
"Presentation (titles, labels, SQL presets) lives in web/src/lib/datasets.js",
"keyed by id; that file throws at load if the two lists disagree."
],
"datasets": [
{
+2 -2
View File
@@ -1,6 +1,6 @@
# Docs
- [`project-overview.md`](./project-overview.md) — goal, scope, constraints, the datasets, history
- [`system-architecture.md`](./system-architecture.md) — data flow, canonical schema, routing, how one frontend serves both exam years
- [`system-architecture.md`](./system-architecture.md) — data flow, the download and how it is kept, canonical schema, routing, how one frontend serves both exam years
- [`data-pipeline.md`](./data-pipeline.md) — Excel parse quirks, per-dataset formats, overflow-sheet gotcha, expected row counts, verifying a rebuild
- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, adding a dataset, rollback, troubleshooting
- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, what the databases publish as, adding a dataset, rollback, troubleshooting
+9 -4
View File
@@ -1,6 +1,6 @@
# Data Pipeline
From raw Excel files to a compressed SQLite file the browser can load.
From raw Excel files to the SQLite file the browser downloads and queries.
One Go binary (`parser/`) builds every dataset. What differs per dataset is
parse rules only — sheet strategy, column layout, validation guards — declared
@@ -182,9 +182,14 @@ silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what
## Expected row counts
The databases are written with 4 KiB pages (`PRAGMA page_size` in
`parser/internal/writer/writer.go`) because the browser fetches them a page per
HTTP request, and those requests are serial. `CHUNK_BYTES` in
`web/src/lib/sqlite.svelte.js` must match.
`parser/internal/writer/writer.go`) — SQLite's own default, set explicitly so
the published file does not change shape if that default ever moves. Nothing on
the client depends on the figure any more. It did under the previous design,
which read the file over HTTP a page at a time and had to be told the page
size; the browser now downloads the file whole.
The pragma runs before the DDL, because a page size cannot change once a table
exists.
| id | Source rows | Skipped | DB rows |
| --- | --- | --- | --- |
+34 -17
View File
@@ -7,13 +7,24 @@ One-time setup: **Settings → Pages → Source: GitHub Actions**.
## What the workflow does
1. Checkout, Go toolchain, Node 24, `npm ci` in `web/`
2. Parser, crawler and assembler test suites, web lint, `govulncheck` over all three modules
1. Checkout, Go toolchain (version taken from `parser/go.mod`), Node 24,
`npm ci` in `web/`
2. Parser, crawler and assembler test suites, `npm test` and `npm run lint` in
`web/`, `govulncheck` over all three Go modules
3. `go -C assembler run ./cmd/assemble` — the whole pipeline: compile the
parser, build and verify each database into `.build/public/db/`, run the web
build, assemble `_site/`. The databases are restored from the Actions cache
when nothing that determines them has changed, and only the site is built
4. `actions/upload-pages-artifact` + `actions/deploy-pages`
5. After deploying, each published database is read back from the live URL and
must begin `SQLite format 3` — proof the site serves a database rather than
an error page or a truncated upload
Pull requests run steps 1–3 and stop. The deploy job is guarded to `main`, so a
branch is verified end to end without touching the live site. The concurrency
group is keyed by ref rather than shared, because a shared lane let a
pull-request run cancel an in-flight `main` deploy while every check stayed
green.
The database build dominates the runtime: roughly 348 MB of Excel to parse. It
is cached in Actions, keyed on `data/**`, `parser/**` and `datasets.json`, so
@@ -62,39 +73,44 @@ shows up as a blank page with 404s on `/_app/...`.
indices, SBD validation, header tokens, blank-row stripping. No SQL: the
schema is canonical and lives in `parser/internal/schema/schema.go`
3. Add an entry to `datasets.json` — id, `expectedRows`, `dbSizeMb`
4. Add its presentation to `CONTENT` in `web/src/datasets.js`
4. Add its presentation to `CONTENT` in `web/src/lib/datasets.js`
Nothing else. The assembler and the router both read the registry, and the
frontend adapts to whichever columns the dataset populates. The last two steps
check each other, so forgetting either fails rather than half-working.
## Why no uncompressed database can ship
## What the databases are published as
The assembler deletes the source once compression succeeds, so the raw file
does not survive the build, and it then fails the job if any `.db`,
`.db-journal`, `.db-wal` or `.db-shm` reached the output.
One uncompressed `<id>.sqlite3` per dataset. Pages gzips it on the wire anyway,
so a `.gz` artifact would only mean decompressing twice.
Both guards exist because the previous pipeline wrote a 100+ MB uncompressed
database into the source tree and relied on an `rm` step to keep it out of the
artifact — one missing line away from publishing it.
`.sqlite3` is the only accepted name, and that is deliberate: it leaves `.db`
free, so site assembly can treat any stray `.db`, `.db.gz`, `.sqlite30` or
SQLite journal (`-journal`, `-wal`, `-shm`) in the output as the mistake it is
and fail the job. An earlier pipeline wrote a database into the source tree and
relied on an `rm` step to keep it out of the artifact — one missing line away
from publishing something it did not mean to.
## Notes
- **The database is not cacheable across deploys.** Every rebuild lays SQLite
pages out differently, so the file changes even when the data does not. Only
the pages a query touches are fetched, so this costs far less than it used
to, but a deploy does invalidate what a returning visitor had cached.
- **A deploy invalidates what returning visitors kept.** Every rebuild lays
SQLite pages out differently, so the file — and the ETag the browser stored it
under — changes even when the data does not. The visitor downloads once more
and keeps that copy until the next deploy.
- **The 100 MB file limit is a Git limit, not a Pages one.** It applies to
files committed to a repository; the databases are built in CI and uploaded
as a Pages artifact, and the documented Pages limits are a 1 GB published
site and 100 GB/month of bandwidth, with no per-file figure. The two
databases are 142 MB and 119 MB.
- **Total artifact is about 262 MB**, comfortably inside the 1 GB site limit.
- **Total artifact is about 263 MB**, comfortably inside the 1 GB site limit.
Dropping the search index the range-request design needed halved both files.
- **Pages compresses the databases on the wire, which is a benefit here.** The
extension is unknown to Pages, so the file is served as
`application/octet-stream`, which `mime-db` marks compressible: 142 MB stored
becomes about 31 MB delivered, and the browser decompresses it transparently.
`application/octet-stream`, which `mime-db` marks compressible, and the
browser decompresses it transparently: 142 MB stored becomes about 31 MB
delivered for 2016, and 119 MB becomes about 36 MB for 2017. Note that the
smaller database is the larger download: the two compress differently, so
neither transfer figure can be derived from the stored size.
- **Verify the bytes, not the headers.** A bare `curl -sI` reports success
whatever the host does. Read the file's first bytes instead:
@@ -119,3 +135,4 @@ run rebuilds the older state. There is no data to migrate.
| Tab crashes on a phone | The database needs 142 MB of memory for 2016, 119 MB for 2017; a low-memory device may have the tab killed |
| Deploy fails on assembly | A stray database artefact reached the output — a journal, a `.gz`, or a name from an earlier design; the error names the files |
| Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints |
| Visitor still sees old data after a deploy | Their stored copy is keyed by ETag, so this should not happen. Confirm the deployed file's `ETag` header actually changed |
+4 -2
View File
@@ -21,8 +21,10 @@ running entirely in the browser and hosted for free on GitHub Pages. Covers the
## Constraints
- **Zero backend.** The database (238–289 MB per dataset) stays on the server
and the browser downloads it once and queries it in memory.
- **Zero backend.** The browser downloads the whole database (142 MB for 2016,
119 MB for 2017; about 31 MB and 36 MB on the wire, since the host gzips it)
and queries it in memory. The copy is kept on the device, so the transfer is
paid once rather than once per visit.
- **Read-only.** `INSERT`/`UPDATE`/`DELETE` are rejected, so nobody is misled
into thinking edits persist. The file is fetched, never written.
- **Row caps.** 100 rows for lookups, 1000 for custom SQL, to prevent browser
+47 -13
View File
@@ -1,10 +1,13 @@
# System Architecture
Static site, no backend. The browser downloads the whole SQLite file once and
queries it in memory with `sql.js` (SQLite compiled to WebAssembly). A dataset
costs about 31 MB to fetch and 142 MB of memory to hold; every query after that
is local — an exam-number lookup is immediate, and a name search scans all
877,460 rows in a few hundred milliseconds.
queries it in memory with `sql.js` (SQLite compiled to WebAssembly). 2016 costs
about 31 MB to fetch and 142 MB of memory to hold; 2017, 36 MB and 119 MB. Every
query after that is local — an exam-number lookup is immediate, and a name
search scans all 877,460 rows in a few hundred milliseconds.
The download is kept on the device, so the transfer is paid once per device
rather than once per visit.
That is a reversal of the previous design, which read the file where it lay
over HTTP range requests. See [Considered and not taken](#considered-and-not-taken).
@@ -34,8 +37,35 @@ data/<id>/*.xls(x)
│
▼ browser
download gate → whole file → sql.js in memory → queries never leave the tab
↕
Cache Storage, keyed by ETag — a second visit skips the gate
```
## The download, and keeping it
`web/src/lib/database.svelte.js` owns the transfer; `web/src/lib/db-cache.js`
owns what survives it. Opening a dataset page:
1. A conditional request asks the server for the database's current ETag.
2. If a copy of exactly that version is in Cache Storage, it opens without
asking — consent was given once and reuse costs no network.
3. Otherwise the gate states the transfer and the memory cost, and downloads
only when the visitor accepts. The response is stored under
`<url>?v=<etag>`, and every other version of that database is dropped, so a
dataset never occupies the disk twice.
4. If the server cannot be reached at all, any stored version is used. Slightly
old frozen exam results beat an error page, and this is what lets the site
answer offline.
Cache Storage rather than IndexedDB: the thing being stored is an HTTP
response, the browser can stream it to disk instead of holding it as one
buffer, and eviction under storage pressure is the browser's business.
Progress is measured against `dbSizeMb` from the registry, not
`Content-Length`. The server sends the file gzipped, so `Content-Length` counts
compressed bytes while the stream yields decompressed ones; the compressed
figure is reported separately as the transfer size.
## The dataset id
One identifier ties the whole pipeline together:
@@ -84,18 +114,18 @@ CREATE TABLE student (
lich_su, dia_ly, gdcd, khxh,
tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
```
One table and no secondary indexes. The primary key is the only one, and it
comes free with the table. Everything else is a scan, which in memory costs a
few hundred milliseconds — an index would buy that back at the price of
megabytes on every visitor's network. `TestDDLIsFrozen` in `parser` holds an
independent copy of this DDL, so changing the shape has to be deliberate.
Every dataset gets all 22 columns; ones it has no data for are NULL, costing
about a byte per row. `khtn`, `khxh` and `gdcd` are empty on 2016;
`ten_cum_thi` and `gioi_tinh` are empty on 2017.
`idx_ten_cum_thi` is partial, so it holds zero entries where the column is
always NULL.
## Routing
URLs are flat, one segment per dataset, and the segment is the id:
@@ -176,6 +206,8 @@ total descending.
| --- | --- | --- |
| Storage | Static SQLite file, downloaded whole | No backend; the datasets are frozen, and one large transfer is something browsers and CDNs are both good at |
| Download gate | Blocking, no dismiss | The page has no answers before the file arrives, and 31 MB of someone's mobile data should be asked for rather than spent silently |
| Keeping the download | Cache Storage, versioned by ETag | The transfer is paid once per device instead of once per visit; the ETag is what stops a redeploy being answered with last week's data |
| Searching | On submit, never while typing | A name search scans every row, so a keystroke-per-search would run hundreds of full scans to answer one question |
| Compression | None published | The host gzips on the wire, so a `.gz` artifact would only be decompressed twice |
| WASM hosting | Bundled with the app | `sql.js` ships its own build; one less third-party runtime dependency |
| Diacritics search | Pre-computed `ho_ten_ascii` | `LOWER(REPLACE(...))` at query time is far slower over 877,460 rows than a column computed once at build |
@@ -214,9 +246,11 @@ total descending.
WebAssembly memory for as long as the page is open: 142 MB for 2016, 119 MB
for 2017. A low-memory phone may have the tab killed, which is why the gate
states the figure before the download starts.
- **The download is repeated every visit.** Nothing is persisted — a reload
fetches the file again. Caching it in IndexedDB or the Cache API would fix
that and has not been done.
- **A deploy costs returning visitors the transfer again.** Every rebuild lays
SQLite pages out differently, so the file — and its ETag — changes even when
the data does not. The stored copy is then a stale version and is replaced.
- **The stored copy can be evicted.** Cache Storage is subject to the browser's
own storage pressure, so a device short on disk falls back to downloading.
- **A visitor who will not download cannot use the site.** That is the
deliberate shape of the gate, and it makes the first impression a 31 MB ask.
- **Hosted size.** 526 MB for both datasets against the 1 GB GitHub Pages
+3 -3
View File
@@ -13,9 +13,9 @@ xlsxread build --schema parser/configs/<id>.yml --input data/<id> --output <db>
xlsxread audit --schema parser/configs/<id>.yml --input data/<id> --db <db>
```
This stage only produces a database. Verifying it against the expected row
count, compressing it and publishing it belong to `assembler/`, which compiles
this binary and drives it per dataset:
This stage only produces a database. Verifying it against the expected row count
and size, and publishing it, belong to `assembler/`, which compiles this binary
and drives it per dataset:
```bash
go -C assembler run ./cmd/assemble db
+1 -1
View File
@@ -11,7 +11,7 @@ import { vitePreprocess } from "@sveltejs/vite-plugin-svelte";
* depth — GitHub Pages answers /thptqg/anything/at/all/ with that one file.
*
* `files.assets` points outside this workspace because the assembler stages the
* gzipped databases there (assembler/internal/databases).
* built databases there (assembler/internal/databases).
*
* @type {import('@sveltejs/kit').Config}
*/