mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
docs: describe the design the code has, not the two before it
Three reversals had landed without the documentation following them, so the docs described a pipeline that compresses its output, a schema with three secondary indexes, and a browser that re-downloads the file on every visit. None of those are true any more. - Compression: the assembler stopped producing .gz when the databases began shipping as .sqlite3. The deployment guide's "why no uncompressed database can ship" section explained a guard that now exists for the opposite reason — to keep .db, .gz and journals out, so .sqlite3 stays the only name. - Indexes: the architecture printed a DDL with three CREATE INDEX statements and a paragraph on the partial one. schema.go carries none. - Persistence: "the download is repeated every visit ... has not been done" was listed as an open risk after db-cache.js closed it. Replaced with the ETag flow, the offline fallback, and the risks that did replace it. Measured both transfers rather than scaling one from the other, which would have been wrong: 2016 is 142 MB stored and 31 MB delivered, 2017 is 119 MB and 36 MB. The smaller database is the larger download, so neither figure follows from the stored size. Also corrects a CHUNK_BYTES reference to a module that no longer exists, the 238-289 MB per-dataset figure, two paths to web/src/lib/datasets.js, and the CI step list, which omitted npm test and the post-deploy header check. The two code comments that said the same outdated things go with them.
This commit is contained in:
1 parent
d8b5b66fcf
commit
2c24943cc6
11 files changed
+112
-52
No files matched your search
@@ -2,7 +2,7 @@
|
||||
|
||||
Exam-score lookup for Vietnam's national high school graduation exam. Four
|
||||
stages, each independent: `crawler/` (Go) fetches spreadsheets into `data/<id>/`,
|
||||
`parser/` (Go) turns them into `<id>.db`, `assembler/` (Go) verifies, compresses
|
||||
`parser/` (Go) turns them into `<id>.sqlite3`, `assembler/` (Go) verifies them
|
||||
and builds `_site/`, and `web/` (SvelteKit + Tailwind) serves every dataset from
|
||||
one app.
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
|
||||
school graduation exam. Client-side SQL over a SQLite database the browser
|
||||
downloads once and queries in memory, built
|
||||
downloads once, keeps, and queries in memory, built
|
||||
from the published `.xls`/`.xlsx` score files by the Go `parser` module. Where
|
||||
those files come from: [data pipeline](./docs/data-pipeline.md#sources).
|
||||
|
||||
@@ -24,8 +24,8 @@ pass between them.
|
||||
|
||||
```
|
||||
crawler/ Go — re-fetches the source spreadsheets → data/
|
||||
parser/ Go — Excel to SQLite data/ → .db
|
||||
assembler/ Go — verifies, compresses, builds, assembles .db + web/ → _site/
|
||||
parser/ Go — Excel to SQLite data/ → .sqlite3
|
||||
assembler/ Go — verifies, builds and assembles .sqlite3 + web/ → _site/
|
||||
web/ npm — the frontend, one SvelteKit app for every dataset
|
||||
data/<id>/ raw Excel files, one directory per dataset
|
||||
datasets.json the registry: which datasets exist, and their expected size
|
||||
@@ -54,7 +54,7 @@ npx serve _site
|
||||
```
|
||||
|
||||
That one command compiles the parser, builds and verifies each database against
|
||||
its registry row count, compresses it, builds the web app and assembles `_site` —
|
||||
its registry row count and size, builds the web app and assembles `_site` —
|
||||
refusing to continue if a database is short, an artifact looks truncated, or one
|
||||
is missing altogether. Sub-steps when iterating:
|
||||
|
||||
|
||||
@@ -73,10 +73,12 @@ func BuildParser(p Paths) (string, error) {
|
||||
return bin, nil
|
||||
}
|
||||
|
||||
// Build runs the parser for one dataset, verifies the result and compresses it.
|
||||
// Build runs the parser for one dataset and verifies the result.
|
||||
//
|
||||
// Only the .gz survives: shipping a 100+ MB uncompressed database is made
|
||||
// structurally impossible rather than left to a cleanup step.
|
||||
// Two guards, both refusing to publish rather than warning: the row count must
|
||||
// equal the registry's exactly, and the file must be at least minSizeRatio of
|
||||
// the size the registry records. A truncated or short database is the failure
|
||||
// that would otherwise reach the site with a green pipeline.
|
||||
func Build(p Paths, bin string, d registry.Dataset) error {
|
||||
if err := os.MkdirAll(p.OutDir, 0o755); err != nil {
|
||||
return err
|
||||
|
||||
+2
-2
@@ -13,8 +13,8 @@
|
||||
" is also the memory figure the download gate shows, since the",
|
||||
" browser holds the whole file while the tab is open.",
|
||||
"",
|
||||
"Presentation (titles, labels, SQL presets) lives in web/src/datasets.js keyed",
|
||||
"by id; that file throws at load if the two lists disagree."
|
||||
"Presentation (titles, labels, SQL presets) lives in web/src/lib/datasets.js",
|
||||
"keyed by id; that file throws at load if the two lists disagree."
|
||||
],
|
||||
"datasets": [
|
||||
{
|
||||
|
||||
+2
-2
@@ -1,6 +1,6 @@
|
||||
# Docs
|
||||
|
||||
- [`project-overview.md`](./project-overview.md) — goal, scope, constraints, the datasets, history
|
||||
- [`system-architecture.md`](./system-architecture.md) — data flow, canonical schema, routing, how one frontend serves both exam years
|
||||
- [`system-architecture.md`](./system-architecture.md) — data flow, the download and how it is kept, canonical schema, routing, how one frontend serves both exam years
|
||||
- [`data-pipeline.md`](./data-pipeline.md) — Excel parse quirks, per-dataset formats, overflow-sheet gotcha, expected row counts, verifying a rebuild
|
||||
- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, adding a dataset, rollback, troubleshooting
|
||||
- [`deployment-guide.md`](./deployment-guide.md) — GitHub Pages workflow, what the databases publish as, adding a dataset, rollback, troubleshooting
|
||||
@@ -1,6 +1,6 @@
|
||||
# Data Pipeline
|
||||
|
||||
From raw Excel files to a compressed SQLite file the browser can load.
|
||||
From raw Excel files to the SQLite file the browser downloads and queries.
|
||||
|
||||
One Go binary (`parser/`) builds every dataset. What differs per dataset is
|
||||
parse rules only — sheet strategy, column layout, validation guards — declared
|
||||
@@ -182,9 +182,14 @@ silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what
|
||||
## Expected row counts
|
||||
|
||||
The databases are written with 4 KiB pages (`PRAGMA page_size` in
|
||||
`parser/internal/writer/writer.go`) because the browser fetches them a page per
|
||||
HTTP request, and those requests are serial. `CHUNK_BYTES` in
|
||||
`web/src/lib/sqlite.svelte.js` must match.
|
||||
`parser/internal/writer/writer.go`) — SQLite's own default, set explicitly so
|
||||
the published file does not change shape if that default ever moves. Nothing on
|
||||
the client depends on the figure any more. It did under the previous design,
|
||||
which read the file over HTTP a page at a time and had to be told the page
|
||||
size; the browser now downloads the file whole.
|
||||
|
||||
The pragma runs before the DDL, because a page size cannot change once a table
|
||||
exists.
|
||||
|
||||
| id | Source rows | Skipped | DB rows |
|
||||
| --- | --- | --- | --- |
|
||||
|
||||
+34
-17
@@ -7,13 +7,24 @@ One-time setup: **Settings → Pages → Source: GitHub Actions**.
|
||||
|
||||
## What the workflow does
|
||||
|
||||
1. Checkout, Go toolchain, Node 24, `npm ci` in `web/`
|
||||
2. Parser, crawler and assembler test suites, web lint, `govulncheck` over all three modules
|
||||
1. Checkout, Go toolchain (version taken from `parser/go.mod`), Node 24,
|
||||
`npm ci` in `web/`
|
||||
2. Parser, crawler and assembler test suites, `npm test` and `npm run lint` in
|
||||
`web/`, `govulncheck` over all three Go modules
|
||||
3. `go -C assembler run ./cmd/assemble` — the whole pipeline: compile the
|
||||
parser, build and verify each database into `.build/public/db/`, run the web
|
||||
build, assemble `_site/`. The databases are restored from the Actions cache
|
||||
when nothing that determines them has changed, and only the site is built
|
||||
4. `actions/upload-pages-artifact` + `actions/deploy-pages`
|
||||
5. After deploying, each published database is read back from the live URL and
|
||||
must begin `SQLite format 3` — proof the site serves a database rather than
|
||||
an error page or a truncated upload
|
||||
|
||||
Pull requests run steps 1–3 and stop. The deploy job is guarded to `main`, so a
|
||||
branch is verified end to end without touching the live site. The concurrency
|
||||
group is keyed by ref rather than shared, because a shared lane let a
|
||||
pull-request run cancel an in-flight `main` deploy while every check stayed
|
||||
green.
|
||||
|
||||
The database build dominates the runtime: roughly 348 MB of Excel to parse. It
|
||||
is cached in Actions, keyed on `data/**`, `parser/**` and `datasets.json`, so
|
||||
@@ -62,39 +73,44 @@ shows up as a blank page with 404s on `/_app/...`.
|
||||
indices, SBD validation, header tokens, blank-row stripping. No SQL: the
|
||||
schema is canonical and lives in `parser/internal/schema/schema.go`
|
||||
3. Add an entry to `datasets.json` — id, `expectedRows`, `dbSizeMb`
|
||||
4. Add its presentation to `CONTENT` in `web/src/datasets.js`
|
||||
4. Add its presentation to `CONTENT` in `web/src/lib/datasets.js`
|
||||
|
||||
Nothing else. The assembler and the router both read the registry, and the
|
||||
frontend adapts to whichever columns the dataset populates. The last two steps
|
||||
check each other, so forgetting either fails rather than half-working.
|
||||
|
||||
## Why no uncompressed database can ship
|
||||
## What the databases are published as
|
||||
|
||||
The assembler deletes the source once compression succeeds, so the raw file
|
||||
does not survive the build, and it then fails the job if any `.db`,
|
||||
`.db-journal`, `.db-wal` or `.db-shm` reached the output.
|
||||
One uncompressed `<id>.sqlite3` per dataset. Pages gzips it on the wire anyway,
|
||||
so a `.gz` artifact would only mean decompressing twice.
|
||||
|
||||
Both guards exist because the previous pipeline wrote a 100+ MB uncompressed
|
||||
database into the source tree and relied on an `rm` step to keep it out of the
|
||||
artifact — one missing line away from publishing it.
|
||||
`.sqlite3` is the only accepted name, and that is deliberate: it leaves `.db`
|
||||
free, so site assembly can treat any stray `.db`, `.db.gz`, `.sqlite30` or
|
||||
SQLite journal (`-journal`, `-wal`, `-shm`) in the output as the mistake it is
|
||||
and fail the job. An earlier pipeline wrote a database into the source tree and
|
||||
relied on an `rm` step to keep it out of the artifact — one missing line away
|
||||
from publishing something it did not mean to.
|
||||
|
||||
## Notes
|
||||
|
||||
- **The database is not cacheable across deploys.** Every rebuild lays SQLite
|
||||
pages out differently, so the file changes even when the data does not. Only
|
||||
the pages a query touches are fetched, so this costs far less than it used
|
||||
to, but a deploy does invalidate what a returning visitor had cached.
|
||||
- **A deploy invalidates what returning visitors kept.** Every rebuild lays
|
||||
SQLite pages out differently, so the file — and the ETag the browser stored it
|
||||
under — changes even when the data does not. The visitor downloads once more
|
||||
and keeps that copy until the next deploy.
|
||||
- **The 100 MB file limit is a Git limit, not a Pages one.** It applies to
|
||||
files committed to a repository; the databases are built in CI and uploaded
|
||||
as a Pages artifact, and the documented Pages limits are a 1 GB published
|
||||
site and 100 GB/month of bandwidth, with no per-file figure. The two
|
||||
databases are 142 MB and 119 MB.
|
||||
- **Total artifact is about 262 MB**, comfortably inside the 1 GB site limit.
|
||||
- **Total artifact is about 263 MB**, comfortably inside the 1 GB site limit.
|
||||
Dropping the search index the range-request design needed halved both files.
|
||||
- **Pages compresses the databases on the wire, which is a benefit here.** The
|
||||
extension is unknown to Pages, so the file is served as
|
||||
`application/octet-stream`, which `mime-db` marks compressible: 142 MB stored
|
||||
becomes about 31 MB delivered, and the browser decompresses it transparently.
|
||||
`application/octet-stream`, which `mime-db` marks compressible, and the
|
||||
browser decompresses it transparently: 142 MB stored becomes about 31 MB
|
||||
delivered for 2016, and 119 MB becomes about 36 MB for 2017. Note that the
|
||||
smaller database is the larger download: the two compress differently, so
|
||||
neither transfer figure can be derived from the stored size.
|
||||
- **Verify the bytes, not the headers.** A bare `curl -sI` reports success
|
||||
whatever the host does. Read the file's first bytes instead:
|
||||
|
||||
@@ -119,3 +135,4 @@ run rebuilds the older state. There is no data to migrate.
|
||||
| Tab crashes on a phone | The database needs 142 MB of memory for 2016, 119 MB for 2017; a low-memory device may have the tab killed |
|
||||
| Deploy fails on assembly | A stray database artefact reached the output — a journal, a `.gz`, or a name from an earlier design; the error names the files |
|
||||
| Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints |
|
||||
| Visitor still sees old data after a deploy | Their stored copy is keyed by ETag, so this should not happen. Confirm the deployed file's `ETag` header actually changed |
|
||||
@@ -21,8 +21,10 @@ running entirely in the browser and hosted for free on GitHub Pages. Covers the
|
||||
|
||||
## Constraints
|
||||
|
||||
- **Zero backend.** The database (238–289 MB per dataset) stays on the server
|
||||
and the browser downloads it once and queries it in memory.
|
||||
- **Zero backend.** The browser downloads the whole database (142 MB for 2016,
|
||||
119 MB for 2017; about 31 MB and 36 MB on the wire, since the host gzips it)
|
||||
and queries it in memory. The copy is kept on the device, so the transfer is
|
||||
paid once rather than once per visit.
|
||||
- **Read-only.** `INSERT`/`UPDATE`/`DELETE` are rejected, so nobody is misled
|
||||
into thinking edits persist. The file is fetched, never written.
|
||||
- **Row caps.** 100 rows for lookups, 1000 for custom SQL, to prevent browser
|
||||
|
||||
+47
-13
@@ -1,10 +1,13 @@
|
||||
# System Architecture
|
||||
|
||||
Static site, no backend. The browser downloads the whole SQLite file once and
|
||||
queries it in memory with `sql.js` (SQLite compiled to WebAssembly). A dataset
|
||||
costs about 31 MB to fetch and 142 MB of memory to hold; every query after that
|
||||
is local — an exam-number lookup is immediate, and a name search scans all
|
||||
877,460 rows in a few hundred milliseconds.
|
||||
queries it in memory with `sql.js` (SQLite compiled to WebAssembly). 2016 costs
|
||||
about 31 MB to fetch and 142 MB of memory to hold; 2017, 36 MB and 119 MB. Every
|
||||
query after that is local — an exam-number lookup is immediate, and a name
|
||||
search scans all 877,460 rows in a few hundred milliseconds.
|
||||
|
||||
The download is kept on the device, so the transfer is paid once per device
|
||||
rather than once per visit.
|
||||
|
||||
That is a reversal of the previous design, which read the file where it lay
|
||||
over HTTP range requests. See [Considered and not taken](#considered-and-not-taken).
|
||||
@@ -34,8 +37,35 @@ data/<id>/*.xls(x)
|
||||
│
|
||||
▼ browser
|
||||
download gate → whole file → sql.js in memory → queries never leave the tab
|
||||
↕
|
||||
Cache Storage, keyed by ETag — a second visit skips the gate
|
||||
```
|
||||
|
||||
## The download, and keeping it
|
||||
|
||||
`web/src/lib/database.svelte.js` owns the transfer; `web/src/lib/db-cache.js`
|
||||
owns what survives it. Opening a dataset page:
|
||||
|
||||
1. A conditional request asks the server for the database's current ETag.
|
||||
2. If a copy of exactly that version is in Cache Storage, it opens without
|
||||
asking — consent was given once and reuse costs no network.
|
||||
3. Otherwise the gate states the transfer and the memory cost, and downloads
|
||||
only when the visitor accepts. The response is stored under
|
||||
`<url>?v=<etag>`, and every other version of that database is dropped, so a
|
||||
dataset never occupies the disk twice.
|
||||
4. If the server cannot be reached at all, any stored version is used. Slightly
|
||||
old frozen exam results beat an error page, and this is what lets the site
|
||||
answer offline.
|
||||
|
||||
Cache Storage rather than IndexedDB: the thing being stored is an HTTP
|
||||
response, the browser can stream it to disk instead of holding it as one
|
||||
buffer, and eviction under storage pressure is the browser's business.
|
||||
|
||||
Progress is measured against `dbSizeMb` from the registry, not
|
||||
`Content-Length`. The server sends the file gzipped, so `Content-Length` counts
|
||||
compressed bytes while the stream yields decompressed ones; the compressed
|
||||
figure is reported separately as the transfer size.
|
||||
|
||||
## The dataset id
|
||||
|
||||
One identifier ties the whole pipeline together:
|
||||
@@ -84,18 +114,18 @@ CREATE TABLE student (
|
||||
lich_su, dia_ly, gdcd, khxh,
|
||||
tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung REAL
|
||||
);
|
||||
CREATE INDEX idx_ho_ten ON student(ho_ten);
|
||||
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
|
||||
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
|
||||
```
|
||||
|
||||
One table and no secondary indexes. The primary key is the only one, and it
|
||||
comes free with the table. Everything else is a scan, which in memory costs a
|
||||
few hundred milliseconds — an index would buy that back at the price of
|
||||
megabytes on every visitor's network. `TestDDLIsFrozen` in `parser` holds an
|
||||
independent copy of this DDL, so changing the shape has to be deliberate.
|
||||
|
||||
Every dataset gets all 22 columns; ones it has no data for are NULL, costing
|
||||
about a byte per row. `khtn`, `khxh` and `gdcd` are empty on 2016;
|
||||
`ten_cum_thi` and `gioi_tinh` are empty on 2017.
|
||||
|
||||
`idx_ten_cum_thi` is partial, so it holds zero entries where the column is
|
||||
always NULL.
|
||||
|
||||
## Routing
|
||||
|
||||
URLs are flat, one segment per dataset, and the segment is the id:
|
||||
@@ -176,6 +206,8 @@ total descending.
|
||||
| --- | --- | --- |
|
||||
| Storage | Static SQLite file, downloaded whole | No backend; the datasets are frozen, and one large transfer is something browsers and CDNs are both good at |
|
||||
| Download gate | Blocking, no dismiss | The page has no answers before the file arrives, and 31 MB of someone's mobile data should be asked for rather than spent silently |
|
||||
| Keeping the download | Cache Storage, versioned by ETag | The transfer is paid once per device instead of once per visit; the ETag is what stops a redeploy being answered with last week's data |
|
||||
| Searching | On submit, never while typing | A name search scans every row, so a keystroke-per-search would run hundreds of full scans to answer one question |
|
||||
| Compression | None published | The host gzips on the wire, so a `.gz` artifact would only be decompressed twice |
|
||||
| WASM hosting | Bundled with the app | `sql.js` ships its own build; one less third-party runtime dependency |
|
||||
| Diacritics search | Pre-computed `ho_ten_ascii` | `LOWER(REPLACE(...))` at query time is far slower over 877,460 rows than a column computed once at build |
|
||||
@@ -214,9 +246,11 @@ total descending.
|
||||
WebAssembly memory for as long as the page is open: 142 MB for 2016, 119 MB
|
||||
for 2017. A low-memory phone may have the tab killed, which is why the gate
|
||||
states the figure before the download starts.
|
||||
- **The download is repeated every visit.** Nothing is persisted — a reload
|
||||
fetches the file again. Caching it in IndexedDB or the Cache API would fix
|
||||
that and has not been done.
|
||||
- **A deploy costs returning visitors the transfer again.** Every rebuild lays
|
||||
SQLite pages out differently, so the file — and its ETag — changes even when
|
||||
the data does not. The stored copy is then a stale version and is replaced.
|
||||
- **The stored copy can be evicted.** Cache Storage is subject to the browser's
|
||||
own storage pressure, so a device short on disk falls back to downloading.
|
||||
- **A visitor who will not download cannot use the site.** That is the
|
||||
deliberate shape of the gate, and it makes the first impression a 31 MB ask.
|
||||
- **Hosted size.** 526 MB for both datasets against the 1 GB GitHub Pages
|
||||
|
||||
+3
-3
@@ -13,9 +13,9 @@ xlsxread build --schema parser/configs/<id>.yml --input data/<id> --output <db>
|
||||
xlsxread audit --schema parser/configs/<id>.yml --input data/<id> --db <db>
|
||||
```
|
||||
|
||||
This stage only produces a database. Verifying it against the expected row
|
||||
count, compressing it and publishing it belong to `assembler/`, which compiles
|
||||
this binary and drives it per dataset:
|
||||
This stage only produces a database. Verifying it against the expected row count
|
||||
and size, and publishing it, belong to `assembler/`, which compiles this binary
|
||||
and drives it per dataset:
|
||||
|
||||
```bash
|
||||
go -C assembler run ./cmd/assemble db
|
||||
|
||||
@@ -11,7 +11,7 @@ import { vitePreprocess } from "@sveltejs/vite-plugin-svelte";
|
||||
* depth — GitHub Pages answers /thptqg/anything/at/all/ with that one file.
|
||||
*
|
||||
* `files.assets` points outside this workspace because the assembler stages the
|
||||
* gzipped databases there (assembler/internal/databases).
|
||||
* built databases there (assembler/internal/databases).
|
||||
*
|
||||
* @type {import('@sveltejs/kit').Config}
|
||||
*/
|
||||
|
||||
Reference in new issue
Block a user