Files
thptqg/README.md
T
tiennm99 c988bfafcf fix: correct the 2016 source attribution and link full article URLs
The site credited 2016 to Bộ GD&ĐT, which those files were never fetched
from. Both datasets come from published articles: 2016 from an aggregator
on dtnt.bacninh.edu.vn listing one spreadsheet per exam cluster, 2017 from
baotintuc.vn. The README, the architecture table and the web footer all
repeated the ministry claim.

The footer now shows each dataset's full article URL as a link rather than
a bare host, so the citation can be checked. That needs overflow-wrap on
the footer: the 2016 URL is 110 characters with no break opportunity and
would otherwise scroll the page sideways on a phone.

Also records that a full 2016 crawl has been run successfully. The host was
marked unconfirmed and data/2016/ described as the only recoverable copy;
both datasets are now rebuildable from source.
2026-08-14 09:26:27 +07:00

3.9 KiB

thptqg

Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high school graduation exam. Client-side SQL (sql.js) over a SQLite database built from the published .xls/.xlsx score files by the Go parser module. Where those files come from: data pipeline.

Live at tiennm99.github.io/thptqg.

Dataset Exam Candidates Site
2016 2016 877,461 /2016/
2017 2017 861,068 /2017/

Two earlier 2017 publications (2017-old, 2017-old2) were kept for a while because they disagreed with the current one. They have been removed; they remain in git history.

Layout

The repository is one directory per pipeline stage, plus the two stores they pass between them.

crawler/      Go   — re-fetches the source spreadsheets      → data/
parser/       Go   — Excel to SQLite                          data/ → .db
assembler/    Go   — verifies, compresses, builds, assembles  .db + web/ → _site/
web/          npm  — the frontend, one Vite app for every dataset
data/<id>/         raw Excel files, one directory per dataset
datasets.json      the registry: which datasets exist, and their expected size
docs/              architecture, data pipeline, deployment

Each stage runs on its own and hands its output to the next through the stores. web/ is the only npm project; the three stages are independent Go modules.

datasets.json is the contract between them. It is JSON because Go and the Vite app both read it and neither needs a dependency to do so; presentation stays in web/src/datasets.js, keyed by id, which fails loudly if the two disagree.

The dataset id is one identifier end to end:

data/2017/ → parser/configs/2017.yml → db/2017.db.gz → /thptqg/2017/

Build

(cd web && npm ci)
go -C assembler run ./cmd/assemble        # databases, then the site, into _site/
npx serve _site

That one command compiles the parser, builds and verifies each database against its registry row count, compresses it, builds the web app and assembles _site — refusing to continue if a database is short, an artifact looks truncated, or one is missing altogether. Sub-steps when iterating:

go -C assembler run ./cmd/assemble db 2017   # one database
go -C assembler run ./cmd/assemble site      # web build and _site only
go -C assembler run ./cmd/assemble verify A B  # compare two sets of databases
(cd web && npm run dev)                      # the app against staged databases

The source spreadsheets are committed, so a crawl is only needed to refresh them:

go -C crawler run ./cmd/crawl 2016
go -C crawler run ./cmd/crawl 2017

Each reads the download links out of the article that published the dataset, so no link list is kept in the repository. Crawling is idempotent — files already present are skipped — and is never part of the build.

Pushing to main runs the same steps in .github/workflows/deploy-pages.yml and publishes to GitHub Pages.

Adding a dataset

  1. Put the Excel files in data/<id>/
  2. Add parser/configs/<id>.yml — sheet mode, column indices, validation guards. No SQL; the schema is canonical.
  3. Add an entry to datasets.json with its expected row count and size
  4. Add the matching presentation to CONTENT in web/src/datasets.js

Everything else follows: the assembler, the router and the hub all read the registry, and the UI adapts to whichever columns the dataset fills. Steps 3 and 4 check each other, so forgetting either one fails rather than half-working.

Docs

See docs/ — overview, architecture, data pipeline, deployment.