Files
thptqg/.github/workflows/deploy-pages.yml
T
tiennm99 ceb694a747 refactor: split into web/crawler/go-parser and drop the 2017 archives
Move the frontend into web/, the repo's only npm workspace, and replace the
JS crawler with a Go module covering both remaining datasets. The crawler
writes to a .part file and renames on completion: writing straight to the
destination left truncated files that the skip-if-present check would then
skip forever.

Remove the 2017-old and 2017-old2 datasets. They were successive publications
of the same exam, kept side by side so the disagreement stayed inspectable;
the current 2017 supersedes them and they remain in git history.

Recover the 2016 crawler source from the Internet Archive's copy of the
aggregator article, whose original host no longer resolves. All 119 filenames
are verified against data/2016 in both directions, but no archive captured the
spreadsheets themselves, so the host still serving them is unconfirmed and
data/2016 remains the only confirmed copy.

Filenames are load-bearing throughout: go-parser sorts inputs bytewise and
inserts last-wins, so they decide which row survives a duplicate exam number.
2026-08-13 21:45:05 +07:00

99 lines
3.3 KiB
YAML

name: Deploy to GitHub Pages
on:
push:
branches: [main]
# Pull requests run the build job only. Before this, the workflow triggered on
# main-push and workflow_dispatch alone, so "verify on a branch first" was not
# actually possible: pushing to a branch ran nothing, and dispatching from one
# published that branch straight to the live site.
pull_request:
workflow_dispatch:
permissions:
contents: read
pages: write
id-token: write
concurrency:
group: pages
cancel-in-progress: true
jobs:
build:
runs-on: ubuntu-latest
env:
# The parser is pure Go — grate, excelize, yaml.v3 and modernc.org/sqlite
# are all cgo-free — so no C toolchain is needed. Set explicitly rather
# than relying on the default.
CGO_ENABLED: '0'
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version: '1.26'
cache-dependency-path: go-parser/go.sum
- uses: actions/setup-node@v4
with:
node-version: '24'
cache: 'npm'
cache-dependency-path: package-lock.json
- run: npm ci
# The reader-fidelity suite compares all 299 real input files against a
# committed hash oracle, so it is the regression guard for the whole
# reader. Runs before anything is built.
- name: Test parser
run: npm run test:go
# The crawler is not part of the build — it only refreshes data/ by hand.
# It is still compiled and tested here so it cannot rot unnoticed, and
# because its filename test guards the parser: input filenames decide
# which row survives a duplicate exam number.
- name: Test crawler
run: npm run test:crawler
# excelize carries an open advisory, and the 2017 refresh runbook feeds
# network-downloaded spreadsheets straight into the parser.
- name: Vulnerability scan
run: |
go install golang.org/x/vuln/cmd/govulncheck@latest
GOVULNCHECK="$(go env GOPATH)/bin/govulncheck"
(cd go-parser && "$GOVULNCHECK" ./...)
(cd crawler && "$GOVULNCHECK" ./...)
# One parser binary builds every dataset; build-db.js reads the dataset
# list from web/src/datasets.js, verifies each database against its row
# count, then gzips it in place leaving no uncompressed file behind.
- name: Build databases
run: |
npm run build:go
npm run build:db
# One Vite build produces every page. web/scripts/assemble-site.js copies
# the emitted index.html to each dataset path (and the legacy nested
# URLs), then fails the job if an uncompressed database reached _site.
- name: Build and assemble site
run: npm run build:site
- uses: actions/upload-pages-artifact@v3
with:
path: _site
deploy:
# Guarded to main. Without this, a workflow_dispatch from any branch would
# publish that branch's output to the live site, and concurrency
# cancel-in-progress would kill an in-flight good deploy on the way.
if: github.ref == 'refs/heads/main'
needs: build
runs-on: ubuntu-latest
environment:
name: github-pages
url: ${{ steps.deployment.outputs.page_url }}
steps:
- id: deployment
uses: actions/deploy-pages@v4