3 Commits
Author SHA1 Message Date
tiennm99 9710524ced fix(ci): require the Go release that patches encoding/xml
govulncheck fails the pipeline on GO-2026-6088: encoding/xml decodes
without a recursion depth guard, reachable from excelize's OpenFile,
GetRows and GetSheetList and from buildCRFixups directly. The parser is
fed spreadsheets downloaded over the network by the crawler, so the path
is real.

Raising the go directive to 1.26.6 in all three modules puts the fix
below every build rather than leaving it to whichever patch release the
runner happens to install.

govulncheck is clean on all three modules, and every suite passes on the
new toolchain — including the reader fidelity sweep, which matters here
because buildCRFixups depends on how encoding/xml normalises line endings.
2026-08-14 10:44:26 +07:00
tiennm99 b04f9844f9 refactor(crawler): read the file lists from the source articles
Both sources carried their download links as a hardcoded array, which is not a
crawl: the lists could drift from what the articles actually published, and
nothing would say so. A source now names the article and how to name what it
finds there, and internal/article reads the links out of that page at run time.

2016 takes its filenames straight from the URL. 2017 cannot — the CDN names are
inconsistent (Angiang.xls, 1BaRiaVungTau.xls, 23HaiPhong.xls) — so it derives
them from the province in the link text, transliterated to ASCII the same way
go-parser builds ho_ten_ascii.

Filenames stay load-bearing: go-parser sorts inputs bytewise and inserts
last-wins, so they decide which row survives a duplicate exam number. Saved
copies of both articles are committed as fixtures, and a test asserts that
reading them and applying each naming rule reproduces data/<id> exactly, in both
directions. Resolve also rejects a page that yields the wrong number of links or
two links that would write the same file, since either silently costs the
dataset files that only the row-count guard would notice afterwards.

Verified against the live 2017 article: a from-scratch crawl of all 63 files
leaves the committed data unchanged.
2026-08-13 22:05:07 +07:00
tiennm99 ceb694a747 refactor: split into web/crawler/go-parser and drop the 2017 archives
Move the frontend into web/, the repo's only npm workspace, and replace the
JS crawler with a Go module covering both remaining datasets. The crawler
writes to a .part file and renames on completion: writing straight to the
destination left truncated files that the skip-if-present check would then
skip forever.

Remove the 2017-old and 2017-old2 datasets. They were successive publications
of the same exam, kept side by side so the disagreement stayed inspectable;
the current 2017 supersedes them and they remain in git history.

Recover the 2016 crawler source from the Internet Archive's copy of the
aggregator article, whose original host no longer resolves. All 119 filenames
are verified against data/2016 in both directions, but no archive captured the
spreadsheets themselves, so the host still serving them is unconfirmed and
data/2016 remains the only confirmed copy.

Filenames are load-bearing throughout: go-parser sorts inputs bytewise and
inserts last-wins, so they decide which row survives a duplicate exam number.
2026-08-13 21:45:05 +07:00