mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
Both sources carried their download links as a hardcoded array, which is not a crawl: the lists could drift from what the articles actually published, and nothing would say so. A source now names the article and how to name what it finds there, and internal/article reads the links out of that page at run time. 2016 takes its filenames straight from the URL. 2017 cannot — the CDN names are inconsistent (Angiang.xls, 1BaRiaVungTau.xls, 23HaiPhong.xls) — so it derives them from the province in the link text, transliterated to ASCII the same way go-parser builds ho_ten_ascii. Filenames stay load-bearing: go-parser sorts inputs bytewise and inserts last-wins, so they decide which row survives a duplicate exam number. Saved copies of both articles are committed as fixtures, and a test asserts that reading them and applying each naming rule reproduces data/<id> exactly, in both directions. Resolve also rejects a page that yields the wrong number of links or two links that would write the same file, since either silently costs the dataset files that only the row-count guard would notice afterwards. Verified against the live 2017 article: a from-scratch crawl of all 63 files leaves the committed data unchanged.
9 lines
119 B
AMPL
9 lines
119 B
AMPL
module github.com/tiennm99/thptqg/crawler
|
|
|
|
go 1.26.5
|
|
|
|
require (
|
|
golang.org/x/net v0.58.0
|
|
golang.org/x/text v0.41.0
|
|
)
|