refactor: split into web/crawler/go-parser and drop the 2017 archives

Move the frontend into web/, the repo's only npm workspace, and replace the
JS crawler with a Go module covering both remaining datasets. The crawler
writes to a .part file and renames on completion: writing straight to the
destination left truncated files that the skip-if-present check would then
skip forever.

Remove the 2017-old and 2017-old2 datasets. They were successive publications
of the same exam, kept side by side so the disagreement stayed inspectable;
the current 2017 supersedes them and they remain in git history.

Recover the 2016 crawler source from the Internet Archive's copy of the
aggregator article, whose original host no longer resolves. All 119 filenames
are verified against data/2016 in both directions, but no archive captured the
spreadsheets themselves, so the host still serving them is unconfirmed and
data/2016 remains the only confirmed copy.

Filenames are load-bearing throughout: go-parser sorts inputs bytewise and
inserts last-wins, so they decide which row survives a duplicate exam number.
This commit is contained in:
tiennm99 committed 2026-08-13 21:45:05 +07:00
1 parent 8530df0154
commit ceb694a747
163 files changed
+1317 -510

No files matched your search

+14 -6
View File
@@ -49,25 +49,33 @@ jobs:
- name: Test parser
run: npm run test:go
# The crawler is not part of the build — it only refreshes data/ by hand.
# It is still compiled and tested here so it cannot rot unnoticed, and
# because its filename test guards the parser: input filenames decide
# which row survives a duplicate exam number.
- name: Test crawler
run: npm run test:crawler
# excelize carries an open advisory, and the 2017 refresh runbook feeds
# network-downloaded spreadsheets straight into the parser.
- name: Vulnerability scan
working-directory: go-parser
run: |
go install golang.org/x/vuln/cmd/govulncheck@latest
"$(go env GOPATH)/bin/govulncheck" ./...
GOVULNCHECK="$(go env GOPATH)/bin/govulncheck"
(cd go-parser && "$GOVULNCHECK" ./...)
(cd crawler && "$GOVULNCHECK" ./...)
# One parser binary builds every dataset; build-db.js reads the dataset
# list from src/datasets.js, verifies each database against its known row
# list from web/src/datasets.js, verifies each database against its row
# count, then gzips it in place leaving no uncompressed file behind.
- name: Build databases
run: |
npm run build:go
npm run build:db
# One Vite build produces every page. scripts/assemble-site.js copies the
# emitted index.html to each dataset path (and the legacy nested URLs),
# then fails the job if any uncompressed database reached the artifact.
# One Vite build produces every page. web/scripts/assemble-site.js copies
# the emitted index.html to each dataset path (and the legacy nested
# URLs), then fails the job if an uncompressed database reached _site.
- name: Build and assemble site
run: npm run build:site
+34 -17
View File
@@ -2,7 +2,7 @@
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
school graduation exam. Client-side SQL (sql.js) over a SQLite database built
from the ministry's raw `.xls` score files by the Rust `xlsxread` parser.
from the ministry's raw `.xls` score files by the Go `xlsxread` parser.
Live at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
@@ -10,43 +10,60 @@ Live at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
| --- | --- | --- | --- |
| `2016` | 2016 | 877,461 | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
| `2017` | 2017 | 861,068 | [/2017/](https://tiennm99.github.io/thptqg/2017/) |
| `2017-old` | 2017 | 847,348 | [/2017-old/](https://tiennm99.github.io/thptqg/2017-old/) |
| `2017-old2` | 2017 | 679,764 | [/2017-old2/](https://tiennm99.github.io/thptqg/2017-old2/) |
The three 2017 datasets are successive publications of the same exam and they
disagree; all three are kept so the differences stay inspectable.
Two earlier 2017 publications (`2017-old`, `2017-old2`) were kept for a while
because they disagreed with the current one. They have been removed; they remain
in git history.
## Layout
```
index.html + src/ the frontend — one app serving all four datasets and the hub
datasets.js the four dataset ids and their per-dataset content
router.js pathname → dataset
web/ the frontend — one Vite app serving both datasets and the hub
src/datasets.js the dataset ids and their per-dataset content
src/router.js pathname → dataset
scripts/ site assembly
crawler/ Go — re-fetches the source spreadsheets
internal/sources/ one file per data source: its links and local filenames
internal/fetch/ concurrent, resumable downloading
go-parser/ Go — Excel to SQLite
internal/schema/ canonical 22-column table: DDL, INSERT, subject regexes
configs/<id>.yml per-dataset parse rules only, no SQL
scripts/ database build, parity verification
data/<id>/ raw Excel files, one directory per dataset
parser/ the Rust parser
src/schema.rs canonical 22-column table: DDL, INSERT, subject regexes
configs/<id>.yml per-dataset parse rules only, no SQL
scripts/ database build, crawler, parity verification
scripts/ site assembly
docs/ architecture, data pipeline, deployment
```
`web/` is the only npm workspace; `crawler/` and `go-parser/` are independent Go
modules. The one cross-boundary import is `web/src/datasets.js`, which
`go-parser/scripts/build-db.js` reads for the dataset list and expected sizes.
The dataset id is one identifier end to end:
```
data/2017-old/ → go-parser/configs/2017-old.yml → db/2017-old.db.gz → /thptqg/2017-old/
data/2017/ → go-parser/configs/2017.yml → db/2017.db.gz → /thptqg/2017/
```
## Build
```bash
npm ci
npm run build:go # compile the parser
npm run build:db # build + gzip all four databases (add an id for just one)
npm run build:go # compile the parser
npm run build:db # build + gzip both databases (add an id for just one)
npm run build:site # one Vite build, then assemble into _site/
npx serve _site
```
The source spreadsheets are committed, so a crawl is only needed to refresh
them:
```bash
npm run crawl:2016 # re-fetch data/2016/
npm run crawl:2017 # re-fetch data/2017/ from the baotintuc.vn CDN
```
Crawling is idempotent — files already present are skipped — and is never part
of the build.
Pushing to `main` runs the same steps in
`.github/workflows/deploy-pages.yml` and publishes to GitHub Pages.
@@ -55,7 +72,7 @@ Pushing to `main` runs the same steps in
1. Put the Excel files in `data/<id>/`
2. Add `go-parser/configs/<id>.yml` — sheet mode, column indices, validation
guards. No SQL; the schema is canonical.
3. Add an entry to `DATASETS` in `src/datasets.js`
3. Add an entry to `DATASETS` in `web/src/datasets.js`
Everything else follows: the build script, the site assembly and the router all
read that one list, and the UI adapts to whichever columns the dataset fills.
+145
View File
@@ -0,0 +1,145 @@
// Command crawl downloads a dataset's source spreadsheets into data/<id>/.
//
// The argument is the dataset id, the same one used by go-parser's configs and
// the published site paths.
//
// crawl 2017 # from the baotintuc.vn CDN
// crawl 2016 # source not yet configured
// crawl 2017 --list # show what would be downloaded
//
// Runs are idempotent: a file already present and non-empty is skipped, so an
// interrupted crawl can simply be re-run.
package main
import (
"context"
"flag"
"fmt"
"os"
"os/signal"
"path/filepath"
"syscall"
"time"
"github.com/tiennm99/thptqg/crawler/internal/fetch"
"github.com/tiennm99/thptqg/crawler/internal/sources"
)
func main() {
if err := run(os.Args[1:]); err != nil {
fmt.Fprintf(os.Stderr, "crawl: %v\n", err)
os.Exit(1)
}
}
func usage() {
fmt.Fprint(os.Stderr, "usage: crawl <dataset> [flags]\n\nDatasets:\n")
for _, s := range sources.All() {
fmt.Fprintf(os.Stderr, " %-6s %s\n", s.ID, s.Summary)
}
fmt.Fprint(os.Stderr, "\nFlags:\n")
fmt.Fprint(os.Stderr, " --out string output directory (default ../data/<dataset>)\n")
fmt.Fprint(os.Stderr, " --concurrency int parallel downloads (default 6)\n")
fmt.Fprint(os.Stderr, " --timeout duration per-file timeout (default 2m)\n")
fmt.Fprint(os.Stderr, " --list print the file list and exit\n")
}
func run(args []string) error {
if len(args) == 0 || args[0] == "-h" || args[0] == "--help" {
usage()
if len(args) == 0 {
return fmt.Errorf("no dataset given")
}
return nil
}
src, err := sources.Lookup(args[0])
if err != nil {
usage()
return err
}
fs := flag.NewFlagSet(src.ID, flag.ContinueOnError)
// The default is relative to the crawler module directory, which is where
// both `go -C crawler run ./cmd/crawl` and a manual `cd crawler` land.
out := fs.String("out", filepath.Join("..", "data", src.ID), "output directory")
concurrency := fs.Int("concurrency", 6, "parallel downloads")
timeout := fs.Duration("timeout", 2*time.Minute, "per-file timeout")
list := fs.Bool("list", false, "print the file list and exit")
if err := fs.Parse(args[1:]); err != nil {
return err
}
files, err := src.Files()
if err != nil {
return err
}
if len(files) == 0 {
return fmt.Errorf("source %q produced no files", src.ID)
}
outDir, err := filepath.Abs(*out)
if err != nil {
return err
}
items := make([]fetch.Item, 0, len(files))
for _, f := range files {
items = append(items, fetch.Item{
Name: f.Name,
URL: f.URL,
Path: filepath.Join(outDir, f.Dest),
})
}
if *list {
for _, it := range items {
fmt.Printf("%s\t%s\n", filepath.Base(it.Path), it.URL)
}
return nil
}
// Ctrl-C cancels in-flight requests; each worker deletes its partial file
// on the way out, so an interrupted run leaves nothing half-written.
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
fmt.Printf("Downloading %d files to %s...\n", len(items), outDir)
results, runErr := fetch.Run(ctx, items, fetch.Options{
Concurrency: *concurrency,
Timeout: *timeout,
Headers: src.Headers,
OnResult: printResult,
})
if runErr != nil {
return runErr
}
ok, skip, failed := fetch.Tally(results)
fmt.Printf("\nDone. ok=%d skip=%d fail=%d\n", ok, skip, len(failed))
if len(failed) > 0 {
for _, r := range failed {
fmt.Fprintf(os.Stderr, " %s: %v\n", r.Item.Name, r.Err)
}
return fmt.Errorf("%d file(s) failed", len(failed))
}
return nil
}
func printResult(done, total int, r fetch.Result) {
tag := map[fetch.Status]string{
fetch.StatusOK: "✓",
fetch.StatusSkip: "·",
fetch.StatusFail: "✗",
}[r.Status]
detail := ""
switch r.Status {
case fetch.StatusFail:
detail = r.Err.Error()
default:
detail = fmt.Sprintf("%.0f KB", float64(r.Size)/1024)
}
fmt.Printf(" %s [%d/%d] %-20s %s\n", tag, done, total, r.Item.Name, detail)
}
+3
View File
@@ -0,0 +1,3 @@
module github.com/tiennm99/thptqg/crawler
go 1.26.5
+205
View File
@@ -0,0 +1,205 @@
// Package fetch downloads a list of files concurrently, skipping any that are
// already on disk.
//
// It is deliberately source-agnostic: it knows nothing about provinces, exam
// years or spreadsheet formats. Each source in internal/sources produces a
// []Item and this package moves the bytes.
package fetch
import (
"context"
"fmt"
"io"
"net/http"
"os"
"path/filepath"
"sync"
"time"
)
// Item is one file to download.
type Item struct {
// Name is the human label used in progress output, e.g. "An Giang".
Name string
URL string
// Path is the absolute destination, including filename.
Path string
}
// Status is the outcome of one item.
type Status int
const (
// StatusOK means the file was downloaded during this run.
StatusOK Status = iota
// StatusSkip means a non-empty file was already present.
StatusSkip
// StatusFail means the download did not complete; Result.Err says why.
StatusFail
)
// Result records what happened to one item.
type Result struct {
Item Item
Status Status
Size int64
Err error
}
// Options configures a run. The zero value is usable: Concurrency and Timeout
// fall back to defaults.
type Options struct {
// Concurrency is the number of files downloaded at once.
Concurrency int
// Timeout bounds each individual request, headers and body together.
// The original JS crawler had no timeout, so one stalled connection could
// hang the whole run indefinitely.
Timeout time.Duration
// Headers are sent with every request. The CDN this was written against
// rejects requests without a browser User-Agent and a matching Referer.
Headers map[string]string
// OnResult, if set, is called once per completed item. Calls are
// serialised, so it does not need its own locking, but they arrive in
// completion order rather than list order.
OnResult func(done, total int, r Result)
}
const (
defaultConcurrency = 6
defaultTimeout = 2 * time.Minute
)
// Run downloads every item, returning one Result each. It does not return an
// error for a failed download — that is reported per item — only for a problem
// that stops the run as a whole, such as an unusable output directory.
//
// Results come back in completion order, not input order.
func Run(ctx context.Context, items []Item, opts Options) ([]Result, error) {
if opts.Concurrency <= 0 {
opts.Concurrency = defaultConcurrency
}
if opts.Timeout <= 0 {
opts.Timeout = defaultTimeout
}
// Every item's parent directory must exist before any worker starts, so a
// missing output directory fails once here rather than N times in parallel.
for _, it := range items {
if err := os.MkdirAll(filepath.Dir(it.Path), 0o755); err != nil {
return nil, fmt.Errorf("cannot create output directory: %w", err)
}
}
client := &http.Client{Timeout: opts.Timeout}
var (
mu sync.Mutex
results = make([]Result, 0, len(items))
next int
)
work := make(chan Item)
var wg sync.WaitGroup
for range opts.Concurrency {
wg.Add(1)
go func() {
defer wg.Done()
for it := range work {
r := download(ctx, client, it, opts.Headers)
mu.Lock()
results = append(results, r)
next++
if opts.OnResult != nil {
opts.OnResult(next, len(items), r)
}
mu.Unlock()
}
}()
}
for _, it := range items {
select {
case work <- it:
case <-ctx.Done():
close(work)
wg.Wait()
return results, ctx.Err()
}
}
close(work)
wg.Wait()
return results, nil
}
// download fetches one item, or reports that it was already present.
//
// The body is streamed to a .part file and renamed into place only once it is
// complete. Without that, an interrupted run leaves a truncated file at the
// final path — and because the skip check only tests for a non-empty file,
// every later run would skip it and the corruption would persist silently.
func download(ctx context.Context, client *http.Client, it Item, headers map[string]string) Result {
if st, err := os.Stat(it.Path); err == nil && st.Size() > 0 {
return Result{Item: it, Status: StatusSkip, Size: st.Size()}
}
fail := func(err error) Result {
return Result{Item: it, Status: StatusFail, Err: err}
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, it.URL, nil)
if err != nil {
return fail(err)
}
for k, v := range headers {
req.Header.Set(k, v)
}
resp, err := client.Do(req)
if err != nil {
return fail(err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return fail(fmt.Errorf("HTTP %d", resp.StatusCode))
}
part := it.Path + ".part"
f, err := os.Create(part)
if err != nil {
return fail(err)
}
n, err := io.Copy(f, resp.Body)
if cerr := f.Close(); err == nil {
err = cerr
}
if err != nil {
os.Remove(part)
return fail(err)
}
if n == 0 {
os.Remove(part)
return fail(fmt.Errorf("empty response body"))
}
if err := os.Rename(part, it.Path); err != nil {
os.Remove(part)
return fail(err)
}
return Result{Item: it, Status: StatusOK, Size: n}
}
// Tally counts results by status.
func Tally(results []Result) (ok, skip int, failed []Result) {
for _, r := range results {
switch r.Status {
case StatusOK:
ok++
case StatusSkip:
skip++
case StatusFail:
failed = append(failed, r)
}
}
return ok, skip, failed
}
+184
View File
@@ -0,0 +1,184 @@
package fetch
import (
"context"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"testing"
)
func serve(t *testing.T, h http.HandlerFunc) string {
t.Helper()
s := httptest.NewServer(h)
t.Cleanup(s.Close)
return s.URL
}
func TestDownloadsAndSkips(t *testing.T) {
var hits int
url := serve(t, func(w http.ResponseWriter, _ *http.Request) {
hits++
w.Write([]byte("payload"))
})
dir := t.TempDir()
items := []Item{{Name: "one", URL: url, Path: filepath.Join(dir, "one.xls")}}
results, err := Run(context.Background(), items, Options{})
if err != nil {
t.Fatal(err)
}
if results[0].Status != StatusOK {
t.Fatalf("status = %v, err = %v", results[0].Status, results[0].Err)
}
if got, _ := os.ReadFile(items[0].Path); string(got) != "payload" {
t.Errorf("content = %q", got)
}
// Second run must not re-fetch.
results, err = Run(context.Background(), items, Options{})
if err != nil {
t.Fatal(err)
}
if results[0].Status != StatusSkip {
t.Errorf("status = %v, want StatusSkip", results[0].Status)
}
if hits != 1 {
t.Errorf("server hit %d times, want 1", hits)
}
}
func TestNon200Fails(t *testing.T) {
url := serve(t, func(w http.ResponseWriter, _ *http.Request) {
http.Error(w, "gone", http.StatusNotFound)
})
dir := t.TempDir()
path := filepath.Join(dir, "missing.xls")
results, err := Run(context.Background(), []Item{{Name: "x", URL: url, Path: path}}, Options{})
if err != nil {
t.Fatal(err)
}
if results[0].Status != StatusFail {
t.Fatalf("status = %v, want StatusFail", results[0].Status)
}
if _, err := os.Stat(path); !os.IsNotExist(err) {
t.Error("a failed download must leave no file at the destination")
}
assertNoPartFiles(t, dir)
}
// TestFailureLeavesNoPartial is the reason downloads land on a .part file
// first. Writing straight to the destination would leave a truncated file
// there, and the skip check — which only tests for a non-empty file — would
// skip it on every later run, so the corruption would never be re-fetched.
func TestFailureLeavesNoPartial(t *testing.T) {
url := serve(t, func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Length", "1000")
w.Write([]byte("short"))
// Closing early makes the body read fail mid-copy.
if f, ok := w.(http.Flusher); ok {
f.Flush()
}
panic(http.ErrAbortHandler)
})
dir := t.TempDir()
path := filepath.Join(dir, "truncated.xls")
results, err := Run(context.Background(), []Item{{Name: "x", URL: url, Path: path}}, Options{})
if err != nil {
t.Fatal(err)
}
if results[0].Status != StatusFail {
t.Fatalf("status = %v, want StatusFail", results[0].Status)
}
if _, err := os.Stat(path); !os.IsNotExist(err) {
t.Error("a truncated download must not be left at the destination")
}
assertNoPartFiles(t, dir)
}
func TestEmptyBodyFails(t *testing.T) {
url := serve(t, func(w http.ResponseWriter, _ *http.Request) {})
dir := t.TempDir()
path := filepath.Join(dir, "empty.xls")
results, err := Run(context.Background(), []Item{{Name: "x", URL: url, Path: path}}, Options{})
if err != nil {
t.Fatal(err)
}
if results[0].Status != StatusFail {
t.Errorf("an empty body must fail, got %v", results[0].Status)
}
assertNoPartFiles(t, dir)
}
func TestHeadersAreSent(t *testing.T) {
var gotUA, gotRef string
url := serve(t, func(w http.ResponseWriter, r *http.Request) {
gotUA, gotRef = r.Header.Get("User-Agent"), r.Header.Get("Referer")
w.Write([]byte("ok"))
})
_, err := Run(context.Background(),
[]Item{{Name: "x", URL: url, Path: filepath.Join(t.TempDir(), "x.xls")}},
Options{Headers: map[string]string{"User-Agent": "test-agent", "Referer": "https://example.test/"}})
if err != nil {
t.Fatal(err)
}
if gotUA != "test-agent" || gotRef != "https://example.test/" {
t.Errorf("headers not sent: ua=%q referer=%q", gotUA, gotRef)
}
}
func TestAllItemsRun(t *testing.T) {
url := serve(t, func(w http.ResponseWriter, _ *http.Request) {
w.Write([]byte("x"))
})
dir := t.TempDir()
var items []Item
for i := range 20 {
items = append(items, Item{
Name: "f",
URL: url,
Path: filepath.Join(dir, "sub", string(rune('a'+i))+".xls"),
})
}
results, err := Run(context.Background(), items, Options{Concurrency: 4})
if err != nil {
t.Fatal(err)
}
if len(results) != len(items) {
t.Fatalf("got %d results, want %d", len(results), len(items))
}
ok, _, failed := Tally(results)
if ok != len(items) {
t.Errorf("ok = %d, want %d (failures: %v)", ok, len(items), failed)
}
}
func TestTally(t *testing.T) {
ok, skip, failed := Tally([]Result{
{Status: StatusOK}, {Status: StatusOK}, {Status: StatusSkip}, {Status: StatusFail},
})
if ok != 2 || skip != 1 || len(failed) != 1 {
t.Errorf("ok=%d skip=%d fail=%d", ok, skip, len(failed))
}
}
func assertNoPartFiles(t *testing.T, dir string) {
t.Helper()
entries, err := os.ReadDir(dir)
if err != nil {
t.Fatal(err)
}
for _, e := range entries {
if filepath.Ext(e.Name()) == ".part" {
t.Errorf("leftover partial file: %s", e.Name())
}
}
}
+182
View File
@@ -0,0 +1,182 @@
package sources
// source2016 fetches the 2016 dataset: one spreadsheet per exam cluster
// (cụm thi), 4 .xls and 115 .xlsx.
//
// # Provenance
//
// The list below was recovered from the aggregator article that originally
// published these files:
//
// cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html
//
// The site that first carried it (dtntbacgiang.edu.vn) no longer resolves, so
// the list was read from the Internet Archive's copy. It is verified, not
// guessed: all 119 filenames match data/2016/ exactly in both directions, which
// TestSource2016MatchesFilesOnDisk asserts.
//
// # Filenames are load-bearing — do not "tidy" them
//
// Dest is the server-assigned name, a 32-hex content hash followed by the
// cluster slug and a millisecond timestamp. It is kept verbatim because
// go-parser sorts input files bytewise and inserts last-wins, so filenames
// decide which row survives a duplicate exam number. That is live for 2016, not
// hypothetical: 877,464 source rows collapse to 877,461, so three rows' contents
// depend on this sort order. The hashes in
// go-parser/testdata/reader-fidelity-hashes.tsv are keyed by full path as well,
// and they are frozen — they were produced by the Rust reader, which no longer
// exists.
//
// # The host is not verified
//
// Only the paths are known to be right; the host that still serves them is not.
// The original is gone and the Internet Archive captured the article but none of
// the spreadsheets, so this could not be confirmed by fetching. baseURL points
// at a mirror of the same article that is still online. If it serves these files
// from a different directory, uploadPath is the single constant to change.
const (
baseURL = "https://dtnt.bacninh.edu.vn"
uploadPath = "/upload/s/20180102/"
)
var source2016 = Source{
ID: "2016",
Summary: "119 exam-cluster files (4 .xls + 115 .xlsx)",
Files: func() ([]File, error) {
out := make([]File, 0, len(clusters2016))
for _, c := range clusters2016 {
out = append(out, File{
Name: c.name,
URL: baseURL + uploadPath + c.file,
Dest: c.file,
})
}
return out, nil
},
}
type cluster struct{ name, file string }
// clusters2016 is in the article's publication order. That order is incidental —
// go-parser re-sorts by filename — but it is kept as published so the list can
// be diffed against the source article.
var clusters2016 = []cluster{
{"Trường Đại học Bách khoa Hà Nội", "b34c777942ca8a4de91f23a35cce1c6bdhbachkhoahn-1468901491486.xlsx"},
{"Trường Đại học Sư phạm Hà Nội", "76f19098881e9c79f66ac4f1b42a8a55dhsuphamhanoi-1468901144203.xlsx"},
{"Trường Đại học Thuỷ lợi * Cơ sở 1 ở phía Bắc", "4a3fa95a83fe9b3b2fafab309323fd22dhthuyloi-1468901840919.xlsx"},
{"Học viện Kỹ thuật Quân sự * Cơ sở 1 ở phía Bắc (Quân đội)", "4334ea476a2b8196d54ba40a341831echvkythuatquansu-1468920342994.xlsx"},
{"Trường Đại học Lâm nghiệp", "c5ecc58a731fd1a54b7e6cc3bc99449ddhlamnghiep-1468939300790.xlsx"},
{"Trường Đại học Bách khoa - Đại học Quốc gia Thành phố Hồ Chí Minh", "357c6aaf59aeaa6c11c3e83c595d38cfdhbachkhoa-tphcm-1468939678829.xlsx"},
{"Trường Đại học Khoa học tự nhiên - Đại học Quốc gia Thành phố Hồ Chí Minh", "64e44466ffeff9a587260795644805d0dhkhtunhien-dhqgtphcm-1468921015146.xlsx"},
{"Trường Đại học Khoa học xã hội và Nhân văn - Đại học Quốc gia Thành phố Hồ Chí Minh", "5fd13ae5287cbf0149bd733b507e0608dhkhxhnv-dhqgtphcm-1468920630193.xlsx"},
{"Trường Đại học Sư phạm Tp.HCM", "a995de3f7a7055cf10b2f002f0e194f7dhsptphcm-1468920792291.xlsx"},
{"Trường Đại học Hàng Hải", "c6ef16bc9f75ce0fd16c1229afad4e71dhhanghai-1468906457479.xlsx"},
{"Trường Đại học Sư phạm - Đại học Thái Nguyên", "ee09c723da4e86cdbe26213413450af2dhsupham-dhthainguyen-1468939061872.xlsx"},
{"Trường Đại học Kỹ thuật công nghiệp - Đại học Thái Nguyên", "8af8f1b29b6a40e26ccb4a2e19e8eededhkythuatcongnghiep-dhthainguyen-1468901919568.xlsx"},
{"Trường Đại học Nông lâm - Đại học Thái Nguyên", "53e5e20ebd74d84741513b703fbd13bfdhnonglam-dhthainguyen-1468939023637.xlsx"},
{"Học viện Ngân hàng", "f2846894bd5533e02b122a1e3b4333f6hvnganhang-1468939472850.xlsx"},
{"Trường Đại học Luật Hà Nội", "9e0e564ae801e347bd7649da8b89ad90dhluathn-1468920377978.xlsx"},
{"Trường Đại học Tân Trào", "20db5eaace927596f1806f151649ee6edhtantrao-1468901296040.xlsx"},
{"Trường Đại học Xây dựng Hà Nội", "7b5d74564ab9f6e0e137f6d161fab6f0dhxaydung-1468940561682.xlsx"},
{"Trường Đại học Khoa học - Đại học Thái Nguyên", "c2cfaaf287fb1e7120894e545771b027dhkhoahoc-dhthainguyen-1468900365545.xlsx"},
{"Đại học Thái Nguyên", "3daa6ee33fdafadc3c112c18bff18c2ddhthainguyen-1468940268065.xlsx"},
{"Học viện Tài chính", "96c95c85c5549a06f2e92ffcacf8ff21hvtaichinh-1468934470924.xlsx"},
{"Trường Đại học Tây Bắc", "74e6cd0dc78a3fb7dcf56ae57283b9b9dhtaybac-1468940302052.xlsx"},
{"Trường Đại học Hùng Vương", "778ab074ed6a95d209fa901fd5d68843dhhungvuong-1468901964127.xlsx"},
{"Trường Đại học Sư phạm Hà Nội 2", "a11e707e6ed5fe4d29a8e32380d7558cdhsphanoi-2-1468902006284.xlsx"},
{"Trường Đại học Ngoại thương * Cơ sở 1 ở phía Bắc", "ec84067268f47d523b1e80b493248377dhngoaithuong-1468920480753.xlsx"},
{"Trường Đại học Kinh tế Quốc dân", "8f948c2a5ab49a3c5fd03b6ea04887a7dhkinhtequocdanfix-1468902083782.xlsx"},
{"Trường Đại học Giao thông Vận tải", "51bb2f9b6e9f565c2d705606fd56958bdhgiaothongvantai-1468939107463.xlsx"},
{"Học viện Nông Nghiệp Việt Nam", "5aede4951e7af8c1e1056c747a553386hvnongnghiepvn-1468900610577.xlsx"},
{"Trường Đại học Sư phạm Kỹ thuật Hưng Yên", "f85f04bb2bbd84457b83ac846728a2a6dhspkythuathungyen-1468920667091.xlsx"},
{"Trường Đại học Hải Phòng", "deca197a916a8f633f6894b4e13c3ca4dhhaiphong-1468901178368.xlsx"},
{"Trường Đại học Thương mại", "b2552e20c9ff4b237244dfbf2e278a47dhthuongmai-1468902876264.xlsx"},
{"Trường Đại học Công nghiệp Hà Nội", "82338a4a93ea1e0cde26e0061066ca8adhcongnghiephanoi-1468920154757.xlsx"},
{"Y Dược Thái Bình", "22e8617886412abd90ae3d33cc6d7db4dhyduocthaibinh-1468941020145.xlsx"},
{"Trường Đại học Mỏ Địa chất", "c174631d90303276ac63ff38b31157bbdhmodiachat-1468900697349.xlsx"},
{"Trường Đại học Hồng Đức", "bd857add0cfe92c1c247469ff2104527dh-hong-duc-1468900427119.xlsx"},
{"Trường Đại học Vinh", "64d2635efd30a0d0db4f7c30b9735261dhvinh-1468940230027.xlsx"},
{"Trường ĐH Sư phạm - ĐH Huế", "ab2f259ef55cb08b3a436d3f15d42cc7dhsupham-dhhue-1468934431828.xlsx"},
{"Trường ĐH Khoa học - ĐH Huế", "87ffaf3197ba20f0919f98a0b3f33914dhkhoahoc-dhhue-1468938729472.xlsx"},
{"Trường ĐH Kinh tế - ĐH Huế", "8916da7b3c4d58cffc1dd29dc7a8aabfdhkinhte-dhhue-1468938560205.xlsx"},
{"Đại học Huế", "919276a0348634459f8937dfcd6c7129dhhue-1468938792115.xlsx"},
{"ĐH Đà Nẵng", "b42fd58ead4c54279d7df5f53160a851dhdanang-1468920187203.xlsx"},
{"Trường ĐH Bách khoa - ĐH Đà Nẵng", "43213f61074068d9177685eb6c3f2d3bdhbachkhoa-dhdanang-1468938494437.xlsx"},
{"Trường ĐH Sư phạm - ĐH Đà Nẵng", "54bbcf865c137d58cf62af15733bd376dhsupham-dhdanang-1468938526558.xlsx"},
{"Trường Đại học Quy Nhơn", "70bf8cd2b6e87523219e55149a8524c5dhquynhon-1468938878328.xlsx"},
{"Trường ĐH Xây dựng miền Trung", "4e9bb192d56c9fa92b3e0ec13d8236c8dhxaydungmientrung-1468940978346.xlsx"},
{"Trường ĐH Nông Lâm TPHCM", "b2a739584a8861502573d1c6372fd6e9dhnonglamtphcm-1468939555326.xlsx"},
{"Trường ĐH Ngoại ngữ - ĐH Đà Nẵng", "9ba098a7913ad837a66e173b3fc41fd7dhngoaingu-dhdanang-1468902750424.xlsx"},
{"Trường ĐH Tây Nguyên", "216ee4e385dbc236767be899b068bf98dhtaynguyen-1468940490240.xlsx"},
{"Trường ĐH Tài chính Marketing", "af6828f837a80e03d77a65f3d310ec60dh-tai-chinh-marketing-1468924711596.xlsx"},
{"Trường ĐH Nha Trang", "6375fd13557bee08d20f1f3e46dac130dhnhatrang-coso1-1468901350425.xlsx"},
{"Trường ĐH Giao thông vận tải TP.HCM", "c5c28959fcd0b4e39dd9fd538db5448adhgtvttphcm-1468939150821.xlsx"},
{"Trường ĐH Sư phạm Kĩ thuật TP.HCM", "514877c1cfebdc9c2251f651b5f02c7edhspkythuattphcm-1468906568655.xlsx"},
{"Trường ĐH Đà Lạt", "2b33640b0dd154490f2809d96be81797dhdalat-1468902130616.xlsx"},
{"Trường ĐH Kinh tế TP.HCM", "f0aa5a19c2c8c1fde8211341f0e98e26dhkinhtetphcm-1468939229043.xlsx"},
{"Trường ĐH Kinh tế - Luật - ĐHQG TP.HCM", "ff877788e43bd84f0119dae026996892dhkinhteluat-dhqgtphcm-1468939860894.xlsx"},
{"Trường ĐH Công nghiệp Thực phẩm TP.HCM", "167116af6c8a096c2a9307edfa115f77dhcnthucpham-tphcm-1468910578187.xls"},
{"Trường ĐH Công nghiệp TP.HCM", "487c7287c1c9735722b31d877942becddhcongnghieptphcm-1468983373303.xlsx"},
{"Trường ĐH Sài Gòn", "d75282bff0a4d8cd1421c3fb74f789c3dhsaigon-1468900867561.xlsx"},
{"Trường ĐH Đồng Tháp", "388fcdfdac18ef6eee6125a2a015714ddhdongthap-1468934545359.xlsx"},
{"Trường Đại học An Giang", "fc511852659150f6305f6f624445300ddhangiang-1468940189321.xlsx"},
{"Trường ĐH Tôn Đức Thắng", "9d78f6dce7965fb4b7a337279920d4b1dhtonducthang-1468902168623.xlsx"},
{"Trường ĐH Tiền Giang", "bacd1b2133b37d0ea9fbf46445265036dhtiengiang-1468940391597.xlsx"},
{"Trường ĐH Cần Thơ", "b35a3f26a162bec3740e2da5639327afdhcantho-1468907131146.xls"},
{"Trường ĐH Cần Thơ - Hậu Giang", "7f1cbe7f35d1f539ca817a0ab8bd00c2dhcantho-haugiang-1468907167921.xls"},
{"Trường ĐH Luật TP.HCM", "d652d24e8bd4fe6a48d634b116f37c27dhluattphcm-1468939398928.xlsx"},
{"Trường ĐH Sư phạm Kỹ thuật Vĩnh Long", "4c2fd846ed2a0670ba6d1200d1d68e1dspkythuatvinhlong-1468902640438.xlsx"},
{"Trường ĐH Trà Vinh", "7418e83ff9d07b146216a3beafa927bbdhtravinh-1468902716174.xlsx"},
{"Trường ĐH Ngân hàng TP.HCM", "86a692bebf9b0b7c58adb9876b9befd5dhnganhangtphcm-1468902244566.xlsx"},
{"Trường ĐH Cần Thơ - Bạc Liêu", "5ca241aa1de8d25a3e6ff6db75234025dhcantho-baclieu-1468907071460.xls"},
{"Trường ĐH Kiên Giang", "cdc912664d7217c75eb9b67c6ef3797adhkiengiang-1468901222446.xlsx"},
{"Trường ĐH Y Dược Cần Thơ", "023718c7d3cf7ace3a7116fabb12bd9cdhyduoccantho-1468920829104.xlsx"},
{"Sở GDĐT Hà Nội", "1c66db0265bf45df237b37a7bcaeaef3hanoi-1468932790746.xlsx"},
{"Sở GDĐT Hà Giang", "178814c33c8b85a77d0fc295fa6d2be1hagiang-1468932847230.xlsx"},
{"Sở GDĐT Cao Bằng", "ddb631f833af894200ef10cd408d2ef0caobang-1468932877769.xlsx"},
{"Sở GDĐT Lai Châu", "a3aa106104f114c6426deb2d5e5da5fdlaichau-1468932954079.xlsx"},
{"Sở GDĐT Lào Cai", "d035f470c13adbf5785797c6ccdcb204laocai-1468919805337.xlsx"},
{"xem TẠI ĐÂY ) 07 Sở GDĐT Tuyên Quang", "d5ecd2574b5ab36c62dffad435262cb8tuyenquang-1468902296807.xlsx"},
{"Sở GDĐT Lạng Sơn", "7afd5d7fdff7c170a6342cd8f9474b22langson-1468932984587.xlsx"},
{"Sở GDĐT Bắc Kạn", "17843bccde10f362c2b9600bc49cd3f8backan-1468933016076.xlsx"},
{"Sở GDĐT Thái Nguyên", "5fb3a1b66aff077ee3da482be08cd391thainguyen-1468902797020.xlsx"},
{"Sở GDĐT Yên Bái", "97aa741a4c2b1276559eb021779daa32yenbai-1468933116055.xlsx"},
{"Sở GDĐT Sơn La", "4b98fdd7300b86a94636e4d15394fc19sonla-1468899319124.xlsx"},
{"Sở GDĐT Phú Thọ", "5465cca59495dcd9a48f689e2d38b85ephutho-1468902409682.xlsx"},
{"xem TẠI ĐÂY ) 14 Sở GDĐT Vĩnh Phúc", "4e0ccee19eb64513d23a6dff96ff7f9cvinhphuc-1468919874814.xlsx"},
{"Sở GDĐT Quảng Ninh", "a31b3446bae5a17c517348d8c38c5b7equangninh-1468919920197.xlsx"},
{"Sở GDĐT Bắc Giang", "b7a0c1ecce7c8445b3a1021091a5e654bacgiang-1468983153932.xlsx"},
{"Sở GDĐT Bắc Ninh", "7e0d0b6fb981d407cffdb79d724e8aecbacninh-1468919984998.xlsx"},
{"Sở GDĐT Hải Dương", "7cfc91bd6a1f5ae9d7917e920f469b90haiduong-1468899640007.xlsx"},
{"Sở GDĐT Hưng Yên", "92045214159cb0335b2836aa1c4bb25ahungyen-1468899732002.xlsx"},
{"Sở GDĐT Hoà Bình", "505a876dfd1ea321291e29192aebe95ehoabinh-1468983220125.xlsx"},
{"Sở GDĐT Hà Nam", "e693bc9de67acc916032c6a3b2fc9eedhanam-1468933146526.xlsx"},
{"Sở GDĐT Nam Định", "dfa63c5b68a8f638379ca0e9c5a1c556namdinh-1468933234061.xlsx"},
{"Sở GDĐT Thái Bình", "6960b6fa493e1365a183af461e945389thaibinh-1468920019273.xlsx"},
{"xem TẠI ĐÂY ) 24 Sở GDĐT Ninh Bình", "0b583c9875443e65ffa449d5ca76fae8ninhbinh-1468899778557.xlsx"},
{"xem TẠI ĐÂY ) 25 Sở GDĐT Thanh Hoá", "238adc19e80c0daef089a307af23edcethanhhoa-1468933304999.xlsx"},
{"Sở GDĐT Nghệ An", "7f12b90225e91f9d00aac3bcdf9974e1nghean-1468920043121.xlsx"},
{"Sở GDĐT Quảng Bình", "9c730c43bf4309171528c3853ed53b19quangbinh-1468933376032.xlsx"},
{"Sở GDĐT Quảng Trị", "2208903e1fa12220d100f93e40d5ff3bquangtri-1468933440097.xlsx"},
{"Sở GDĐT Thừa Thiên -Huế", "6badc7cfa258b59ad67e2b64f713b95fthuathienhue-1468983258438.xlsx"},
{"Sở GDĐT Quảng Nam", "ec765c32190773f5a18cd8202a5eaf0dquangnam-1468933501467.xlsx"},
{"Sở GDĐT Quảng Ngãi", "e80dc028d53fa49ee247391a0653dee5quangngai-1468933592400.xlsx"},
{"Sở GDĐT Kon Tum", "cc57dc986acb525086b86e85df54949dkontum-1468933632113.xlsx"},
{"Sở GDĐT Bình Định", "412e62265f35ca0078a0d94af5d1ed8abinhdinh-1468899884192.xlsx"},
{"Sở GDĐT Gia Lai", "698b930ab7697a0672bbc39168024c9cgialai-1468933753098.xlsx"},
{"Sở GDĐT Đắk Lắk", "2004a8d225fd87524e324d2d519e87a4daclac-1468899960505.xlsx"},
{"Sở GDĐT Khánh Hoà", "4ad416ccf014bd5d01f415238566b210khanhhoa-1468933789968.xlsx"},
{"Sở GDĐT Lâm Đồng", "2fa6d9a308bd11bc09f02d787eef2c15lamdong-1468933821082.xlsx"},
{"Sở GDĐT Ninh Thuận", "15c85d901a73dd49f2bae71aadcdbb5cninhthuan-1468920071763.xlsx"},
{"Sở GDĐT Đồng Nai", "538c9676b70f2509573245d58fba93bcdongnai-1468938241935.xlsx"},
{"Sở GDĐT Đồng Tháp", "575df8a4df74e87abb8f55d807324374dongthap-1468933902699.xlsx"},
{"Sở GDĐT Kiên Giang", "7435bb054e8e4ec34c173af1b6ad8402kiengiang-1468900147395.xlsx"},
{"Sở GDĐT Cần Thơ", "5b16d7156a03bb7ed0618796bae87bc3cantho-1468933954714.xlsx"},
{"Sở GDĐT Bến Tre", "74145979ee30186bf4e40ee3e6a3ac74bentre-1468934010040.xlsx"},
{"Sở GDĐT Vĩnh Long", "7f4362ef567cdf74495368e0fc11072avinhlong-1468934067981.xlsx"},
{"Sở GDĐT Trà Vinh", "08cdeb3636dcb2adaef829d62968274atravinh-1468902443859.xlsx"},
{"Sở GDĐT Sóc Trăng", "a8ff80d9478747e36d21ec95cc9ecb8fsoctrang-1468902686409.xlsx"},
{"Sở GDĐT Bạc Liêu", "18055d17e46854ecc6c8be585bcb6e5cbaclieu-1468934144150.xlsx"},
{"Sở GDĐT Điện Biên", "3c0e56abd8c5334f4e1ff7df3ac23a15dienbien-1468920113820.xlsx"},
{"Sở GDĐT Đăk Nông", "bf259a908ff8a8dcc067a761204459f4daknong-1468934173803.xlsx"},
{"Sở GDĐT Hậu Giang", "964c4368131f6be6058fabb1caefba1chaugiang-1468934247324.xlsx"}}
+121
View File
@@ -0,0 +1,121 @@
package sources
import (
"regexp"
"strings"
)
// article is the page the links were taken from. The CDN rejects requests that
// do not carry it as a Referer.
const article = "https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm"
// browserUA is required too: the CDN 403s an unrecognised User-Agent.
const browserUA = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0 Safari/537.36"
// source2017 fetches the 2017 dataset from the baotintuc.vn CDN: one .xls per
// province.
//
// This is the only dataset that can be re-fetched from its original source.
// The CDN filenames are inconsistent (Angiang.xls, 1BaRiaVungTau.xls,
// 23HaiPhong.xls, gia-lai.xls), so local names are derived from the province
// name instead — which is what produced the data/2017/ files currently on disk,
// and what any re-crawl must keep producing.
var source2017 = Source{
ID: "2017",
Summary: "63 province .xls files from the baotintuc.vn CDN",
Headers: map[string]string{
"User-Agent": browserUA,
"Referer": article,
},
Files: func() ([]File, error) {
out := make([]File, 0, len(provinces2017))
for _, p := range provinces2017 {
out = append(out, File{
Name: p.name,
URL: p.url,
Dest: slug(p.name) + ".xls",
})
}
return out, nil
},
}
var nonAlphanumeric = regexp.MustCompile(`[^a-z0-9]+`)
// slug converts a province name to its local filename stem: lowercase, with
// every run of non-alphanumeric characters collapsed to a single hyphen.
//
// "Ba Ria - Vung Tau" -> "ba-ria-vung-tau"
//
// The names are already unaccented ASCII, so no transliteration is involved.
func slug(name string) string {
return strings.Trim(nonAlphanumeric.ReplaceAllString(strings.ToLower(name), "-"), "-")
}
type province struct{ name, url string }
var provinces2017 = []province{
{"An Giang", "https://cdnmedia.baotintuc.vn/2017/07/06/17/57/Angiang.xls"},
{"Bac Lieu", "https://cdnmedia.baotintuc.vn/2017/07/06/17/59/Baclieu.xls"},
{"Ba Ria - Vung Tau", "https://cdnmedia.baotintuc.vn/2017/07/06/08/17/1BaRiaVungTau.xls"},
{"Bac Giang", "https://cdnmedia.baotintuc.vn/2017/07/06/09/33/BacGiang.xls"},
{"Bac Kan", "https://cdnmedia.baotintuc.vn/2017/07/06/08/27/BacKan.xls"},
{"Bac Ninh", "https://cdnmedia.baotintuc.vn/2017/07/06/09/33/BacNinh.xls"},
{"Ben Tre", "https://cdnmedia.baotintuc.vn/2017/07/06/09/34/BenTre.xls"},
{"Binh Duong", "https://cdnmedia.baotintuc.vn/2017/07/06/09/34/BinhDuong.xls"},
{"Binh Thuan", "https://cdnmedia.baotintuc.vn/2017/07/06/09/35/BinhThuan.xls"},
{"Binh Phuoc", "https://cdnmedia.baotintuc.vn/2017/07/06/09/09/BinhPhuoc.xls"},
{"Binh Dinh", "https://cdnmedia.baotintuc.vn/2017/07/06/13/05/Binhdinh.xls"},
{"Ca Mau", "https://cdnmedia.baotintuc.vn/2017/07/06/13/05/Camau.xls"},
{"Cao Bang", "https://cdnmedia.baotintuc.vn/2017/07/06/18/02/Caobang.xls"},
{"Can Tho", "https://cdnmedia.baotintuc.vn/2017/07/06/13/07/Cantho.xls"},
{"Da Nang", "https://cdnmedia.baotintuc.vn/2017/07/06/18/03/Danang.xls"},
{"Dak Nong", "https://cdnmedia.baotintuc.vn/2017/07/06/09/35/DakNong.xls"},
{"Dak Lak", "https://cdnmedia.baotintuc.vn/2017/07/06/18/02/Daklak.xls"},
{"Dong Nai", "https://cdnmedia.baotintuc.vn/2017/07/06/13/08/dongnai.xls"},
{"Dong Thap", "https://cdnmedia.baotintuc.vn/2017/07/06/17/59/Dongthap.xls"},
{"Dien Bien", "https://cdnmedia.baotintuc.vn/2017/07/06/09/08/DienBien.xls"},
{"Gia Lai", "https://cdnmedia.baotintuc.vn/2017/07/06/13/09/Gia-Lai.xls"},
{"Ha Giang", "https://cdnmedia.baotintuc.vn/2017/07/06/18/18/Hagiang.xls"},
{"Ha Noi", "https://cdnmedia.baotintuc.vn/2017/07/07/08/16/HaNoi.xls"},
{"Ha Nam", "https://cdnmedia.baotintuc.vn/2017/07/06/09/07/Hanam.xls"},
{"Ha Tinh", "https://cdnmedia.baotintuc.vn/2017/07/06/09/36/HaTinh.xls"},
{"Hai Phong", "https://cdnmedia.baotintuc.vn/2017/07/06/08/18/23HaiPhong.xls"},
{"Hai Duong", "https://cdnmedia.baotintuc.vn/2017/07/06/09/36/HaiDuong.xls"},
{"Hau Giang", "https://cdnmedia.baotintuc.vn/2017/07/06/18/00/Haugiang.xls"},
{"Ho Chi Minh", "https://cdnmedia.baotintuc.vn/2017/07/06/08/26/HCM.xls"},
{"Hoa Binh", "https://cdnmedia.baotintuc.vn/2017/07/06/09/06/HoaBinh.xls"},
{"Hung Yen", "https://cdnmedia.baotintuc.vn/2017/07/06/08/25/HungYen.xls"},
{"Khanh Hoa", "https://cdnmedia.baotintuc.vn/2017/07/06/18/04/Khanhhoa.xls"},
{"Kien Giang", "https://cdnmedia.baotintuc.vn/2017/07/06/08/28/KienGiang.xls"},
{"Kon Tum", "https://cdnmedia.baotintuc.vn/2017/07/06/18/05/KonTum.xls"},
{"Nam Dinh", "https://cdnmedia.baotintuc.vn/2017/07/06/07/55/13NamDinh.xls"},
{"Nghe An", "https://cdnmedia.baotintuc.vn/2017/07/06/07/55/14NgheAn.xls"},
{"Ninh Binh", "https://cdnmedia.baotintuc.vn/2017/07/06/09/37/NinhBinh.xls"},
{"Ninh Thuan", "https://cdnmedia.baotintuc.vn/2017/07/06/18/00/Ninhthuan.xls"},
{"Lao Cai", "https://cdnmedia.baotintuc.vn/2017/07/06/08/52/11LaoCai.xls"},
{"Lai Chau", "https://cdnmedia.baotintuc.vn/2017/07/06/13/33/LaiChau.xls"},
{"Lang Son", "https://cdnmedia.baotintuc.vn/2017/07/06/18/05/Langson.xls"},
{"Lam Dong", "https://cdnmedia.baotintuc.vn/2017/07/06/09/00/LamDong.xls"},
{"Long An", "https://cdnmedia.baotintuc.vn/2017/07/06/08/53/12LongAn.xls"},
{"Quang Binh", "https://cdnmedia.baotintuc.vn/2017/07/06/09/23/QuangBinh.xls"},
{"Quang Nam", "https://cdnmedia.baotintuc.vn/2017/07/06/09/24/QuangNam.xls"},
{"Quang Ninh", "https://cdnmedia.baotintuc.vn/2017/07/06/18/06/Quangninh.xls"},
{"Quang Ngai", "https://cdnmedia.baotintuc.vn/2017/07/06/09/24/QuangNgai.xls"},
{"Quang Tri", "https://cdnmedia.baotintuc.vn/2017/07/06/09/25/QuangTri.xls"},
{"Phu Tho", "https://cdnmedia.baotintuc.vn/2017/07/06/09/23/PhuTho.xls"},
{"Phu Yen", "https://cdnmedia.baotintuc.vn/2017/07/06/09/38/PhuYen.xls"},
{"Son La", "https://cdnmedia.baotintuc.vn/2017/07/06/18/01/Sonla.xls"},
{"Soc Trang", "https://cdnmedia.baotintuc.vn/2017/07/06/18/06/Soctrang.xls"},
{"Tay Ninh", "https://cdnmedia.baotintuc.vn/2017/07/06/09/26/TayNinh.xls"},
{"Thai Binh", "https://cdnmedia.baotintuc.vn/2017/07/06/09/09/ThaiBinh.xls"},
{"Thai Nguyen", "https://cdnmedia.baotintuc.vn/2017/07/06/13/35/ThaiNguyen.xls"},
{"Thanh Hoa", "https://cdnmedia.baotintuc.vn/2017/07/06/13/10/thanhoa.xls"},
{"Tra Vinh", "https://cdnmedia.baotintuc.vn/2017/07/06/09/38/TraVinh.xls"},
{"Thua Thien Hue", "https://cdnmedia.baotintuc.vn/2017/07/06/13/10/thuathienhue.xls"},
{"Tien Giang", "https://cdnmedia.baotintuc.vn/2017/07/06/13/11/tiengiang.xls"},
{"Tuyen Quang", "https://cdnmedia.baotintuc.vn/2017/07/06/09/39/TuyenQuang.xls"},
{"Vinh Phuc", "https://cdnmedia.baotintuc.vn/2017/07/06/09/40/VinhPhuc.xls"},
{"Vinh Long", "https://cdnmedia.baotintuc.vn/2017/07/06/13/13/vinhlong.xls"},
{"Yen Bai", "https://cdnmedia.baotintuc.vn/2017/07/06/18/07/Yenbai.xls"},
}
+53
View File
@@ -0,0 +1,53 @@
// Package sources describes where each dataset's spreadsheets come from.
//
// A source is a dataset id plus the list of remote files that fill it.
// Downloading is internal/fetch's job; deciding what to download and what to
// call it locally is this package's.
package sources
import "fmt"
// File is one remote spreadsheet.
type File struct {
// Name is the human label used in progress output, e.g. "An Giang".
Name string
URL string
// Dest is the filename within the dataset directory.
//
// This is not cosmetic. go-parser sorts its input files and inserts with
// INSERT OR REPLACE, which is last-wins, so filenames decide which row
// survives a duplicate exam number. A re-crawl that names files differently
// can produce a database with the same row count and different content.
Dest string
}
// Source is one crawlable dataset.
type Source struct {
// ID is the dataset id: the subcommand, the directory under data/, and the
// go-parser config name, all at once. Keeping it single means a source
// cannot be pointed at the wrong dataset's directory.
ID string
Summary string
// Headers are sent with every request for this source.
Headers map[string]string
// Files lists what to download. It returns an error rather than an empty
// list when a source is known but not yet configured, so "nothing to do"
// can never be mistaken for success.
Files func() ([]File, error)
}
// registry is ordered: it drives the help output.
var registry = []Source{source2016, source2017}
// All returns every known source, in help order.
func All() []Source { return registry }
// Lookup finds a source by ID.
func Lookup(id string) (Source, error) {
for _, s := range registry {
if s.ID == id {
return s, nil
}
}
return Source{}, fmt.Errorf("unknown dataset %q", id)
}
+156
View File
@@ -0,0 +1,156 @@
package sources
import (
"os"
"path/filepath"
"strings"
"testing"
)
func TestSlug(t *testing.T) {
for in, want := range map[string]string{
"An Giang": "an-giang",
"Ba Ria - Vung Tau": "ba-ria-vung-tau",
"Ho Chi Minh": "ho-chi-minh",
"Thua Thien Hue": "thua-thien-hue",
} {
if got := slug(in); got != want {
t.Errorf("slug(%q) = %q, want %q", in, got, want)
}
}
}
func TestSource2017IsComplete(t *testing.T) {
files, err := source2017.Files()
if err != nil {
t.Fatal(err)
}
if len(files) != 63 {
t.Errorf("got %d provinces, want 63", len(files))
}
// A duplicate Dest would silently overwrite one province with another and
// leave the dataset one file short, with no error anywhere.
seenDest := map[string]string{}
seenURL := map[string]string{}
for _, f := range files {
if prev, dup := seenDest[f.Dest]; dup {
t.Errorf("%q and %q both write to %s", prev, f.Name, f.Dest)
}
seenDest[f.Dest] = f.Name
if prev, dup := seenURL[f.URL]; dup {
t.Errorf("%q and %q share a URL: %s", prev, f.Name, f.URL)
}
seenURL[f.URL] = f.Name
if !strings.HasPrefix(f.URL, "https://") {
t.Errorf("%s: URL is not https: %s", f.Name, f.URL)
}
if !strings.HasSuffix(f.Dest, ".xls") {
t.Errorf("%s: Dest %q does not end in .xls", f.Name, f.Dest)
}
}
}
// TestMatchesFilesOnDisk is the guard that matters.
//
// go-parser sorts its inputs and inserts last-wins, so filenames decide which
// row survives a duplicate exam number. If a re-crawl started naming files
// differently, the rebuilt database could differ in content while still passing
// the row-count guard in build-db.js. The committed data/<id>/ is the oracle:
// whatever a source would write must match it exactly, in both directions.
//
// For 2016 this is also the only evidence that the recovered link list is the
// right one — the host that served those files is gone, so the list cannot be
// checked by fetching it.
func TestMatchesFilesOnDisk(t *testing.T) {
for _, src := range All() {
t.Run(src.ID, func(t *testing.T) {
dir := filepath.Join("..", "..", "..", "data", src.ID)
entries, err := os.ReadDir(dir)
if os.IsNotExist(err) {
t.Skipf("%s not present", dir)
}
if err != nil {
t.Fatal(err)
}
onDisk := map[string]bool{}
for _, e := range entries {
if !e.IsDir() {
onDisk[e.Name()] = true
}
}
files, err := src.Files()
if err != nil {
t.Fatal(err)
}
for _, f := range files {
if !onDisk[f.Dest] {
t.Errorf("crawler would write %q, which is not in %s — "+
"a renamed input can change which duplicate row survives", f.Dest, dir)
}
delete(onDisk, f.Dest)
}
for name := range onDisk {
t.Errorf("%s/%s exists but no source produces it", dir, name)
}
})
}
}
func TestLookup(t *testing.T) {
for _, s := range All() {
got, err := Lookup(s.ID)
if err != nil {
t.Errorf("Lookup(%q): %v", s.ID, err)
}
if got.ID != s.ID || got.Summary == "" {
t.Errorf("Lookup(%q) returned %+v", s.ID, got)
}
}
if _, err := Lookup("nope"); err == nil {
t.Error("Lookup of an unknown source must fail")
}
}
// TestIDsAreDatasetIDs: the ID doubles as the directory under data/, so a typo
// would send a crawl into a new directory the parser never reads.
func TestIDsAreDatasetIDs(t *testing.T) {
for _, s := range All() {
dir := filepath.Join("..", "..", "..", "data", s.ID)
if _, err := os.Stat(dir); os.IsNotExist(err) {
t.Errorf("source %q has no dataset directory at %s", s.ID, dir)
}
}
}
// TestSource2016IsComplete: 4 .xls + 115 .xlsx, matching what the source
// article published and what data/2016/ holds.
func TestSource2016IsComplete(t *testing.T) {
files, err := source2016.Files()
if err != nil {
t.Fatal(err)
}
if len(files) != 119 {
t.Errorf("got %d clusters, want 119", len(files))
}
var xls, xlsx int
for _, f := range files {
switch filepath.Ext(f.Dest) {
case ".xls":
xls++
case ".xlsx":
xlsx++
default:
t.Errorf("%s: unexpected extension in %q", f.Name, f.Dest)
}
if !strings.HasPrefix(f.URL, baseURL+uploadPath) {
t.Errorf("%s: URL does not use the configured base: %s", f.Name, f.URL)
}
}
if xls != 4 || xlsx != 115 {
t.Errorf("got %d .xls and %d .xlsx, want 4 and 115", xls, xlsx)
}
}
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Loaded 100 of 163 files, more files were not shown because too many files have changed in this diff. Show more