fix: correct the 2016 source attribution and link full article URLs

The site credited 2016 to Bộ GD&ĐT, which those files were never fetched
from. Both datasets come from published articles: 2016 from an aggregator
on dtnt.bacninh.edu.vn listing one spreadsheet per exam cluster, 2017 from
baotintuc.vn. The README, the architecture table and the web footer all
repeated the ministry claim.

The footer now shows each dataset's full article URL as a link rather than
a bare host, so the citation can be checked. That needs overflow-wrap on
the footer: the 2016 URL is 110 characters with no break opportunity and
would otherwise scroll the page sideways on a phone.

Also records that a full 2016 crawl has been run successfully. The host was
marked unconfirmed and data/2016/ described as the only recoverable copy;
both datasets are now rebuildable from source.
This commit is contained in:
tiennm99 committed 2026-08-14 09:26:27 +07:00
1 parent 3995e9d864
commit c988bfafcf
9 files changed
+52 -31

No files matched your search

+2 -1
View File
@@ -2,7 +2,8 @@
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
school graduation exam. Client-side SQL (sql.js) over a SQLite database built
from the ministry's raw `.xls` score files by the Go `xlsxread` parser.
from the published `.xls`/`.xlsx` score files by the Go `parser` module. Where
those files come from: [data pipeline](./docs/data-pipeline.md#sources).
Live at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
+1 -2
View File
@@ -9,8 +9,7 @@ import (
// source2016 fetches the 2016 dataset: one spreadsheet per exam cluster
// (cụm thi), 4 .xls and 115 .xlsx.
//
// The article is a mirror. The site that first published this list
// (dtntbacgiang.edu.vn) no longer resolves; this copy of the same article is
// The article is a school site that aggregated every cluster's file. It is
// still online, and its links are read at crawl time like any other source.
//
// Dest keeps the server's own filename verbatim — a 32-hex content hash, the
+4 -5
View File
@@ -13,11 +13,10 @@ import (
// The fixtures are saved copies of the two source articles, kept so the link
// extraction and the naming rules can be exercised without the network.
//
// article-2017 is the live page. article-2016 came from the Internet Archive's
// copy of the site that first published that list (dtntbacgiang.edu.vn, which
// no longer resolves); the mirror the crawler actually reads carries the same
// article. Its hrefs are relative, so they resolve against whichever host the
// source names, and the filenames are identical either way.
// article-2017 is the live page. article-2016 came from an archived copy of the
// same article the crawler reads. Its hrefs are relative, so they resolve
// against whichever host the source names, and the filenames are identical
// either way.
func fixture(t *testing.T, id string) *gzip.Reader {
t.Helper()
f, err := os.Open(filepath.Join("testdata", "article-"+id+".html.gz"))
+13 -10
View File
@@ -11,8 +11,10 @@ regexes are canonical and live in `parser/internal/schema/schema.go`.
| id | Files | Origin | Host live? |
| --- | --- | --- | --- |
| `2016` | 4 `.xls` + 115 `.xlsx` | aggregator article, 119 exam clusters | unconfirmed |
| `2017` | 63 `.xls` | baotintuc.vn CDN | **yes** |
| `2016` | 4 `.xls` + 115 `.xlsx` | `dtnt.bacninh.edu.vn` aggregator article, 119 exam clusters | **yes** |
| `2017` | 63 `.xls` | `baotintuc.vn` article, files on its CDN | **yes** |
Full article URLs are below.
Crawling lives in `crawler/`, a separate Go module. It is never part of the
build — the source files are committed, so a crawl only refreshes them. Both
@@ -29,15 +31,16 @@ go -C crawler run ./cmd/crawl 2017 --list # list only, download nothing
and its CDN is still serving the files.
**2016** comes from the aggregator article
`cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html`, served from
a mirror — the site that first published it (`dtntbacgiang.edu.vn`) no longer
resolves.
`https://dtnt.bacninh.edu.vn/tin-tuc/tin-tuc-su-kien/cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html`,
which lists one spreadsheet per exam cluster. A full crawl against it has been
run successfully and reproduces `data/2016/`, so the dataset is recoverable from
source like 2017 is.
That mirror is **not reachable from every network.** It resolves to a Vietnamese
address that times out from at least some hosts abroad, in which case the crawl
stops with a connection error before downloading anything. `data/2016/` is
therefore still the only confirmed copy: do not delete it on the assumption that
a crawl can restore it.
That host is **not reachable from every network**, though: it resolves to a
Vietnamese address that times out from at least some hosts abroad, in which case
the crawl stops with a connection error before downloading anything. That is a
connectivity problem, not a missing dataset — retry from a network that can
reach the host.
## How a source is defined
+5 -6
View File
@@ -35,18 +35,17 @@ running entirely in the browser and hosted for free on GitHub Pages. Covers the
| id | Exam | Candidates | Notes |
| --- | --- | --- | --- |
| `2016` | 2016 | 877,461 | 119 files, three column layouts |
| `2017` | 2017 | 861,068 | current generation, reproducible from source |
| `2017` | 2017 | 861,068 | current generation of three publications |
Two further 2017 datasets (`2017-old`, `2017-old2`) were kept alongside these
because the three publications disagreed. They have been removed; git history
still has them.
Both datasets now have a crawler source. 2017 comes from the baotintuc.vn CDN,
which is still live. 2016 comes from an aggregator article whose original host
(`dtntbacgiang.edu.vn`) no longer resolves — its link list was recovered from
the Internet Archive and is pointed at a mirror that is still online. See
[data-pipeline](./data-pipeline.md#sources) for what that does and does not
guarantee.
which is still live. 2016 comes from an aggregator article on
`dtnt.bacninh.edu.vn`, also still online. Both datasets have been crawled
successfully, so either can be rebuilt from source — see
[data-pipeline](./data-pipeline.md#sources) for the full article URLs.
## History
+6 -3
View File
@@ -51,8 +51,11 @@ that does not exist.
| id | Exam | Rows | Source |
| --- | --- | --- | --- |
| `2016` | 2016 | 877,461 | Bộ GD&ĐT |
| `2017` | 2017 | 861,068 | baotintuc.vn |
| `2016` | 2016 | 877,461 | `dtnt.bacninh.edu.vn` |
| `2017` | 2017 | 861,068 | `baotintuc.vn` |
Full source URLs are in [data-pipeline](./data-pipeline.md#sources); the web
footer links to them per dataset.
## Canonical schema
@@ -182,6 +185,6 @@ total descending.
run out.
- **`sql.js.org` dependency.** If that CDN is unreachable, the WASM fails to
load. Self-hosting `sql-wasm.wasm` and updating `SQL_WASM_URL` in
`use-sqlite.js` is the fix.
`web/src/hooks/use-sqlite.js` is the fix.
- **Excel format drift.** A new source file with an unseen header layout needs a
new branch in `parser/internal/ingest/detect2016.go` or a new config.
+3
View File
@@ -484,6 +484,9 @@ footer {
border-top: 1px solid var(--border);
color: var(--text-subtle);
font-size: 0.85rem;
/* The source is a full article URL and some are long enough to overflow a
phone screen, which would scroll the whole page sideways. */
overflow-wrap: anywhere;
}
/* ---- Student detail card ---- */
+4 -1
View File
@@ -236,7 +236,10 @@ function DatasetApp({ dataset }) {
<footer>
<p>
Nguồn: {dataset.source}
Nguồn:{" "}
<a href={dataset.source} target="_blank" rel="noopener noreferrer">
{dataset.source}
</a>
{totalCount !== null && ` · ${totalCount.toLocaleString("vi-VN")} thí sinh`}
{" · Dữ liệu chỉ mang tính tham khảo"}
</p>
+14 -3
View File
@@ -20,13 +20,23 @@ import { PRESETS_2016, PRESETS_2017 } from "./lib/sql-presets.js";
const SUBTITLE = "Dữ liệu thí sinh toàn quốc · Hỗ trợ truy vấn SQL tùy chỉnh";
/** Presentation, keyed by the ids declared in datasets.json. */
/**
* Presentation, keyed by the ids declared in datasets.json.
*
* `source` is the full article URL the dataset's spreadsheets come from, shown
* in the footer as a link. The canonical copy of each URL is the crawler's
* `Article` field (crawler/internal/sources/source_<id>.go); it is repeated here
* because a Go module and a Vite app cannot share a constant.
*/
const CONTENT = {
2016: {
label: "Kỳ thi 2016",
title: "Tra cứu điểm thi THPT Quốc gia 2016",
subtitle: SUBTITLE,
source: "Bộ GD&ĐT",
// The article the 119 cluster spreadsheets are fetched from — a school
// site that aggregated every cluster's file, not the ministry.
source:
"https://dtnt.bacninh.edu.vn/tin-tuc/tin-tuc-su-kien/cong-bo-diem-thi-thptqg-2016-toan-bo-120-cum-thi-da-co-diem.html",
examples: ["17006021", "Nguyễn Thị Hoa"],
presets: PRESETS_2016,
},
@@ -34,7 +44,8 @@ const CONTENT = {
label: "Kỳ thi 2017",
title: "Tra cứu điểm thi THPT Quốc gia 2017",
subtitle: SUBTITLE,
source: "baotintuc.vn",
source:
"https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm",
examples: ["49008235", "Nguyễn Minh Tiến"],
presets: PRESETS_2017,
},