Files
thptqg/go-parser/configs/2016.yml
T
tiennm99 00a08d5fab refactor(parser): remove the Rust crate now that Go is at parity
The Go parser has matched the Rust one field-by-field across all four datasets,
so the Rust crate is retired and CI builds the Go binary instead.

crawl-baotintuc.js moves to go-parser/scripts/ — it is the only mechanism for
refreshing data/2017 and has a documented runbook. check-duplicates.js and
diff-datasets.js are dropped: both had been broken since before the repo was
unified, and neither had a caller. db-stats.js and verify-parity.js are dropped
as superseded by differential-parity.mjs, which compares more and cannot
silently skip a dataset.

The reader-fidelity oracle is kept and marked frozen. It was produced by the
Rust reader so it can no longer be regenerated, but it still fails if any single
cell of any of the 299 input files reads differently.

Source comments cite the original Rust by file and line; those paths resolve at
tag pre-go-parser-removal, recorded in go-parser/README.md.
2026-08-13 20:38:25 +07:00

33 lines
1.2 KiB
YAML

# thptqg2016 data/ — 4 .xls + 115 .xlsx mixed files.
#
# Three column layouts exist across the 119 files; the binary selects the
# right one per-file at runtime via format_detection = "thptqg2016":
#
# separate-scores header SBD(0)/HOTEN(1)/TOAN(2)... — dhhanghai files
# mapped header SOBAODANH|SBD + DIEM_THI — most provinces
# default no header; positional 6-col layout — remaining files
#
# This is the only dataset that populates ten_cum_thi and gioi_tinh, and the
# only one whose files carry Tiếng Đức / Tiếng Nhật scores.
#
# sheet_mode = "all": several provinces overflow into Sheet2 (65k Excel row cap).
# strip_blank_rows = false: no blank-row anomaly observed in this dataset.
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
format_detection: thptqg2016
reader:
sheet_mode: all
strip_blank_rows: false
validation:
require_numeric_sbd: false
require_nonempty_name: true
require_nonempty_sbd: true
header:
# Tokens that identify a header row by first-cell content (uppercased).
# Covers both SOBAODANH-style and SBD-style headers.
tokens: ["SOBAODANH", "SBD", "HO_TEN", "HOTEN", "HỌ TÊN", "STT"]