mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
The Go parser has matched the Rust one field-by-field across all four datasets, so the Rust crate is retired and CI builds the Go binary instead. crawl-baotintuc.js moves to go-parser/scripts/ — it is the only mechanism for refreshing data/2017 and has a documented runbook. check-duplicates.js and diff-datasets.js are dropped: both had been broken since before the repo was unified, and neither had a caller. db-stats.js and verify-parity.js are dropped as superseded by differential-parity.mjs, which compares more and cannot silently skip a dataset. The reader-fidelity oracle is kept and marked frozen. It was produced by the Rust reader so it can no longer be regenerated, but it still fails if any single cell of any of the 299 input files reads differently. Source comments cite the original Rust by file and line; those paths resolve at tag pre-go-parser-removal, recorded in go-parser/README.md.
33 lines
1.2 KiB
YAML
33 lines
1.2 KiB
YAML
# thptqg2016 data/ — 4 .xls + 115 .xlsx mixed files.
|
|
#
|
|
# Three column layouts exist across the 119 files; the binary selects the
|
|
# right one per-file at runtime via format_detection = "thptqg2016":
|
|
#
|
|
# separate-scores header SBD(0)/HOTEN(1)/TOAN(2)... — dhhanghai files
|
|
# mapped header SOBAODANH|SBD + DIEM_THI — most provinces
|
|
# default no header; positional 6-col layout — remaining files
|
|
#
|
|
# This is the only dataset that populates ten_cum_thi and gioi_tinh, and the
|
|
# only one whose files carry Tiếng Đức / Tiếng Nhật scores.
|
|
#
|
|
# sheet_mode = "all": several provinces overflow into Sheet2 (65k Excel row cap).
|
|
# strip_blank_rows = false: no blank-row anomaly observed in this dataset.
|
|
#
|
|
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
|
|
|
|
format_detection: thptqg2016
|
|
|
|
reader:
|
|
sheet_mode: all
|
|
strip_blank_rows: false
|
|
|
|
validation:
|
|
require_numeric_sbd: false
|
|
require_nonempty_name: true
|
|
require_nonempty_sbd: true
|
|
|
|
header:
|
|
# Tokens that identify a header row by first-cell content (uppercased).
|
|
# Covers both SOBAODANH-style and SBD-style headers.
|
|
tokens: ["SOBAODANH", "SBD", "HO_TEN", "HOTEN", "HỌ TÊN", "STT"]
|