mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 12:28:58 +00:00
Adds go-parser/, a Go reimplementation of the xlsxread parser, verified byte-for-byte against the Rust original before any cutover. Reader fidelity is exact across all 299 input files: the canonical cell dump of every sheet matches calamine's, locked in as a test against a committed hash oracle. Reaching that required replacing extrame/xls, which corrupted 69% of cells and dropped a further 28% on the BIFF corpus, with pbnjay/grate; correcting excelize's number-format application and trailing-cell trimming; restoring carriage returns that XML line-ending normalisation strips from 2,233 ten_cum_thi values; and gating numeric re-rendering on cell type so shared strings that merely look numeric keep their leading zeros. The differential gate compares both parsers over all four datasets: 3,265,641 rows with identical full-table SHA-256, identical per-column non-NULL counts, identical schema metadata and identical stdout. Config moves from TOML to YAML for both parsers, so they keep reading the same files and the gate stays meaningful. Verified by rebuilding 2016 and 2017-old2 with Rust under the new configs and matching the recorded counts. build-db.js now refuses to publish a database whose row count does not match the known figure, closing a path where an under-producing parser could ship a truncated public dataset with green CI. The deploy workflow gains a pull_request trigger and guards deploy to main, so branch verification can no longer publish to production.
33 lines
1.2 KiB
YAML
33 lines
1.2 KiB
YAML
# thptqg2016 data/ — 4 .xls + 115 .xlsx mixed files.
|
|
#
|
|
# Three column layouts exist across the 119 files; the binary selects the
|
|
# right one per-file at runtime via format_detection = "thptqg2016":
|
|
#
|
|
# separate-scores header SBD(0)/HOTEN(1)/TOAN(2)... — dhhanghai files
|
|
# mapped header SOBAODANH|SBD + DIEM_THI — most provinces
|
|
# default no header; positional 6-col layout — remaining files
|
|
#
|
|
# This is the only dataset that populates ten_cum_thi and gioi_tinh, and the
|
|
# only one whose files carry Tiếng Đức / Tiếng Nhật scores.
|
|
#
|
|
# sheet_mode = "all": several provinces overflow into Sheet2 (65k Excel row cap).
|
|
# strip_blank_rows = false: no blank-row anomaly observed in this dataset.
|
|
#
|
|
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
|
|
|
|
format_detection: thptqg2016
|
|
|
|
reader:
|
|
sheet_mode: all
|
|
strip_blank_rows: false
|
|
|
|
validation:
|
|
require_numeric_sbd: false
|
|
require_nonempty_name: true
|
|
require_nonempty_sbd: true
|
|
|
|
header:
|
|
# Tokens that identify a header row by first-cell content (uppercased).
|
|
# Covers both SOBAODANH-style and SBD-style headers.
|
|
tokens: ["SOBAODANH", "SBD", "HO_TEN", "HOTEN", "HỌ TÊN", "STT"]
|