Files
thptqg/parser/internal/schema/schema.go
T
tiennm99 219c7c6a69 fix(parser): read the two 2016 layouts that were column-shifted
Four of the 119 files in data/2016 publish one score column per subject
instead of a DIEM_THI sentence, and none of them was being read correctly.

The ĐH Công nghiệp Thực phẩm file puts a three-row ministry title block
above its header, so no header was recognised and the positional fallback
shifted every column by one: the serial number became so_bao_danh, the
exam number became ho_ten, the name became ngay_sinh, and the national ID
became the score cell. All 7,833 rows were unusable. The three ĐH Cần Thơ
files name an SBD column but no DIEM_THI, so they fell to the same
fallback: surname into ngay_sinh, given name into ten_cum_thi, birth date
into the score cell, and 12,152 candidates with no scores at all.

Both are now read by FormatSubjectColumns, which resolves identity and one
column per subject from the header. The header is searched for in the
first five rows, so a title block no longer hides it.

The Cần Thơ score columns are numbered rather than named. They follow the
order the exam was sat — each morning an essay paper, each afternoon a
multiple-choice one — which is what identifies them: columns 1/3/5/7
quantise to 0.25 and 2/4/6/8 do not, and each column's mean lands within
0.5 of the same subject's mean across the rest of the dataset. The
foreign language is filed under the subject its N1..N6 code names.

Gender now accepts the 0/1 encoding those files use: of the rows marked
1, 53% carry "Thị" in the name against 1% of those marked 0. Birth dates
in the compact ddmmyy form are expanded so the column holds one format.

A score of 0 is stored rather than dropped, recovering 302 real scores
that a JavaScript falsy check had been turning into NULL.

Row count falls by one, to 877,460: the removed row is the title line
"ĐƠN VỊ: / TRƯỜNG ĐẠI HỌC CÔNG NGHIỆP THỰC PHẨM TP. HỒ CHÍ MINH", which
had been stored as a student. The dataset has no duplicate exam numbers;
the three rows previously described as collapsing were that same file's
title and header lines being counted and then rejected.

Also drops behaviour that existed only to match the parser this one
replaced: the inert "SINH " header token, the untrimmed diem_thi cell, an
unreachable blank-row branch, and a cross-check test against a database
that can no longer exist. None of them changes output.

Verified by rebuilding both datasets: 877,460 and 861,068 rows, both
artifacts through the assembler's row and size guards, and the reader
fidelity suite unchanged across all 182 files.
2026-08-14 10:38:59 +07:00

141 lines
4.5 KiB
Go

// Package schema is the single source of truth for the SQL shape of every
// dataset.
//
// Every dataset writes into the same 22-column student table. Columns a dataset has no data for bind NULL.
//
// Column provenance:
//
// ten_cum_thi, gioi_tinh -> 2016 only
// khtn, khxh, gdcd -> 2017 only
// everything else -> both, all six languages included
//
// The DDL, the INSERT and the subject regexes belong here and nowhere else. One
// copy per dataset is what let the 2016 and 2017 schemas drift apart; the
// per-dataset configs carry only parse rules.
package schema
import "regexp"
// DDL is executed verbatim after the output database is (re)created.
//
// idx_ten_cum_thi is partial, so it holds zero entries on the 2017 dataset —
// where the column is always NULL — while staying useful for the 2016
// cluster-grouping queries. Partial indexes are SQLite-specific.
//
// This text is frozen: it decides the shape of every database the parser
// produces. TestDDLIsFrozen holds an independent copy so any edit has to be
// deliberate.
const DDL = `
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
ngay_sinh TEXT,
ten_cum_thi TEXT,
gioi_tinh TEXT,
toan REAL,
ngu_van REAL,
vat_ly REAL,
hoa_hoc REAL,
sinh_hoc REAL,
khtn REAL,
lich_su REAL,
dia_ly REAL,
gdcd REAL,
khxh REAL,
tieng_anh REAL,
tieng_phap REAL,
tieng_nga REAL,
tieng_duc REAL,
tieng_nhat REAL,
tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
`
// IdentityFields are the identity columns, in INSERT parameter order.
var IdentityFields = []string{
"so_bao_danh",
"ho_ten",
"ho_ten_ascii",
"ngay_sinh",
"ten_cum_thi",
"gioi_tinh",
}
// ScoreFields are the subject columns, in INSERT parameter order. Bound NULL
// when a row has no score for that subject.
var ScoreFields = []string{
"toan",
"ngu_van",
"vat_ly",
"hoa_hoc",
"sinh_hoc",
"khtn",
"lich_su",
"dia_ly",
"gdcd",
"khxh",
"tieng_anh",
"tieng_phap",
"tieng_nga",
"tieng_duc",
"tieng_nhat",
"tieng_trung",
}
// ParamCount is the total bound parameters per row.
const ParamCount = 22
// InsertSQL is a positional INSERT matching IdentityFields then ScoreFields.
//
// OR REPLACE is a behavioural contract, not an optimisation: a repeated SBD
// overwrites the earlier row rather than aborting the transaction, so the last
// file to supply a duplicate wins.
const InsertSQL = `
INSERT OR REPLACE INTO student
(so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh, ten_cum_thi, gioi_tinh,
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn,
lich_su, dia_ly, gdcd, khxh,
tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung)
VALUES
(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
`
// scorePatternSources holds the regex per subject, applied to the DIEM_THI cell
// text. The literals contain Vietnamese subject names exactly as they appear in
// the source files — copy them, never retype them.
//
// Every pattern runs against every dataset. A subject absent from a given exam
// year simply never matches and stays NULL: 2016 files contain no "KHTN:",
// "KHXH:" or "GDCD:" tokens, since those combined papers did not exist yet.
var scorePatternSources = map[string]string{
"toan": `Toán:\s*(\d+(?:\.\d+)?)`,
"ngu_van": `Ngữ văn:\s*(\d+(?:\.\d+)?)`,
"vat_ly": `Vật lí:\s*(\d+(?:\.\d+)?)`,
"hoa_hoc": `Hóa học:\s*(\d+(?:\.\d+)?)`,
"sinh_hoc": `Sinh học:\s*(\d+(?:\.\d+)?)`,
"khtn": `KHTN:\s*(\d+(?:\.\d+)?)`,
"lich_su": `Lịch sử:\s*(\d+(?:\.\d+)?)`,
"dia_ly": `Địa lí:\s*(\d+(?:\.\d+)?)`,
"gdcd": `GDCD:\s*(\d+(?:\.\d+)?)`,
"khxh": `KHXH:\s*(\d+(?:\.\d+)?)`,
"tieng_anh": `Tiếng Anh:\s*(\d+(?:\.\d+)?)`,
"tieng_phap": `Tiếng Pháp:\s*(\d+(?:\.\d+)?)`,
"tieng_nga": `Tiếng Nga:\s*(\d+(?:\.\d+)?)`,
"tieng_duc": `Tiếng Đức:\s*(\d+(?:\.\d+)?)`,
"tieng_nhat": `Tiếng Nhật:\s*(\d+(?:\.\d+)?)`,
"tieng_trung": `Tiếng Trung:\s*(\d+(?:\.\d+)?)`,
}
// ScorePatterns holds the subject regexes, compiled once at package init.
var ScorePatterns = func() map[string]*regexp.Regexp {
out := make(map[string]*regexp.Regexp, len(scorePatternSources))
for field, src := range scorePatternSources {
out[field] = regexp.MustCompile(src)
}
return out
}()