fix(parser): read the two 2016 layouts that were column-shifted

Four of the 119 files in data/2016 publish one score column per subject
instead of a DIEM_THI sentence, and none of them was being read correctly.

The ĐH Công nghiệp Thực phẩm file puts a three-row ministry title block
above its header, so no header was recognised and the positional fallback
shifted every column by one: the serial number became so_bao_danh, the
exam number became ho_ten, the name became ngay_sinh, and the national ID
became the score cell. All 7,833 rows were unusable. The three ĐH Cần Thơ
files name an SBD column but no DIEM_THI, so they fell to the same
fallback: surname into ngay_sinh, given name into ten_cum_thi, birth date
into the score cell, and 12,152 candidates with no scores at all.

Both are now read by FormatSubjectColumns, which resolves identity and one
column per subject from the header. The header is searched for in the
first five rows, so a title block no longer hides it.

The Cần Thơ score columns are numbered rather than named. They follow the
order the exam was sat — each morning an essay paper, each afternoon a
multiple-choice one — which is what identifies them: columns 1/3/5/7
quantise to 0.25 and 2/4/6/8 do not, and each column's mean lands within
0.5 of the same subject's mean across the rest of the dataset. The
foreign language is filed under the subject its N1..N6 code names.

Gender now accepts the 0/1 encoding those files use: of the rows marked
1, 53% carry "Thị" in the name against 1% of those marked 0. Birth dates
in the compact ddmmyy form are expanded so the column holds one format.

A score of 0 is stored rather than dropped, recovering 302 real scores
that a JavaScript falsy check had been turning into NULL.

Row count falls by one, to 877,460: the removed row is the title line
"ĐƠN VỊ: / TRƯỜNG ĐẠI HỌC CÔNG NGHIỆP THỰC PHẨM TP. HỒ CHÍ MINH", which
had been stored as a student. The dataset has no duplicate exam numbers;
the three rows previously described as collapsing were that same file's
title and header lines being counted and then rejected.

Also drops behaviour that existed only to match the parser this one
replaced: the inert "SINH " header token, the untrimmed diem_thi cell, an
unreachable blank-row branch, and a cross-check test against a database
that can no longer exist. None of them changes output.

Verified by rebuilding both datasets: 877,460 and 861,068 rows, both
artifacts through the assembler's row and size guards, and the reader
fidelity suite unchanged across all 182 files.
This commit is contained in:
tiennm99 committed 2026-08-14 10:38:59 +07:00
1 parent 26d0ba81ce
commit 219c7c6a69
15 files changed
+535 -237

No files matched your search

+8 -4
View File
@@ -34,16 +34,20 @@ hashes every real input file. That is the point of it; do not skip it.
## Things that look wrong but are not
- **`data/2016/` filenames are content hashes and must stay verbatim.** The
parser sorts inputs bytewise and inserts last-wins, so filenames decide which
row survives a duplicate exam number — 877,464 source rows collapse to
877,461. Renaming them also breaks `parser/testdata/reader-fidelity-hashes.tsv`,
which is keyed by full path.
parser sorts inputs bytewise and inserts last-wins, so filenames would decide
which row survives a duplicate exam number. Renaming them also breaks
`parser/testdata/reader-fidelity-hashes.tsv`, which is keyed by full path.
- **`parser/testdata/reader-fidelity-hashes.tsv` is frozen and cannot be
regenerated.** A mismatch is a reader bug until proven otherwise, never a cue
to refresh the file. `parser/cmd/dumpcells` narrows a failure to the cell.
- **`ToAscii` filters the literal range U+0300..U+036F, not `unicode.Mn`.** It
must match `toAscii` in `web/src/App.jsx`, or accent-insensitive search
silently misses rows.
- **The 2016 files use four different layouts, and detection is per sheet.**
Two of them publish scores in one column per subject instead of a `DIEM_THI`
sentence, and one puts a three-row ministry title block above its header.
`parser/internal/ingest/detect2016.go` holds the header tables; they are
observations about 119 specific files, not a rule to generalise.
- **`base` in `web/vite.config.js` is absolute.** One emitted `index.html` is
copied to every dataset path, so every URL is a real static file and no SPA
404-fallback is needed — a fallback would break the `?q=` deep links.
+1 -1
View File
@@ -9,7 +9,7 @@ Live at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
| Dataset | Exam | Candidates | Site |
| --- | --- | --- | --- |
| `2016` | 2016 | 877,461 | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
| `2016` | 2016 | 877,460 | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
| `2017` | 2017 | 861,068 | [/2017/](https://tiennm99.github.io/thptqg/2017/) |
Two earlier 2017 publications (`2017-old`, `2017-old2`) were kept for a while
+2 -2
View File
@@ -18,8 +18,8 @@
"datasets": [
{
"id": "2016",
"expectedRows": 877461,
"dbSizeMb": 44
"expectedRows": 877460,
"dbSizeMb": 45
},
{
"id": "2017",
+16 -4
View File
@@ -92,17 +92,29 @@ server-assigned names verbatim, since it did not choose them.
| 2 | SOBAODANH | 8-digit string, first 2 digits = province |
| 3 | DIEM_THI | concatenated per-subject scores, e.g. `"Toán: 6.80 Ngữ văn: 5.25 …"` |
2016 has **three** layouts across its 119 files, chosen per file at runtime via
2016 has **four** layouts across its 119 files, chosen per sheet at runtime via
`format_detection: thptqg2016` in its config:
| Format | Detected by | Notes |
| --- | --- | --- |
| `separate-scores` | `row[0] == "SBD"` and `row[2] == "TOAN"` | one column per subject; scores read directly, no regex |
| `mapped` | header contains `SOBAODANH`/`SBD` plus `DIEM_THI` | column indices derived from the header |
| `subject-columns` | header names three or more subject columns | university-cluster files; identity and scores both resolved by header name |
| `default` | no recognised header | positional 6-column layout |
The `mapped` and `default` layouts also carry `TEN_CUMTHI` and `GIOI_TINH`,
which is why only 2016 populates those columns.
The header may sit below a title block, so the first five rows are searched for
it; rows above it are not data.
`subject-columns` covers two spellings. The ĐH Công nghiệp Thực phẩm file names
its columns `TO VA LI HO SI SU DI NN` with the language code in `Môn NN`; the
three ĐH Cần Thơ files publish `cdiem1..cdiem8`, numbered in the order the 2016
exam was sat — Toán, Ngoại ngữ, Ngữ văn, Vật lí, Địa lí, Hóa học, Lịch sử, Sinh
học — with the language code in `ngoaingu`. The language score is filed under
the subject its `N1`..`N6` code names.
The `mapped`, `default` and `subject-columns` layouts also carry `GIOI_TINH`
(and `mapped`/`default` `TEN_CUMTHI`), which is why only 2016 populates those
columns.
## Score text parsing
@@ -171,7 +183,7 @@ silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what
| id | Source rows | Skipped | DB rows |
| --- | --- | --- | --- |
| `2016` | 877,464 | 3 duplicate SBDs collapsed | **877,461** |
| `2016` | 877,460 | 0 | **877,460** |
| `2017` | 861,068 | 0 | **861,068** |
## Verifying a rebuild
+1 -1
View File
@@ -34,7 +34,7 @@ running entirely in the browser and hosted for free on GitHub Pages. Covers the
| id | Exam | Candidates | Notes |
| --- | --- | --- | --- |
| `2016` | 2016 | 877,461 | 119 files, three column layouts |
| `2016` | 2016 | 877,460 | 119 files, four column layouts |
| `2017` | 2017 | 861,068 | current generation of three publications |
Two further 2017 datasets (`2017-old`, `2017-old2`) were kept alongside these
+4 -4
View File
@@ -51,7 +51,7 @@ that does not exist.
| id | Exam | Rows | Source |
| --- | --- | --- | --- |
| `2016` | 2016 | 877,461 | `dtnt.bacninh.edu.vn` |
| `2016` | 2016 | 877,460 | `dtnt.bacninh.edu.vn` |
| `2017` | 2017 | 861,068 | `baotintuc.vn` |
Full source URLs are in [data-pipeline](./data-pipeline.md#sources); the web
@@ -139,9 +139,9 @@ and had drifted apart on exactly this rule).
| 9 digits with leading zero | `017006021` | 2016 |
| 2-4 letters then digits | `BAL000001` | 2016 — exam cluster code |
The letter-prefixed form is the majority case for 2016: 616,593 of 877,461
candidates (70.3%). Letter prefixes are upper-cased before lookup, so
`bal000001` resolves.
The letter-prefixed form is the majority case for 2016: 624,424 of 877,460
candidates (71.2%); the remaining 253,036 are all 9-digit. Letter prefixes are
upper-cased before lookup, so `bal000001` resolves.
## Score tiers
+23 -28
View File
@@ -26,7 +26,7 @@ go -C assembler run ./cmd/assemble db
| path | role |
|---|---|
| `internal/reader` | spreadsheet reading; the only place that knows about file formats |
| `internal/ingest` | dataset policy — sheet selection, header skipping, blank rows, the build loop, and the 2016 per-sheet format detection |
| `internal/ingest` | dataset policy — sheet selection, header skipping, blank rows, the build loop, and the 2016 per-sheet layout detection |
| `internal/transform` | `ToAscii`, score-regex parsing, row validation |
| `internal/schema` | the canonical 22-column table: DDL, INSERT, subject regexes |
| `internal/config` | per-dataset YAML parse rules |
@@ -37,40 +37,35 @@ The reader deliberately knows nothing about datasets: it reports every sheet and
every row verbatim. All policy lives in `ingest`. That split is what made the
reader independently verifiable against a hash oracle.
## Provenance
## Behaviour worth knowing
This is a port of a Rust crate that occupied this same path until the Go
implementation reached full parity, when it was built alongside as `go-parser/`
and moved back here once the Rust was removed. Source comments cite the original
by file and line (`parser/src/transform.rs:56` and similar) — those refer to the
Rust tree and resolve at the tag **`pre-go-parser-removal`**, the last commit
containing it.
The port was gated on a field-by-field comparison of both implementations across
the four datasets that existed then — 3,265,641 rows with identical full-table
SHA-256, identical per-column non-NULL counts, identical schema metadata and
identical build stdout. `assemble verify` is that comparator today and
still runs against any two sets of databases.
Behaviour was matched bug-for-bug, deliberately. Several quirks look like
defects and are load-bearing for the published data:
- a parsed score of `0` becomes NULL in the 2016 separate-scores layout,
replicating a JavaScript falsy check;
- `ToAscii` strips combining marks in the literal range U+0300–U+036F rather
than by Unicode category, which is narrower;
- gender is a two-value allowlist, and anything else becomes NULL;
- `diem_thi` is read untrimmed while the other three fields are trimmed;
- the `"SINH "` header token carries a trailing space.
than by Unicode category. That covers every Vietnamese diacritic and must
stay identical to `toAscii` in `web/src/App.jsx`, or accent-insensitive
search misses rows.
- Gender is normalised to `Nam`/`Nữ`; the Cần Thơ files write `0`/`1` instead
and are translated. Anything else becomes NULL.
- A score of `0` is a real score — the candidate sat the paper and scored
nothing — and is stored, not dropped.
- Birth dates are stored as `dd/mm/yyyy`. The Cần Thơ files' compact `ddmmyy`
is expanded; the century is always 19xx, since a 2016 candidate born later
would have sat the exam under age.
Each has a test naming it, so none can be tidied away by accident.
This code began as a port of a Rust crate that occupied the same path, and was
gated on a field-by-field comparison against it. That comparison is over:
correctness against the source spreadsheets decides behaviour now, not
agreement with the old implementation. `assemble verify` is the comparator
that gated it, and still compares any two sets of built databases.
## Verification
`testdata/reader-fidelity-hashes.tsv` holds a SHA-256 per input file over a
canonical dump of every cell of every sheet. It is **frozen**: it was produced
by the Rust reader, which no longer exists, so it cannot be regenerated. It
still fails if any single cell of any input file reads differently.
canonical dump of every cell of every sheet. It is **frozen**: the tool that
produced it no longer exists, so it cannot be regenerated. It still fails if
any single cell of any input file reads differently, which is what makes it
useful — cell rendering and sheet geometry are settled, and a change there is
a regression until proven otherwise. Dataset policy lives in `ingest`, so
fixing a layout never touches it.
A mismatch names the file but not the cell. `cmd/dumpcells` prints the stream
the hash is taken over, so two runs can be diffed:
+3 -2
View File
@@ -1,10 +1,11 @@
# thptqg2016 data/ — 4 .xls + 115 .xlsx mixed files.
#
# Three column layouts exist across the 119 files; the binary selects the
# right one per-file at runtime via format_detection: thptqg2016:
# Four column layouts exist across the 119 files; the binary selects the
# right one per-sheet at runtime via format_detection: thptqg2016:
#
# separate-scores header SBD(0)/HOTEN(1)/TOAN(2)... — dhhanghai files
# mapped header SOBAODANH|SBD + DIEM_THI — most provinces
# subject-columns header names >=3 subject columns — dhcantho, dhcnthucpham
# default no header; positional 6-col layout — remaining files
#
# This is the only dataset that populates ten_cum_thi and gioi_tinh, and the
+288 -70
View File
@@ -14,14 +14,11 @@ import (
)
// The 2016 dataset's 119 files were produced by inconsistent tooling and use
// three different column layouts, chosen per sheet at runtime. The literals
// four different column layouts, chosen per sheet at runtime. The literals
// below are institutional knowledge about those files, with no rule to derive
// them from — extend them, but do not tidy them into something more regular.
// KnownHeaders are the upper-cased first-cell values that identify a header row.
//
// NOTE "SINH " carries a TRAILING SPACE. Trimming it would change which rows
// are recognised as headers.
var KnownHeaders = []string{
"SOBAODANH",
"SBD",
@@ -37,7 +34,7 @@ var KnownHeaders = []string{
"VAN",
"LY",
"HOA",
"SINH ",
"SINH",
"SU",
"DIA",
}
@@ -61,7 +58,7 @@ func IsHeaderRow2016(row []reader.Cell) bool {
return isKnownHeader(strings.ToUpper(strings.TrimSpace(row[0].Str)))
}
// FormatKind is one of the three 2016 layouts.
// FormatKind is one of the 2016 layouts.
type FormatKind int
const (
@@ -73,9 +70,13 @@ const (
// FormatDefault is the positional 6-column layout used when no header is
// recognised. It is FormatMapped with fixed indices, not a separate path.
FormatDefault
// FormatSubjectColumns resolves BOTH identity and one score column per
// subject from the header by name — the university-cluster files, which
// publish scores in columns instead of a DIEM_THI sentence.
FormatSubjectColumns
)
// Format is a detected layout plus, for the mapped case, its column indices.
// Format is a detected layout plus its resolved column indices.
type Format struct {
Kind FormatKind
Sbd int
@@ -84,6 +85,14 @@ type Format struct {
TenCumThi *int
GioiTinh *int
DiemThi int
// Subjects maps a schema score field to its column, for
// FormatSubjectColumns only.
Subjects map[string]int
// LangScore is the foreign-language score column; the subject it counts
// towards is named by the code in LangCode.
LangScore *int
LangCode *int
}
// defaultFormat is FormatMapped with the fixed positional indices.
@@ -95,6 +104,90 @@ func defaultFormat() Format {
}
}
// identity holds the identity columns resolved from a header row. A nil field
// means the header does not name that column.
type identity struct {
Sbd, HoTen, NgaySinh, TenCumThi, GioiTinh, DiemThi *int
}
// resolveIdentity maps header names to identity columns, order-independent.
// Each column has several spellings across the corpus, including the accented
// and unseparated ones the university files use.
func resolveIdentity(cols []string) identity {
var id identity
for i := range cols {
idx := i
switch cols[i] {
case "SOBAODANH", "SBD":
id.Sbd = &idx
case "HO_TEN", "HOTEN", "HỌ TÊN":
id.HoTen = &idx
case "NGAY_SINH", "NGAYSINH", "NGÀY SINH":
id.NgaySinh = &idx
case "TEN_CUMTHI":
id.TenCumThi = &idx
case "GIOI_TINH", "GIỚI TÍNH", "PHAI":
id.GioiTinh = &idx
case "DIEM_THI":
id.DiemThi = &idx
}
}
return id
}
// subjectColumnFamily is one university-cluster spelling of the per-subject
// score columns. The two families are told apart by which names resolve.
type subjectColumnFamily struct {
scores map[string]string // header name -> schema score field
langScore string // header of the foreign-language score column
langCode string // header of the N1..N6 language code column
}
// subjectColumnFamilies are tried in order. The Cần Thơ family goes first
// because its CDIEM<n> names are unambiguous, while the two-letter family
// shares "HO" with that layout's surname column.
var subjectColumnFamilies = []subjectColumnFamily{
{scores: canThoScoreHeaders, langScore: "CDIEM2", langCode: "NGOAINGU"},
{scores: subjectCodeHeaders, langScore: "NN", langCode: "MÔN NN"},
}
// canThoScoreHeaders maps the Cần Thơ cluster's published-score columns to
// schema fields. The columns are numbered in the order the 2016 exam was sat —
// each morning an essay paper (quarter-point scores), each afternoon a
// multiple-choice one — which is what identifies them: CDIEM1/3/5/7 quantise to
// 0.25 and CDIEM2/4/6/8 do not, and the per-column means match the published
// national averages for those subjects.
var canThoScoreHeaders = map[string]string{
"CDIEM1": "toan",
"CDIEM3": "ngu_van",
"CDIEM4": "vat_ly",
"CDIEM5": "dia_ly",
"CDIEM6": "hoa_hoc",
"CDIEM7": "lich_su",
"CDIEM8": "sinh_hoc",
}
// subjectCodeHeaders maps the two-letter subject columns to schema fields.
var subjectCodeHeaders = map[string]string{
"TO": "toan",
"VA": "ngu_van",
"LI": "vat_ly",
"HO": "hoa_hoc",
"SI": "sinh_hoc",
"SU": "lich_su",
"DI": "dia_ly",
}
// languageFields maps the ministry's foreign-language codes to schema fields.
var languageFields = map[string]string{
"N1": "tieng_anh",
"N2": "tieng_nga",
"N3": "tieng_phap",
"N4": "tieng_trung",
"N5": "tieng_duc",
"N6": "tieng_nhat",
}
// DetectFormat inspects a header row and decides which layout applies.
func DetectFormat(headerRow []reader.Cell) Format {
cols := make([]string, len(headerRow))
@@ -107,54 +200,85 @@ func DetectFormat(headerRow []reader.Cell) Format {
return Format{Kind: FormatSeparateScores}
}
// Format 2: resolve indices by header name, order-independent.
var sbdIdx, hoTenIdx, ngaySinhIdx, tenCumThiIdx, gioiTinhIdx, diemThiIdx *int
for i := range cols {
idx := i
switch cols[i] {
case "SOBAODANH", "SBD":
sbdIdx = &idx
case "HO_TEN", "HOTEN", "HỌ TÊN":
hoTenIdx = &idx
case "NGAY_SINH":
ngaySinhIdx = &idx
case "TEN_CUMTHI":
tenCumThiIdx = &idx
case "GIOI_TINH":
gioiTinhIdx = &idx
case "DIEM_THI":
diemThiIdx = &idx
id := resolveIdentity(cols)
// Format 2: a free-text DIEM_THI cell alongside an SBD.
if id.Sbd != nil && id.DiemThi != nil {
hoTen := 1 // fallback: col 1, present in all known files
if id.HoTen != nil {
hoTen = *id.HoTen
}
return Format{
Kind: FormatMapped, Sbd: *id.Sbd, HoTen: hoTen,
NgaySinh: id.NgaySinh, TenCumThi: id.TenCumThi, GioiTinh: id.GioiTinh,
DiemThi: *id.DiemThi,
}
}
if sbdIdx != nil && diemThiIdx != nil {
hoTen := 1 // fallback: col 1, present in all known files
if hoTenIdx != nil {
hoTen = *hoTenIdx
}
return Format{
Kind: FormatMapped, Sbd: *sbdIdx, HoTen: hoTen,
NgaySinh: ngaySinhIdx, TenCumThi: tenCumThiIdx, GioiTinh: gioiTinhIdx,
DiemThi: *diemThiIdx,
}
// Format 3: one column per subject.
if f, ok := subjectColumnsFormat(cols, id); ok {
return f
}
// Unrecognised header, or none at all.
return defaultFormat()
}
// parseFloatCell parses a per-subject score cell.
// subjectColumnsFormat resolves a per-subject-column layout, if the header
// names one.
//
// A parsed 0.0 becomes "no score", so a genuine zero is indistinguishable from
// a blank. Not obviously correct, but it is the shipped behaviour and the
// published data depends on it.
func parseFloatCell(row []reader.Cell, idx int) (float64, bool) {
// Three resolved subjects are required: a stray one- or two-letter header in an
// unrelated file must not turn that file into this layout.
func subjectColumnsFormat(cols []string, id identity) (Format, bool) {
if id.Sbd == nil {
return Format{}, false
}
for _, fam := range subjectColumnFamilies {
subjects := make(map[string]int)
for i, name := range cols {
if field, ok := fam.scores[name]; ok {
subjects[field] = i
}
}
if len(subjects) < 3 {
continue
}
hoTen := 1
if id.HoTen != nil {
hoTen = *id.HoTen
}
return Format{
Kind: FormatSubjectColumns, Sbd: *id.Sbd, HoTen: hoTen,
NgaySinh: id.NgaySinh, TenCumThi: id.TenCumThi, GioiTinh: id.GioiTinh,
DiemThi: -1,
Subjects: subjects,
LangScore: indexOf(cols, fam.langScore),
LangCode: indexOf(cols, fam.langCode),
}, true
}
return Format{}, false
}
// indexOf returns the column holding name, or nil when the header lacks it.
func indexOf(cols []string, name string) *int {
for i := range cols {
if cols[i] == name {
idx := i
return &idx
}
}
return nil
}
// parseScoreCell parses a per-subject score cell. A parsed 0 is a real score:
// the candidate sat the paper and scored nothing.
func parseScoreCell(row []reader.Cell, idx int) (float64, bool) {
s := cellAt(row, idx)
if s == "" {
return 0, false
}
v, err := strconv.ParseFloat(s, 64)
if err != nil || v == 0.0 {
if err != nil {
return 0, false
}
return v, true
@@ -182,7 +306,7 @@ func processSeparateScoresRow(row []reader.Cell) *transform.ParsedRow {
"sinh_hoc": 6, "lich_su": 7, "dia_ly": 8,
"tieng_anh": 11, // NGOAINGU total
} {
if v, ok := parseFloatCell(row, idx); ok {
if v, ok := parseScoreCell(row, idx); ok {
scores[field] = v
}
}
@@ -209,47 +333,146 @@ func processMappedRow(row []reader.Cell, f Format) *transform.ParsedRow {
return nil
}
optional := func(idx *int) *string {
if idx == nil {
return nil
}
if s := cellAt(row, *idx); s != "" {
return &s
}
return &transform.ParsedRow{
SoBaoDanh: sbd,
HoTen: hoTen,
HoTenAscii: transform.ToAscii(hoTen),
NgaySinh: birthDate(row, f.NgaySinh),
TenCumThi: optionalCell(row, f.TenCumThi),
GioiTinh: gender(row, f.GioiTinh),
Scores: transform.ParseScores(cellAt(row, f.DiemThi)),
}
}
// processSubjectColumnsRow reads a row whose scores sit in one column per
// subject, with the foreign language filed under the subject its code names.
func processSubjectColumnsRow(row []reader.Cell, f Format) *transform.ParsedRow {
sbd := cellAt(row, f.Sbd)
hoTen := cellAt(row, f.HoTen)
if sbd == "" || hoTen == "" {
return nil
}
if isKnownHeader(strings.ToUpper(sbd)) || isKnownHeader(strings.ToUpper(hoTen)) {
return nil
}
// Gender is a two-value allowlist, not a general enum: anything other than
// exactly "Nam" or "Nữ" becomes nil.
var gioiTinh *string
if f.GioiTinh != nil {
if s := cellAt(row, *f.GioiTinh); s == "Nam" || s == "Nữ" {
gioiTinh = &s
scores := make(map[string]float64)
for field, idx := range f.Subjects {
if v, ok := parseScoreCell(row, idx); ok {
scores[field] = v
}
}
// The free-text score cell is read untrimmed, unlike every other cell here.
diemThi := ""
if f.DiemThi >= 0 && f.DiemThi < len(row) {
diemThi = row[f.DiemThi].Str
if f.LangScore != nil {
if v, ok := parseScoreCell(row, *f.LangScore); ok {
code := "N1" // an unlabelled language paper is English in this corpus
if f.LangCode != nil {
if c := strings.ToUpper(cellAt(row, *f.LangCode)); c != "" {
code = c
}
}
// An unknown code is left out rather than guessed at.
if field, ok := languageFields[code]; ok {
scores[field] = v
}
}
}
return &transform.ParsedRow{
SoBaoDanh: sbd,
HoTen: hoTen,
HoTenAscii: transform.ToAscii(hoTen),
NgaySinh: optional(f.NgaySinh),
TenCumThi: optional(f.TenCumThi),
GioiTinh: gioiTinh,
Scores: transform.ParseScores(diemThi),
NgaySinh: birthDate(row, f.NgaySinh),
TenCumThi: optionalCell(row, f.TenCumThi),
GioiTinh: gender(row, f.GioiTinh),
Scores: scores,
}
}
// optionalCell returns the trimmed cell at idx, or nil when the column is
// absent or its cell empty.
func optionalCell(row []reader.Cell, idx *int) *string {
if idx == nil {
return nil
}
if s := cellAt(row, *idx); s != "" {
return &s
}
return nil
}
// gender normalises the gender cell to "Nam"/"Nữ", or nil when it says neither.
//
// The Cần Thơ files encode it numerically instead: of the rows marked 1, 53%
// carry the female marker "Thị" in the name, against 1% of those marked 0.
func gender(row []reader.Cell, idx *int) *string {
s := optionalCell(row, idx)
if s == nil {
return nil
}
switch *s {
case "Nam", "Nữ":
return s
case "0":
nam := "Nam"
return &nam
case "1":
nu := "Nữ"
return &nu
}
return nil
}
// birthDate returns the birth-date cell as dd/mm/yyyy.
//
// Most files already write it that way; the Cần Thơ ones use a compact ddmmyy,
// expanded here so the column holds one format. The century is always 19xx: a
// 2016 candidate born after 1999 would have sat the exam under age.
func birthDate(row []reader.Cell, idx *int) *string {
s := optionalCell(row, idx)
if s == nil || len(*s) != 6 || !allDigits(*s) {
return s
}
out := (*s)[0:2] + "/" + (*s)[2:4] + "/19" + (*s)[4:6]
return &out
}
func allDigits(s string) bool {
for i := 0; i < len(s); i++ {
if s[i] < '0' || s[i] > '9' {
return false
}
}
return true
}
// headerScanRows is how far into a sheet the header may sit. One file opens
// with a three-row ministry title block above its header; a sheet with no
// header at all has data from row 0 and simply finds nothing in that window.
const headerScanRows = 5
// sheetFormat picks the layout for one sheet and returns the row its data
// starts on. Rows above the header are a title block and are not data.
func sheetFormat(rows [][]reader.Cell) (Format, int) {
limit := headerScanRows
if len(rows) < limit {
limit = len(rows)
}
for i := 0; i < limit; i++ {
if IsHeaderRow2016(rows[i]) {
return DetectFormat(rows[i]), i + 1
}
}
return defaultFormat(), 0
}
// ProcessRow2016 dispatches a data row through the detected layout. A nil return
// means the row is empty or invalid and should be skipped.
func ProcessRow2016(row []reader.Cell, f Format) *transform.ParsedRow {
if f.Kind == FormatSeparateScores {
switch f.Kind {
case FormatSeparateScores:
return processSeparateScoresRow(row)
case FormatSubjectColumns:
return processSubjectColumnsRow(row, f)
}
// FormatMapped and FormatDefault share one implementation; Default is just a
// fixed index tuple.
@@ -336,12 +559,7 @@ func processFile2016(path string, cfg *config.DatasetConfig, ins *writer.Inserte
continue
}
f := defaultFormat()
startIdx := 0
if IsHeaderRow2016(rows[0]) {
f = DetectFormat(rows[0])
startIdx = 1
}
f, startIdx := sheetFormat(rows)
for _, row := range rows[startIdx:] {
// Rows shorter than 2 cells are dropped BEFORE the counter, so they
+177 -25
View File
@@ -1,6 +1,7 @@
package ingest
import (
"strings"
"testing"
"github.com/tiennm99/thptqg/parser/internal/reader"
@@ -31,20 +32,45 @@ func TestIsHeaderRow2016(t *testing.T) {
}
}
// TestKnownHeadersHasTrailingSpaceToken pins the "SINH " literal. Trimming it
// would silently change which rows count as headers.
func TestKnownHeadersHasTrailingSpaceToken(t *testing.T) {
var found bool
// TestKnownHeadersAreTrimmed: the cell is trimmed before comparison, so a token
// carrying surrounding space could never match anything.
func TestKnownHeadersAreTrimmed(t *testing.T) {
for _, h := range KnownHeaders {
if h == "SINH " {
found = true
if h != strings.TrimSpace(h) {
t.Errorf("KnownHeaders token %q has surrounding space and can never match", h)
}
}
if !found {
t.Error(`KnownHeaders must contain "SINH " WITH its trailing space`)
}
// TestSheetFormatSkipsTitleBlock: one file opens with a ministry title block
// above its header row, and the rows above the header are not data.
func TestSheetFormatSkipsTitleBlock(t *testing.T) {
rows := [][]reader.Cell{
cells("BỘ GIÁO DỤC VÀ ĐÀO TẠO", "", ""),
cells("ĐƠN VỊ:", "TRƯỜNG ĐẠI HỌC X", ""),
cells("DANH SÁCH ĐIỂM THÍ SINH", "", ""),
cells("STT", "SBD", "Họ tên", "Ngày sinh", "Giới tính", "CMND", "Tỉnh", "TO", "VA", "LI"),
cells("1", "DCT000001", "Nguyễn Văn A", "01/01/1998", "Nam", "123", "46", "8", "7", "6"),
}
if len(KnownHeaders) != 17 {
t.Errorf("KnownHeaders has %d tokens, want 17", len(KnownHeaders))
f, start := sheetFormat(rows)
if start != 4 {
t.Errorf("data starts at row %d, want 4", start)
}
if f.Kind != FormatSubjectColumns {
t.Errorf("Kind = %v, want FormatSubjectColumns", f.Kind)
}
}
// TestSheetFormatHeaderlessSheet: a sheet whose first row is already data keeps
// the positional layout and starts at row 0.
func TestSheetFormatHeaderlessSheet(t *testing.T) {
rows := [][]reader.Cell{
cells("YTB000001", "PHẠM THỊ HỒNG ÁI", "21/08/1998", "Y Dược Thái Bình", "Nữ", "Toán: 7.75"),
cells("YTB000002", "TRẦN VĂN B", "22/08/1998", "Y Dược Thái Bình", "Nam", "Toán: 5"),
}
f, start := sheetFormat(rows)
if start != 0 || f.Kind != FormatDefault {
t.Errorf("got kind=%v start=%d, want FormatDefault at 0", f.Kind, start)
}
}
@@ -133,36 +159,37 @@ func TestProcessSeparateScoresRow(t *testing.T) {
}
}
// TestZeroScoreBecomesNull pins the quirk that a literal 0 is indistinguishable
// from "no score".
func TestZeroScoreBecomesNull(t *testing.T) {
// TestZeroScoreIsKept: a candidate who sat a paper and scored nothing has a
// score of 0, not a missing one.
func TestZeroScoreIsKept(t *testing.T) {
row := cells("1000", "A", "0", "0.0", "1", "", "", "", "", "", "", "")
got := ProcessRow2016(row, Format{Kind: FormatSeparateScores})
if got == nil {
t.Fatal("row rejected")
}
if _, ok := got.Scores["toan"]; ok {
t.Error(`a "0" score must become NULL, not 0`)
for field, want := range map[string]float64{"toan": 0, "ngu_van": 0, "vat_ly": 1} {
v, ok := got.Scores[field]
if !ok || v != want {
t.Errorf("%s = %v (present=%v), want %v", field, v, ok, want)
}
}
if _, ok := got.Scores["ngu_van"]; ok {
t.Error(`a "0.0" score must become NULL, not 0`)
}
if got.Scores["vat_ly"] != 1 {
t.Error("a non-zero score must survive")
if _, ok := got.Scores["hoa_hoc"]; ok {
t.Error("an empty cell must stay absent")
}
}
// TestGenderAllowlist: exactly two accepted values, matched case-sensitively.
func TestGenderAllowlist(t *testing.T) {
// TestGenderNormalisation: "Nam"/"Nữ" verbatim, the Cần Thơ files' 1/0 encoding
// translated, everything else NULL.
func TestGenderNormalisation(t *testing.T) {
f := defaultFormat()
for in, want := range map[string]string{"Nam": "Nam", "Nữ": "Nữ"} {
for in, want := range map[string]string{"Nam": "Nam", "Nữ": "Nữ", "1": "Nữ", "0": "Nam"} {
row := cells("1", "A", "", "", in, "")
got := ProcessRow2016(row, f)
if got == nil || got.GioiTinh == nil || *got.GioiTinh != want {
t.Errorf("gioi_tinh for %q not preserved", in)
t.Errorf("gioi_tinh for %q = %v, want %q", in, got.GioiTinh, want)
}
}
for _, in := range []string{"nam", "NAM", "Unknown", "M", "F", "nữ", ""} {
for _, in := range []string{"nam", "NAM", "Unknown", "M", "F", "nữ", "", "2"} {
row := cells("1", "A", "", "", in, "")
got := ProcessRow2016(row, f)
if got == nil {
@@ -174,6 +201,131 @@ func TestGenderAllowlist(t *testing.T) {
}
}
// TestBirthDateExpandsCompactForm: ddmmyy becomes dd/mm/19yy so the column
// holds one format; anything else is passed through.
func TestBirthDateExpandsCompactForm(t *testing.T) {
f := defaultFormat()
for in, want := range map[string]string{
"140798": "14/07/1998",
"01/01/1998": "01/01/1998",
"1998": "1998",
} {
got := ProcessRow2016(cells("1", "A", in, "", "", ""), f)
if got == nil || got.NgaySinh == nil || *got.NgaySinh != want {
t.Errorf("ngay_sinh for %q = %v, want %q", in, got.NgaySinh, want)
}
}
}
// --- per-subject score columns ---
// TestDetectFormatSubjectColumnsTwoLetter covers the two-letter subject columns,
// whose foreign-language score is filed under the subject its code names.
func TestDetectFormatSubjectColumnsTwoLetter(t *testing.T) {
header := cells("STT", "SBD", "Họ tên", "Ngày sinh", "Giới tính", "CMND", "Tỉnh",
"TO", "VA", "LI", "HO", "SI", "SU", "DI", "NN", "Môn NN")
f := DetectFormat(header)
if f.Kind != FormatSubjectColumns {
t.Fatalf("Kind = %v, want FormatSubjectColumns", f.Kind)
}
if f.Sbd != 1 || f.HoTen != 2 {
t.Errorf("identity: sbd=%d ho_ten=%d", f.Sbd, f.HoTen)
}
if f.NgaySinh == nil || *f.NgaySinh != 3 || f.GioiTinh == nil || *f.GioiTinh != 4 {
t.Error("accented identity headers not resolved")
}
if f.Subjects["toan"] != 7 || f.Subjects["dia_ly"] != 13 {
t.Errorf("subject columns = %v", f.Subjects)
}
row := cells("1", "DCT000073", "Trần Thị Phước An", "23/11/1998", "Nữ", "291183999", "46",
"6.25", "4.5", "5.8", "", "", "", "", "5.38", "N1")
got := ProcessRow2016(row, f)
if got == nil {
t.Fatal("row rejected")
}
if got.SoBaoDanh != "DCT000073" || got.HoTen != "Trần Thị Phước An" {
t.Errorf("identity = %q / %q", got.SoBaoDanh, got.HoTen)
}
want := map[string]float64{"toan": 6.25, "ngu_van": 4.5, "vat_ly": 5.8, "tieng_anh": 5.38}
if len(got.Scores) != len(want) {
t.Errorf("scores = %v, want %v", got.Scores, want)
}
for k, v := range want {
if got.Scores[k] != v {
t.Errorf("%s = %v, want %v", k, got.Scores[k], v)
}
}
}
// TestDetectFormatSubjectColumnsCanTho covers the CDIEM<n> columns, which are
// numbered in exam-timetable order rather than named.
func TestDetectFormatSubjectColumnsCanTho(t *testing.T) {
header := cells("sbd", "hoten", "ho", "ten", "phai", "ngaysinh", "socmnd",
"cdiem1", "cdiem2", "cdiem3", "cdiem4", "cdiem5", "cdiem6", "cdiem7", "cdiem8", "ngoaingu")
f := DetectFormat(header)
if f.Kind != FormatSubjectColumns {
t.Fatalf("Kind = %v, want FormatSubjectColumns", f.Kind)
}
// "ho" is the surname column here, and must not be read as hoa_hoc.
if got, ok := f.Subjects["hoa_hoc"]; ok && got == 2 {
t.Error("surname column resolved as a score column")
}
if f.Subjects["toan"] != 7 || f.Subjects["ngu_van"] != 9 || f.Subjects["sinh_hoc"] != 14 {
t.Errorf("subject columns = %v", f.Subjects)
}
if f.LangScore == nil || *f.LangScore != 8 || f.LangCode == nil || *f.LangCode != 15 {
t.Error("language score/code columns not resolved")
}
row := cells("TCT000001", "Dương Diễm ái", "Dương Diễm", "ái", "1", "140798", "362539228",
"07.50", "05.60", "06.50", "", "", "08.20", "", "07.60", "N3")
got := ProcessRow2016(row, f)
if got == nil {
t.Fatal("row rejected")
}
if got.GioiTinh == nil || *got.GioiTinh != "Nữ" {
t.Errorf("gioi_tinh = %v, want Nữ", got.GioiTinh)
}
if got.NgaySinh == nil || *got.NgaySinh != "14/07/1998" {
t.Errorf("ngay_sinh = %v", got.NgaySinh)
}
want := map[string]float64{"toan": 7.5, "tieng_phap": 5.6, "ngu_van": 6.5, "hoa_hoc": 8.2, "sinh_hoc": 7.6}
if len(got.Scores) != len(want) {
t.Errorf("scores = %v, want %v", got.Scores, want)
}
for k, v := range want {
if got.Scores[k] != v {
t.Errorf("%s = %v, want %v", k, got.Scores[k], v)
}
}
}
// TestSubjectColumnsUnknownLanguageCode: an unrecognised code drops the score
// rather than filing it under a guess.
func TestSubjectColumnsUnknownLanguageCode(t *testing.T) {
header := cells("sbd", "hoten", "ho", "ten", "phai", "ngaysinh", "socmnd",
"cdiem1", "cdiem2", "cdiem3", "cdiem4", "cdiem5", "cdiem6", "cdiem7", "cdiem8", "ngoaingu")
f := DetectFormat(header)
row := cells("TCT000002", "B", "", "", "0", "", "", "5", "6", "", "", "", "", "", "", "N9")
got := ProcessRow2016(row, f)
if got == nil {
t.Fatal("row rejected")
}
if len(got.Scores) != 1 || got.Scores["toan"] != 5 {
t.Errorf("scores = %v, want only toan", got.Scores)
}
}
// TestSubjectColumnsNeedsThreeSubjects: a stray two-letter header elsewhere must
// not turn an unrelated file into this layout.
func TestSubjectColumnsNeedsThreeSubjects(t *testing.T) {
f := DetectFormat(cells("SBD", "HOTEN", "DI", "SU"))
if f.Kind == FormatSubjectColumns {
t.Error("two subject columns must not be enough to select the layout")
}
}
// TestLeakedHeaderRowSkipped: a header repeated inside the data must not become
// a student.
func TestLeakedHeaderRowSkipped(t *testing.T) {
+1 -9
View File
@@ -174,15 +174,7 @@ func Standard(cfg *config.DatasetConfig, inputDir, outputPath string) error {
soBaoDanh = cellAt(row, cols.SoBaoDanh)
}
switch transform.ValidateRow(hoTen, soBaoDanh, &cfg.Validation, stripBlank, allBlank) {
case transform.SkipBlankRow:
// Falls through to transform and insert. Unreachable here:
// BlankRow requires stripBlank && allBlank, which returned
// above. Kept so the two call sites with opposite outcomes stay
// visibly distinct.
case transform.SkipNone:
// proceed
default:
if transform.ValidateRow(hoTen, soBaoDanh, &cfg.Validation, stripBlank, allBlank) != transform.SkipNone {
fileSkipped++
return
}
+5 -5
View File
@@ -5,9 +5,9 @@
//
// Column provenance:
//
// ten_cum_thi, gioi_tinh, tieng_duc, tieng_nhat -> 2016 only
// khtn, khxh, gdcd, tieng_nga -> 2017 datasets only
// everything else -> both
// ten_cum_thi, gioi_tinh -> 2016 only
// khtn, khxh, gdcd -> 2017 only
// everything else -> both, all six languages included
//
// The DDL, the INSERT and the subject regexes belong here and nowhere else. One
// copy per dataset is what let the 2016 and 2017 schemas drift apart; the
@@ -109,8 +109,8 @@ VALUES
// the source files — copy them, never retype them.
//
// Every pattern runs against every dataset. A subject absent from a given exam
// year simply never matches and stays NULL: 2016 files contain no "KHTN:" or
// "Tiếng Nga:" tokens, and 2017 files contain no "Tiếng Đức:" or "Tiếng Nhật:".
// year simply never matches and stays NULL: 2016 files contain no "KHTN:",
// "KHXH:" or "GDCD:" tokens, since those combined papers did not exist yet.
var scorePatternSources = map[string]string{
"toan": `Toán:\s*(\d+(?:\.\d+)?)`,
"ngu_van": `Ngữ văn:\s*(\d+(?:\.\d+)?)`,
@@ -1,66 +0,0 @@
package transform_test
import (
"database/sql"
"os"
"testing"
_ "github.com/tiennm99/thptqg/parser/internal/sqlitedb"
"github.com/tiennm99/thptqg/parser/internal/transform"
)
// TestToAsciiAgainstRustOutput cross-checks ToAscii against a reference database
// on real data.
//
// Such a database is its own oracle: every row carries ho_ten alongside the
// ho_ten_ascii derived from it, so the whole table is a name -> expected-slug
// corpus far broader than the 20 hand-picked unit cases.
//
// GO_PARSER_RUST_DB=/tmp/rust-2016.db go test ./internal/transform/
//
// Skips when unset, so the default suite stays hermetic. The reference parser is
// gone from the tree, so no new oracle database can be produced and this always
// skips today.
func TestToAsciiAgainstRustOutput(t *testing.T) {
path := os.Getenv("GO_PARSER_RUST_DB")
if path == "" {
t.Skip("GO_PARSER_RUST_DB not set; skipping cross-check against Rust output")
}
if _, err := os.Stat(path); err != nil {
t.Skipf("GO_PARSER_RUST_DB=%s not readable: %v", path, err)
}
db, err := sql.Open("sqlite", "file:"+path+"?mode=ro")
if err != nil {
t.Fatalf("open %s: %v", path, err)
}
defer db.Close()
rows, err := db.Query("SELECT ho_ten, ho_ten_ascii FROM student")
if err != nil {
t.Fatalf("query: %v", err)
}
defer rows.Close()
var checked, bad int
for rows.Next() {
var name, rustAscii string
if err := rows.Scan(&name, &rustAscii); err != nil {
t.Fatalf("scan: %v", err)
}
checked++
if got := transform.ToAscii(name); got != rustAscii {
bad++
if bad <= 5 {
t.Errorf("ToAscii(%q)\n rust = %q\n go = %q", name, rustAscii, got)
}
}
}
if err := rows.Err(); err != nil {
t.Fatalf("iterate: %v", err)
}
if checked == 0 {
t.Fatal("database contained no rows")
}
t.Logf("compared %d real names, %d mismatches", checked, bad)
}
+2 -11
View File
@@ -68,9 +68,7 @@ type ParsedRow struct {
//
// The distinction is load-bearing for the printed counters: the two non-blank
// reasons count as source rows while SkipBlankRow does not. That split lives in
// the CALLER, not here — the build loop drops blank rows before the source-row
// counter and, at the later switch, lets SkipBlankRow fall through to transform
// and insert. Both call sites have to stay as they are.
// the CALLER — the build loop drops blank rows before the source-row counter.
type SkipReason int
const (
@@ -188,14 +186,7 @@ func TransformRow(raw []reader.Cell, cfg *config.DatasetConfig) (*ParsedRow, err
hoTen := get(cols.HoTen)
ngaySinh := get(cols.NgaySinh)
soBaoDanh := get(cols.SoBaoDanh)
// diem_thi is read WITHOUT trimming, unlike the three fields above. Harmless
// because the score patterns are unanchored, but it is the shipped behaviour
// — do not "tidy" it.
diemThi := ""
if cols.DiemThi >= 0 && cols.DiemThi < len(raw) {
diemThi = raw[cols.DiemThi].Str
}
diemThi := get(cols.DiemThi)
var ngaySinhOpt *string
if ngaySinh != "" {
+4 -5
View File
@@ -242,10 +242,10 @@ func TestTransformRowShortRowYieldsEmptyFields(t *testing.T) {
}
}
// TestTransformRowDiemThiIsNotTrimmed pins an asymmetry that is easy to
// "tidy away": ho_ten, ngay_sinh and so_bao_danh are trimmed, but diem_thi is
// read raw.
func TestTransformRowDiemThiIsNotTrimmed(t *testing.T) {
// TestTransformRowTrimsEveryField: every column is read trimmed, diem_thi
// included. Source cells routinely carry padding — 850k of the 861k 2017 rows
// have whitespace around their score cell.
func TestTransformRowTrimsEveryField(t *testing.T) {
row := cells(" A ", " 01/01/2000 ", " 123 ", " Toán: 5 ")
got, err := TransformRow(row, fixedColumnCfg())
if err != nil {
@@ -257,7 +257,6 @@ func TestTransformRowDiemThiIsNotTrimmed(t *testing.T) {
if got.NgaySinh == nil || *got.NgaySinh != "01/01/2000" {
t.Errorf("ngay_sinh = %v, want trimmed", got.NgaySinh)
}
// Untrimmed diem_thi still parses — the regexes are unanchored.
if got.Scores["toan"] != 5 {
t.Errorf("scores = %v", got.Scores)
}