mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
fix(parser): read the two 2016 layouts that were column-shifted
Four of the 119 files in data/2016 publish one score column per subject instead of a DIEM_THI sentence, and none of them was being read correctly. The ĐH Công nghiệp Thực phẩm file puts a three-row ministry title block above its header, so no header was recognised and the positional fallback shifted every column by one: the serial number became so_bao_danh, the exam number became ho_ten, the name became ngay_sinh, and the national ID became the score cell. All 7,833 rows were unusable. The three ĐH Cần Thơ files name an SBD column but no DIEM_THI, so they fell to the same fallback: surname into ngay_sinh, given name into ten_cum_thi, birth date into the score cell, and 12,152 candidates with no scores at all. Both are now read by FormatSubjectColumns, which resolves identity and one column per subject from the header. The header is searched for in the first five rows, so a title block no longer hides it. The Cần Thơ score columns are numbered rather than named. They follow the order the exam was sat — each morning an essay paper, each afternoon a multiple-choice one — which is what identifies them: columns 1/3/5/7 quantise to 0.25 and 2/4/6/8 do not, and each column's mean lands within 0.5 of the same subject's mean across the rest of the dataset. The foreign language is filed under the subject its N1..N6 code names. Gender now accepts the 0/1 encoding those files use: of the rows marked 1, 53% carry "Thị" in the name against 1% of those marked 0. Birth dates in the compact ddmmyy form are expanded so the column holds one format. A score of 0 is stored rather than dropped, recovering 302 real scores that a JavaScript falsy check had been turning into NULL. Row count falls by one, to 877,460: the removed row is the title line "ĐƠN VỊ: / TRƯỜNG ĐẠI HỌC CÔNG NGHIỆP THỰC PHẨM TP. HỒ CHÍ MINH", which had been stored as a student. The dataset has no duplicate exam numbers; the three rows previously described as collapsing were that same file's title and header lines being counted and then rejected. Also drops behaviour that existed only to match the parser this one replaced: the inert "SINH " header token, the untrimmed diem_thi cell, an unreachable blank-row branch, and a cross-check test against a database that can no longer exist. None of them changes output. Verified by rebuilding both datasets: 877,460 and 861,068 rows, both artifacts through the assembler's row and size guards, and the reader fidelity suite unchanged across all 182 files.
This commit is contained in:
1 parent
26d0ba81ce
commit
219c7c6a69
15 files changed
+535
-237
No files matched your search
@@ -34,16 +34,20 @@ hashes every real input file. That is the point of it; do not skip it.
|
||||
## Things that look wrong but are not
|
||||
|
||||
- **`data/2016/` filenames are content hashes and must stay verbatim.** The
|
||||
parser sorts inputs bytewise and inserts last-wins, so filenames decide which
|
||||
row survives a duplicate exam number — 877,464 source rows collapse to
|
||||
877,461. Renaming them also breaks `parser/testdata/reader-fidelity-hashes.tsv`,
|
||||
which is keyed by full path.
|
||||
parser sorts inputs bytewise and inserts last-wins, so filenames would decide
|
||||
which row survives a duplicate exam number. Renaming them also breaks
|
||||
`parser/testdata/reader-fidelity-hashes.tsv`, which is keyed by full path.
|
||||
- **`parser/testdata/reader-fidelity-hashes.tsv` is frozen and cannot be
|
||||
regenerated.** A mismatch is a reader bug until proven otherwise, never a cue
|
||||
to refresh the file. `parser/cmd/dumpcells` narrows a failure to the cell.
|
||||
- **`ToAscii` filters the literal range U+0300..U+036F, not `unicode.Mn`.** It
|
||||
must match `toAscii` in `web/src/App.jsx`, or accent-insensitive search
|
||||
silently misses rows.
|
||||
- **The 2016 files use four different layouts, and detection is per sheet.**
|
||||
Two of them publish scores in one column per subject instead of a `DIEM_THI`
|
||||
sentence, and one puts a three-row ministry title block above its header.
|
||||
`parser/internal/ingest/detect2016.go` holds the header tables; they are
|
||||
observations about 119 specific files, not a rule to generalise.
|
||||
- **`base` in `web/vite.config.js` is absolute.** One emitted `index.html` is
|
||||
copied to every dataset path, so every URL is a real static file and no SPA
|
||||
404-fallback is needed — a fallback would break the `?q=` deep links.
|
||||
|
||||
@@ -9,7 +9,7 @@ Live at **[tiennm99.github.io/thptqg](https://tiennm99.github.io/thptqg/)**.
|
||||
|
||||
| Dataset | Exam | Candidates | Site |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 2016 | 877,461 | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
|
||||
| `2016` | 2016 | 877,460 | [/2016/](https://tiennm99.github.io/thptqg/2016/) |
|
||||
| `2017` | 2017 | 861,068 | [/2017/](https://tiennm99.github.io/thptqg/2017/) |
|
||||
|
||||
Two earlier 2017 publications (`2017-old`, `2017-old2`) were kept for a while
|
||||
|
||||
+2
-2
@@ -18,8 +18,8 @@
|
||||
"datasets": [
|
||||
{
|
||||
"id": "2016",
|
||||
"expectedRows": 877461,
|
||||
"dbSizeMb": 44
|
||||
"expectedRows": 877460,
|
||||
"dbSizeMb": 45
|
||||
},
|
||||
{
|
||||
"id": "2017",
|
||||
|
||||
+16
-4
@@ -92,17 +92,29 @@ server-assigned names verbatim, since it did not choose them.
|
||||
| 2 | SOBAODANH | 8-digit string, first 2 digits = province |
|
||||
| 3 | DIEM_THI | concatenated per-subject scores, e.g. `"Toán: 6.80 Ngữ văn: 5.25 …"` |
|
||||
|
||||
2016 has **three** layouts across its 119 files, chosen per file at runtime via
|
||||
2016 has **four** layouts across its 119 files, chosen per sheet at runtime via
|
||||
`format_detection: thptqg2016` in its config:
|
||||
|
||||
| Format | Detected by | Notes |
|
||||
| --- | --- | --- |
|
||||
| `separate-scores` | `row[0] == "SBD"` and `row[2] == "TOAN"` | one column per subject; scores read directly, no regex |
|
||||
| `mapped` | header contains `SOBAODANH`/`SBD` plus `DIEM_THI` | column indices derived from the header |
|
||||
| `subject-columns` | header names three or more subject columns | university-cluster files; identity and scores both resolved by header name |
|
||||
| `default` | no recognised header | positional 6-column layout |
|
||||
|
||||
The `mapped` and `default` layouts also carry `TEN_CUMTHI` and `GIOI_TINH`,
|
||||
which is why only 2016 populates those columns.
|
||||
The header may sit below a title block, so the first five rows are searched for
|
||||
it; rows above it are not data.
|
||||
|
||||
`subject-columns` covers two spellings. The ĐH Công nghiệp Thực phẩm file names
|
||||
its columns `TO VA LI HO SI SU DI NN` with the language code in `Môn NN`; the
|
||||
three ĐH Cần Thơ files publish `cdiem1..cdiem8`, numbered in the order the 2016
|
||||
exam was sat — Toán, Ngoại ngữ, Ngữ văn, Vật lí, Địa lí, Hóa học, Lịch sử, Sinh
|
||||
học — with the language code in `ngoaingu`. The language score is filed under
|
||||
the subject its `N1`..`N6` code names.
|
||||
|
||||
The `mapped`, `default` and `subject-columns` layouts also carry `GIOI_TINH`
|
||||
(and `mapped`/`default` `TEN_CUMTHI`), which is why only 2016 populates those
|
||||
columns.
|
||||
|
||||
## Score text parsing
|
||||
|
||||
@@ -171,7 +183,7 @@ silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what
|
||||
|
||||
| id | Source rows | Skipped | DB rows |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 877,464 | 3 duplicate SBDs collapsed | **877,461** |
|
||||
| `2016` | 877,460 | 0 | **877,460** |
|
||||
| `2017` | 861,068 | 0 | **861,068** |
|
||||
|
||||
## Verifying a rebuild
|
||||
|
||||
@@ -34,7 +34,7 @@ running entirely in the browser and hosted for free on GitHub Pages. Covers the
|
||||
|
||||
| id | Exam | Candidates | Notes |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 2016 | 877,461 | 119 files, three column layouts |
|
||||
| `2016` | 2016 | 877,460 | 119 files, four column layouts |
|
||||
| `2017` | 2017 | 861,068 | current generation of three publications |
|
||||
|
||||
Two further 2017 datasets (`2017-old`, `2017-old2`) were kept alongside these
|
||||
|
||||
@@ -51,7 +51,7 @@ that does not exist.
|
||||
|
||||
| id | Exam | Rows | Source |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 2016 | 877,461 | `dtnt.bacninh.edu.vn` |
|
||||
| `2016` | 2016 | 877,460 | `dtnt.bacninh.edu.vn` |
|
||||
| `2017` | 2017 | 861,068 | `baotintuc.vn` |
|
||||
|
||||
Full source URLs are in [data-pipeline](./data-pipeline.md#sources); the web
|
||||
@@ -139,9 +139,9 @@ and had drifted apart on exactly this rule).
|
||||
| 9 digits with leading zero | `017006021` | 2016 |
|
||||
| 2-4 letters then digits | `BAL000001` | 2016 — exam cluster code |
|
||||
|
||||
The letter-prefixed form is the majority case for 2016: 616,593 of 877,461
|
||||
candidates (70.3%). Letter prefixes are upper-cased before lookup, so
|
||||
`bal000001` resolves.
|
||||
The letter-prefixed form is the majority case for 2016: 624,424 of 877,460
|
||||
candidates (71.2%); the remaining 253,036 are all 9-digit. Letter prefixes are
|
||||
upper-cased before lookup, so `bal000001` resolves.
|
||||
|
||||
## Score tiers
|
||||
|
||||
|
||||
+23
-28
@@ -26,7 +26,7 @@ go -C assembler run ./cmd/assemble db
|
||||
| path | role |
|
||||
|---|---|
|
||||
| `internal/reader` | spreadsheet reading; the only place that knows about file formats |
|
||||
| `internal/ingest` | dataset policy — sheet selection, header skipping, blank rows, the build loop, and the 2016 per-sheet format detection |
|
||||
| `internal/ingest` | dataset policy — sheet selection, header skipping, blank rows, the build loop, and the 2016 per-sheet layout detection |
|
||||
| `internal/transform` | `ToAscii`, score-regex parsing, row validation |
|
||||
| `internal/schema` | the canonical 22-column table: DDL, INSERT, subject regexes |
|
||||
| `internal/config` | per-dataset YAML parse rules |
|
||||
@@ -37,40 +37,35 @@ The reader deliberately knows nothing about datasets: it reports every sheet and
|
||||
every row verbatim. All policy lives in `ingest`. That split is what made the
|
||||
reader independently verifiable against a hash oracle.
|
||||
|
||||
## Provenance
|
||||
## Behaviour worth knowing
|
||||
|
||||
This is a port of a Rust crate that occupied this same path until the Go
|
||||
implementation reached full parity, when it was built alongside as `go-parser/`
|
||||
and moved back here once the Rust was removed. Source comments cite the original
|
||||
by file and line (`parser/src/transform.rs:56` and similar) — those refer to the
|
||||
Rust tree and resolve at the tag **`pre-go-parser-removal`**, the last commit
|
||||
containing it.
|
||||
|
||||
The port was gated on a field-by-field comparison of both implementations across
|
||||
the four datasets that existed then — 3,265,641 rows with identical full-table
|
||||
SHA-256, identical per-column non-NULL counts, identical schema metadata and
|
||||
identical build stdout. `assemble verify` is that comparator today and
|
||||
still runs against any two sets of databases.
|
||||
|
||||
Behaviour was matched bug-for-bug, deliberately. Several quirks look like
|
||||
defects and are load-bearing for the published data:
|
||||
|
||||
- a parsed score of `0` becomes NULL in the 2016 separate-scores layout,
|
||||
replicating a JavaScript falsy check;
|
||||
- `ToAscii` strips combining marks in the literal range U+0300–U+036F rather
|
||||
than by Unicode category, which is narrower;
|
||||
- gender is a two-value allowlist, and anything else becomes NULL;
|
||||
- `diem_thi` is read untrimmed while the other three fields are trimmed;
|
||||
- the `"SINH "` header token carries a trailing space.
|
||||
than by Unicode category. That covers every Vietnamese diacritic and must
|
||||
stay identical to `toAscii` in `web/src/App.jsx`, or accent-insensitive
|
||||
search misses rows.
|
||||
- Gender is normalised to `Nam`/`Nữ`; the Cần Thơ files write `0`/`1` instead
|
||||
and are translated. Anything else becomes NULL.
|
||||
- A score of `0` is a real score — the candidate sat the paper and scored
|
||||
nothing — and is stored, not dropped.
|
||||
- Birth dates are stored as `dd/mm/yyyy`. The Cần Thơ files' compact `ddmmyy`
|
||||
is expanded; the century is always 19xx, since a 2016 candidate born later
|
||||
would have sat the exam under age.
|
||||
|
||||
Each has a test naming it, so none can be tidied away by accident.
|
||||
This code began as a port of a Rust crate that occupied the same path, and was
|
||||
gated on a field-by-field comparison against it. That comparison is over:
|
||||
correctness against the source spreadsheets decides behaviour now, not
|
||||
agreement with the old implementation. `assemble verify` is the comparator
|
||||
that gated it, and still compares any two sets of built databases.
|
||||
|
||||
## Verification
|
||||
|
||||
`testdata/reader-fidelity-hashes.tsv` holds a SHA-256 per input file over a
|
||||
canonical dump of every cell of every sheet. It is **frozen**: it was produced
|
||||
by the Rust reader, which no longer exists, so it cannot be regenerated. It
|
||||
still fails if any single cell of any input file reads differently.
|
||||
canonical dump of every cell of every sheet. It is **frozen**: the tool that
|
||||
produced it no longer exists, so it cannot be regenerated. It still fails if
|
||||
any single cell of any input file reads differently, which is what makes it
|
||||
useful — cell rendering and sheet geometry are settled, and a change there is
|
||||
a regression until proven otherwise. Dataset policy lives in `ingest`, so
|
||||
fixing a layout never touches it.
|
||||
|
||||
A mismatch names the file but not the cell. `cmd/dumpcells` prints the stream
|
||||
the hash is taken over, so two runs can be diffed:
|
||||
|
||||
@@ -1,10 +1,11 @@
|
||||
# thptqg2016 data/ — 4 .xls + 115 .xlsx mixed files.
|
||||
#
|
||||
# Three column layouts exist across the 119 files; the binary selects the
|
||||
# right one per-file at runtime via format_detection: thptqg2016:
|
||||
# Four column layouts exist across the 119 files; the binary selects the
|
||||
# right one per-sheet at runtime via format_detection: thptqg2016:
|
||||
#
|
||||
# separate-scores header SBD(0)/HOTEN(1)/TOAN(2)... — dhhanghai files
|
||||
# mapped header SOBAODANH|SBD + DIEM_THI — most provinces
|
||||
# subject-columns header names >=3 subject columns — dhcantho, dhcnthucpham
|
||||
# default no header; positional 6-col layout — remaining files
|
||||
#
|
||||
# This is the only dataset that populates ten_cum_thi and gioi_tinh, and the
|
||||
|
||||
@@ -14,14 +14,11 @@ import (
|
||||
)
|
||||
|
||||
// The 2016 dataset's 119 files were produced by inconsistent tooling and use
|
||||
// three different column layouts, chosen per sheet at runtime. The literals
|
||||
// four different column layouts, chosen per sheet at runtime. The literals
|
||||
// below are institutional knowledge about those files, with no rule to derive
|
||||
// them from — extend them, but do not tidy them into something more regular.
|
||||
|
||||
// KnownHeaders are the upper-cased first-cell values that identify a header row.
|
||||
//
|
||||
// NOTE "SINH " carries a TRAILING SPACE. Trimming it would change which rows
|
||||
// are recognised as headers.
|
||||
var KnownHeaders = []string{
|
||||
"SOBAODANH",
|
||||
"SBD",
|
||||
@@ -37,7 +34,7 @@ var KnownHeaders = []string{
|
||||
"VAN",
|
||||
"LY",
|
||||
"HOA",
|
||||
"SINH ",
|
||||
"SINH",
|
||||
"SU",
|
||||
"DIA",
|
||||
}
|
||||
@@ -61,7 +58,7 @@ func IsHeaderRow2016(row []reader.Cell) bool {
|
||||
return isKnownHeader(strings.ToUpper(strings.TrimSpace(row[0].Str)))
|
||||
}
|
||||
|
||||
// FormatKind is one of the three 2016 layouts.
|
||||
// FormatKind is one of the 2016 layouts.
|
||||
type FormatKind int
|
||||
|
||||
const (
|
||||
@@ -73,9 +70,13 @@ const (
|
||||
// FormatDefault is the positional 6-column layout used when no header is
|
||||
// recognised. It is FormatMapped with fixed indices, not a separate path.
|
||||
FormatDefault
|
||||
// FormatSubjectColumns resolves BOTH identity and one score column per
|
||||
// subject from the header by name — the university-cluster files, which
|
||||
// publish scores in columns instead of a DIEM_THI sentence.
|
||||
FormatSubjectColumns
|
||||
)
|
||||
|
||||
// Format is a detected layout plus, for the mapped case, its column indices.
|
||||
// Format is a detected layout plus its resolved column indices.
|
||||
type Format struct {
|
||||
Kind FormatKind
|
||||
Sbd int
|
||||
@@ -84,6 +85,14 @@ type Format struct {
|
||||
TenCumThi *int
|
||||
GioiTinh *int
|
||||
DiemThi int
|
||||
|
||||
// Subjects maps a schema score field to its column, for
|
||||
// FormatSubjectColumns only.
|
||||
Subjects map[string]int
|
||||
// LangScore is the foreign-language score column; the subject it counts
|
||||
// towards is named by the code in LangCode.
|
||||
LangScore *int
|
||||
LangCode *int
|
||||
}
|
||||
|
||||
// defaultFormat is FormatMapped with the fixed positional indices.
|
||||
@@ -95,6 +104,90 @@ func defaultFormat() Format {
|
||||
}
|
||||
}
|
||||
|
||||
// identity holds the identity columns resolved from a header row. A nil field
|
||||
// means the header does not name that column.
|
||||
type identity struct {
|
||||
Sbd, HoTen, NgaySinh, TenCumThi, GioiTinh, DiemThi *int
|
||||
}
|
||||
|
||||
// resolveIdentity maps header names to identity columns, order-independent.
|
||||
// Each column has several spellings across the corpus, including the accented
|
||||
// and unseparated ones the university files use.
|
||||
func resolveIdentity(cols []string) identity {
|
||||
var id identity
|
||||
for i := range cols {
|
||||
idx := i
|
||||
switch cols[i] {
|
||||
case "SOBAODANH", "SBD":
|
||||
id.Sbd = &idx
|
||||
case "HO_TEN", "HOTEN", "HỌ TÊN":
|
||||
id.HoTen = &idx
|
||||
case "NGAY_SINH", "NGAYSINH", "NGÀY SINH":
|
||||
id.NgaySinh = &idx
|
||||
case "TEN_CUMTHI":
|
||||
id.TenCumThi = &idx
|
||||
case "GIOI_TINH", "GIỚI TÍNH", "PHAI":
|
||||
id.GioiTinh = &idx
|
||||
case "DIEM_THI":
|
||||
id.DiemThi = &idx
|
||||
}
|
||||
}
|
||||
return id
|
||||
}
|
||||
|
||||
// subjectColumnFamily is one university-cluster spelling of the per-subject
|
||||
// score columns. The two families are told apart by which names resolve.
|
||||
type subjectColumnFamily struct {
|
||||
scores map[string]string // header name -> schema score field
|
||||
langScore string // header of the foreign-language score column
|
||||
langCode string // header of the N1..N6 language code column
|
||||
}
|
||||
|
||||
// subjectColumnFamilies are tried in order. The Cần Thơ family goes first
|
||||
// because its CDIEM<n> names are unambiguous, while the two-letter family
|
||||
// shares "HO" with that layout's surname column.
|
||||
var subjectColumnFamilies = []subjectColumnFamily{
|
||||
{scores: canThoScoreHeaders, langScore: "CDIEM2", langCode: "NGOAINGU"},
|
||||
{scores: subjectCodeHeaders, langScore: "NN", langCode: "MÔN NN"},
|
||||
}
|
||||
|
||||
// canThoScoreHeaders maps the Cần Thơ cluster's published-score columns to
|
||||
// schema fields. The columns are numbered in the order the 2016 exam was sat —
|
||||
// each morning an essay paper (quarter-point scores), each afternoon a
|
||||
// multiple-choice one — which is what identifies them: CDIEM1/3/5/7 quantise to
|
||||
// 0.25 and CDIEM2/4/6/8 do not, and the per-column means match the published
|
||||
// national averages for those subjects.
|
||||
var canThoScoreHeaders = map[string]string{
|
||||
"CDIEM1": "toan",
|
||||
"CDIEM3": "ngu_van",
|
||||
"CDIEM4": "vat_ly",
|
||||
"CDIEM5": "dia_ly",
|
||||
"CDIEM6": "hoa_hoc",
|
||||
"CDIEM7": "lich_su",
|
||||
"CDIEM8": "sinh_hoc",
|
||||
}
|
||||
|
||||
// subjectCodeHeaders maps the two-letter subject columns to schema fields.
|
||||
var subjectCodeHeaders = map[string]string{
|
||||
"TO": "toan",
|
||||
"VA": "ngu_van",
|
||||
"LI": "vat_ly",
|
||||
"HO": "hoa_hoc",
|
||||
"SI": "sinh_hoc",
|
||||
"SU": "lich_su",
|
||||
"DI": "dia_ly",
|
||||
}
|
||||
|
||||
// languageFields maps the ministry's foreign-language codes to schema fields.
|
||||
var languageFields = map[string]string{
|
||||
"N1": "tieng_anh",
|
||||
"N2": "tieng_nga",
|
||||
"N3": "tieng_phap",
|
||||
"N4": "tieng_trung",
|
||||
"N5": "tieng_duc",
|
||||
"N6": "tieng_nhat",
|
||||
}
|
||||
|
||||
// DetectFormat inspects a header row and decides which layout applies.
|
||||
func DetectFormat(headerRow []reader.Cell) Format {
|
||||
cols := make([]string, len(headerRow))
|
||||
@@ -107,54 +200,85 @@ func DetectFormat(headerRow []reader.Cell) Format {
|
||||
return Format{Kind: FormatSeparateScores}
|
||||
}
|
||||
|
||||
// Format 2: resolve indices by header name, order-independent.
|
||||
var sbdIdx, hoTenIdx, ngaySinhIdx, tenCumThiIdx, gioiTinhIdx, diemThiIdx *int
|
||||
for i := range cols {
|
||||
idx := i
|
||||
switch cols[i] {
|
||||
case "SOBAODANH", "SBD":
|
||||
sbdIdx = &idx
|
||||
case "HO_TEN", "HOTEN", "HỌ TÊN":
|
||||
hoTenIdx = &idx
|
||||
case "NGAY_SINH":
|
||||
ngaySinhIdx = &idx
|
||||
case "TEN_CUMTHI":
|
||||
tenCumThiIdx = &idx
|
||||
case "GIOI_TINH":
|
||||
gioiTinhIdx = &idx
|
||||
case "DIEM_THI":
|
||||
diemThiIdx = &idx
|
||||
id := resolveIdentity(cols)
|
||||
|
||||
// Format 2: a free-text DIEM_THI cell alongside an SBD.
|
||||
if id.Sbd != nil && id.DiemThi != nil {
|
||||
hoTen := 1 // fallback: col 1, present in all known files
|
||||
if id.HoTen != nil {
|
||||
hoTen = *id.HoTen
|
||||
}
|
||||
return Format{
|
||||
Kind: FormatMapped, Sbd: *id.Sbd, HoTen: hoTen,
|
||||
NgaySinh: id.NgaySinh, TenCumThi: id.TenCumThi, GioiTinh: id.GioiTinh,
|
||||
DiemThi: *id.DiemThi,
|
||||
}
|
||||
}
|
||||
|
||||
if sbdIdx != nil && diemThiIdx != nil {
|
||||
hoTen := 1 // fallback: col 1, present in all known files
|
||||
if hoTenIdx != nil {
|
||||
hoTen = *hoTenIdx
|
||||
}
|
||||
return Format{
|
||||
Kind: FormatMapped, Sbd: *sbdIdx, HoTen: hoTen,
|
||||
NgaySinh: ngaySinhIdx, TenCumThi: tenCumThiIdx, GioiTinh: gioiTinhIdx,
|
||||
DiemThi: *diemThiIdx,
|
||||
}
|
||||
// Format 3: one column per subject.
|
||||
if f, ok := subjectColumnsFormat(cols, id); ok {
|
||||
return f
|
||||
}
|
||||
|
||||
// Unrecognised header, or none at all.
|
||||
return defaultFormat()
|
||||
}
|
||||
|
||||
// parseFloatCell parses a per-subject score cell.
|
||||
// subjectColumnsFormat resolves a per-subject-column layout, if the header
|
||||
// names one.
|
||||
//
|
||||
// A parsed 0.0 becomes "no score", so a genuine zero is indistinguishable from
|
||||
// a blank. Not obviously correct, but it is the shipped behaviour and the
|
||||
// published data depends on it.
|
||||
func parseFloatCell(row []reader.Cell, idx int) (float64, bool) {
|
||||
// Three resolved subjects are required: a stray one- or two-letter header in an
|
||||
// unrelated file must not turn that file into this layout.
|
||||
func subjectColumnsFormat(cols []string, id identity) (Format, bool) {
|
||||
if id.Sbd == nil {
|
||||
return Format{}, false
|
||||
}
|
||||
for _, fam := range subjectColumnFamilies {
|
||||
subjects := make(map[string]int)
|
||||
for i, name := range cols {
|
||||
if field, ok := fam.scores[name]; ok {
|
||||
subjects[field] = i
|
||||
}
|
||||
}
|
||||
if len(subjects) < 3 {
|
||||
continue
|
||||
}
|
||||
hoTen := 1
|
||||
if id.HoTen != nil {
|
||||
hoTen = *id.HoTen
|
||||
}
|
||||
return Format{
|
||||
Kind: FormatSubjectColumns, Sbd: *id.Sbd, HoTen: hoTen,
|
||||
NgaySinh: id.NgaySinh, TenCumThi: id.TenCumThi, GioiTinh: id.GioiTinh,
|
||||
DiemThi: -1,
|
||||
Subjects: subjects,
|
||||
LangScore: indexOf(cols, fam.langScore),
|
||||
LangCode: indexOf(cols, fam.langCode),
|
||||
}, true
|
||||
}
|
||||
return Format{}, false
|
||||
}
|
||||
|
||||
// indexOf returns the column holding name, or nil when the header lacks it.
|
||||
func indexOf(cols []string, name string) *int {
|
||||
for i := range cols {
|
||||
if cols[i] == name {
|
||||
idx := i
|
||||
return &idx
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// parseScoreCell parses a per-subject score cell. A parsed 0 is a real score:
|
||||
// the candidate sat the paper and scored nothing.
|
||||
func parseScoreCell(row []reader.Cell, idx int) (float64, bool) {
|
||||
s := cellAt(row, idx)
|
||||
if s == "" {
|
||||
return 0, false
|
||||
}
|
||||
v, err := strconv.ParseFloat(s, 64)
|
||||
if err != nil || v == 0.0 {
|
||||
if err != nil {
|
||||
return 0, false
|
||||
}
|
||||
return v, true
|
||||
@@ -182,7 +306,7 @@ func processSeparateScoresRow(row []reader.Cell) *transform.ParsedRow {
|
||||
"sinh_hoc": 6, "lich_su": 7, "dia_ly": 8,
|
||||
"tieng_anh": 11, // NGOAINGU total
|
||||
} {
|
||||
if v, ok := parseFloatCell(row, idx); ok {
|
||||
if v, ok := parseScoreCell(row, idx); ok {
|
||||
scores[field] = v
|
||||
}
|
||||
}
|
||||
@@ -209,47 +333,146 @@ func processMappedRow(row []reader.Cell, f Format) *transform.ParsedRow {
|
||||
return nil
|
||||
}
|
||||
|
||||
optional := func(idx *int) *string {
|
||||
if idx == nil {
|
||||
return nil
|
||||
}
|
||||
if s := cellAt(row, *idx); s != "" {
|
||||
return &s
|
||||
}
|
||||
return &transform.ParsedRow{
|
||||
SoBaoDanh: sbd,
|
||||
HoTen: hoTen,
|
||||
HoTenAscii: transform.ToAscii(hoTen),
|
||||
NgaySinh: birthDate(row, f.NgaySinh),
|
||||
TenCumThi: optionalCell(row, f.TenCumThi),
|
||||
GioiTinh: gender(row, f.GioiTinh),
|
||||
Scores: transform.ParseScores(cellAt(row, f.DiemThi)),
|
||||
}
|
||||
}
|
||||
|
||||
// processSubjectColumnsRow reads a row whose scores sit in one column per
|
||||
// subject, with the foreign language filed under the subject its code names.
|
||||
func processSubjectColumnsRow(row []reader.Cell, f Format) *transform.ParsedRow {
|
||||
sbd := cellAt(row, f.Sbd)
|
||||
hoTen := cellAt(row, f.HoTen)
|
||||
if sbd == "" || hoTen == "" {
|
||||
return nil
|
||||
}
|
||||
if isKnownHeader(strings.ToUpper(sbd)) || isKnownHeader(strings.ToUpper(hoTen)) {
|
||||
return nil
|
||||
}
|
||||
|
||||
// Gender is a two-value allowlist, not a general enum: anything other than
|
||||
// exactly "Nam" or "Nữ" becomes nil.
|
||||
var gioiTinh *string
|
||||
if f.GioiTinh != nil {
|
||||
if s := cellAt(row, *f.GioiTinh); s == "Nam" || s == "Nữ" {
|
||||
gioiTinh = &s
|
||||
scores := make(map[string]float64)
|
||||
for field, idx := range f.Subjects {
|
||||
if v, ok := parseScoreCell(row, idx); ok {
|
||||
scores[field] = v
|
||||
}
|
||||
}
|
||||
|
||||
// The free-text score cell is read untrimmed, unlike every other cell here.
|
||||
diemThi := ""
|
||||
if f.DiemThi >= 0 && f.DiemThi < len(row) {
|
||||
diemThi = row[f.DiemThi].Str
|
||||
if f.LangScore != nil {
|
||||
if v, ok := parseScoreCell(row, *f.LangScore); ok {
|
||||
code := "N1" // an unlabelled language paper is English in this corpus
|
||||
if f.LangCode != nil {
|
||||
if c := strings.ToUpper(cellAt(row, *f.LangCode)); c != "" {
|
||||
code = c
|
||||
}
|
||||
}
|
||||
// An unknown code is left out rather than guessed at.
|
||||
if field, ok := languageFields[code]; ok {
|
||||
scores[field] = v
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return &transform.ParsedRow{
|
||||
SoBaoDanh: sbd,
|
||||
HoTen: hoTen,
|
||||
HoTenAscii: transform.ToAscii(hoTen),
|
||||
NgaySinh: optional(f.NgaySinh),
|
||||
TenCumThi: optional(f.TenCumThi),
|
||||
GioiTinh: gioiTinh,
|
||||
Scores: transform.ParseScores(diemThi),
|
||||
NgaySinh: birthDate(row, f.NgaySinh),
|
||||
TenCumThi: optionalCell(row, f.TenCumThi),
|
||||
GioiTinh: gender(row, f.GioiTinh),
|
||||
Scores: scores,
|
||||
}
|
||||
}
|
||||
|
||||
// optionalCell returns the trimmed cell at idx, or nil when the column is
|
||||
// absent or its cell empty.
|
||||
func optionalCell(row []reader.Cell, idx *int) *string {
|
||||
if idx == nil {
|
||||
return nil
|
||||
}
|
||||
if s := cellAt(row, *idx); s != "" {
|
||||
return &s
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// gender normalises the gender cell to "Nam"/"Nữ", or nil when it says neither.
|
||||
//
|
||||
// The Cần Thơ files encode it numerically instead: of the rows marked 1, 53%
|
||||
// carry the female marker "Thị" in the name, against 1% of those marked 0.
|
||||
func gender(row []reader.Cell, idx *int) *string {
|
||||
s := optionalCell(row, idx)
|
||||
if s == nil {
|
||||
return nil
|
||||
}
|
||||
switch *s {
|
||||
case "Nam", "Nữ":
|
||||
return s
|
||||
case "0":
|
||||
nam := "Nam"
|
||||
return &nam
|
||||
case "1":
|
||||
nu := "Nữ"
|
||||
return &nu
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// birthDate returns the birth-date cell as dd/mm/yyyy.
|
||||
//
|
||||
// Most files already write it that way; the Cần Thơ ones use a compact ddmmyy,
|
||||
// expanded here so the column holds one format. The century is always 19xx: a
|
||||
// 2016 candidate born after 1999 would have sat the exam under age.
|
||||
func birthDate(row []reader.Cell, idx *int) *string {
|
||||
s := optionalCell(row, idx)
|
||||
if s == nil || len(*s) != 6 || !allDigits(*s) {
|
||||
return s
|
||||
}
|
||||
out := (*s)[0:2] + "/" + (*s)[2:4] + "/19" + (*s)[4:6]
|
||||
return &out
|
||||
}
|
||||
|
||||
func allDigits(s string) bool {
|
||||
for i := 0; i < len(s); i++ {
|
||||
if s[i] < '0' || s[i] > '9' {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// headerScanRows is how far into a sheet the header may sit. One file opens
|
||||
// with a three-row ministry title block above its header; a sheet with no
|
||||
// header at all has data from row 0 and simply finds nothing in that window.
|
||||
const headerScanRows = 5
|
||||
|
||||
// sheetFormat picks the layout for one sheet and returns the row its data
|
||||
// starts on. Rows above the header are a title block and are not data.
|
||||
func sheetFormat(rows [][]reader.Cell) (Format, int) {
|
||||
limit := headerScanRows
|
||||
if len(rows) < limit {
|
||||
limit = len(rows)
|
||||
}
|
||||
for i := 0; i < limit; i++ {
|
||||
if IsHeaderRow2016(rows[i]) {
|
||||
return DetectFormat(rows[i]), i + 1
|
||||
}
|
||||
}
|
||||
return defaultFormat(), 0
|
||||
}
|
||||
|
||||
// ProcessRow2016 dispatches a data row through the detected layout. A nil return
|
||||
// means the row is empty or invalid and should be skipped.
|
||||
func ProcessRow2016(row []reader.Cell, f Format) *transform.ParsedRow {
|
||||
if f.Kind == FormatSeparateScores {
|
||||
switch f.Kind {
|
||||
case FormatSeparateScores:
|
||||
return processSeparateScoresRow(row)
|
||||
case FormatSubjectColumns:
|
||||
return processSubjectColumnsRow(row, f)
|
||||
}
|
||||
// FormatMapped and FormatDefault share one implementation; Default is just a
|
||||
// fixed index tuple.
|
||||
@@ -336,12 +559,7 @@ func processFile2016(path string, cfg *config.DatasetConfig, ins *writer.Inserte
|
||||
continue
|
||||
}
|
||||
|
||||
f := defaultFormat()
|
||||
startIdx := 0
|
||||
if IsHeaderRow2016(rows[0]) {
|
||||
f = DetectFormat(rows[0])
|
||||
startIdx = 1
|
||||
}
|
||||
f, startIdx := sheetFormat(rows)
|
||||
|
||||
for _, row := range rows[startIdx:] {
|
||||
// Rows shorter than 2 cells are dropped BEFORE the counter, so they
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
package ingest
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/tiennm99/thptqg/parser/internal/reader"
|
||||
@@ -31,20 +32,45 @@ func TestIsHeaderRow2016(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestKnownHeadersHasTrailingSpaceToken pins the "SINH " literal. Trimming it
|
||||
// would silently change which rows count as headers.
|
||||
func TestKnownHeadersHasTrailingSpaceToken(t *testing.T) {
|
||||
var found bool
|
||||
// TestKnownHeadersAreTrimmed: the cell is trimmed before comparison, so a token
|
||||
// carrying surrounding space could never match anything.
|
||||
func TestKnownHeadersAreTrimmed(t *testing.T) {
|
||||
for _, h := range KnownHeaders {
|
||||
if h == "SINH " {
|
||||
found = true
|
||||
if h != strings.TrimSpace(h) {
|
||||
t.Errorf("KnownHeaders token %q has surrounding space and can never match", h)
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
t.Error(`KnownHeaders must contain "SINH " WITH its trailing space`)
|
||||
}
|
||||
|
||||
// TestSheetFormatSkipsTitleBlock: one file opens with a ministry title block
|
||||
// above its header row, and the rows above the header are not data.
|
||||
func TestSheetFormatSkipsTitleBlock(t *testing.T) {
|
||||
rows := [][]reader.Cell{
|
||||
cells("BỘ GIÁO DỤC VÀ ĐÀO TẠO", "", ""),
|
||||
cells("ĐƠN VỊ:", "TRƯỜNG ĐẠI HỌC X", ""),
|
||||
cells("DANH SÁCH ĐIỂM THÍ SINH", "", ""),
|
||||
cells("STT", "SBD", "Họ tên", "Ngày sinh", "Giới tính", "CMND", "Tỉnh", "TO", "VA", "LI"),
|
||||
cells("1", "DCT000001", "Nguyễn Văn A", "01/01/1998", "Nam", "123", "46", "8", "7", "6"),
|
||||
}
|
||||
if len(KnownHeaders) != 17 {
|
||||
t.Errorf("KnownHeaders has %d tokens, want 17", len(KnownHeaders))
|
||||
f, start := sheetFormat(rows)
|
||||
if start != 4 {
|
||||
t.Errorf("data starts at row %d, want 4", start)
|
||||
}
|
||||
if f.Kind != FormatSubjectColumns {
|
||||
t.Errorf("Kind = %v, want FormatSubjectColumns", f.Kind)
|
||||
}
|
||||
}
|
||||
|
||||
// TestSheetFormatHeaderlessSheet: a sheet whose first row is already data keeps
|
||||
// the positional layout and starts at row 0.
|
||||
func TestSheetFormatHeaderlessSheet(t *testing.T) {
|
||||
rows := [][]reader.Cell{
|
||||
cells("YTB000001", "PHẠM THỊ HỒNG ÁI", "21/08/1998", "Y Dược Thái Bình", "Nữ", "Toán: 7.75"),
|
||||
cells("YTB000002", "TRẦN VĂN B", "22/08/1998", "Y Dược Thái Bình", "Nam", "Toán: 5"),
|
||||
}
|
||||
f, start := sheetFormat(rows)
|
||||
if start != 0 || f.Kind != FormatDefault {
|
||||
t.Errorf("got kind=%v start=%d, want FormatDefault at 0", f.Kind, start)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -133,36 +159,37 @@ func TestProcessSeparateScoresRow(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestZeroScoreBecomesNull pins the quirk that a literal 0 is indistinguishable
|
||||
// from "no score".
|
||||
func TestZeroScoreBecomesNull(t *testing.T) {
|
||||
// TestZeroScoreIsKept: a candidate who sat a paper and scored nothing has a
|
||||
// score of 0, not a missing one.
|
||||
func TestZeroScoreIsKept(t *testing.T) {
|
||||
row := cells("1000", "A", "0", "0.0", "1", "", "", "", "", "", "", "")
|
||||
got := ProcessRow2016(row, Format{Kind: FormatSeparateScores})
|
||||
if got == nil {
|
||||
t.Fatal("row rejected")
|
||||
}
|
||||
if _, ok := got.Scores["toan"]; ok {
|
||||
t.Error(`a "0" score must become NULL, not 0`)
|
||||
for field, want := range map[string]float64{"toan": 0, "ngu_van": 0, "vat_ly": 1} {
|
||||
v, ok := got.Scores[field]
|
||||
if !ok || v != want {
|
||||
t.Errorf("%s = %v (present=%v), want %v", field, v, ok, want)
|
||||
}
|
||||
}
|
||||
if _, ok := got.Scores["ngu_van"]; ok {
|
||||
t.Error(`a "0.0" score must become NULL, not 0`)
|
||||
}
|
||||
if got.Scores["vat_ly"] != 1 {
|
||||
t.Error("a non-zero score must survive")
|
||||
if _, ok := got.Scores["hoa_hoc"]; ok {
|
||||
t.Error("an empty cell must stay absent")
|
||||
}
|
||||
}
|
||||
|
||||
// TestGenderAllowlist: exactly two accepted values, matched case-sensitively.
|
||||
func TestGenderAllowlist(t *testing.T) {
|
||||
// TestGenderNormalisation: "Nam"/"Nữ" verbatim, the Cần Thơ files' 1/0 encoding
|
||||
// translated, everything else NULL.
|
||||
func TestGenderNormalisation(t *testing.T) {
|
||||
f := defaultFormat()
|
||||
for in, want := range map[string]string{"Nam": "Nam", "Nữ": "Nữ"} {
|
||||
for in, want := range map[string]string{"Nam": "Nam", "Nữ": "Nữ", "1": "Nữ", "0": "Nam"} {
|
||||
row := cells("1", "A", "", "", in, "")
|
||||
got := ProcessRow2016(row, f)
|
||||
if got == nil || got.GioiTinh == nil || *got.GioiTinh != want {
|
||||
t.Errorf("gioi_tinh for %q not preserved", in)
|
||||
t.Errorf("gioi_tinh for %q = %v, want %q", in, got.GioiTinh, want)
|
||||
}
|
||||
}
|
||||
for _, in := range []string{"nam", "NAM", "Unknown", "M", "F", "nữ", ""} {
|
||||
for _, in := range []string{"nam", "NAM", "Unknown", "M", "F", "nữ", "", "2"} {
|
||||
row := cells("1", "A", "", "", in, "")
|
||||
got := ProcessRow2016(row, f)
|
||||
if got == nil {
|
||||
@@ -174,6 +201,131 @@ func TestGenderAllowlist(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestBirthDateExpandsCompactForm: ddmmyy becomes dd/mm/19yy so the column
|
||||
// holds one format; anything else is passed through.
|
||||
func TestBirthDateExpandsCompactForm(t *testing.T) {
|
||||
f := defaultFormat()
|
||||
for in, want := range map[string]string{
|
||||
"140798": "14/07/1998",
|
||||
"01/01/1998": "01/01/1998",
|
||||
"1998": "1998",
|
||||
} {
|
||||
got := ProcessRow2016(cells("1", "A", in, "", "", ""), f)
|
||||
if got == nil || got.NgaySinh == nil || *got.NgaySinh != want {
|
||||
t.Errorf("ngay_sinh for %q = %v, want %q", in, got.NgaySinh, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// --- per-subject score columns ---
|
||||
|
||||
// TestDetectFormatSubjectColumnsTwoLetter covers the two-letter subject columns,
|
||||
// whose foreign-language score is filed under the subject its code names.
|
||||
func TestDetectFormatSubjectColumnsTwoLetter(t *testing.T) {
|
||||
header := cells("STT", "SBD", "Họ tên", "Ngày sinh", "Giới tính", "CMND", "Tỉnh",
|
||||
"TO", "VA", "LI", "HO", "SI", "SU", "DI", "NN", "Môn NN")
|
||||
f := DetectFormat(header)
|
||||
if f.Kind != FormatSubjectColumns {
|
||||
t.Fatalf("Kind = %v, want FormatSubjectColumns", f.Kind)
|
||||
}
|
||||
if f.Sbd != 1 || f.HoTen != 2 {
|
||||
t.Errorf("identity: sbd=%d ho_ten=%d", f.Sbd, f.HoTen)
|
||||
}
|
||||
if f.NgaySinh == nil || *f.NgaySinh != 3 || f.GioiTinh == nil || *f.GioiTinh != 4 {
|
||||
t.Error("accented identity headers not resolved")
|
||||
}
|
||||
if f.Subjects["toan"] != 7 || f.Subjects["dia_ly"] != 13 {
|
||||
t.Errorf("subject columns = %v", f.Subjects)
|
||||
}
|
||||
|
||||
row := cells("1", "DCT000073", "Trần Thị Phước An", "23/11/1998", "Nữ", "291183999", "46",
|
||||
"6.25", "4.5", "5.8", "", "", "", "", "5.38", "N1")
|
||||
got := ProcessRow2016(row, f)
|
||||
if got == nil {
|
||||
t.Fatal("row rejected")
|
||||
}
|
||||
if got.SoBaoDanh != "DCT000073" || got.HoTen != "Trần Thị Phước An" {
|
||||
t.Errorf("identity = %q / %q", got.SoBaoDanh, got.HoTen)
|
||||
}
|
||||
want := map[string]float64{"toan": 6.25, "ngu_van": 4.5, "vat_ly": 5.8, "tieng_anh": 5.38}
|
||||
if len(got.Scores) != len(want) {
|
||||
t.Errorf("scores = %v, want %v", got.Scores, want)
|
||||
}
|
||||
for k, v := range want {
|
||||
if got.Scores[k] != v {
|
||||
t.Errorf("%s = %v, want %v", k, got.Scores[k], v)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestDetectFormatSubjectColumnsCanTho covers the CDIEM<n> columns, which are
|
||||
// numbered in exam-timetable order rather than named.
|
||||
func TestDetectFormatSubjectColumnsCanTho(t *testing.T) {
|
||||
header := cells("sbd", "hoten", "ho", "ten", "phai", "ngaysinh", "socmnd",
|
||||
"cdiem1", "cdiem2", "cdiem3", "cdiem4", "cdiem5", "cdiem6", "cdiem7", "cdiem8", "ngoaingu")
|
||||
f := DetectFormat(header)
|
||||
if f.Kind != FormatSubjectColumns {
|
||||
t.Fatalf("Kind = %v, want FormatSubjectColumns", f.Kind)
|
||||
}
|
||||
// "ho" is the surname column here, and must not be read as hoa_hoc.
|
||||
if got, ok := f.Subjects["hoa_hoc"]; ok && got == 2 {
|
||||
t.Error("surname column resolved as a score column")
|
||||
}
|
||||
if f.Subjects["toan"] != 7 || f.Subjects["ngu_van"] != 9 || f.Subjects["sinh_hoc"] != 14 {
|
||||
t.Errorf("subject columns = %v", f.Subjects)
|
||||
}
|
||||
if f.LangScore == nil || *f.LangScore != 8 || f.LangCode == nil || *f.LangCode != 15 {
|
||||
t.Error("language score/code columns not resolved")
|
||||
}
|
||||
|
||||
row := cells("TCT000001", "Dương Diễm ái", "Dương Diễm", "ái", "1", "140798", "362539228",
|
||||
"07.50", "05.60", "06.50", "", "", "08.20", "", "07.60", "N3")
|
||||
got := ProcessRow2016(row, f)
|
||||
if got == nil {
|
||||
t.Fatal("row rejected")
|
||||
}
|
||||
if got.GioiTinh == nil || *got.GioiTinh != "Nữ" {
|
||||
t.Errorf("gioi_tinh = %v, want Nữ", got.GioiTinh)
|
||||
}
|
||||
if got.NgaySinh == nil || *got.NgaySinh != "14/07/1998" {
|
||||
t.Errorf("ngay_sinh = %v", got.NgaySinh)
|
||||
}
|
||||
want := map[string]float64{"toan": 7.5, "tieng_phap": 5.6, "ngu_van": 6.5, "hoa_hoc": 8.2, "sinh_hoc": 7.6}
|
||||
if len(got.Scores) != len(want) {
|
||||
t.Errorf("scores = %v, want %v", got.Scores, want)
|
||||
}
|
||||
for k, v := range want {
|
||||
if got.Scores[k] != v {
|
||||
t.Errorf("%s = %v, want %v", k, got.Scores[k], v)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestSubjectColumnsUnknownLanguageCode: an unrecognised code drops the score
|
||||
// rather than filing it under a guess.
|
||||
func TestSubjectColumnsUnknownLanguageCode(t *testing.T) {
|
||||
header := cells("sbd", "hoten", "ho", "ten", "phai", "ngaysinh", "socmnd",
|
||||
"cdiem1", "cdiem2", "cdiem3", "cdiem4", "cdiem5", "cdiem6", "cdiem7", "cdiem8", "ngoaingu")
|
||||
f := DetectFormat(header)
|
||||
row := cells("TCT000002", "B", "", "", "0", "", "", "5", "6", "", "", "", "", "", "", "N9")
|
||||
got := ProcessRow2016(row, f)
|
||||
if got == nil {
|
||||
t.Fatal("row rejected")
|
||||
}
|
||||
if len(got.Scores) != 1 || got.Scores["toan"] != 5 {
|
||||
t.Errorf("scores = %v, want only toan", got.Scores)
|
||||
}
|
||||
}
|
||||
|
||||
// TestSubjectColumnsNeedsThreeSubjects: a stray two-letter header elsewhere must
|
||||
// not turn an unrelated file into this layout.
|
||||
func TestSubjectColumnsNeedsThreeSubjects(t *testing.T) {
|
||||
f := DetectFormat(cells("SBD", "HOTEN", "DI", "SU"))
|
||||
if f.Kind == FormatSubjectColumns {
|
||||
t.Error("two subject columns must not be enough to select the layout")
|
||||
}
|
||||
}
|
||||
|
||||
// TestLeakedHeaderRowSkipped: a header repeated inside the data must not become
|
||||
// a student.
|
||||
func TestLeakedHeaderRowSkipped(t *testing.T) {
|
||||
|
||||
@@ -174,15 +174,7 @@ func Standard(cfg *config.DatasetConfig, inputDir, outputPath string) error {
|
||||
soBaoDanh = cellAt(row, cols.SoBaoDanh)
|
||||
}
|
||||
|
||||
switch transform.ValidateRow(hoTen, soBaoDanh, &cfg.Validation, stripBlank, allBlank) {
|
||||
case transform.SkipBlankRow:
|
||||
// Falls through to transform and insert. Unreachable here:
|
||||
// BlankRow requires stripBlank && allBlank, which returned
|
||||
// above. Kept so the two call sites with opposite outcomes stay
|
||||
// visibly distinct.
|
||||
case transform.SkipNone:
|
||||
// proceed
|
||||
default:
|
||||
if transform.ValidateRow(hoTen, soBaoDanh, &cfg.Validation, stripBlank, allBlank) != transform.SkipNone {
|
||||
fileSkipped++
|
||||
return
|
||||
}
|
||||
|
||||
@@ -5,9 +5,9 @@
|
||||
//
|
||||
// Column provenance:
|
||||
//
|
||||
// ten_cum_thi, gioi_tinh, tieng_duc, tieng_nhat -> 2016 only
|
||||
// khtn, khxh, gdcd, tieng_nga -> 2017 datasets only
|
||||
// everything else -> both
|
||||
// ten_cum_thi, gioi_tinh -> 2016 only
|
||||
// khtn, khxh, gdcd -> 2017 only
|
||||
// everything else -> both, all six languages included
|
||||
//
|
||||
// The DDL, the INSERT and the subject regexes belong here and nowhere else. One
|
||||
// copy per dataset is what let the 2016 and 2017 schemas drift apart; the
|
||||
@@ -109,8 +109,8 @@ VALUES
|
||||
// the source files — copy them, never retype them.
|
||||
//
|
||||
// Every pattern runs against every dataset. A subject absent from a given exam
|
||||
// year simply never matches and stays NULL: 2016 files contain no "KHTN:" or
|
||||
// "Tiếng Nga:" tokens, and 2017 files contain no "Tiếng Đức:" or "Tiếng Nhật:".
|
||||
// year simply never matches and stays NULL: 2016 files contain no "KHTN:",
|
||||
// "KHXH:" or "GDCD:" tokens, since those combined papers did not exist yet.
|
||||
var scorePatternSources = map[string]string{
|
||||
"toan": `Toán:\s*(\d+(?:\.\d+)?)`,
|
||||
"ngu_van": `Ngữ văn:\s*(\d+(?:\.\d+)?)`,
|
||||
|
||||
@@ -1,66 +0,0 @@
|
||||
package transform_test
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"os"
|
||||
"testing"
|
||||
|
||||
_ "github.com/tiennm99/thptqg/parser/internal/sqlitedb"
|
||||
"github.com/tiennm99/thptqg/parser/internal/transform"
|
||||
)
|
||||
|
||||
// TestToAsciiAgainstRustOutput cross-checks ToAscii against a reference database
|
||||
// on real data.
|
||||
//
|
||||
// Such a database is its own oracle: every row carries ho_ten alongside the
|
||||
// ho_ten_ascii derived from it, so the whole table is a name -> expected-slug
|
||||
// corpus far broader than the 20 hand-picked unit cases.
|
||||
//
|
||||
// GO_PARSER_RUST_DB=/tmp/rust-2016.db go test ./internal/transform/
|
||||
//
|
||||
// Skips when unset, so the default suite stays hermetic. The reference parser is
|
||||
// gone from the tree, so no new oracle database can be produced and this always
|
||||
// skips today.
|
||||
func TestToAsciiAgainstRustOutput(t *testing.T) {
|
||||
path := os.Getenv("GO_PARSER_RUST_DB")
|
||||
if path == "" {
|
||||
t.Skip("GO_PARSER_RUST_DB not set; skipping cross-check against Rust output")
|
||||
}
|
||||
if _, err := os.Stat(path); err != nil {
|
||||
t.Skipf("GO_PARSER_RUST_DB=%s not readable: %v", path, err)
|
||||
}
|
||||
|
||||
db, err := sql.Open("sqlite", "file:"+path+"?mode=ro")
|
||||
if err != nil {
|
||||
t.Fatalf("open %s: %v", path, err)
|
||||
}
|
||||
defer db.Close()
|
||||
|
||||
rows, err := db.Query("SELECT ho_ten, ho_ten_ascii FROM student")
|
||||
if err != nil {
|
||||
t.Fatalf("query: %v", err)
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
var checked, bad int
|
||||
for rows.Next() {
|
||||
var name, rustAscii string
|
||||
if err := rows.Scan(&name, &rustAscii); err != nil {
|
||||
t.Fatalf("scan: %v", err)
|
||||
}
|
||||
checked++
|
||||
if got := transform.ToAscii(name); got != rustAscii {
|
||||
bad++
|
||||
if bad <= 5 {
|
||||
t.Errorf("ToAscii(%q)\n rust = %q\n go = %q", name, rustAscii, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
t.Fatalf("iterate: %v", err)
|
||||
}
|
||||
if checked == 0 {
|
||||
t.Fatal("database contained no rows")
|
||||
}
|
||||
t.Logf("compared %d real names, %d mismatches", checked, bad)
|
||||
}
|
||||
@@ -68,9 +68,7 @@ type ParsedRow struct {
|
||||
//
|
||||
// The distinction is load-bearing for the printed counters: the two non-blank
|
||||
// reasons count as source rows while SkipBlankRow does not. That split lives in
|
||||
// the CALLER, not here — the build loop drops blank rows before the source-row
|
||||
// counter and, at the later switch, lets SkipBlankRow fall through to transform
|
||||
// and insert. Both call sites have to stay as they are.
|
||||
// the CALLER — the build loop drops blank rows before the source-row counter.
|
||||
type SkipReason int
|
||||
|
||||
const (
|
||||
@@ -188,14 +186,7 @@ func TransformRow(raw []reader.Cell, cfg *config.DatasetConfig) (*ParsedRow, err
|
||||
hoTen := get(cols.HoTen)
|
||||
ngaySinh := get(cols.NgaySinh)
|
||||
soBaoDanh := get(cols.SoBaoDanh)
|
||||
|
||||
// diem_thi is read WITHOUT trimming, unlike the three fields above. Harmless
|
||||
// because the score patterns are unanchored, but it is the shipped behaviour
|
||||
// — do not "tidy" it.
|
||||
diemThi := ""
|
||||
if cols.DiemThi >= 0 && cols.DiemThi < len(raw) {
|
||||
diemThi = raw[cols.DiemThi].Str
|
||||
}
|
||||
diemThi := get(cols.DiemThi)
|
||||
|
||||
var ngaySinhOpt *string
|
||||
if ngaySinh != "" {
|
||||
|
||||
@@ -242,10 +242,10 @@ func TestTransformRowShortRowYieldsEmptyFields(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestTransformRowDiemThiIsNotTrimmed pins an asymmetry that is easy to
|
||||
// "tidy away": ho_ten, ngay_sinh and so_bao_danh are trimmed, but diem_thi is
|
||||
// read raw.
|
||||
func TestTransformRowDiemThiIsNotTrimmed(t *testing.T) {
|
||||
// TestTransformRowTrimsEveryField: every column is read trimmed, diem_thi
|
||||
// included. Source cells routinely carry padding — 850k of the 861k 2017 rows
|
||||
// have whitespace around their score cell.
|
||||
func TestTransformRowTrimsEveryField(t *testing.T) {
|
||||
row := cells(" A ", " 01/01/2000 ", " 123 ", " Toán: 5 ")
|
||||
got, err := TransformRow(row, fixedColumnCfg())
|
||||
if err != nil {
|
||||
@@ -257,7 +257,6 @@ func TestTransformRowDiemThiIsNotTrimmed(t *testing.T) {
|
||||
if got.NgaySinh == nil || *got.NgaySinh != "01/01/2000" {
|
||||
t.Errorf("ngay_sinh = %v, want trimmed", got.NgaySinh)
|
||||
}
|
||||
// Untrimmed diem_thi still parses — the regexes are unanchored.
|
||||
if got.Scores["toan"] != 5 {
|
||||
t.Errorf("scores = %v", got.Scores)
|
||||
}
|
||||
|
||||
Reference in new issue
Block a user