refactor(parser): reimplement the parser in Go alongside the Rust crate

Adds go-parser/, a Go reimplementation of the xlsxread parser, verified
byte-for-byte against the Rust original before any cutover.

Reader fidelity is exact across all 299 input files: the canonical cell dump
of every sheet matches calamine's, locked in as a test against a committed
hash oracle. Reaching that required replacing extrame/xls, which corrupted
69% of cells and dropped a further 28% on the BIFF corpus, with pbnjay/grate;
correcting excelize's number-format application and trailing-cell trimming;
restoring carriage returns that XML line-ending normalisation strips from
2,233 ten_cum_thi values; and gating numeric re-rendering on cell type so
shared strings that merely look numeric keep their leading zeros.

The differential gate compares both parsers over all four datasets:
3,265,641 rows with identical full-table SHA-256, identical per-column
non-NULL counts, identical schema metadata and identical stdout.

Config moves from TOML to YAML for both parsers, so they keep reading the
same files and the gate stays meaningful. Verified by rebuilding 2016 and
2017-old2 with Rust under the new configs and matching the recorded counts.

build-db.js now refuses to publish a database whose row count does not match
the known figure, closing a path where an under-producing parser could ship a
truncated public dataset with green CI. The deploy workflow gains a
pull_request trigger and guards deploy to main, so branch verification can no
longer publish to production.
This commit is contained in:
tiennm99 committed 2026-08-13 20:27:44 +07:00
1 parent 8902232747
commit 0eb174721d
60 files changed
+6050 -189

No files matched your search

+34 -7
View File
@@ -3,6 +3,11 @@ name: Deploy to GitHub Pages
on:
push:
branches: [main]
# Pull requests run the build job only. Before this, the workflow triggered on
# main-push and workflow_dispatch alone, so "verify on a branch first" was not
# actually possible: pushing to a branch ran nothing, and dispatching from one
# published that branch straight to the live site.
pull_request:
workflow_dispatch:
permissions:
@@ -17,14 +22,18 @@ concurrency:
jobs:
build:
runs-on: ubuntu-latest
env:
# The parser is pure Go — grate, excelize, yaml.v3 and modernc.org/sqlite
# are all cgo-free — so no C toolchain is needed. Set explicitly rather
# than relying on the default.
CGO_ENABLED: '0'
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
- uses: actions/setup-go@v5
with:
workspaces: parser
go-version: '1.26'
cache-dependency-path: go-parser/go.sum
- uses: actions/setup-node@v4
with:
@@ -34,12 +43,26 @@ jobs:
- run: npm ci
# The reader-fidelity suite compares all 299 real input files against a
# committed hash oracle, so it is the regression guard for the whole
# reader. Runs before anything is built.
- name: Test parser
run: npm run test:go
# excelize carries an open advisory, and the 2017 refresh runbook feeds
# network-downloaded spreadsheets straight into the parser.
- name: Vulnerability scan
working-directory: go-parser
run: |
go install golang.org/x/vuln/cmd/govulncheck@latest
"$(go env GOPATH)/bin/govulncheck" ./...
# One parser binary builds every dataset; build-db.js reads the dataset
# list from src/datasets.js and gzips each database in place, leaving no
# uncompressed file behind.
# list from src/datasets.js, verifies each database against its known row
# count, then gzips it in place leaving no uncompressed file behind.
- name: Build databases
run: |
npm run build:rust
npm run build:go
npm run build:db
# One Vite build produces every page. scripts/assemble-site.js copies the
@@ -53,6 +76,10 @@ jobs:
path: _site
deploy:
# Guarded to main. Without this, a workflow_dispatch from any branch would
# publish that branch's output to the live site, and concurrency
# cancel-in-progress would kill an in-flight good deploy on the way.
if: github.ref == 'refs/heads/main'
needs: build
runs-on: ubuntu-latest
environment:
+4
View File
@@ -22,3 +22,7 @@ parser/target/
### Assembled Pages artifact ###
_site/
# Go parser build output and regenerable ground-truth dumps
go-parser/bin/
go-parser/testdata/dumps/
+4 -4
View File
@@ -25,7 +25,7 @@ index.html + src/ the frontend — one app serving all four datasets and the
data/<id>/ raw Excel files, one directory per dataset
parser/ the Rust parser
src/schema.rs canonical 22-column table: DDL, INSERT, subject regexes
configs/<id>.toml per-dataset parse rules only, no SQL
configs/<id>.yml per-dataset parse rules only, no SQL
scripts/ database build, crawler, parity verification
scripts/ site assembly
docs/ architecture, data pipeline, deployment
@@ -34,14 +34,14 @@ docs/ architecture, data pipeline, deployment
The dataset id is one identifier end to end:
```
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
data/2017-old/ → parser/configs/2017-old.yml → db/2017-old.db.gz → /thptqg/2017-old/
```
## Build
```bash
npm ci
npm run build:rust # compile the parser
npm run build:go # compile the parser
npm run build:db # build + gzip all four databases (add an id for just one)
npm run build:site # one Vite build, then assemble into _site/
npx serve _site
@@ -53,7 +53,7 @@ Pushing to `main` runs the same steps in
## Adding a dataset
1. Put the Excel files in `data/<id>/`
2. Add `parser/configs/<id>.toml` — sheet mode, column indices, validation
2. Add `parser/configs/<id>.yml` — sheet mode, column indices, validation
guards. No SQL; the schema is canonical.
3. Add an entry to `DATASETS` in `src/datasets.js`
+5 -5
View File
@@ -4,8 +4,8 @@ From raw Excel files to a compressed SQLite file the browser can load.
One Rust binary (`parser/`) builds every dataset. What differs per dataset is
parse rules only — sheet strategy, column layout, validation guards — declared
in `parser/configs/<id>.toml`. The table shape, the INSERT and the subject
regexes are canonical and live in `parser/src/schema.rs`.
in `parser/configs/<id>.yml`. The table shape, the INSERT and the subject
regexes are canonical and live in `go-parser/internal/schema/schema.go`.
## Sources
@@ -54,7 +54,7 @@ which is why only 2016 populates those columns.
## Score text parsing
`SCORE_PATTERNS` in `parser/src/schema.rs` defines one regex per subject, and
`SCORE_PATTERNS` in `go-parser/internal/schema/schema.go` defines one regex per subject, and
**all 16 run against every dataset**. A subject a given exam year did not offer
simply never matches and stays NULL.
@@ -110,7 +110,7 @@ silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what
| id | Source rows | Skipped | DB rows |
| --- | --- | --- | --- |
| `2016` | 877,464 | 3 duplicate SBDs collapsed | **877,461** |
| `2017` | 861,131 | 63 empty | **861,068** |
| `2017` | 861,068 | 0 | **861,068** |
| `2017-old` | 847,349 | 1 header leak | **847,348** |
| `2017-old2` | 679,764 | 0 | **679,764** |
@@ -135,7 +135,7 @@ regenerated — the two old crates no longer exist. Both scripts use the built-i
```bash
rm data/2017/*.xls
node parser/scripts/crawl-baotintuc.js
node parser/scripts/build-db.js 2017
node go-parser/scripts/build-db.js 2017
```
Then re-run the parity check above and confirm the row count still matches.
+6 -6
View File
@@ -7,8 +7,8 @@ One-time setup: **Settings → Pages → Source: GitHub Actions**.
## What the workflow does
1. Checkout, Rust toolchain, Node 24, `npm ci`
2. `npm run build:rust` — one parser binary
1. Checkout, Go toolchain, Node 24, `npm ci`
2. `npm run build:go` — one parser binary
3. `npm run build:db` — builds and gzips all four databases into
`.build/public/db/`
4. `npm run build:site` — one Vite build, then `scripts/assemble-site.js`
@@ -35,7 +35,7 @@ string intact.
```bash
npm ci
npm run build:rust
npm run build:go
npm run build:db # all four; pass an id to build just one
npm run build:site # vite build + assemble into _site/
npx serve _site
@@ -44,7 +44,7 @@ npx serve _site
To rebuild a single dataset:
```bash
node parser/scripts/build-db.js 2017-old
node go-parser/scripts/build-db.js 2017-old
```
## Base path
@@ -56,9 +56,9 @@ up as a blank page with 404s on `/assets/...`.
## Adding a dataset
1. Put the Excel files in `data/<id>/`
2. Add `parser/configs/<id>.toml` with the parse rules — sheet mode, column
2. Add `parser/configs/<id>.yml` with the parse rules — sheet mode, column
indices, SBD validation, header tokens, blank-row stripping. No SQL: the
schema is canonical and lives in `parser/src/schema.rs`
schema is canonical and lives in `go-parser/internal/schema/schema.go`
3. Add an entry to `DATASETS` in `src/datasets.js`
Nothing else. The build script, the site assembly and the router all read that
+3 -3
View File
@@ -31,11 +31,11 @@ data/<id>/*.xls(x)
One identifier ties the whole pipeline together:
```
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
data/2017-old/ → parser/configs/2017-old.yml → db/2017-old.db.gz → /thptqg/2017-old/
```
`src/datasets.js` declares the four ids once. The frontend, the database build
(`parser/scripts/build-db.js`) and the site assembly all import that list, so
(`go-parser/scripts/build-db.js`) and the site assembly all import that list, so
adding a dataset means adding one entry and one config file.
| id | Exam | Rows | Source |
@@ -47,7 +47,7 @@ adding a dataset means adding one entry and one config file.
## Canonical schema
Defined once in `parser/src/schema.rs` — DDL, INSERT, column order and the 16
Defined once in `go-parser/internal/schema/schema.go` — DDL, INSERT, column order and the 16
subject regexes. The four TOML configs carry no SQL at all, only per-dataset
parse rules. Config parsing uses `deny_unknown_fields`, so a leftover `[schema]`
block fails loudly instead of looking effective while `schema.rs` drives the
+1 -1
View File
@@ -28,7 +28,7 @@ export default defineConfig([
},
{
// Node-executed files (Vite config, parser tooling) run with Node globals.
files: ['vite.config.js', 'scripts/**/*.js', 'parser/scripts/**/*.js'],
files: ['vite.config.js', 'scripts/**/*.js', 'go-parser/scripts/**/*.js', 'parser/scripts/**/*.js'],
languageOptions: {
globals: { ...globals.node },
},
+88
View File
@@ -0,0 +1,88 @@
// Command dumpcells emits the canonical cell rendering of a spreadsheet, for
// comparison against the Rust/calamine ground truth produced by
// parser/examples/dump_cells.rs.
//
// The canonical stream carries geometry and rendered cell values only. The
// calamine Data variant is deliberately excluded: Data::Empty and
// Data::String("") both render "" and both count as blank everywhere
// downstream, so the distinction cannot affect the database.
//
// Usage: dumpcells <spreadsheet> [out-file]
package main
import (
"bufio"
"fmt"
"os"
"strings"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
)
// escape mirrors the Rust dumper so field separators can never break the format.
func escape(s string) string {
var b strings.Builder
b.Grow(len(s))
for _, ch := range s {
switch ch {
case '\\':
b.WriteString(`\\`)
case '\t':
b.WriteString(`\t`)
case '\n':
b.WriteString(`\n`)
case '\r':
b.WriteString(`\r`)
default:
b.WriteRune(ch)
}
}
return b.String()
}
func main() {
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, "usage: dumpcells <spreadsheet> [out-file]")
os.Exit(2)
}
path := os.Args[1]
out := os.Stdout
if len(os.Args) > 2 {
f, err := os.Create(os.Args[2])
if err != nil {
fmt.Fprintf(os.Stderr, "create: %v\n", err)
os.Exit(1)
}
defer f.Close()
out = f
}
w := bufio.NewWriterSize(out, 1<<20)
defer w.Flush()
wb, err := reader.Open(path)
if err != nil {
fmt.Fprintf(os.Stderr, "open: %v\n", err)
os.Exit(1)
}
defer wb.Close()
sheets := wb.Sheets()
fmt.Fprintf(w, "FILE\t%s\n", escape(path))
fmt.Fprintf(w, "SHEETCOUNT\t%d\n", len(sheets))
for _, sh := range sheets {
fmt.Fprintf(w, "SHEET\t%d\t%s\t%d\t%d\n", sh.Index, escape(sh.Name), sh.Height, sh.Width)
err := wb.EachRow(sh.Index, func(s reader.Sheet, rowIdx int, row []reader.Cell) error {
fmt.Fprintf(w, "ROW\t%d\t%d\t%d\n", s.Index, rowIdx, len(row))
for c, cell := range row {
fmt.Fprintf(w, "CELL\t%d\t%d\t%d\t%s\n", s.Index, rowIdx, c, escape(cell.Str))
}
return nil
})
if err != nil {
fmt.Fprintf(os.Stderr, "rows: %v\n", err)
os.Exit(1)
}
}
}
+115
View File
@@ -0,0 +1,115 @@
// Command xlsxread reads .xls/.xlsx files and builds SQLite databases for the
// thptqg datasets.
//
// The CLI contract is fixed by parser/scripts/build-db.js and must match the
// Rust binary exactly:
//
// xlsxread build --schema <config.yml> --input <dir> --output <db>
// xlsxread audit --schema <config.yml> --input <dir> --db <db>
//
// Implemented with the standard flag package rather than a CLI framework: two
// subcommands with three flags each do not justify a dependency.
package main
import (
"flag"
"fmt"
"os"
"github.com/tiennm99/thptqg/go-parser/internal/audit"
"github.com/tiennm99/thptqg/go-parser/internal/config"
"github.com/tiennm99/thptqg/go-parser/internal/ingest"
)
func usage() {
fmt.Fprint(os.Stderr, `xlsxread — read .xls/.xlsx files and build SQLite databases for thptqg datasets
Usage:
xlsxread build --schema <config.yml> --input <dir> --output <db>
xlsxread audit --schema <config.yml> --input <dir> --db <db>
`)
}
func main() {
if len(os.Args) < 2 {
usage()
os.Exit(2)
}
switch os.Args[1] {
case "build":
runBuild(os.Args[2:])
case "audit":
runAudit(os.Args[2:])
case "-h", "--help", "help":
usage()
default:
fmt.Fprintf(os.Stderr, "unknown subcommand %q\n\n", os.Args[1])
usage()
os.Exit(2)
}
}
func runBuild(args []string) {
fs := flag.NewFlagSet("build", flag.ExitOnError)
schemaPath := fs.String("schema", "", "path to the dataset YAML config file")
inputDir := fs.String("input", "", "directory containing the .xls / .xlsx source files")
outputPath := fs.String("output", "", "output SQLite database path")
fs.Parse(args)
if *schemaPath == "" || *inputDir == "" || *outputPath == "" {
fmt.Fprintln(os.Stderr, "build requires --schema, --input and --output")
os.Exit(2)
}
cfg, err := config.Load(*schemaPath)
if err != nil {
fatalf("Failed to load config: %v", err)
}
// The 2016 dataset selects its column layout per file at runtime; every other
// dataset uses the fixed columns: mapping (main.rs:63-67).
if cfg.FormatDetection != nil && *cfg.FormatDetection == "thptqg2016" {
if err := ingest.Detect2016(cfg, *inputDir, *outputPath); err != nil {
fatalf("%v", err)
}
return
}
if err := ingest.Standard(cfg, *inputDir, *outputPath); err != nil {
fatalf("%v", err)
}
}
func runAudit(args []string) {
fs := flag.NewFlagSet("audit", flag.ExitOnError)
schemaPath := fs.String("schema", "", "path to the dataset YAML config file")
inputDir := fs.String("input", "", "directory containing the .xlsx source files")
dbPath := fs.String("db", "", "SQLite database to compare against")
fs.Parse(args)
if *schemaPath == "" || *inputDir == "" || *dbPath == "" {
fmt.Fprintln(os.Stderr, "audit requires --schema, --input and --db")
os.Exit(2)
}
cfg, err := config.Load(*schemaPath)
if err != nil {
fatalf("Failed to load config: %v", err)
}
res, err := audit.Run(*inputDir, *dbPath, cfg)
if err != nil {
fatalf("%v", err)
}
audit.PrintReport(res)
// A mismatch is a non-zero exit even though the audit itself succeeded
// (main.rs:46-48) — CI treats it as a failure signal.
if !res.Matched {
os.Exit(1)
}
}
func fatalf(format string, a ...any) {
fmt.Fprintf(os.Stderr, "Error: "+format+"\n", a...)
os.Exit(1)
}
+30
View File
@@ -0,0 +1,30 @@
module github.com/tiennm99/thptqg/go-parser
go 1.26.5
require (
github.com/pbnjay/grate v0.0.0-20231006022435-3f8e65d74a14
github.com/xuri/excelize/v2 v2.11.0
golang.org/x/text v0.41.0
gopkg.in/yaml.v3 v3.0.1
modernc.org/sqlite v1.56.0
)
require (
github.com/dustin/go-humanize v1.0.1 // indirect
github.com/google/uuid v1.6.0 // indirect
github.com/mattn/go-isatty v0.0.24 // indirect
github.com/ncruces/go-strftime v1.0.0 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/richardlehane/mscfb v1.0.7 // indirect
github.com/richardlehane/msoleps v1.0.6 // indirect
github.com/tiendc/go-deepcopy v1.7.2 // indirect
github.com/xuri/efp v0.0.1 // indirect
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9 // indirect
golang.org/x/crypto v0.53.0 // indirect
golang.org/x/net v0.56.0 // indirect
golang.org/x/sys v0.47.0 // indirect
modernc.org/libc v1.74.4 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
)
+82
View File
@@ -0,0 +1,82 @@
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3 h1:LMLX+LgTNWpfvCBdFebv6EsYotImrt/Ppc5cXIriCSo=
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3/go.mod h1:jl5iWTm0/hd5PjEYEOuwAJ57L/CibdZfrqZ5XA5GrCk=
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
github.com/mattn/go-isatty v0.0.24/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
github.com/ncruces/go-strftime v1.0.0 h1:HMFp8mLCTPp341M/ZnA4qaf7ZlsbTc+miZjCLOFAw7w=
github.com/ncruces/go-strftime v1.0.0/go.mod h1:Fwc5htZGVVkseilnfgOVb9mKy6w1naJmn9CehxcKcls=
github.com/pbnjay/grate v0.0.0-20231006022435-3f8e65d74a14 h1:ZfXdW7GIVZT3Z9oejLJ+GHrrQv/ezU2Bwqn0BF37s4g=
github.com/pbnjay/grate v0.0.0-20231006022435-3f8e65d74a14/go.mod h1:VaZEKQrYbYr2untVA/EFNdC6hM7GyARRNM+k4+5CmA0=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo=
github.com/richardlehane/mscfb v1.0.7 h1:oeoiM0WE79vHwE8RpIYYvIAc8ajTH2mb6UZm55/+EB0=
github.com/richardlehane/mscfb v1.0.7/go.mod h1:pe0+IUIc0AHh0+teNzBlJCtSyZdFOGgV4ZK9bsoV+Jo=
github.com/richardlehane/msoleps v1.0.6 h1:9BvkpjvD+iUBalUY4esMwv6uBkfOip/Lzvd93jvR9gg=
github.com/richardlehane/msoleps v1.0.6/go.mod h1:BWev5JBpU9Ko2WAgmZEuiz4/u3ZYTKbjLycmwiWUfWg=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/tiendc/go-deepcopy v1.7.2 h1:Ut2yYR7W9tWjTQitganoIue4UGxZwCcJy3orjrrIj44=
github.com/tiendc/go-deepcopy v1.7.2/go.mod h1:4bKjNC2r7boYOkD2IOuZpYjmlDdzjbpTRyCx+goBCJQ=
github.com/xuri/efp v0.0.1 h1:fws5Rv3myXyYni8uwj2qKjVaRP30PdjeYe2Y6FDsCL8=
github.com/xuri/efp v0.0.1/go.mod h1:ybY/Jr0T0GTCnYjKqmdwxyxn2BQf2RcQIIvex5QldPI=
github.com/xuri/excelize/v2 v2.11.0 h1:HxaEFl6sRN2+8J5a8HaKq+0M4FsjBGMnWWtjOCPSG88=
github.com/xuri/excelize/v2 v2.11.0/go.mod h1:jxFLbzaIwGQ5ufFNvYfUOHqXhfPaNmP14KWfmNz2Uak=
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9 h1:+C0TIdyyYmzadGaL/HBLbf3WdLgC29pgyhTjAT/0nuE=
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9/go.mod h1:WwHg+CVyzlv/TX9xqBFXEZAuxOPxn2k1GNHwG41IIUQ=
golang.org/x/crypto v0.53.0 h1:QZ4Muo8THX6CizN2vPPd5fBGHyogrdK9fG4wLPFUsto=
golang.org/x/crypto v0.53.0/go.mod h1:DNLU434OwVakk9PzuwV8w62mAJpRJL3vsgcfp4Qnsio=
golang.org/x/image v0.38.0 h1:5l+q+Y9JDC7mBOMjo4/aPhMDcxEptsX+Tt3GgRQRPuE=
golang.org/x/image v0.38.0/go.mod h1:/3f6vaXC+6CEanU4KJxbcUZyEePbyKbaLoDOe4ehFYY=
golang.org/x/mod v0.38.0 h1:MECBjubtXD7yj4HrhIUcywNaGeNVUdfVnxmPajOk4yk=
golang.org/x/mod v0.38.0/go.mod h1:V6Xz0pq8TQ3dGqVQ1FVHuelZpAL0uNhSkk9ogYP3c40=
golang.org/x/net v0.56.0 h1:Rw8j/hFzGvJUZwNBXnAtf5sVDVt+65SK2C7IxCxZt5o=
golang.org/x/net v0.56.0/go.mod h1:D3Ku6r+V6JROoZK144D2XfMHFcMq/0zSfLelVTCFKec=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE=
golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI=
modernc.org/cc/v4 v4.29.1/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
modernc.org/ccgo/v4 v4.34.6 h1:sBgfIwyN0TQ9C5hwIeuqyeAKyMWnbvj2fvpF4L11uzU=
modernc.org/ccgo/v4 v4.34.6/go.mod h1:SZ8YcN9NG7XVsQYdm6jYBvi8PQP1qi+kqB6OhjqI3Fk=
modernc.org/fileutil v1.4.0 h1:j6ZzNTftVS054gi281TyLjHPp6CPHr2KCxEXjEbD6SM=
modernc.org/fileutil v1.4.0/go.mod h1:EqdKFDxiByqxLk8ozOxObDSfcVOv/54xDs/DUHdvCUU=
modernc.org/gc/v2 v2.6.5 h1:nyqdV8q46KvTpZlsw66kWqwXRHdjIlJOhG6kxiV/9xI=
modernc.org/gc/v2 v2.6.5/go.mod h1:YgIahr1ypgfe7chRuJi2gD7DBQiKSLMPgBQe9oIiito=
modernc.org/gc/v3 v3.1.4 h1:2g65LGVSmFQrXeITAw97x7hCRvZFcyE1uDP+7Vng7JI=
modernc.org/gc/v3 v3.1.4/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
modernc.org/goabi0 v0.2.0 h1:HvEowk7LxcPd0eq6mVOAEMai46V+i7Jrj13t4AzuNks=
modernc.org/goabi0 v0.2.0/go.mod h1:CEFRnnJhKvWT1c1JTI3Avm+tgOWbkOu5oPA8eH8LnMI=
modernc.org/libc v1.74.4 h1:fX1Omw4o2/1C2iRkkIsrQTasJQldLhRmuPreXLoWs9k=
modernc.org/libc v1.74.4/go.mod h1:eeQAS9W3sZeKYMFubydxJpII9ybHWshk+7or7bLG9co=
modernc.org/mathutil v1.7.1 h1:GCZVGXdaN8gTqB1Mf/usp1Y/hSqgI2vAGGP4jZMCxOU=
modernc.org/mathutil v1.7.1/go.mod h1:4p5IwJITfppl0G4sUEDtCr4DthTaT47/N3aT6MhfgJg=
modernc.org/memory v1.11.0 h1:o4QC8aMQzmcwCK3t3Ux/ZHmwFPzE6hf2Y5LbkRs+hbI=
modernc.org/memory v1.11.0/go.mod h1:/JP4VbVC+K5sU2wZi9bHoq2MAkCnrt2r98UGeSK7Mjw=
modernc.org/opt v0.2.0 h1:tGyef5ApycA7FSEOMraay9SaTk5zmbx7Tu+cJs4QKZg=
modernc.org/opt v0.2.0/go.mod h1:03fq9lsNfvkYSfxrfUhZCWPk1lm4cq4N+Bh//bEtgns=
modernc.org/sortutil v1.2.1 h1:+xyoGf15mM3NMlPDnFqrteY07klSFxLElE2PVuWIJ7w=
modernc.org/sortutil v1.2.1/go.mod h1:7ZI3a3REbai7gzCLcotuw9AC4VZVpYMjDzETGsSMqJE=
modernc.org/sqlite v1.56.0 h1:/D8e2RfFqoy/Zc6PuC76U28zFwmI/sYx1Kjm4yEn9e0=
modernc.org/sqlite v1.56.0/go.mod h1:yCJ2cmAaIkHQ25oXWrF8H4O1lIfPYPR26yCEDj2P3pQ=
modernc.org/strutil v1.2.1 h1:UneZBkQA+DX2Rp35KcM69cSsNES9ly8mQWD71HKlOA0=
modernc.org/strutil v1.2.1/go.mod h1:EHkiggD70koQxjVdSBM3JKM7k6L0FbGE5eymy9i3B9A=
modernc.org/token v1.1.0 h1:Xl7Ap9dKaEs5kLoOQeQmPWevfnk/DM5qcLcYlA8ys6Y=
modernc.org/token v1.1.0/go.mod h1:UGzOrNV1mAFSEB63lOFHIpNRUVMvYTc6yu1SMY/XTDM=
+148
View File
@@ -0,0 +1,148 @@
// Package audit compares distinct SBDs in the source spreadsheets against the
// row count in a built database — a port of parser/src/audit.rs.
//
// Two deliberate divergences from the build path are preserved, both inherited
// from audit-row-counts.js:
//
// - only .xlsx files are considered (audit.rs:53-59)
// - only sheet 0 is read, regardless of sheet_mode (audit.rs:81-84)
//
// These are intentional, not oversights. Do not "fix" them.
package audit
import (
"database/sql"
"fmt"
"os"
"path/filepath"
"sort"
"strings"
"github.com/tiennm99/thptqg/go-parser/internal/config"
"github.com/tiennm99/thptqg/go-parser/internal/ingest"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
"github.com/tiennm99/thptqg/go-parser/internal/sqlitedb"
)
// Result carries the audit counters.
type Result struct {
TotalDataRows uint64
BothEmpty uint64
EmptyName uint64
EmptySbd uint64
DistinctSbds int
DBCount int64
Matched bool
}
// Run collects distinct SBDs from the .xlsx files in inputDir and compares the
// total against the student row count in dbPath.
func Run(inputDir, dbPath string, cfg *config.DatasetConfig) (*Result, error) {
entries, err := os.ReadDir(inputDir)
if err != nil {
return nil, fmt.Errorf("read input dir %s: %w", inputDir, err)
}
var files []string
for _, e := range entries {
if e.IsDir() {
continue
}
if strings.EqualFold(filepath.Ext(e.Name()), ".xlsx") {
files = append(files, filepath.Join(inputDir, e.Name()))
}
}
sort.Strings(files)
res := &Result{}
seen := make(map[string]struct{})
// The audit uses positional defaults for format-detection configs, mirroring
// audit-row-counts.js's fixed column assumption (audit.rs:104-111).
hoTenCol, sbdCol := 1, 0
if cfg.Columns != nil {
hoTenCol, sbdCol = cfg.Columns.HoTen, cfg.Columns.SoBaoDanh
}
for _, file := range files {
wb, err := reader.Open(file)
if err != nil {
return nil, err
}
sheets := wb.Sheets()
if len(sheets) == 0 {
wb.Close()
continue
}
firstRow := true
err = wb.EachRow(sheets[0].Index, func(_ reader.Sheet, _ int, row []reader.Cell) error {
if firstRow {
firstRow = false
if ingest.IsHeaderRow(row, cfg.Header.Tokens) {
return nil
}
}
res.TotalDataRows++
hoTen := cellAt(row, hoTenCol)
sbd := cellAt(row, sbdCol)
if hoTen == "" && sbd == "" {
res.BothEmpty++
return nil
}
if hoTen == "" {
res.EmptyName++
}
if sbd == "" {
res.EmptySbd++
}
if sbd != "" {
seen[sbd] = struct{}{}
}
return nil
})
wb.Close()
if err != nil {
return nil, err
}
}
// Opened read-only: the audit must never mutate the database it inspects
// (audit.rs:139).
db, err := sql.Open(sqlitedb.DriverName, "file:"+dbPath+"?mode=ro")
if err != nil {
return nil, fmt.Errorf("open db %s: %w", dbPath, err)
}
defer db.Close()
if err := db.QueryRow("SELECT COUNT(*) FROM student").Scan(&res.DBCount); err != nil {
return nil, fmt.Errorf("count rows: %w", err)
}
res.DistinctSbds = len(seen)
res.Matched = int64(res.DistinctSbds) == res.DBCount
return res, nil
}
// PrintReport mirrors audit-row-counts.js:54-62 exactly.
func PrintReport(r *Result) {
fmt.Println("=== Source vs DB ===")
fmt.Printf("Source: total data rows across all files: %d\n", r.TotalDataRows)
fmt.Printf("Source: rows with empty name AND sbd (skipped): %d\n", r.BothEmpty)
fmt.Printf("Source: rows with missing name only: %d\n", r.EmptyName)
fmt.Printf("Source: rows with missing sbd only: %d\n", r.EmptySbd)
fmt.Printf("Source: distinct SBDs: %d\n", r.DistinctSbds)
fmt.Printf("DB: row count: %d\n", r.DBCount)
if r.Matched {
fmt.Println("Match: YES — all unique SBDs accounted for")
} else {
fmt.Printf("Match: NO — gap of %d\n", int64(r.DistinctSbds)-r.DBCount)
}
}
func cellAt(row []reader.Cell, idx int) string {
if idx < 0 || idx >= len(row) {
return ""
}
return strings.TrimSpace(row[idx].Str)
}
+113
View File
@@ -0,0 +1,113 @@
// Package config loads per-dataset parse rules — a port of parser/src/config.rs.
//
// The config deliberately carries no SQL. The table shape, the INSERT and the
// subject regexes are identical for every dataset and live in internal/schema;
// keeping them here meant four copies of the same DDL, which is how the 2016 and
// 2017 schemas drifted apart.
package config
import (
"bytes"
"fmt"
"os"
"gopkg.in/yaml.v3"
)
// SheetMode selects which sheets of a workbook are read.
type SheetMode string
const (
// SheetModeAll iterates every sheet, which is what recovers the Hanoi and
// HCM rows that overflow past Excel's 65,536-row cap into a second sheet.
SheetModeAll SheetMode = "all"
// SheetModeFirst reads sheet 0 only.
SheetModeFirst SheetMode = "first"
)
// valid reports whether m is one of the two values the Rust enum accepts.
//
// Checked after decoding rather than via a custom unmarshaler: the decoder
// assigns named string types directly, so a typo would otherwise decode
// silently and be read as "not all" downstream.
func (m SheetMode) valid() bool {
return m == SheetModeAll || m == SheetModeFirst
}
// DatasetConfig is the per-dataset parse rule set.
type DatasetConfig struct {
Reader ReaderCfg `yaml:"reader"`
// Columns holds fixed column indices. Nil when FormatDetection handles
// per-file mapping — a pointer rather than a value because a zero ColumnMap
// would silently mean "every column is index 0".
Columns *ColumnMap `yaml:"columns"`
Validation ValidationCfg `yaml:"validation"`
Header HeaderCfg `yaml:"header"`
// FormatDetection, when set to "thptqg2016", enables per-file format
// auto-detection: each file's header row is inspected at runtime to choose
// the column layout (separate-scores / mapped / default-positional).
FormatDetection *string `yaml:"format_detection"`
}
type ReaderCfg struct {
SheetMode SheetMode `yaml:"sheet_mode"`
// StripBlankRows skips rows where every cell is empty before counting them
// as source rows (a 2017-old2 quirk).
StripBlankRows bool `yaml:"strip_blank_rows"`
}
// ColumnMap holds zero-indexed column positions in the source row. Used by the
// 2017 configs; 2016 uses runtime format detection instead.
type ColumnMap struct {
HoTen int `yaml:"ho_ten"`
NgaySinh int `yaml:"ngay_sinh"`
SoBaoDanh int `yaml:"so_bao_danh"`
DiemThi int `yaml:"diem_thi"`
}
type ValidationCfg struct {
// RequireNumericSbd mirrors build-database-old.js / -old2.js, which require
// so_bao_danh to match ^\d+$.
RequireNumericSbd bool `yaml:"require_numeric_sbd"`
RequireNonemptyName bool `yaml:"require_nonempty_name"`
RequireNonemptySbd bool `yaml:"require_nonempty_sbd"`
}
type HeaderCfg struct {
// Tokens are matched against the uppercased first cell to detect a header row.
Tokens []string `yaml:"tokens"`
}
// Parse decodes a config, rejecting any key the struct does not declare.
//
// Strictness is load-bearing and has its own test. Rust gets it from serde's
// deny_unknown_fields; yaml.v3 needs KnownFields(true) explicitly, and Go YAML
// decoders ignore unknown keys by default. Without it a leftover `schema:`
// mapping would look effective while internal/schema actually drove the build —
// the drift that produced two divergent schemas before the unification.
func Parse(src []byte) (*DatasetConfig, error) {
var cfg DatasetConfig
dec := yaml.NewDecoder(bytes.NewReader(src))
dec.KnownFields(true)
if err := dec.Decode(&cfg); err != nil {
return nil, fmt.Errorf("decode config: %w", err)
}
if !cfg.Reader.SheetMode.valid() {
return nil, fmt.Errorf("invalid sheet_mode %q (want %q or %q)",
cfg.Reader.SheetMode, SheetModeAll, SheetModeFirst)
}
return &cfg, nil
}
// Load reads and parses a config file.
func Load(path string) (*DatasetConfig, error) {
src, err := os.ReadFile(path)
if err != nil {
return nil, fmt.Errorf("read config %s: %w", path, err)
}
cfg, err := Parse(src)
if err != nil {
return nil, fmt.Errorf("%s: %w", path, err)
}
return cfg, nil
}
+227
View File
@@ -0,0 +1,227 @@
package config
import (
"os"
"path/filepath"
"testing"
)
// sampleYAML mirrors SAMPLE_YAML in parser/src/config.rs.
const sampleYAML = `
reader:
sheet_mode: all
strip_blank_rows: false
columns:
ho_ten: 0
ngay_sinh: 1
so_bao_danh: 2
diem_thi: 3
validation:
require_numeric_sbd: false
require_nonempty_name: true
require_nonempty_sbd: true
header:
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
`
// TestConfigRoundTrip ports config_round_trip (config.rs:115).
func TestConfigRoundTrip(t *testing.T) {
cfg, err := Parse([]byte(sampleYAML))
if err != nil {
t.Fatalf("parse failed: %v", err)
}
if cfg.Reader.SheetMode != SheetModeAll {
t.Errorf("sheet_mode = %q, want all", cfg.Reader.SheetMode)
}
if cfg.Reader.StripBlankRows {
t.Error("strip_blank_rows should be false")
}
if cfg.Columns == nil {
t.Fatal("columns should be present")
}
if cfg.Columns.HoTen != 0 {
t.Errorf("ho_ten = %d, want 0", cfg.Columns.HoTen)
}
if cfg.Columns.DiemThi != 3 {
t.Errorf("diem_thi = %d, want 3", cfg.Columns.DiemThi)
}
if cfg.Validation.RequireNumericSbd {
t.Error("require_numeric_sbd should be false")
}
if !cfg.Validation.RequireNonemptyName {
t.Error("require_nonempty_name should be true")
}
if len(cfg.Header.Tokens) != 3 {
t.Errorf("tokens = %d, want 3", len(cfg.Header.Tokens))
}
if cfg.FormatDetection != nil {
t.Errorf("format_detection = %v, want nil", *cfg.FormatDetection)
}
}
// TestConfigRejectsLeftoverSQLSections ports config_rejects_leftover_sql_sections
// (config.rs:132). This is the load-bearing one: Rust gets the behaviour free
// from serde's deny_unknown_fields, whereas Go YAML decoders ignore unknown keys
// unless KnownFields(true) is set. Without it a stale schema: mapping would look effective while
// internal/schema silently drove the build — exactly the drift that produced two
// divergent schemas before the unification.
func TestConfigRejectsLeftoverSQLSections(t *testing.T) {
withDDL := sampleYAML + "\nschema:\n ddl: \"CREATE TABLE student (so_bao_danh TEXT);\"\n"
if _, err := Parse([]byte(withDDL)); err == nil {
t.Fatal("config with a leftover schema: mapping must be rejected")
}
}
// TestConfigRejectsUnknownScalarKey guards the same property for a stray scalar,
// not just a stray mapping.
func TestConfigRejectsUnknownScalarKey(t *testing.T) {
withKey := sampleYAML + "\nunexpected_key: 1\n"
if _, err := Parse([]byte(withKey)); err == nil {
t.Fatal("config with an unknown scalar key must be rejected")
}
}
// TestConfigFirstSheetMode ports config_first_sheet_mode (config.rs:140).
func TestConfigFirstSheetMode(t *testing.T) {
src := []byte(replaceAll(sampleYAML, "sheet_mode: all", "sheet_mode: first"))
cfg, err := Parse(src)
if err != nil {
t.Fatalf("parse failed: %v", err)
}
if cfg.Reader.SheetMode != SheetModeFirst {
t.Errorf("sheet_mode = %q, want first", cfg.Reader.SheetMode)
}
}
// TestConfigRejectsUnknownSheetMode: the Rust enum accepts only "all"/"first",
// so anything else must fail rather than defaulting.
func TestConfigRejectsUnknownSheetMode(t *testing.T) {
src := []byte(replaceAll(sampleYAML, "sheet_mode: all", "sheet_mode: second"))
if _, err := Parse(src); err == nil {
t.Fatal("unknown sheet_mode must be rejected")
}
}
// TestConfigFormatDetectionField ports config_format_detection_field
// (config.rs:147): a 2016-style config has no columns: mapping at all.
func TestConfigFormatDetectionField(t *testing.T) {
const src = `
format_detection: thptqg2016
reader:
sheet_mode: all
strip_blank_rows: false
validation:
require_numeric_sbd: false
require_nonempty_name: true
require_nonempty_sbd: true
header:
tokens: ["SBD", "SOBAODANH", "STT"]
`
cfg, err := Parse([]byte(src))
if err != nil {
t.Fatalf("parse failed: %v", err)
}
if cfg.FormatDetection == nil || *cfg.FormatDetection != "thptqg2016" {
t.Errorf("format_detection = %v, want thptqg2016", cfg.FormatDetection)
}
if cfg.Columns != nil {
t.Error("columns must be nil when format_detection drives the layout")
}
}
// TestLoadRealConfigs loads the four shipped configs — the Go binary reads the
// same files as Rust, never a fork — and asserts the per-dataset differences
// recorded during scouting.
func TestLoadRealConfigs(t *testing.T) {
root := repoRoot(t)
want := map[string]struct {
sheetMode SheetMode
stripBlank bool
numericSbd bool
hasColumns bool
formatDet string
tokenCount int
}{
"2016": {SheetModeAll, false, false, false, "thptqg2016", 6},
"2017": {SheetModeAll, false, false, true, "", 3},
"2017-old": {SheetModeFirst, false, true, true, "", 3},
"2017-old2": {SheetModeAll, true, true, true, "", 3},
}
for id, w := range want {
t.Run(id, func(t *testing.T) {
cfg, err := Load(filepath.Join(root, "parser", "configs", id+".yml"))
if err != nil {
t.Fatalf("load: %v", err)
}
if cfg.Reader.SheetMode != w.sheetMode {
t.Errorf("sheet_mode = %q, want %q", cfg.Reader.SheetMode, w.sheetMode)
}
if cfg.Reader.StripBlankRows != w.stripBlank {
t.Errorf("strip_blank_rows = %v, want %v", cfg.Reader.StripBlankRows, w.stripBlank)
}
if cfg.Validation.RequireNumericSbd != w.numericSbd {
t.Errorf("require_numeric_sbd = %v, want %v", cfg.Validation.RequireNumericSbd, w.numericSbd)
}
if (cfg.Columns != nil) != w.hasColumns {
t.Errorf("columns present = %v, want %v", cfg.Columns != nil, w.hasColumns)
}
got := ""
if cfg.FormatDetection != nil {
got = *cfg.FormatDetection
}
if got != w.formatDet {
t.Errorf("format_detection = %q, want %q", got, w.formatDet)
}
if len(cfg.Header.Tokens) != w.tokenCount {
t.Errorf("tokens = %d, want %d", len(cfg.Header.Tokens), w.tokenCount)
}
// Every dataset requires a non-empty name and SBD.
if !cfg.Validation.RequireNonemptyName || !cfg.Validation.RequireNonemptySbd {
t.Error("both non-empty validations should be true for every dataset")
}
})
}
}
func replaceAll(s, old, new string) string {
out := ""
for {
i := indexOf(s, old)
if i < 0 {
return out + s
}
out += s[:i] + new
s = s[i+len(old):]
}
}
func indexOf(s, sub string) int {
for i := 0; i+len(sub) <= len(s); i++ {
if s[i:i+len(sub)] == sub {
return i
}
}
return -1
}
func repoRoot(t *testing.T) string {
t.Helper()
dir, err := os.Getwd()
if err != nil {
t.Fatalf("getwd: %v", err)
}
for i := 0; i < 6; i++ {
if _, err := os.Stat(filepath.Join(dir, "parser", "configs")); err == nil {
return dir
}
dir = filepath.Dir(dir)
}
t.Fatal("could not locate repo root")
return ""
}
+382
View File
@@ -0,0 +1,382 @@
package ingest
import (
"fmt"
"os"
"path/filepath"
"strconv"
"strings"
"github.com/tiennm99/thptqg/go-parser/internal/config"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
"github.com/tiennm99/thptqg/go-parser/internal/transform"
"github.com/tiennm99/thptqg/go-parser/internal/writer"
)
// The 2016 dataset's 119 files were produced by inconsistent tooling and use
// three different column layouts, chosen per sheet at runtime. This is a port of
// parser/src/format_detect_2016.rs — institutional knowledge encoded as
// literals, with no abstraction to derive it from, so everything here is copied
// verbatim rather than rationalised.
// KnownHeaders are the upper-cased first-cell values that identify a header row
// (format_detect_2016.rs:36-54, mirroring KNOWN_HEADERS in build-database.js).
//
// NOTE "SINH " carries a TRAILING SPACE, exactly as in the Rust source. Trimming
// it would change which rows are recognised as headers.
var KnownHeaders = []string{
"SOBAODANH",
"SBD",
"HO_TEN",
"HOTEN",
"HỌ TÊN",
"NGAY_SINH",
"TEN_CUMTHI",
"GIOI_TINH",
"DIEM_THI",
"STT",
"TOAN",
"VAN",
"LY",
"HOA",
"SINH ",
"SU",
"DIA",
}
func isKnownHeader(s string) bool {
for _, h := range KnownHeaders {
if h == s {
return true
}
}
return false
}
// IsHeaderRow2016 reports whether row[0] is a known header token
// (format_detect_2016.rs:58-64).
//
// The guard here is len < 2, not the len < 3 used by the 2017 header check.
func IsHeaderRow2016(row []reader.Cell) bool {
if len(row) < 2 {
return false
}
return isKnownHeader(strings.ToUpper(strings.TrimSpace(row[0].Str)))
}
// FormatKind is one of the three 2016 layouts.
type FormatKind int
const (
// FormatSeparateScores has one column per subject rather than a free-text
// DIEM_THI cell — the dhhanghai-style files.
FormatSeparateScores FormatKind = iota
// FormatMapped resolves column indices from the header by name.
FormatMapped
// FormatDefault is the positional 6-column layout used when no header is
// recognised. It is FormatMapped with fixed indices, not a separate path.
FormatDefault
)
// Format is a detected layout plus, for the mapped case, its column indices.
type Format struct {
Kind FormatKind
Sbd int
HoTen int
NgaySinh *int
TenCumThi *int
GioiTinh *int
DiemThi int
}
// defaultFormat is FormatMapped with the fixed positional indices
// (format_detect_2016.rs:295-306).
func defaultFormat() Format {
two, three, four := 2, 3, 4
return Format{
Kind: FormatDefault, Sbd: 0, HoTen: 1,
NgaySinh: &two, TenCumThi: &three, GioiTinh: &four, DiemThi: 5,
}
}
// DetectFormat inspects a header row and decides which layout applies
// (format_detect_2016.rs:95-145).
func DetectFormat(headerRow []reader.Cell) Format {
cols := make([]string, len(headerRow))
for i, c := range headerRow {
cols[i] = strings.ToUpper(strings.TrimSpace(c.Str))
}
// Format 1: SBD in col 0 AND TOAN in col 2.
if len(cols) > 2 && cols[0] == "SBD" && cols[2] == "TOAN" {
return Format{Kind: FormatSeparateScores}
}
// Format 2: resolve indices by header name, order-independent.
var sbdIdx, hoTenIdx, ngaySinhIdx, tenCumThiIdx, gioiTinhIdx, diemThiIdx *int
for i := range cols {
idx := i
switch cols[i] {
case "SOBAODANH", "SBD":
sbdIdx = &idx
case "HO_TEN", "HOTEN", "HỌ TÊN":
hoTenIdx = &idx
case "NGAY_SINH":
ngaySinhIdx = &idx
case "TEN_CUMTHI":
tenCumThiIdx = &idx
case "GIOI_TINH":
gioiTinhIdx = &idx
case "DIEM_THI":
diemThiIdx = &idx
}
}
if sbdIdx != nil && diemThiIdx != nil {
hoTen := 1 // fallback: col 1, present in all known files
if hoTenIdx != nil {
hoTen = *hoTenIdx
}
return Format{
Kind: FormatMapped, Sbd: *sbdIdx, HoTen: hoTen,
NgaySinh: ngaySinhIdx, TenCumThi: tenCumThiIdx, GioiTinh: gioiTinhIdx,
DiemThi: *diemThiIdx,
}
}
// Unrecognised header, or none at all.
return defaultFormat()
}
// parseFloatCell parses a per-subject score cell.
//
// A parsed 0.0 becomes "no score" — this replicates JavaScript's
// `parseFloat(row[N]) || null`, where 0 is falsy (format_detect_2016.rs:165).
// It means a genuine zero is indistinguishable from a blank. Not obviously
// correct, but it is the shipped behaviour and the published data depends on it.
func parseFloatCell(row []reader.Cell, idx int) (float64, bool) {
s := cellAt(row, idx)
if s == "" {
return 0, false
}
v, err := strconv.ParseFloat(s, 64)
if err != nil || v == 0.0 {
return 0, false
}
return v, true
}
// processSeparateScoresRow handles the fixed 12-column layout
// (format_detect_2016.rs:176-216):
//
// 0=SBD 1=HOTEN 2=TOAN 3=VAN 4=LY 5=HOA 6=SINH 7=SU 8=DIA
// 9=NGOAINGUTN 10=NGOAINGUTL 11=NGOAINGU(total -> tieng_anh)
//
// tieng_phap / tieng_duc / tieng_nhat / tieng_trung are structurally unreachable
// in this format, and ngay_sinh / ten_cum_thi / gioi_tinh are always nil.
// Scores are read as floats directly; this layout has no free-text score cell,
// so the subject regexes never run.
func processSeparateScoresRow(row []reader.Cell) *transform.ParsedRow {
sbd := cellAt(row, 0)
hoTen := cellAt(row, 1)
if sbd == "" || hoTen == "" {
return nil
}
scores := make(map[string]float64)
for field, idx := range map[string]int{
"toan": 2, "ngu_van": 3, "vat_ly": 4, "hoa_hoc": 5,
"sinh_hoc": 6, "lich_su": 7, "dia_ly": 8,
"tieng_anh": 11, // NGOAINGU total
} {
if v, ok := parseFloatCell(row, idx); ok {
scores[field] = v
}
}
return &transform.ParsedRow{
SoBaoDanh: sbd,
HoTen: hoTen,
HoTenAscii: transform.ToAscii(hoTen),
Scores: scores,
}
}
// processMappedRow handles header-derived column indices
// (format_detect_2016.rs:226-289).
func processMappedRow(row []reader.Cell, f Format) *transform.ParsedRow {
sbd := cellAt(row, f.Sbd)
hoTen := cellAt(row, f.HoTen)
if sbd == "" || hoTen == "" {
return nil
}
// Leaked-header guard: a row whose SBD or name cell is itself a header token
// is a repeated header, not data (format_detect_2016.rs:244-250).
if isKnownHeader(strings.ToUpper(sbd)) || isKnownHeader(strings.ToUpper(hoTen)) {
return nil
}
optional := func(idx *int) *string {
if idx == nil {
return nil
}
if s := cellAt(row, *idx); s != "" {
return &s
}
return nil
}
// Gender is a two-value allowlist, not a general enum: anything other than
// exactly "Nam" or "Nữ" becomes nil (format_detect_2016.rs:263-271).
var gioiTinh *string
if f.GioiTinh != nil {
if s := cellAt(row, *f.GioiTinh); s == "Nam" || s == "Nữ" {
gioiTinh = &s
}
}
// Read untrimmed, matching format_detect_2016.rs:273-276.
diemThi := ""
if f.DiemThi >= 0 && f.DiemThi < len(row) {
diemThi = row[f.DiemThi].Str
}
return &transform.ParsedRow{
SoBaoDanh: sbd,
HoTen: hoTen,
HoTenAscii: transform.ToAscii(hoTen),
NgaySinh: optional(f.NgaySinh),
TenCumThi: optional(f.TenCumThi),
GioiTinh: gioiTinh,
Scores: transform.ParseScores(diemThi),
}
}
// ProcessRow2016 dispatches a data row through the detected layout. A nil return
// means the row is empty or invalid and should be skipped.
func ProcessRow2016(row []reader.Cell, f Format) *transform.ParsedRow {
if f.Kind == FormatSeparateScores {
return processSeparateScoresRow(row)
}
// FormatMapped and FormatDefault share one implementation; Default is just a
// fixed index tuple.
return processMappedRow(row, f)
}
// Detect2016 ingests the 2016 dataset — the port of run_build_2016 and
// process_file_2016 (main.rs:211-377).
func Detect2016(cfg *config.DatasetConfig, inputDir, outputPath string) error {
files, err := InputFiles(inputDir)
if err != nil {
return err
}
label := DatasetLabel(inputDir)
fmt.Printf("[build:2016] %s/ → %s (%d files)\n", label, outputPath, len(files))
db, err := writer.OpenDB(outputPath)
if err != nil {
return err
}
defer db.Close()
tx, err := db.Begin()
if err != nil {
return fmt.Errorf("begin: %w", err)
}
ins, err := writer.Prepare(tx)
if err != nil {
tx.Rollback()
return err
}
var st writer.Stats // Skipped stays 0 on this path, matching main.rs:251
for _, file := range files {
base := filepath.Base(file)
fileRows, err := processFile2016(file, cfg, ins, base, &st)
if err != nil {
fmt.Fprintf(os.Stderr, " [error] %s: %v\n", base, err)
st.Errors++
continue
}
fmt.Printf(" %s: %d rows\n", base, fileRows)
}
if err := ins.Close(); err != nil {
tx.Rollback()
return err
}
if err := tx.Commit(); err != nil {
return fmt.Errorf("commit: %w", err)
}
return writer.Finish(db, outputPath, st, label, false)
}
// processFile2016 detects the layout per SHEET and processes that sheet's rows.
//
// Detection is per sheet, not per file (main.rs:344-349): within one workbook,
// sheet 2 may legitimately detect differently from sheet 1, which matters for
// the province files that overflow past Excel's row cap.
func processFile2016(path string, cfg *config.DatasetConfig, ins *writer.Inserter, base string, st *writer.Stats) (uint64, error) {
wb, err := reader.Open(path)
if err != nil {
return 0, err
}
defer wb.Close()
sheets := wb.Sheets()
if len(sheets) == 0 {
return 0, nil
}
if cfg.Reader.SheetMode == config.SheetModeFirst {
sheets = sheets[:1]
}
var fileRows uint64
for _, sh := range sheets {
var rows [][]reader.Cell
if err := wb.EachRow(sh.Index, func(_ reader.Sheet, _ int, row []reader.Cell) error {
rows = append(rows, row)
return nil
}); err != nil {
return fileRows, err
}
if len(rows) == 0 {
continue
}
f := defaultFormat()
startIdx := 0
if IsHeaderRow2016(rows[0]) {
f = DetectFormat(rows[0])
startIdx = 1
}
for _, row := range rows[startIdx:] {
// Rows shorter than 2 cells are dropped BEFORE the counter
// (main.rs:351-353), so they never appear in the source total.
if len(row) < 2 {
continue
}
st.SourceRows++
parsed := ProcessRow2016(row, f)
if parsed == nil {
// Empty or invalid. Note this is NOT counted as skipped — the
// 2016 path leaves that counter at zero (main.rs:251), so the
// stats block reports insertable == source rows and the Audit
// line absorbs the difference.
continue
}
if err := ins.Insert(parsed); err != nil {
st.Errors++
if st.Errors <= 5 {
fmt.Fprintf(os.Stderr, " [warn] %s: %v\n", base, err)
}
continue
}
fileRows++
}
}
return fileRows, nil
}
@@ -0,0 +1,252 @@
package ingest
import (
"testing"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
)
// Ports the 11 tests in parser/src/format_detect_2016.rs, plus a guard per quirk.
//
// Phase 1 settled the Data -> string translation these fixtures need, so no
// guessing: Data::String(s) is s verbatim, Data::Float(8.0) renders "8" (Rust's
// f64 Display drops the .0, matching Go's FormatFloat(v,'f',-1,64)),
// Data::Float(8.5) renders "8.5", and Data::Empty is "" with IsEmpty true.
// --- header detection ---
func TestIsHeaderRow2016(t *testing.T) {
if !IsHeaderRow2016(cells("SBD", "HOTEN", "TOAN")) {
t.Error("SBD header not detected")
}
if !IsHeaderRow2016(cells("sobaodanh", "x")) {
t.Error("lowercase SOBAODANH not detected")
}
if IsHeaderRow2016(cells("Nguyen Van A", "01/01/2000")) {
t.Error("data row wrongly detected as header")
}
// The 2016 guard is len < 2, unlike the 2017 check's len < 3.
if IsHeaderRow2016(cells("SBD")) {
t.Error("single-cell row must not be a header")
}
if !IsHeaderRow2016(cells("SBD", "x")) {
t.Error("two-cell row with a header token must be a header")
}
}
// TestKnownHeadersHasTrailingSpaceToken pins the "SINH " literal. Trimming it
// would silently change which rows count as headers.
func TestKnownHeadersHasTrailingSpaceToken(t *testing.T) {
var found bool
for _, h := range KnownHeaders {
if h == "SINH " {
found = true
}
}
if !found {
t.Error(`KnownHeaders must contain "SINH " WITH its trailing space`)
}
if len(KnownHeaders) != 17 {
t.Errorf("KnownHeaders has %d tokens, want 17", len(KnownHeaders))
}
}
// --- format detection ---
func TestDetectFormatSeparateScores(t *testing.T) {
f := DetectFormat(cells("SBD", "HOTEN", "TOAN", "VAN"))
if f.Kind != FormatSeparateScores {
t.Errorf("Kind = %v, want FormatSeparateScores", f.Kind)
}
}
func TestDetectFormatMapped(t *testing.T) {
f := DetectFormat(cells("STT", "SOBAODANH", "HO_TEN", "NGAY_SINH", "TEN_CUMTHI", "GIOI_TINH", "DIEM_THI"))
if f.Kind != FormatMapped {
t.Fatalf("Kind = %v, want FormatMapped", f.Kind)
}
if f.Sbd != 1 || f.HoTen != 2 || f.DiemThi != 6 {
t.Errorf("indices sbd=%d ho_ten=%d diem_thi=%d", f.Sbd, f.HoTen, f.DiemThi)
}
if f.NgaySinh == nil || *f.NgaySinh != 3 || f.TenCumThi == nil || *f.TenCumThi != 4 || f.GioiTinh == nil || *f.GioiTinh != 5 {
t.Error("optional indices not resolved")
}
}
// TestDetectFormatMappedIsOrderIndependent: indices are resolved by name.
func TestDetectFormatMappedIsOrderIndependent(t *testing.T) {
f := DetectFormat(cells("DIEM_THI", "HO_TEN", "SBD"))
if f.Kind != FormatMapped || f.Sbd != 2 || f.DiemThi != 0 || f.HoTen != 1 {
t.Errorf("got %+v", f)
}
}
// TestDetectFormatMappedHoTenFallback ports the col-1 fallback
// (format_detect_2016.rs:132).
func TestDetectFormatMappedHoTenFallback(t *testing.T) {
f := DetectFormat(cells("SBD", "SOMETHING", "DIEM_THI"))
if f.Kind != FormatMapped {
t.Fatalf("Kind = %v, want FormatMapped", f.Kind)
}
if f.HoTen != 1 {
t.Errorf("ho_ten = %d, want fallback 1", f.HoTen)
}
}
// TestDetectFormatDefault: a header lacking SBD or DIEM_THI falls back to the
// positional layout.
func TestDetectFormatDefault(t *testing.T) {
f := DetectFormat(cells("A", "B", "C"))
if f.Kind != FormatDefault {
t.Errorf("Kind = %v, want FormatDefault", f.Kind)
}
if f.Sbd != 0 || f.HoTen != 1 || f.DiemThi != 5 {
t.Errorf("default indices wrong: %+v", f)
}
}
// --- row processing ---
func TestProcessSeparateScoresRow(t *testing.T) {
// 0=SBD 1=HOTEN 2=TOAN 3=VAN 4=LY 5=HOA 6=SINH 7=SU 8=DIA 9,10=NN 11=NN total
row := cells("1000", "Nguyễn Văn Đức", "8", "7.5", "", "6", "", "5", "", "", "", "9.25")
got := ProcessRow2016(row, Format{Kind: FormatSeparateScores})
if got == nil {
t.Fatal("row rejected")
}
if got.SoBaoDanh != "1000" || got.HoTenAscii != "nguyen van duc" {
t.Errorf("sbd=%q ascii=%q", got.SoBaoDanh, got.HoTenAscii)
}
want := map[string]float64{"toan": 8, "ngu_van": 7.5, "hoa_hoc": 6, "lich_su": 5, "tieng_anh": 9.25}
if len(got.Scores) != len(want) {
t.Errorf("scores = %v, want %v", got.Scores, want)
}
for k, v := range want {
if got.Scores[k] != v {
t.Errorf("%s = %v, want %v", k, got.Scores[k], v)
}
}
// These columns are structurally unreachable in this layout.
for _, absent := range []string{"tieng_phap", "tieng_duc", "tieng_nhat", "tieng_trung", "khtn", "khxh", "gdcd"} {
if _, ok := got.Scores[absent]; ok {
t.Errorf("%s must be unreachable in separate-scores", absent)
}
}
if got.NgaySinh != nil || got.TenCumThi != nil || got.GioiTinh != nil {
t.Error("ngay_sinh/ten_cum_thi/gioi_tinh are always nil in separate-scores")
}
}
// TestZeroScoreBecomesNull pins the JS falsy quirk: parseFloat(x) || null means
// a literal 0 is indistinguishable from "no score" (format_detect_2016.rs:165).
func TestZeroScoreBecomesNull(t *testing.T) {
row := cells("1000", "A", "0", "0.0", "1", "", "", "", "", "", "", "")
got := ProcessRow2016(row, Format{Kind: FormatSeparateScores})
if got == nil {
t.Fatal("row rejected")
}
if _, ok := got.Scores["toan"]; ok {
t.Error(`a "0" score must become NULL, not 0`)
}
if _, ok := got.Scores["ngu_van"]; ok {
t.Error(`a "0.0" score must become NULL, not 0`)
}
if got.Scores["vat_ly"] != 1 {
t.Error("a non-zero score must survive")
}
}
// TestGenderAllowlist ports format_detect_2016.rs:263-271 — exactly two values.
func TestGenderAllowlist(t *testing.T) {
f := defaultFormat()
for in, want := range map[string]string{"Nam": "Nam", "Nữ": "Nữ"} {
row := cells("1", "A", "", "", in, "")
got := ProcessRow2016(row, f)
if got == nil || got.GioiTinh == nil || *got.GioiTinh != want {
t.Errorf("gioi_tinh for %q not preserved", in)
}
}
for _, in := range []string{"nam", "NAM", "Unknown", "M", "F", "nữ", ""} {
row := cells("1", "A", "", "", in, "")
got := ProcessRow2016(row, f)
if got == nil {
t.Fatalf("row rejected for %q", in)
}
if got.GioiTinh != nil {
t.Errorf("gioi_tinh for %q = %q, want nil", in, *got.GioiTinh)
}
}
}
// TestLeakedHeaderRowSkipped ports format_detect_2016.rs:244-250.
func TestLeakedHeaderRowSkipped(t *testing.T) {
f := defaultFormat()
if ProcessRow2016(cells("SBD", "HO_TEN", "", "", "", ""), f) != nil {
t.Error("a repeated header row must be skipped")
}
if ProcessRow2016(cells("123", "SOBAODANH", "", "", "", ""), f) != nil {
t.Error("a row whose name cell is a header token must be skipped")
}
if ProcessRow2016(cells("123", "Nguyen Van A", "", "", "", ""), f) == nil {
t.Error("a genuine data row must not be skipped")
}
}
// TestMappedRowEmptyFieldsRejected: either identity field empty rejects the row.
func TestMappedRowEmptyFieldsRejected(t *testing.T) {
f := defaultFormat()
if ProcessRow2016(cells("", "A", "", "", "", ""), f) != nil {
t.Error("empty sbd must reject")
}
if ProcessRow2016(cells("1", "", "", "", "", ""), f) != nil {
t.Error("empty ho_ten must reject")
}
}
// TestMappedRowPopulates2016OnlyColumns: this is the only dataset that fills
// ten_cum_thi and gioi_tinh.
func TestMappedRowPopulates2016OnlyColumns(t *testing.T) {
row := cells("123", "Lê Văn Long", "01/01/1998", "Cụm thi số 1", "Nam", "Toán: 7.5 Tiếng Đức: 6")
got := ProcessRow2016(row, defaultFormat())
if got == nil {
t.Fatal("row rejected")
}
if got.NgaySinh == nil || *got.NgaySinh != "01/01/1998" {
t.Errorf("ngay_sinh = %v", got.NgaySinh)
}
if got.TenCumThi == nil || *got.TenCumThi != "Cụm thi số 1" {
t.Errorf("ten_cum_thi = %v", got.TenCumThi)
}
if got.Scores["toan"] != 7.5 || got.Scores["tieng_duc"] != 6 {
t.Errorf("scores = %v", got.Scores)
}
}
// TestDefaultIsMappedWithFixedIndices: FormatDefault must not be a separate code
// path (format_detect_2016.rs:295-306).
func TestDefaultIsMappedWithFixedIndices(t *testing.T) {
row := cells("123", "A", "01/01/2000", "Cluster", "Nữ", "Toán: 5")
viaDefault := ProcessRow2016(row, defaultFormat())
two, three, four := 2, 3, 4
viaMapped := ProcessRow2016(row, Format{
Kind: FormatMapped, Sbd: 0, HoTen: 1,
NgaySinh: &two, TenCumThi: &three, GioiTinh: &four, DiemThi: 5,
})
if viaDefault == nil || viaMapped == nil {
t.Fatal("row rejected")
}
if viaDefault.SoBaoDanh != viaMapped.SoBaoDanh ||
*viaDefault.TenCumThi != *viaMapped.TenCumThi ||
*viaDefault.GioiTinh != *viaMapped.GioiTinh ||
viaDefault.Scores["toan"] != viaMapped.Scores["toan"] {
t.Error("FormatDefault must behave identically to the equivalent FormatMapped")
}
}
// TestShortRowGuardIsSeparateFromValidation documents that rows under 2 cells
// are dropped by the loop before the counter (main.rs:351-353), not here.
func TestShortRowGuardIsSeparateFromValidation(t *testing.T) {
if got := ProcessRow2016([]reader.Cell{{Str: "123"}}, defaultFormat()); got != nil {
t.Error("a 1-cell row has no name and must be rejected")
}
}
+243
View File
@@ -0,0 +1,243 @@
// Package ingest owns all dataset policy: which sheets to read, which rows are
// headers, which are blank, and how rows are counted.
//
// The reader deliberately has none of this — it reports every sheet and every
// row exactly as calamine would, which is what made its fidelity independently
// testable. This package ports the build loop in parser/src/main.rs plus the
// header/blank helpers in parser/src/reader.rs.
package ingest
import (
"fmt"
"os"
"path/filepath"
"sort"
"strings"
"github.com/tiennm99/thptqg/go-parser/internal/config"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
"github.com/tiennm99/thptqg/go-parser/internal/transform"
"github.com/tiennm99/thptqg/go-parser/internal/writer"
)
// IsHeaderRow reports whether row is a header, by matching its uppercased first
// cell against the configured tokens (reader.rs:28-34).
//
// Rows shorter than 3 cells are never headers (reader.rs:29) — a 1- or 2-cell
// row is a stray fragment, not a real header.
func IsHeaderRow(row []reader.Cell, tokens []string) bool {
if len(row) < 3 {
return false
}
first := strings.ToUpper(strings.TrimSpace(row[0].Str))
for _, t := range tokens {
if strings.ToUpper(t) == first {
return true
}
}
return false
}
// IsAllBlank reports whether every cell is empty or whitespace-only
// (reader.rs:40-43).
//
// Compares on Str only. Cell.IsEmpty is diagnostic: calamine distinguishes
// Data::Empty from an empty string cell, but both render "" and both count as
// blank here, so branching on the flag would invent a distinction the Rust
// original never acts on.
func IsAllBlank(row []reader.Cell) bool {
for _, c := range row {
if strings.TrimSpace(c.Str) != "" {
return false
}
}
return true
}
// InputFiles lists a dataset directory's spreadsheets, sorted.
//
// The sort is load-bearing, not cosmetic: INSERT OR REPLACE is last-wins, so
// file order decides which row survives a duplicate SBD. Rust collects read_dir
// then calls files.sort() (main.rs:82-97) — a bytewise sort on the full path.
func InputFiles(dir string) ([]string, error) {
entries, err := os.ReadDir(dir)
if err != nil {
return nil, fmt.Errorf("cannot read input dir %s: %w", dir, err)
}
var out []string
for _, e := range entries {
if e.IsDir() {
continue
}
switch strings.ToLower(filepath.Ext(e.Name())) {
case ".xls", ".xlsx":
out = append(out, filepath.Join(dir, e.Name()))
}
}
sort.Strings(out)
return out, nil
}
// DatasetLabel derives the stats-wording key from the input directory basename,
// matching main.rs:98-101.
func DatasetLabel(inputDir string) string {
base := filepath.Base(strings.TrimRight(inputDir, string(filepath.Separator)))
if base == "" || base == "." || base == string(filepath.Separator) {
return "data"
}
return base
}
// RowFn consumes one data row of one sheet, after header skipping.
type RowFn func(sheetIdx int, row []reader.Cell)
// ProcessFile applies sheet selection and per-sheet header skipping, invoking fn
// for every remaining row — the port of reader.rs:54-105.
//
// The header check is per SHEET, not per file: first_row resets inside the sheet
// loop (reader.rs:91), so a workbook whose second sheet repeats the header has
// it skipped there too.
func ProcessFile(path string, cfg *config.DatasetConfig, fn RowFn) error {
wb, err := reader.Open(path)
if err != nil {
return err
}
defer wb.Close()
sheets := wb.Sheets()
if len(sheets) == 0 {
return fmt.Errorf("no sheets in %s", path)
}
if cfg.Reader.SheetMode == config.SheetModeFirst {
sheets = sheets[:1]
}
for _, sh := range sheets {
firstRow := true
err := wb.EachRow(sh.Index, func(s reader.Sheet, _ int, row []reader.Cell) error {
if firstRow {
firstRow = false
if IsHeaderRow(row, cfg.Header.Tokens) {
return nil
}
}
fn(s.Index, row)
return nil
})
if err != nil {
return err
}
}
return nil
}
// Standard runs the fixed-column path for the 2017-family datasets — the port of
// run_build_standard (main.rs:74-199).
func Standard(cfg *config.DatasetConfig, inputDir, outputPath string) error {
files, err := InputFiles(inputDir)
if err != nil {
return err
}
label := DatasetLabel(inputDir)
fmt.Printf("[build] %s/ → %s (%d files)\n", label, outputPath, len(files))
db, err := writer.OpenDB(outputPath)
if err != nil {
return err
}
defer db.Close()
isOld2 := strings.Contains(label, "old2")
stripBlank := cfg.Reader.StripBlankRows
// One transaction spans the whole dataset directory (main.rs:120,184).
tx, err := db.Begin()
if err != nil {
return fmt.Errorf("begin: %w", err)
}
ins, err := writer.Prepare(tx)
if err != nil {
tx.Rollback()
return err
}
var st writer.Stats
for _, file := range files {
base := filepath.Base(file)
var fileRows, fileSkipped, fileErrors uint64
procErr := ProcessFile(file, cfg, func(_ int, row []reader.Cell) {
allBlank := IsAllBlank(row)
// 2017-old2: blank rows drop out BEFORE the source-row counter
// (main.rs:134-137, ahead of the increment at :140).
if stripBlank && allBlank {
return
}
st.SourceRows++
hoTen, soBaoDanh := "", ""
if cols := cfg.Columns; cols != nil {
hoTen = cellAt(row, cols.HoTen)
soBaoDanh = cellAt(row, cols.SoBaoDanh)
}
switch transform.ValidateRow(hoTen, soBaoDanh, &cfg.Validation, stripBlank, allBlank) {
case transform.SkipBlankRow:
// Falls through to transform and insert, matching main.rs:150.
// Unreachable here: BlankRow requires stripBlank && allBlank,
// which returned above. Kept so the two call sites with opposite
// outcomes stay visibly distinct.
case transform.SkipNone:
// proceed
default:
fileSkipped++
return
}
parsed, err := transform.TransformRow(row, cfg)
if err != nil {
fileErrors++
return
}
if err := ins.Insert(parsed); err != nil {
fileErrors++
// Only the first five insert warnings print (main.rs:164).
if st.Errors+fileErrors <= 5 {
fmt.Fprintf(os.Stderr, " [warn] %s: %v\n", base, err)
}
return
}
fileRows++
})
if procErr != nil {
// A file that cannot be read is logged and counted, never fatal
// (main.rs:171-177) — one corrupt file must not abandon the batch.
fmt.Fprintf(os.Stderr, " [error] %s: %v\n", base, procErr)
fileErrors++
}
st.Skipped += fileSkipped
st.Errors += fileErrors
fmt.Printf(" %s: %d rows\n", base, fileRows)
}
if err := ins.Close(); err != nil {
tx.Rollback()
return err
}
if err := tx.Commit(); err != nil {
return fmt.Errorf("commit: %w", err)
}
// VACUUM only after COMMIT — SQLite refuses it inside a transaction.
return writer.Finish(db, outputPath, st, label, isOld2)
}
// cellAt returns the trimmed cell at idx, or "" when out of range — the
// unwrap_or_default() behaviour of transform.rs:163-167.
func cellAt(row []reader.Cell, idx int) string {
if idx < 0 || idx >= len(row) {
return ""
}
return strings.TrimSpace(row[idx].Str)
}
+131
View File
@@ -0,0 +1,131 @@
package ingest
import (
"os"
"path/filepath"
"testing"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
)
// Ports the 7 tests in parser/src/reader.rs:116-197. They were listed under
// Phase 1 originally, which was wrong: they exercise header and blank-row
// policy, which lives here rather than in the reader package.
func cells(vals ...string) []reader.Cell {
out := make([]reader.Cell, len(vals))
for i, v := range vals {
out[i] = reader.Cell{Str: v, IsEmpty: v == ""}
}
return out
}
var stdTokens = []string{"HO_TEN", "HỌ TÊN", "STT"}
// header_detects_ho_ten (reader.rs:128)
func TestHeaderDetectsHoTen(t *testing.T) {
if !IsHeaderRow(cells("HO_TEN", "NGAY_SINH", "SBD"), stdTokens) {
t.Error("HO_TEN header not detected")
}
}
// header_detects_stt (reader.rs:139)
func TestHeaderDetectsStt(t *testing.T) {
if !IsHeaderRow(cells("STT", "B", "C"), stdTokens) {
t.Error("STT header not detected")
}
}
// header_detects_ho_ten_unicode (reader.rs:150)
func TestHeaderDetectsHoTenUnicode(t *testing.T) {
if !IsHeaderRow(cells("HỌ TÊN", "B", "C"), stdTokens) {
t.Error("HỌ TÊN header not detected")
}
}
// header_rejects_data_row (reader.rs:161)
func TestHeaderRejectsDataRow(t *testing.T) {
if IsHeaderRow(cells("Nguyen Van A", "01/01/2000", "12345678"), stdTokens) {
t.Error("data row wrongly detected as header")
}
}
// header_rejects_short_row (reader.rs:172) — the <3 cell guard.
func TestHeaderRejectsShortRow(t *testing.T) {
if IsHeaderRow(cells("HO_TEN", ""), stdTokens) {
t.Error("a 2-cell row must never be a header, even with a matching token")
}
}
// header_case_insensitive (reader.rs:179)
func TestHeaderCaseInsensitive(t *testing.T) {
if !IsHeaderRow(cells("ho_ten", "B", "C"), stdTokens) {
t.Error("lowercase header not detected")
}
}
// blank_row_detection (reader.rs:190) — note the third cell is an empty *string*
// cell, not Data::Empty, and must still count as blank.
func TestBlankRowDetection(t *testing.T) {
blank := []reader.Cell{{IsEmpty: true}, {IsEmpty: true}, {Str: ""}}
if !IsAllBlank(blank) {
t.Error("row of empty cells should be blank")
}
if IsAllBlank(cells("Nguyen", "", "")) {
t.Error("row with content should not be blank")
}
}
// TestIsAllBlankIgnoresIsEmptyFlag pins the rule that blankness is decided by
// the rendered string, never by Cell.IsEmpty — calamine emits empty-but-not-Empty
// cells, and treating the flag as authoritative would diverge.
func TestIsAllBlankIgnoresIsEmptyFlag(t *testing.T) {
// Content present but IsEmpty wrongly set: still not blank.
if IsAllBlank([]reader.Cell{{Str: "x", IsEmpty: true}}) {
t.Error("a cell with content must not be blank regardless of IsEmpty")
}
// Whitespace only: blank, matching Rust's trim().is_empty().
if !IsAllBlank([]reader.Cell{{Str: " "}, {Str: "\t"}}) {
t.Error("whitespace-only cells should be blank")
}
}
// TestDatasetLabel covers the basename derivation that drives the stats wording.
func TestDatasetLabel(t *testing.T) {
for in, want := range map[string]string{
"data/2017": "2017",
"data/2017-old2": "2017-old2",
"data/2017-old2/": "2017-old2",
"/abs/path/2016": "2016",
} {
if got := DatasetLabel(in); got != want {
t.Errorf("DatasetLabel(%q) = %q, want %q", in, got, want)
}
}
}
// TestInputFilesSortedAndFiltered: the sort decides which duplicate SBD survives
// INSERT OR REPLACE, so it is behaviour, not presentation.
func TestInputFilesSortedAndFiltered(t *testing.T) {
dir := t.TempDir()
for _, name := range []string{"b.xlsx", "a.xls", "c.XLSX", "notes.txt", "d.csv"} {
if err := writeEmpty(filepath.Join(dir, name)); err != nil {
t.Fatal(err)
}
}
got, err := InputFiles(dir)
if err != nil {
t.Fatal(err)
}
want := []string{"a.xls", "b.xlsx", "c.XLSX"}
if len(got) != len(want) {
t.Fatalf("got %d files %v, want %d", len(got), got, len(want))
}
for i := range want {
if filepath.Base(got[i]) != want[i] {
t.Errorf("file %d = %q, want %q", i, filepath.Base(got[i]), want[i])
}
}
}
func writeEmpty(path string) error { return os.WriteFile(path, nil, 0o644) }
+128
View File
@@ -0,0 +1,128 @@
package reader_test
import (
"bufio"
"crypto/sha256"
"encoding/hex"
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
)
// TestReaderFidelity asserts the Go reader reproduces calamine byte-for-byte on
// every real input file.
//
// The oracle is a committed SHA-256 per file over a canonical cell dump. The
// dumps themselves are real student names and birthdates, so only the hashes are
// committed — regenerate the dumps from the Rust side on demand
// (parser/examples/dump_cells.rs), which is possible because parser/ still
// builds.
//
// The canonical form carries geometry and rendered cell values. The calamine
// Data variant is excluded on purpose: Data::Empty and Data::String("") both
// render "" and both count as blank in is_all_blank and transform, so the
// distinction cannot reach the database.
// Runs by default so CI and `go test ./...` keep the full guarantee; skipped
// under -short, which is how to iterate without paying ~77s to re-read 418 MB.
func TestReaderFidelity(t *testing.T) {
if testing.Short() {
t.Skip("-short: skipping the 299-file corpus sweep")
}
root := repoRoot(t)
manifest := filepath.Join(root, "go-parser", "testdata", "reader-fidelity-hashes.tsv")
f, err := os.Open(manifest)
if err != nil {
t.Fatalf("open manifest: %v", err)
}
defer f.Close()
var checked int
sc := bufio.NewScanner(f)
for sc.Scan() {
line := sc.Text()
if line == "" || strings.HasPrefix(line, "#") {
continue
}
rel, want, ok := strings.Cut(line, "\t")
if !ok {
t.Fatalf("malformed manifest line: %q", line)
}
path := filepath.Join(root, rel)
if _, err := os.Stat(path); err != nil {
t.Skipf("input data not present (%s); skipping fidelity suite", rel)
}
checked++
t.Run(rel, func(t *testing.T) {
t.Parallel()
got, err := canonicalHash(path)
if err != nil {
t.Fatalf("hash %s: %v", rel, err)
}
if got != want {
t.Errorf("cell dump diverges from calamine\n want %s\n got %s", want, got)
}
})
}
if err := sc.Err(); err != nil {
t.Fatalf("read manifest: %v", err)
}
if checked == 0 {
t.Fatal("manifest contained no entries")
}
}
// canonicalHash renders one file in the canonical form and hashes it. Kept
// byte-identical to the awk canonicalisation used to build the manifest.
func canonicalHash(path string) (string, error) {
wb, err := reader.Open(path)
if err != nil {
return "", err
}
defer wb.Close()
h := sha256.New()
sheets := wb.Sheets()
fmt.Fprintf(h, "SHEETCOUNT\t%d\n", len(sheets))
for _, sh := range sheets {
fmt.Fprintf(h, "SHEET\t%d\t%s\t%d\t%d\n", sh.Index, escape(sh.Name), sh.Height, sh.Width)
err := wb.EachRow(sh.Index, func(s reader.Sheet, rowIdx int, row []reader.Cell) error {
fmt.Fprintf(h, "ROW\t%d\t%d\t%d\n", s.Index, rowIdx, len(row))
for c, cell := range row {
fmt.Fprintf(h, "CELL\t%d\t%d\t%d\t%s\n", s.Index, rowIdx, c, escape(cell.Str))
}
return nil
})
if err != nil {
return "", err
}
}
return hex.EncodeToString(h.Sum(nil)), nil
}
func escape(s string) string {
return strings.NewReplacer("\\", `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`).Replace(s)
}
// repoRoot walks up from the test's working directory to the directory holding
// the data/ corpus.
func repoRoot(t *testing.T) string {
t.Helper()
dir, err := os.Getwd()
if err != nil {
t.Fatalf("getwd: %v", err)
}
for i := 0; i < 6; i++ {
if _, err := os.Stat(filepath.Join(dir, "data")); err == nil {
return dir
}
dir = filepath.Dir(dir)
}
t.Fatal("could not locate repo root (no data/ directory found)")
return ""
}
+79
View File
@@ -0,0 +1,79 @@
// Package reader wraps the two spreadsheet libraries behind one streaming
// contract, mirroring parser/src/reader.rs (which wraps calamine's
// open_workbook_auto for both formats).
//
// The contract is deliberately row-streaming rather than whole-workbook
// materialising: data/2017/ha-noi.xls alone holds 72k rows across two sheets,
// and the Rust original keeps at most one sheet range live at a time.
package reader
import (
"fmt"
"path/filepath"
"strings"
)
// Cell is one spreadsheet cell rendered the way calamine's Data::to_string()
// renders it.
//
// IsEmpty tracks calamine's Data::Empty variant separately from a string cell
// that happens to be empty. Rust distinguishes them at reader.rs:42, but both
// render "" and both count as blank in is_all_blank and in transform, so the
// flag is diagnostic only — never compare on it.
type Cell struct {
Str string
IsEmpty bool
}
// Sheet carries a sheet's identity and geometry. Height and Width describe the
// used range, matching calamine's Range::height()/width(); every used range in
// the corpus starts at (0,0), verified across all 299 files.
type Sheet struct {
Index int
Name string
Height int
Width int
}
// RowFunc receives each row of each sheet. Rows are padded to the sheet's used
// width — width is load-bearing because every column read downstream is
// positional with an unwrap_or_default() equivalent, so a short row silently
// NULLs its tail columns.
type RowFunc func(sheet Sheet, rowIdx int, row []Cell) error
// Workbook is one opened spreadsheet.
type Workbook interface {
// Sheets returns sheet identity and geometry in workbook order.
Sheets() []Sheet
// EachRow streams every row of the given sheet in order.
EachRow(sheetIdx int, fn RowFunc) error
Close() error
}
// Open dispatches on file extension, mirroring calamine's open_workbook_auto.
func Open(path string) (Workbook, error) {
switch strings.ToLower(filepath.Ext(path)) {
case ".xls":
return openXLS(path)
case ".xlsx", ".xlsm":
return openXLSX(path)
default:
return nil, fmt.Errorf("unsupported extension: %s", path)
}
}
// padRow extends row to width with empty cells, and truncates if longer.
func padRow(row []Cell, width int) []Cell {
if len(row) == width {
return row
}
if len(row) > width {
return row[:width]
}
out := make([]Cell, width)
copy(out, row)
for i := len(row); i < width; i++ {
out[i] = Cell{IsEmpty: true}
}
return out
}
+119
View File
@@ -0,0 +1,119 @@
package reader
import (
"fmt"
"github.com/pbnjay/grate"
// Registers the BIFF backend with grate.Open.
_ "github.com/pbnjay/grate/xls"
)
// xlsWorkbook reads legacy BIFF through pbnjay/grate.
//
// grate replaced extrame/xls, which was measured against calamine ground truth
// and found to corrupt 69% of cells and drop a further 28%: undecoded UTF-16LE
// and BIFF record framing leaked into cell values, content moved between rows
// and columns, and tail rows came back blank. That was charset-independent and
// unfixable from the outside. grate reproduces calamine exactly on the same
// files.
//
// The one normalisation grate needs is trailing blank rows: it yields rows past
// the end of calamine's used range (one for a populated sheet, two for an empty
// one), so trailing all-blank rows are trimmed. Note this is the opposite of
// the xlsx path, where excelize already trims and calamine keeps a 1x1 empty
// range — in both cases the rule is "match calamine's used range".
type xlsWorkbook struct {
sheets []Sheet
rows [][][]Cell
}
func openXLS(path string) (Workbook, error) {
wb, err := grate.Open(path)
if err != nil {
return nil, fmt.Errorf("grate open %s: %w", path, err)
}
defer wb.Close()
names, err := wb.List()
if err != nil {
return nil, fmt.Errorf("grate list %s: %w", path, err)
}
out := &xlsWorkbook{}
for idx, name := range names {
sh, err := wb.Get(name)
if err != nil {
return nil, fmt.Errorf("grate get %s/%s: %w", path, name, err)
}
var raw [][]string
for sh.Next() {
row := sh.Strings()
cp := make([]string, len(row))
for i, v := range row {
cp[i] = demergeMarker(v)
}
raw = append(raw, cp)
}
// Trim to calamine's used range.
height := len(raw)
for height > 0 && rowAllBlank(raw[height-1]) {
height--
}
raw = raw[:height]
width := 0
for _, r := range raw {
if len(r) > width {
width = len(r)
}
}
cells := make([][]Cell, len(raw))
for i, r := range raw {
row := make([]Cell, len(r))
for j, v := range r {
row[j] = Cell{Str: v, IsEmpty: v == ""}
}
cells[i] = padRow(row, width)
}
out.sheets = append(out.sheets, Sheet{Index: idx, Name: name, Height: height, Width: width})
out.rows = append(out.rows, cells)
}
return out, nil
}
// demergeMarker blanks grate's merged-cell continuation markers.
//
// grate fills the cells covered by a merge with sentinel runes; calamine
// reports them as empty. In this corpus they occur only in the merged title
// block of the 2016 spreadsheets (19 cells in rows 0-2 of one file). Only an
// exact whole-value match is blanked, so a real cell that merely contains an
// arrow is untouched.
func demergeMarker(v string) string {
switch v {
case grate.ContinueColumnMerged, grate.EndColumnMerged,
grate.ContinueRowMerged, grate.EndRowMerged:
return ""
}
return v
}
func (w *xlsWorkbook) Sheets() []Sheet { return w.sheets }
func (w *xlsWorkbook) EachRow(sheetIdx int, fn RowFunc) error {
if sheetIdx < 0 || sheetIdx >= len(w.sheets) {
return fmt.Errorf("sheet index %d out of range", sheetIdx)
}
sh := w.sheets[sheetIdx]
for i, row := range w.rows[sheetIdx] {
if err := fn(sh, i, row); err != nil {
return err
}
}
return nil
}
func (w *xlsWorkbook) Close() error { return nil }
+105
View File
@@ -0,0 +1,105 @@
package reader
import (
"fmt"
"github.com/xuri/excelize/v2"
)
// xlsxWorkbook reads OOXML through excelize.
//
// Two excelize behaviours must be corrected to match calamine:
//
// 1. GetRows applies the cell number format by default, while calamine renders
// the underlying value. RawCellValue: true disables that.
// 2. GetRows trims trailing blank cells, so rows are ragged; calamine returns a
// rectangular used range. Rows are padded back out to the sheet width.
type xlsxWorkbook struct {
f *excelize.File
sheets []Sheet
rows [][][]Cell // [sheetIdx][rowIdx][colIdx]
}
func openXLSX(path string) (Workbook, error) {
f, err := excelize.OpenFile(path)
if err != nil {
return nil, fmt.Errorf("excelize open %s: %w", path, err)
}
wb := &xlsxWorkbook{f: f}
crFixups := buildCRFixups(path)
for idx, name := range f.GetSheetList() {
raw, err := f.GetRows(name, excelize.Options{RawCellValue: true})
if err != nil {
f.Close()
return nil, fmt.Errorf("excelize GetRows %s/%s: %w", path, name, err)
}
// Do NOT trim trailing blank rows: calamine's used range keeps them, and
// excelize's GetRows already drops trailing fully-empty rows itself.
//
// One correction is needed. 63 sheets in 2017-old and 53 in 2017-old2
// hold a single empty shared-string cell at A1; calamine reports those
// as a 1x1 range, while GetRows returns nothing. A genuinely empty sheet
// (230 of them in 2016) is height 0 on both sides. GetCellType tells the
// two apart: the empty-shared-string cell exists in the XML and types as
// CellTypeSharedString, an absent cell types as CellTypeUnset.
if len(raw) == 0 {
if t, terr := f.GetCellType(name, "A1"); terr == nil && t != excelize.CellTypeUnset {
raw = [][]string{{""}}
}
}
height := len(raw)
width := 0
for _, r := range raw {
if len(r) > width {
width = len(r)
}
}
cells := make([][]Cell, len(raw))
for i, r := range raw {
row := make([]Cell, len(r))
for j, v := range r {
if fixed, ok := crFixups[v]; ok {
v = fixed
} else {
v = normalizeNumeric(f, name, j, i, v)
}
row[j] = Cell{Str: v, IsEmpty: v == ""}
}
cells[i] = padRow(row, width)
}
wb.sheets = append(wb.sheets, Sheet{Index: idx, Name: name, Height: height, Width: width})
wb.rows = append(wb.rows, cells)
}
return wb, nil
}
func rowAllBlank(r []string) bool {
for _, v := range r {
if v != "" {
return false
}
}
return true
}
func (w *xlsxWorkbook) Sheets() []Sheet { return w.sheets }
func (w *xlsxWorkbook) EachRow(sheetIdx int, fn RowFunc) error {
if sheetIdx < 0 || sheetIdx >= len(w.sheets) {
return fmt.Errorf("sheet index %d out of range", sheetIdx)
}
sh := w.sheets[sheetIdx]
for i, row := range w.rows[sheetIdx] {
if err := fn(sh, i, row); err != nil {
return err
}
}
return nil
}
func (w *xlsxWorkbook) Close() error { return w.f.Close() }
+155
View File
@@ -0,0 +1,155 @@
package reader
import (
"archive/zip"
"bytes"
"encoding/xml"
"io"
"strconv"
"github.com/xuri/excelize/v2"
)
// normalizeNumeric reproduces calamine's rendering of a numeric cell.
//
// calamine parses a numeric cell to f64 and renders it with Rust's f64 Display,
// so the stored literal "6.0" becomes "6". RawCellValue hands back the literal.
//
// The cell type must be consulted, not guessed: "01063476", "6.00" and "NAN" are
// all shared strings that survive ParseFloat, and renumbering them would drop a
// leading zero, drop a trailing zero, or recase NaN. GetCellType is only called
// when re-rendering would actually change the text, which keeps it off the hot
// path for the ~99% of cells that are already canonical or plainly non-numeric.
func normalizeNumeric(f *excelize.File, sheet string, col, row int, v string) string {
if v == "" {
return v
}
fv, err := strconv.ParseFloat(v, 64)
if err != nil {
return v
}
out := strconv.FormatFloat(fv, 'f', -1, 64)
if out == v {
return v
}
axis, err := excelize.CoordinatesToCellName(col+1, row+1)
if err != nil {
return v
}
// OOXML omits the t attribute on numeric cells, and excelize has no map
// entry for an empty t, so a plain number reports CellTypeUnset rather than
// CellTypeNumber. Unset is only reachable here for a cell that exists and
// parsed as a float — an absent cell is "" and returned above — so both
// values mean "numeric". Shared strings report CellTypeSharedString and are
// left alone, which is what protects "01063476", "6.00" and "NAN".
if t, err := f.GetCellType(sheet, axis); err == nil &&
(t == excelize.CellTypeNumber || t == excelize.CellTypeUnset) {
return out
}
return v
}
// buildCRFixups maps a shared string's line-ending-normalised form back to its
// raw form, for the strings that contain a carriage return.
//
// Go's encoding/xml performs the line-ending normalisation the XML 1.0 spec
// mandates (CRLF and lone CR both become LF), so excelize returns "a\nb" where
// calamine — which reads the raw bytes — returns "a\r\nb". That difference
// reaches the database: in one 2016 file it affects 2,233 TEN_CUMTHI values,
// which populate the ten_cum_thi column.
//
// The trick is that character references are exempt from that normalisation, so
// rewriting literal CR bytes to &#13; before decoding round-trips them intact.
//
// Returns nil when the file has no CR at all, which is the common case.
func buildCRFixups(path string) map[string]string {
zr, err := zip.OpenReader(path)
if err != nil {
return nil
}
defer zr.Close()
var entry *zip.File
for _, zf := range zr.File {
if zf.Name == "xl/sharedStrings.xml" {
entry = zf
break
}
}
if entry == nil {
return nil
}
rc, err := entry.Open()
if err != nil {
return nil
}
defer rc.Close()
data, err := io.ReadAll(rc)
if err != nil || !bytes.ContainsRune(data, '\r') {
return nil
}
data = bytes.ReplaceAll(data, []byte{'\r'}, []byte("&#13;"))
fixups := make(map[string]string)
dec := xml.NewDecoder(bytes.NewReader(data))
var cur bytes.Buffer
inSI, inT := false, false
for {
tok, err := dec.Token()
if err != nil {
break
}
switch t := tok.(type) {
case xml.StartElement:
switch t.Name.Local {
case "si":
inSI, cur = true, bytes.Buffer{}
case "t":
inT = true
}
case xml.CharData:
if inSI && inT {
cur.Write(t)
}
case xml.EndElement:
switch t.Name.Local {
case "t":
inT = false
case "si":
if inSI {
raw := cur.String()
if norm := normalizeEOL(raw); norm != raw {
fixups[norm] = raw
}
}
inSI = false
}
}
}
if len(fixups) == 0 {
return nil
}
return fixups
}
// normalizeEOL applies XML 1.0 line-ending normalisation: CRLF and lone CR
// both collapse to LF. This is what encoding/xml does to the raw bytes, so it
// reproduces the text excelize hands back.
func normalizeEOL(s string) string {
if !bytes.ContainsRune([]byte(s), '\r') {
return s
}
var b bytes.Buffer
b.Grow(len(s))
for i := 0; i < len(s); i++ {
if s[i] == '\r' {
if i+1 < len(s) && s[i+1] == '\n' {
continue // CRLF: the LF is emitted on the next pass
}
b.WriteByte('\n') // lone CR
continue
}
b.WriteByte(s[i])
}
return b.String()
}
+144
View File
@@ -0,0 +1,144 @@
// Package schema is the single source of truth for the SQL shape of every
// dataset — a direct port of parser/src/schema.rs.
//
// All four datasets (2016, 2017, 2017-old, 2017-old2) write into the same
// 22-column student table. Columns a dataset has no data for bind NULL.
//
// Column provenance:
//
// ten_cum_thi, gioi_tinh, tieng_duc, tieng_nhat -> 2016 only
// khtn, khxh, gdcd, tieng_nga -> 2017 datasets only
// everything else -> both
//
// Before this consolidation the DDL, the INSERT and the subject regexes were
// duplicated across four TOML configs, which is how the 2016 and 2017 schemas
// drifted apart. The configs now carry only per-dataset parse rules.
package schema
import "regexp"
// DDL is executed verbatim after the output database is (re)created.
//
// idx_ten_cum_thi is partial, so it holds zero entries on the three 2017
// datasets — where the column is always NULL — while staying useful for the
// 2016 cluster-grouping queries. Partial indexes are SQLite-specific.
//
// Byte-identical to parser/src/schema.rs:26-54; TestDDLMatchesRust enforces it.
const DDL = `
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
ngay_sinh TEXT,
ten_cum_thi TEXT,
gioi_tinh TEXT,
toan REAL,
ngu_van REAL,
vat_ly REAL,
hoa_hoc REAL,
sinh_hoc REAL,
khtn REAL,
lich_su REAL,
dia_ly REAL,
gdcd REAL,
khxh REAL,
tieng_anh REAL,
tieng_phap REAL,
tieng_nga REAL,
tieng_duc REAL,
tieng_nhat REAL,
tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
`
// IdentityFields are the identity columns, in INSERT parameter order.
var IdentityFields = []string{
"so_bao_danh",
"ho_ten",
"ho_ten_ascii",
"ngay_sinh",
"ten_cum_thi",
"gioi_tinh",
}
// ScoreFields are the subject columns, in INSERT parameter order. Bound NULL
// when a row has no score for that subject.
var ScoreFields = []string{
"toan",
"ngu_van",
"vat_ly",
"hoa_hoc",
"sinh_hoc",
"khtn",
"lich_su",
"dia_ly",
"gdcd",
"khxh",
"tieng_anh",
"tieng_phap",
"tieng_nga",
"tieng_duc",
"tieng_nhat",
"tieng_trung",
}
// ParamCount is the total bound parameters per row.
const ParamCount = 22
// InsertSQL is a positional INSERT matching IdentityFields then ScoreFields.
//
// OR REPLACE is a behavioural contract, not an optimisation: a repeated SBD
// overwrites the earlier row rather than aborting the transaction, so the last
// file to supply a duplicate wins.
const InsertSQL = `
INSERT OR REPLACE INTO student
(so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh, ten_cum_thi, gioi_tinh,
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn,
lich_su, dia_ly, gdcd, khxh,
tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung)
VALUES
(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
`
// scorePatternSources holds the regex per subject, applied to the DIEM_THI cell
// text. Copied verbatim from parser/src/schema.rs:126-143 — the literals contain
// Vietnamese text and must never be retyped.
//
// Every pattern runs against every dataset. A subject absent from a given exam
// year simply never matches and stays NULL: 2016 files contain no "KHTN:" or
// "Tiếng Nga:" tokens, and 2017 files contain no "Tiếng Đức:" or "Tiếng Nhật:".
//
// Go's regexp and Rust's regex crate are both RE2, and these patterns use no
// backreferences, lookaround or Unicode classes, so they port with zero risk.
var scorePatternSources = map[string]string{
"toan": `Toán:\s*(\d+(?:\.\d+)?)`,
"ngu_van": `Ngữ văn:\s*(\d+(?:\.\d+)?)`,
"vat_ly": `Vật lí:\s*(\d+(?:\.\d+)?)`,
"hoa_hoc": `Hóa học:\s*(\d+(?:\.\d+)?)`,
"sinh_hoc": `Sinh học:\s*(\d+(?:\.\d+)?)`,
"khtn": `KHTN:\s*(\d+(?:\.\d+)?)`,
"lich_su": `Lịch sử:\s*(\d+(?:\.\d+)?)`,
"dia_ly": `Địa lí:\s*(\d+(?:\.\d+)?)`,
"gdcd": `GDCD:\s*(\d+(?:\.\d+)?)`,
"khxh": `KHXH:\s*(\d+(?:\.\d+)?)`,
"tieng_anh": `Tiếng Anh:\s*(\d+(?:\.\d+)?)`,
"tieng_phap": `Tiếng Pháp:\s*(\d+(?:\.\d+)?)`,
"tieng_nga": `Tiếng Nga:\s*(\d+(?:\.\d+)?)`,
"tieng_duc": `Tiếng Đức:\s*(\d+(?:\.\d+)?)`,
"tieng_nhat": `Tiếng Nhật:\s*(\d+(?:\.\d+)?)`,
"tieng_trung": `Tiếng Trung:\s*(\d+(?:\.\d+)?)`,
}
// ScorePatterns holds the compiled subject regexes, compiled once at init.
// Rust compiles them once per run in CompiledPatterns::new; a package-level map
// is the equivalent for a single-threaded CLI.
var ScorePatterns = func() map[string]*regexp.Regexp {
out := make(map[string]*regexp.Regexp, len(scorePatternSources))
for field, src := range scorePatternSources {
out[field] = regexp.MustCompile(src)
}
return out
}()
+150
View File
@@ -0,0 +1,150 @@
package schema
import (
"strings"
"testing"
)
// Ports the four tests in parser/src/schema.rs:149-213. Their purpose is to stop
// the DDL, the INSERT column list and the field-order constants from drifting
// apart — a drift that silently lands values in the wrong columns.
// TestInsertMatchesFieldOrder ports insert_matches_field_order (schema.rs:156).
func TestInsertMatchesFieldOrder(t *testing.T) {
if ParamCount != 22 {
t.Errorf("ParamCount = %d, want 22", ParamCount)
}
if got := strings.Count(InsertSQL, "?"); got != ParamCount {
t.Errorf("INSERT placeholders = %d, want %d", got, ParamCount)
}
open := strings.Index(InsertSQL, "(")
closeIdx := strings.Index(InsertSQL, ")")
if open < 0 || closeIdx < 0 {
t.Fatal("INSERT must contain a column list")
}
var listed []string
for _, c := range strings.Split(InsertSQL[open+1:closeIdx], ",") {
if c = strings.TrimSpace(c); c != "" {
listed = append(listed, c)
}
}
want := append(append([]string{}, IdentityFields...), ScoreFields...)
if len(listed) != len(want) {
t.Fatalf("INSERT lists %d columns, want %d", len(listed), len(want))
}
for i := range want {
if listed[i] != want[i] {
t.Errorf("column %d: INSERT has %q, field order has %q", i, listed[i], want[i])
}
}
}
// TestScorePatternsCoverScoreFields ports score_patterns_cover_score_fields
// (schema.rs:183).
func TestScorePatternsCoverScoreFields(t *testing.T) {
if len(ScorePatterns) != len(ScoreFields) {
t.Fatalf("%d patterns for %d score columns", len(ScorePatterns), len(ScoreFields))
}
inFields := make(map[string]bool, len(ScoreFields))
for _, f := range ScoreFields {
inFields[f] = true
}
for field := range ScorePatterns {
if !inFields[field] {
t.Errorf("pattern %q has no column", field)
}
}
for _, field := range ScoreFields {
if _, ok := ScorePatterns[field]; !ok {
t.Errorf("column %q has no pattern", field)
}
}
}
// TestDDLColumnsMatchInsert ports ddl_columns_match_insert (schema.rs:201).
func TestDDLColumnsMatchInsert(t *testing.T) {
for _, field := range append(append([]string{}, IdentityFields...), ScoreFields...) {
if !strings.Contains(DDL, field) {
t.Errorf("DDL missing column %q", field)
}
}
}
// TestScorePatternsCompile ports score_patterns_compile (schema.rs:208).
// Compilation happens in the package initialiser, so reaching this point already
// proves it; the explicit checks guard against an empty or partial table.
func TestScorePatternsCompile(t *testing.T) {
for field, re := range ScorePatterns {
if re == nil {
t.Errorf("pattern %q is nil", field)
}
}
}
// TestDDLMatchesRust asserts the DDL is byte-identical to parser/src/schema.rs.
// Anything less and the two parsers can produce structurally different databases
// while every row-level check still passes.
func TestDDLMatchesRust(t *testing.T) {
const want = `
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
ngay_sinh TEXT,
ten_cum_thi TEXT,
gioi_tinh TEXT,
toan REAL,
ngu_van REAL,
vat_ly REAL,
hoa_hoc REAL,
sinh_hoc REAL,
khtn REAL,
lich_su REAL,
dia_ly REAL,
gdcd REAL,
khxh REAL,
tieng_anh REAL,
tieng_phap REAL,
tieng_nga REAL,
tieng_duc REAL,
tieng_nhat REAL,
tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
`
if DDL != want {
t.Errorf("DDL diverges from parser/src/schema.rs:26-54\n--- got ---\n%s\n--- want ---\n%s", DDL, want)
}
}
// TestScorePatternsMatchScores exercises each pattern against the shape the
// DIEM_THI cell actually carries, including the wide runs of spaces seen in the
// real corpus.
func TestScorePatternsMatchScores(t *testing.T) {
const cell = "Toán: 8.50 Ngữ văn: 7.00 Tiếng Đức: 9 KHXH: 5.58 "
cases := map[string]string{
"toan": "8.50",
"ngu_van": "7.00",
"tieng_duc": "9",
"khxh": "5.58",
"tieng_nhat": "", // absent from the cell -> no match
}
for field, want := range cases {
re, ok := ScorePatterns[field]
if !ok {
t.Fatalf("no pattern for %q", field)
}
m := re.FindStringSubmatch(cell)
got := ""
if m != nil {
got = m[1]
}
if got != want {
t.Errorf("%s: matched %q, want %q", field, got, want)
}
}
}
+25
View File
@@ -0,0 +1,25 @@
// Package sqlitedb registers the SQLite driver the parser writes with and names
// it in one place.
//
// modernc.org/sqlite is a pure-Go SQLite, chosen so the whole module stays
// cgo-free — grate, excelize and yaml.v3 are pure Go too, so CI can compile with
// CGO_ENABLED=0 and no C toolchain.
//
// It is a machine-transpiled SQLite rather than the upstream C amalgamation that
// Rust's rusqlite --bundled vendors (libsqlite3-sys 0.30.1). Two things make that
// acceptable: this parser uses only plain SQL — no CTEs, window functions,
// triggers or extensions — and the differential gate compares a full-table
// SHA-256 plus PRAGMA table_info/index_list against live Rust output.
//
// Verified on linux/arm64 with v1.56.0 (SQLite 3.53.3): the full DDL including
// the partial idx_ten_cum_thi index, INSERT OR REPLACE, and VACUUM.
//
// Fallback if the differential gate ever implicates the driver: mattn/go-sqlite3
// is upstream C at the cost of cgo.
package sqlitedb
// Registers "sqlite" with database/sql.
import _ "modernc.org/sqlite"
// DriverName is the database/sql driver name to pass to sql.Open.
const DriverName = "sqlite"
@@ -0,0 +1,65 @@
package transform_test
import (
"database/sql"
"os"
"testing"
_ "github.com/tiennm99/thptqg/go-parser/internal/sqlitedb"
"github.com/tiennm99/thptqg/go-parser/internal/transform"
)
// TestToAsciiAgainstRustOutput cross-checks ToAscii against Rust on real data.
//
// A Rust-built database is its own oracle: every row carries ho_ten alongside
// the ho_ten_ascii that Rust derived from it, so the whole table is a
// name -> expected-slug corpus far broader than the 20 hand-picked unit cases.
//
// Point it at a Rust-built database:
//
// GO_PARSER_RUST_DB=/tmp/rust-2016.db go test ./internal/transform/
//
// Skips when unset, so the default suite stays hermetic.
func TestToAsciiAgainstRustOutput(t *testing.T) {
path := os.Getenv("GO_PARSER_RUST_DB")
if path == "" {
t.Skip("GO_PARSER_RUST_DB not set; skipping cross-check against Rust output")
}
if _, err := os.Stat(path); err != nil {
t.Skipf("GO_PARSER_RUST_DB=%s not readable: %v", path, err)
}
db, err := sql.Open("sqlite", "file:"+path+"?mode=ro")
if err != nil {
t.Fatalf("open %s: %v", path, err)
}
defer db.Close()
rows, err := db.Query("SELECT ho_ten, ho_ten_ascii FROM student")
if err != nil {
t.Fatalf("query: %v", err)
}
defer rows.Close()
var checked, bad int
for rows.Next() {
var name, rustAscii string
if err := rows.Scan(&name, &rustAscii); err != nil {
t.Fatalf("scan: %v", err)
}
checked++
if got := transform.ToAscii(name); got != rustAscii {
bad++
if bad <= 5 {
t.Errorf("ToAscii(%q)\n rust = %q\n go = %q", name, rustAscii, got)
}
}
}
if err := rows.Err(); err != nil {
t.Fatalf("iterate: %v", err)
}
if checked == 0 {
t.Fatal("database contained no rows")
}
t.Logf("compared %d real names, %d mismatches", checked, bad)
}
+223
View File
@@ -0,0 +1,223 @@
// Package transform performs row transformation: ASCII normalisation, score
// regex parsing, and validation — a port of parser/src/transform.rs.
//
// ToAscii replicates build-lib.js toAscii exactly:
//
// str.normalize("NFD").replace(/[̀-ͯ]/g,"").replace(/đ/gi,"d").toLowerCase()
package transform
import (
"errors"
"math"
"strconv"
"strings"
"golang.org/x/text/unicode/norm"
"github.com/tiennm99/thptqg/go-parser/internal/config"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
"github.com/tiennm99/thptqg/go-parser/internal/schema"
)
// ToAscii normalises a Vietnamese name to an ASCII slug.
//
// 1. NFD decompose (splits base + combining diacritics)
// 2. Drop combining marks in U+0300..U+036F
// 3. Replace đ/Đ with d (NFD does not decompose them)
// 4. Lowercase
//
// Step 2 filters a LITERAL CODEPOINT RANGE, not a Unicode category. The inline
// comment at transform.rs:53 says "Unicode category M", but the code at :56 is
// the specification and it checks '\u{0300}'..='\u{036f}'. unicode.Is(unicode.Mn, r)
// is strictly broader and would strip marks Rust keeps, silently changing
// ho_ten_ascii — the column the site's accent-insensitive search runs on.
//
// Step order matters and mirrors transform.rs:54-63: the đ/Đ replacement happens
// before lowercasing.
func ToAscii(s string) string {
if s == "" {
return ""
}
decomposed := norm.NFD.String(s)
var b strings.Builder
b.Grow(len(decomposed))
for _, r := range decomposed {
if r >= 0x0300 && r <= 0x036F {
continue // combining mark, in the range Rust drops
}
switch r {
case 'đ', 'Đ':
b.WriteByte('d')
default:
b.WriteRune(r)
}
}
return strings.ToLower(b.String())
}
// ParsedRow is one row ready for insertion.
type ParsedRow struct {
SoBaoDanh string
HoTen string
HoTenAscii string
NgaySinh *string
// TenCumThi is 2016 only: examination cluster name (TEN_CUMTHI column).
TenCumThi *string
// GioiTinh is 2016 only: gender, normalised to "Nam"/"Nữ" or nil.
GioiTinh *string
// Scores maps subject field -> value. Absent subjects are simply missing and
// bind NULL.
Scores map[string]float64
}
// SkipReason says why a row was skipped, or SkipNone when it passed.
//
// The distinction is load-bearing for the printed counters, and the two
// non-blank reasons are counted as source rows while BlankRow is not — but note
// that split lives in the CALLER, not here. parser/src/main.rs has two call
// sites with opposite outcomes for BlankRow: at :135-137 it returns before the
// counter at :140, while at :151 it matches Err(BlankRow) => {} and falls
// through to transform and insert. The build loop must reproduce both.
type SkipReason int
const (
SkipNone SkipReason = iota
// SkipBlankRow: row is fully blank (2017-old2 only, checked before the
// source-row counter).
SkipBlankRow
// SkipEmptyField: so_bao_danh or ho_ten empty/missing.
SkipEmptyField
// SkipNonNumericSbd: so_bao_danh contains non-digit characters
// (2017-old / 2017-old2 guard).
SkipNonNumericSbd
)
func (s SkipReason) String() string {
switch s {
case SkipNone:
return "none"
case SkipBlankRow:
return "blank_row"
case SkipEmptyField:
return "empty_field"
case SkipNonNumericSbd:
return "non_numeric_sbd"
}
return "unknown"
}
// ValidateRow checks a row against the dataset's validation rules.
//
// Signature mirrors transform.rs:101-107, taking stripBlankRows and allBlank
// explicitly; a shorter signature could not express both blank-row paths.
func ValidateRow(hoTen, soBaoDanh string, cfg *config.ValidationCfg, stripBlankRows, allBlank bool) SkipReason {
// 2017-old2: skip fully blank rows BEFORE counting source rows.
if stripBlankRows && allBlank {
return SkipBlankRow
}
if cfg.RequireNonemptySbd && soBaoDanh == "" {
return SkipEmptyField
}
if cfg.RequireNonemptyName && hoTen == "" {
return SkipEmptyField
}
if cfg.RequireNumericSbd && !allASCIIDigits(soBaoDanh) {
return SkipNonNumericSbd
}
return SkipNone
}
// allASCIIDigits mirrors Rust's chars().all(|c| c.is_ascii_digit()).
//
// Deliberately not strconv.Atoi: Atoi accepts a leading sign, so "+123" would
// pass a check Rust rejects. Empty input returns true, matching Rust's all() on
// an empty iterator — the empty case is caught earlier by RequireNonemptySbd.
func allASCIIDigits(s string) bool {
for i := 0; i < len(s); i++ {
if s[i] < '0' || s[i] > '9' {
return false
}
}
return true
}
// ParseScores extracts subject scores from a DIEM_THI cell.
//
// Every one of the 16 patterns runs against every dataset; a subject absent from
// a given exam year never matches and stays NULL. Matching is unanchored
// first-match, like Rust's Regex::captures.
func ParseScores(diemThi string) map[string]float64 {
out := make(map[string]float64)
if diemThi == "" {
return out
}
for field, re := range schema.ScorePatterns {
m := re.FindStringSubmatch(diemThi)
if m == nil {
continue
}
v, err := strconv.ParseFloat(m[1], 64)
if err != nil {
continue
}
// Unreachable given the pattern shape, but kept for parity with
// transform.rs:136 (is_finite).
if math.IsInf(v, 0) || math.IsNaN(v) {
continue
}
out[field] = v
}
return out
}
// ErrNoColumns is returned when the fixed-column path is used on a config that
// has no columns: mapping. Rust panics here via .expect() (transform.rs:162);
// returning an error is the Go-idiomatic equivalent and is unreachable in
// practice, since only the non-2016 path calls this.
var ErrNoColumns = errors.New("transform: config has no columns mapping")
// TransformRow extracts one row into a ParsedRow using fixed column indices,
// the 2017-family path. 2016 uses runtime format detection instead.
func TransformRow(raw []reader.Cell, cfg *config.DatasetConfig) (*ParsedRow, error) {
cols := cfg.Columns
if cols == nil {
return nil, ErrNoColumns
}
// Trimmed accessor, mirroring the closure at transform.rs:163-167. Out-of-range
// indices yield "" rather than an error, matching unwrap_or_default().
get := func(idx int) string {
if idx < 0 || idx >= len(raw) {
return ""
}
return strings.TrimSpace(raw[idx].Str)
}
hoTen := get(cols.HoTen)
ngaySinh := get(cols.NgaySinh)
soBaoDanh := get(cols.SoBaoDanh)
// diem_thi is read WITHOUT trimming (transform.rs:172-175), unlike the three
// fields above. Harmless because the score patterns are unanchored, but it is
// the shipped behaviour — do not "tidy" it.
diemThi := ""
if cols.DiemThi >= 0 && cols.DiemThi < len(raw) {
diemThi = raw[cols.DiemThi].Str
}
var ngaySinhOpt *string
if ngaySinh != "" {
ngaySinhOpt = &ngaySinh
}
return &ParsedRow{
SoBaoDanh: soBaoDanh,
HoTen: hoTen,
HoTenAscii: ToAscii(hoTen),
NgaySinh: ngaySinhOpt,
TenCumThi: nil,
GioiTinh: nil,
Scores: ParseScores(diemThi),
}, nil
}
@@ -0,0 +1,283 @@
package transform
import (
"testing"
"github.com/tiennm99/thptqg/go-parser/internal/config"
"github.com/tiennm99/thptqg/go-parser/internal/reader"
)
// Ports every test in parser/src/transform.rs's test module (:201-409) — 29 in
// total, not the 20 in the :213-315 range, which covers only ToAscii. The nine
// outside that range are the ParseScores and ValidateRow cases, i.e. exactly the
// behaviours this package's traps concern.
// --- ToAscii: the 20 cases at transform.rs:213-315 ---
func TestToAscii(t *testing.T) {
cases := []struct{ name, in, want string }{
{"plain_latin", "Nguyen Van A", "nguyen van a"},
{"nguyen_thi_hoa", "Nguyễn Thị Hoa", "nguyen thi hoa"},
{"tran_van_duc", "Trần Văn Đức", "tran van duc"},
{"le_thi_my_duyen", "Lê Thị Mỹ Duyên", "le thi my duyen"},
{"pham_thi_lan", "Phạm Thị Lan", "pham thi lan"},
{"bui_thi_thu", "Bùi Thị Thu", "bui thi thu"},
{"hoang_van_truong", "Hoàng Văn Trường", "hoang van truong"},
{"do_thi_ngan", "Đỗ Thị Ngân", "do thi ngan"},
{"nguyen_van_khanh", "Nguyễn Văn Khánh", "nguyen van khanh"},
{"trinh_thi_bich_ngoc", "Trịnh Thị Bích Ngọc", "trinh thi bich ngoc"},
{"vu_thi_dieu", "Vũ Thị Diệu", "vu thi dieu"},
{"nguyen_thi_tuong_vi", "Nguyễn Thị Tường Vi", "nguyen thi tuong vi"},
{"lowercase_d_stroke", "đặng thị hằng", "dang thi hang"},
{"uppercase_d_stroke", "ĐẶNG THỊ HẰNG", "dang thi hang"},
{"mixed_case", "NGUYỄN VĂN AN", "nguyen van an"},
{"tran_thi_kim_anh", "Trần Thị Kim Anh", "tran thi kim anh"},
{"nguyen_thi_phuong_thao", "Nguyễn Thị Phương Thảo", "nguyen thi phuong thao"},
{"le_van_long", "Lê Văn Long", "le van long"},
{"vo_thi_xuan_mai", "Võ Thị Xuân Mai", "vo thi xuan mai"},
{"empty_string", "", ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
if got := ToAscii(c.in); got != c.want {
t.Errorf("ToAscii(%q) = %q, want %q", c.in, got, c.want)
}
})
}
}
// TestToAsciiUsesLiteralRangeNotUnicodeMn guards the highest-value trap in this
// package. transform.rs:56 filters the literal range U+0300..U+036F; the inline
// comment at :53 calls it "Unicode category M", but the code is the spec.
// unicode.Mn is strictly broader, so using it would strip marks Rust keeps.
// U+0654 (ARABIC HAMZA ABOVE) is in Mn but outside the range: Rust keeps it.
func TestToAsciiUsesLiteralRangeNotUnicodeMn(t *testing.T) {
const in = "aٔb"
if got := ToAscii(in); got != in {
t.Errorf("ToAscii(%q) = %q — a combining mark outside U+0300..U+036F must survive; "+
"stripping it means unicode.Mn was used instead of the literal range", in, got)
}
// And a mark inside the range must be stripped.
if got := ToAscii("áb"); got != "ab" {
t.Errorf("ToAscii(\"a\\u0301b\") = %q, want \"ab\"", got)
}
}
// TestToAsciiDStrokeIndependentOfNFD proves the đ/Đ replacement is a separate
// step: NFD does not decompose them, so relying on the mark filter alone loses
// the letter entirely.
func TestToAsciiDStrokeIndependentOfNFD(t *testing.T) {
for _, c := range []struct{ in, want string }{
{"đ", "d"}, {"Đ", "d"}, {"đĐ", "dd"},
} {
if got := ToAscii(c.in); got != c.want {
t.Errorf("ToAscii(%q) = %q, want %q", c.in, got, c.want)
}
}
}
// --- ParseScores: transform.rs:323, :332, :342 ---
func TestParseScoresSingle(t *testing.T) {
s := ParseScores("Toán: 8.5")
if v, ok := s["toan"]; !ok || v != 8.5 {
t.Errorf("toan = %v (present=%v), want 8.5", v, ok)
}
if _, ok := s["ngu_van"]; ok {
t.Error("ngu_van should be absent")
}
}
func TestParseScoresMultiple(t *testing.T) {
s := ParseScores("Toán: 7.25 Ngữ văn: 6.0 Vật lí: 9")
for field, want := range map[string]float64{"toan": 7.25, "ngu_van": 6.0, "vat_ly": 9.0} {
if v, ok := s[field]; !ok || v != want {
t.Errorf("%s = %v (present=%v), want %v", field, v, ok, want)
}
}
}
func TestParseScoresEmptyCell(t *testing.T) {
if s := ParseScores(""); len(s) != 0 {
t.Errorf("ParseScores(\"\") = %v, want empty", s)
}
}
// TestParseScoresRealCellShape uses the wide space runs seen in the corpus.
func TestParseScoresRealCellShape(t *testing.T) {
const cell = "Toán: 4.60 Ngữ văn: 5.50 Lịch sử: 4.50 "
s := ParseScores(cell)
if len(s) != 3 {
t.Fatalf("matched %d subjects, want 3: %v", len(s), s)
}
if s["toan"] != 4.60 || s["ngu_van"] != 5.50 || s["lich_su"] != 4.50 {
t.Errorf("got %v", s)
}
}
// --- ValidateRow: transform.rs:359, :365, :374, :383, :393, :400 ---
func defaultValidation() *config.ValidationCfg {
return &config.ValidationCfg{
RequireNumericSbd: false,
RequireNonemptyName: true,
RequireNonemptySbd: true,
}
}
func TestValidateOK(t *testing.T) {
if r := ValidateRow("Nguyen Van A", "12345678", defaultValidation(), false, false); r != SkipNone {
t.Errorf("got %v, want SkipNone", r)
}
}
func TestValidateEmptySbd(t *testing.T) {
if r := ValidateRow("Nguyen Van A", "", defaultValidation(), false, false); r != SkipEmptyField {
t.Errorf("got %v, want SkipEmptyField", r)
}
}
func TestValidateEmptyName(t *testing.T) {
if r := ValidateRow("", "12345678", defaultValidation(), false, false); r != SkipEmptyField {
t.Errorf("got %v, want SkipEmptyField", r)
}
}
func TestValidateNonNumericSbdRejected(t *testing.T) {
v := defaultValidation()
v.RequireNumericSbd = true
if r := ValidateRow("Nguyen Van A", "12AB5678", v, false, false); r != SkipNonNumericSbd {
t.Errorf("got %v, want SkipNonNumericSbd", r)
}
}
func TestValidateNumericSbdAccepted(t *testing.T) {
v := defaultValidation()
v.RequireNumericSbd = true
if r := ValidateRow("Nguyen Van A", "12345678", v, false, false); r != SkipNone {
t.Errorf("got %v, want SkipNone", r)
}
}
func TestValidateBlankRowSkipped(t *testing.T) {
if r := ValidateRow("", "", defaultValidation(), true, true); r != SkipBlankRow {
t.Errorf("got %v, want SkipBlankRow", r)
}
}
// TestValidateNumericSbdIsDigitScanNotAtoi: Rust uses chars().all(is_ascii_digit),
// which strconv.Atoi does not reproduce — Atoi accepts a leading sign, and would
// wrongly admit "+123".
func TestValidateNumericSbdIsDigitScanNotAtoi(t *testing.T) {
v := defaultValidation()
v.RequireNumericSbd = true
for _, sbd := range []string{"+123", "-123", "12 3", "1.0", "ABC123", "123"} {
if r := ValidateRow("Nguyen Van A", sbd, v, false, false); r != SkipNonNumericSbd {
t.Errorf("ValidateRow(sbd=%q) = %v, want SkipNonNumericSbd", sbd, r)
}
}
}
// TestValidateBlankRowOnlyWhenStripEnabled: with strip_blank_rows false, an
// all-blank row falls through to the empty-field checks instead (transform.rs:109).
func TestValidateBlankRowOnlyWhenStripEnabled(t *testing.T) {
if r := ValidateRow("", "", defaultValidation(), false, true); r != SkipEmptyField {
t.Errorf("got %v, want SkipEmptyField when strip_blank_rows is off", r)
}
}
// --- TransformRow ---
func fixedColumnCfg() *config.DatasetConfig {
return &config.DatasetConfig{
Columns: &config.ColumnMap{HoTen: 0, NgaySinh: 1, SoBaoDanh: 2, DiemThi: 3},
Validation: *defaultValidation(),
}
}
func cells(vals ...string) []reader.Cell {
out := make([]reader.Cell, len(vals))
for i, v := range vals {
out[i] = reader.Cell{Str: v, IsEmpty: v == ""}
}
return out
}
func TestTransformRow(t *testing.T) {
row := cells("Nguyễn Văn Đức", "04/04/1999", "51002167", "Toán: 8.5 Ngữ văn: 7")
got, err := TransformRow(row, fixedColumnCfg())
if err != nil {
t.Fatalf("TransformRow: %v", err)
}
if got.HoTen != "Nguyễn Văn Đức" || got.HoTenAscii != "nguyen van duc" {
t.Errorf("ho_ten=%q ascii=%q", got.HoTen, got.HoTenAscii)
}
if got.SoBaoDanh != "51002167" {
t.Errorf("so_bao_danh = %q", got.SoBaoDanh)
}
if got.NgaySinh == nil || *got.NgaySinh != "04/04/1999" {
t.Errorf("ngay_sinh = %v", got.NgaySinh)
}
// 2016-only columns are never populated on the fixed-column path.
if got.TenCumThi != nil || got.GioiTinh != nil {
t.Error("ten_cum_thi and gioi_tinh must stay nil on the 2017 path")
}
if got.Scores["toan"] != 8.5 || got.Scores["ngu_van"] != 7 {
t.Errorf("scores = %v", got.Scores)
}
}
// TestTransformRowEmptyNgaySinhBecomesNil ports transform.rs:179-183.
func TestTransformRowEmptyNgaySinhBecomesNil(t *testing.T) {
got, err := TransformRow(cells("A", "", "1", ""), fixedColumnCfg())
if err != nil {
t.Fatalf("TransformRow: %v", err)
}
if got.NgaySinh != nil {
t.Errorf("empty ngay_sinh should be nil, got %q", *got.NgaySinh)
}
}
// TestTransformRowShortRowYieldsEmptyFields ports the unwrap_or_default()
// behaviour at transform.rs:163-176: a row shorter than the configured indices
// yields empty strings rather than an error.
func TestTransformRowShortRowYieldsEmptyFields(t *testing.T) {
got, err := TransformRow(cells("OnlyName"), fixedColumnCfg())
if err != nil {
t.Fatalf("TransformRow: %v", err)
}
if got.HoTen != "OnlyName" || got.SoBaoDanh != "" || got.NgaySinh != nil {
t.Errorf("got ho_ten=%q sbd=%q ngay_sinh=%v", got.HoTen, got.SoBaoDanh, got.NgaySinh)
}
}
// TestTransformRowDiemThiIsNotTrimmed pins an asymmetry that is easy to
// "tidy away": ho_ten, ngay_sinh and so_bao_danh are trimmed through the closure
// at transform.rs:164-168, but diem_thi is read raw at :172-175.
func TestTransformRowDiemThiIsNotTrimmed(t *testing.T) {
row := cells(" A ", " 01/01/2000 ", " 123 ", " Toán: 5 ")
got, err := TransformRow(row, fixedColumnCfg())
if err != nil {
t.Fatalf("TransformRow: %v", err)
}
if got.HoTen != "A" || got.SoBaoDanh != "123" {
t.Errorf("trimmed fields wrong: ho_ten=%q sbd=%q", got.HoTen, got.SoBaoDanh)
}
if got.NgaySinh == nil || *got.NgaySinh != "01/01/2000" {
t.Errorf("ngay_sinh = %v, want trimmed", got.NgaySinh)
}
// Untrimmed diem_thi still parses — the regexes are unanchored.
if got.Scores["toan"] != 5 {
t.Errorf("scores = %v", got.Scores)
}
}
// TestTransformRowRequiresColumns: the fixed-column path is only reachable when
// the config has a columns: mapping. Rust panics via .expect() (transform.rs:162);
// Go returns an error instead.
func TestTransformRowRequiresColumns(t *testing.T) {
cfg := &config.DatasetConfig{Validation: *defaultValidation()}
if _, err := TransformRow(cells("A", "B", "C", "D"), cfg); err == nil {
t.Fatal("TransformRow without a columns: mapping must return an error")
}
}
+177
View File
@@ -0,0 +1,177 @@
// Package writer handles SQLite output: DDL setup, INSERT OR REPLACE, VACUUM
// and the stats block — a port of parser/src/writer.rs.
//
// Every dataset writes the same canonical table (internal/schema), so there is
// exactly one insert path. Columns a dataset carries no data for bind NULL.
//
// The stats lines are reproduced verbatim because they are the operator-facing
// output of the build, and docs/deployment-guide.md points at the per-file row
// counts for troubleshooting.
package writer
import (
"database/sql"
"fmt"
"os"
"path/filepath"
"github.com/tiennm99/thptqg/go-parser/internal/schema"
"github.com/tiennm99/thptqg/go-parser/internal/sqlitedb"
"github.com/tiennm99/thptqg/go-parser/internal/transform"
)
// OpenDB deletes any existing database at dbPath, recreates it, and executes the
// canonical DDL.
//
// Deleting the file rather than issuing DROP TABLE mirrors build-lib.js:54 via
// writer.rs:24-30. A consequence worth knowing: a concurrent reader sees the file
// vanish mid-rebuild rather than a transactional swap.
func OpenDB(dbPath string) (*sql.DB, error) {
if _, err := os.Stat(dbPath); err == nil {
if err := os.Remove(dbPath); err != nil {
return nil, fmt.Errorf("remove existing db %s: %w", dbPath, err)
}
}
if parent := filepath.Dir(dbPath); parent != "" && parent != "." {
if err := os.MkdirAll(parent, 0o755); err != nil {
return nil, fmt.Errorf("create %s: %w", parent, err)
}
}
db, err := sql.Open(sqlitedb.DriverName, dbPath)
if err != nil {
return nil, fmt.Errorf("open db %s: %w", dbPath, err)
}
if _, err := db.Exec(schema.DDL); err != nil {
db.Close()
return nil, fmt.Errorf("execute DDL: %w", err)
}
return db, nil
}
// Inserter wraps a prepared INSERT statement.
//
// Rust calls conn.execute(INSERT_SQL, ...) per row (writer.rs:73), re-preparing
// each time. Preparing once is a performance choice, not a parity requirement —
// the SQL and its bindings are identical either way.
type Inserter struct{ stmt *sql.Stmt }
// Prepare compiles the canonical INSERT against tx.
func Prepare(tx *sql.Tx) (*Inserter, error) {
stmt, err := tx.Prepare(schema.InsertSQL)
if err != nil {
return nil, fmt.Errorf("prepare insert: %w", err)
}
return &Inserter{stmt: stmt}, nil
}
// Close releases the prepared statement.
func (i *Inserter) Close() error { return i.stmt.Close() }
// Insert binds one parsed row and executes the INSERT.
//
// Parameter order is IdentityFields then ScoreFields. Subjects absent from
// row.Scores — and the two identity columns only the 2016 layouts populate —
// bind NULL.
func (i *Inserter) Insert(row *transform.ParsedRow) error {
args := make([]any, 0, schema.ParamCount)
args = append(args,
row.SoBaoDanh,
row.HoTen,
row.HoTenAscii,
nullableString(row.NgaySinh),
nullableString(row.TenCumThi),
nullableString(row.GioiTinh),
)
for _, field := range schema.ScoreFields {
if v, ok := row.Scores[field]; ok {
args = append(args, v)
} else {
args = append(args, nil)
}
}
if len(args) != schema.ParamCount {
return fmt.Errorf("built %d params, want %d", len(args), schema.ParamCount)
}
if _, err := i.stmt.Exec(args...); err != nil {
return err
}
return nil
}
func nullableString(s *string) any {
if s == nil {
return nil
}
return *s
}
// Stats carries the counters the build loop accumulates.
type Stats struct {
SourceRows uint64
Skipped uint64
Errors uint64
}
// Finish runs VACUUM and prints the stats block.
//
// VACUUM must run AFTER the transaction commits — SQLite refuses it inside one.
//
// The wording branches on datasetLabel, which Rust derives from the input
// directory's basename (main.rs:98-101). That makes the output depend on a
// filesystem path rather than on config; it is reproduced here for parity, and
// the caller passes the label explicitly so tests are not at the mercy of a
// temp-directory name.
func Finish(db *sql.DB, dbPath string, st Stats, datasetLabel string, isOld2 bool) error {
if _, err := db.Exec("VACUUM"); err != nil {
return fmt.Errorf("vacuum: %w", err)
}
var dbCount int64
if err := db.QueryRow("SELECT COUNT(*) FROM student").Scan(&dbCount); err != nil {
return fmt.Errorf("count rows: %w", err)
}
insertable := st.SourceRows - st.Skipped
fmt.Println()
if isOld2 {
fmt.Printf("Source non-blank data rows: %d\n", st.SourceRows)
fmt.Printf(" skipped (empty/non-numeric SBD): %d\n", st.Skipped)
} else {
fmt.Printf("Source data rows (post-header): %d\n", st.SourceRows)
if containsOld(datasetLabel) {
fmt.Printf(" skipped (empty/non-numeric SBD): %d\n", st.Skipped)
} else {
fmt.Printf(" skipped (empty/invalid): %d\n", st.Skipped)
}
}
fmt.Printf(" insertable: %d\n", insertable)
fmt.Printf(" insert errors: %d\n", st.Errors)
fmt.Printf("DB rows (distinct SBD): %d\n", dbCount)
if !containsOld(datasetLabel) && st.Errors == 0 {
gap := int64(insertable) - dbCount
if gap == 0 {
fmt.Println("Audit: OK — every source row made it in.")
} else {
fmt.Printf("Audit: %d row(s) collapsed (duplicate SBDs overwriting).\n", gap)
}
}
var size int64
if fi, err := os.Stat(dbPath); err == nil {
size = fi.Size()
}
fmt.Printf("Size: %.1f MB\n", float64(size)/1024.0/1024.0)
return nil
}
func containsOld(label string) bool {
for i := 0; i+3 <= len(label); i++ {
if label[i:i+3] == "old" {
return true
}
}
return false
}
+126
View File
@@ -0,0 +1,126 @@
#!/usr/bin/env node
/**
* Build the SQLite database for one or all datasets, verify it, then gzip it.
*
* The dataset list comes from src/datasets.js so it is written in exactly one
* place.
*
* Output goes to .build/public/db/ — the directory Vite copies as its publicDir.
* Only the .gz survives: shipping a 100+ MB uncompressed database is made
* structurally impossible rather than left to a cleanup step.
*
* VERIFICATION IS THE POINT OF THIS SCRIPT, not an extra.
*
* Until now nothing between the parser and the public site asserted that a
* database actually had data in it. The parser logs a file-level failure and
* continues, returns success regardless, and finishes cleanly even at zero rows;
* this script gzipped whatever it got; and scripts/assemble-site.js only greps
* *filenames* for stray .db files. So a reader that silently under-produced
* would publish a truncated dataset with green CI and no red signal anywhere.
*
* The guard below closes that: a build whose row count does not match the known
* figure, or whose artifact is implausibly small, fails the pipeline.
*
* Usage:
* node go-parser/scripts/build-db.js # all four datasets
* node go-parser/scripts/build-db.js 2017-old # just one
*/
import { execFileSync } from "node:child_process";
import { mkdirSync, rmSync, existsSync, statSync } from "node:fs";
import { dirname, resolve } from "node:path";
import { fileURLToPath } from "node:url";
import { DatabaseSync } from "node:sqlite";
import { DATASET_IDS, DATASETS } from "../../src/datasets.js";
const ROOT = resolve(dirname(fileURLToPath(import.meta.url)), "../..");
const BIN = resolve(ROOT, "go-parser/bin/xlsxread");
const OUT_DIR = resolve(ROOT, ".build/public/db");
/**
* Known-good row counts, from docs/data-pipeline.md.
*
* The inputs are frozen historical exam results, so these are exact, not
* approximate. A deviation of even one row means something changed that nobody
* intended — treat it as a build failure, not a warning.
*/
const EXPECTED_ROWS = {
"2016": 877461,
"2017": 861068,
"2017-old": 847348,
"2017-old2": 679764,
};
/** A gzipped database far below its usual size means a truncated build. */
const MIN_SIZE_RATIO = 0.9;
const requested = process.argv.slice(2);
const unknown = requested.filter((id) => !DATASET_IDS.includes(id));
if (unknown.length) {
console.error(`unknown dataset(s): ${unknown.join(", ")}`);
console.error(`known: ${DATASET_IDS.join(", ")}`);
process.exit(2);
}
const targets = requested.length ? requested : DATASET_IDS;
if (!existsSync(BIN)) {
console.error(`parser binary not found at ${BIN}`);
console.error("run: npm run build:go");
process.exit(1);
}
mkdirSync(OUT_DIR, { recursive: true });
for (const id of targets) {
const db = resolve(OUT_DIR, `${id}.db`);
// execFileSync throws on a non-zero exit, so a parser failure aborts the run.
execFileSync(
BIN,
[
"build",
"--schema",
resolve(ROOT, `parser/configs/${id}.yml`),
"--input",
resolve(ROOT, `data/${id}`),
"--output",
db,
],
{ stdio: "inherit" },
);
// --- guard: the database must contain what it is supposed to contain ---
const expected = EXPECTED_ROWS[id];
if (expected === undefined) {
console.error(`no expected row count recorded for ${id}; add one to EXPECTED_ROWS`);
process.exit(1);
}
const conn = new DatabaseSync(db, { readOnly: true });
const actual = conn.prepare("SELECT COUNT(*) c FROM student").get().c;
conn.close();
if (actual !== expected) {
console.error(`\n${id}: row count ${actual}, expected ${expected}`);
console.error("Refusing to publish — the build did not reproduce the known dataset.");
process.exit(1);
}
console.log(` ✓ ${id}: ${actual} rows (matches expected)`);
// -9 without -k: the raw .db must not reach the published artifact.
rmSync(`${db}.gz`, { force: true });
execFileSync("gzip", ["-9", db], { stdio: "inherit" });
const gz = `${db}.gz`;
const sizeMb = statSync(gz).size / 1024 / 1024;
const nominal = DATASETS.find((d) => d.id === id)?.dbSizeMb;
if (nominal && sizeMb < nominal * MIN_SIZE_RATIO) {
console.error(
`\n${id}: ${sizeMb.toFixed(1)} MB is below ${(nominal * MIN_SIZE_RATIO).toFixed(1)} MB ` +
`(${MIN_SIZE_RATIO * 100}% of the expected ${nominal} MB)`,
);
console.error("Refusing to publish — the artifact looks truncated.");
process.exit(1);
}
console.log(` → db/${id}.db.gz (${sizeMb.toFixed(1)} MB)\n`);
}
Binary file not shown.
+34
View File
@@ -0,0 +1,34 @@
#!/usr/bin/env bash
# Regenerates the reader-fidelity oracle from the Rust/calamine ground truth.
#
# Emits SHA-256 per input file over the canonical cell dump. Only hashes are
# committed: the dumps are real student PII, and parser/tests/fixtures/README.md
# establishes that fixtures in this repo carry synthetic data only.
#
# Requires the Rust parser to still build. Run from the repo root.
set -euo pipefail
OUT=go-parser/testdata/reader-fidelity-hashes.tsv
CANON='BEGIN{OFS="\t"}
$1=="FILE"{next}
$1=="SHEETCOUNT"{print;next}
$1=="SHEET"{print $1,$2,$3,$4,$5;next}
$1=="ROW"{print;next}
$1=="CELL"{print $1,$2,$3,$4,$7;next}'
{
echo "# Canonical cell-dump SHA-256 per input file, produced by the Rust/calamine"
echo "# ground truth (parser/examples/dump_cells.rs). The Go reader must reproduce"
echo "# each hash exactly. Hashes only - the dumps themselves are real student PII"
echo "# and are never committed, per parser/tests/fixtures/README.md."
echo "# Regenerate: go-parser/scripts/regen-fidelity-hashes.sh"
} > "$OUT"
for f in data/2016/* data/2017/* data/2017-old/* data/2017-old2/*; do
h=$(cargo run --release --quiet --manifest-path parser/Cargo.toml \
--example dump_cells -- "$f" /dev/stdout \
| awk -F'\t' "$CANON" | sha256sum | cut -d' ' -f1)
printf '%s\t%s\n' "$f" "$h" >> "$OUT"
done
echo "wrote $(grep -cv '^#' "$OUT") hashes to $OUT"
+304
View File
@@ -0,0 +1,304 @@
# Canonical cell-dump SHA-256 per input file, produced by the Rust/calamine
# ground truth (parser/examples/dump_cells.rs). The Go reader must reproduce
# each hash exactly. Hashes only - the dumps themselves are real student PII
# and are never committed, per parser/tests/fixtures/README.md.
# Regenerate: go-parser/scripts/regen-fidelity-hashes.sh
data/2016/023718c7d3cf7ace3a7116fabb12bd9cdhyduoccantho-1468920829104.xlsx fc3afd8da9ed28b0fc177192fab7b19a06535cb45d86841835d8e8828fc61c1c
data/2016/08cdeb3636dcb2adaef829d62968274atravinh-1468902443859.xlsx 71f0e61918c63b95f08a16d502f5e9f2054e3c830bd83c725f1285bb1aabead1
data/2016/0b583c9875443e65ffa449d5ca76fae8ninhbinh-1468899778557.xlsx b0c2cfb9ab419442f151be5a014439f7f477fa7b31a9ae3f37b83073276bdb59
data/2016/15c85d901a73dd49f2bae71aadcdbb5cninhthuan-1468920071763.xlsx f27ceabc10484c0124c7d03f8343248df05719486af81ea2fb387d09a2d11b31
data/2016/167116af6c8a096c2a9307edfa115f77dhcnthucpham-tphcm-1468910578187.xls 9962f437ca047cafb3add0c9eb07f96cf4168d026b7d5947ecd2dc915117ef11
data/2016/17843bccde10f362c2b9600bc49cd3f8backan-1468933016076.xlsx ffdabc85670ef7d5522bb33254d6053682b96deb248406e171bd066749c51cae
data/2016/178814c33c8b85a77d0fc295fa6d2be1hagiang-1468932847230.xlsx 843d83e11b78862598eb77b1a5448f7b09f0a7f1842687b4f0174e69ae4fff8b
data/2016/18055d17e46854ecc6c8be585bcb6e5cbaclieu-1468934144150.xlsx ec257a18650045ec476ebbbdbfed48760b4c51b5aee73fe6456729d7cd8ea365
data/2016/1c66db0265bf45df237b37a7bcaeaef3hanoi-1468932790746.xlsx aeb84a0588ce445b852a9328f8a233444a5e5ace6ea0cdff695d96a85cf0015e
data/2016/2004a8d225fd87524e324d2d519e87a4daclac-1468899960505.xlsx 455f5c3f17397139e50b71329e936ffc1c76a39dc244ee72f49289b3eeb5a47c
data/2016/20db5eaace927596f1806f151649ee6edhtantrao-1468901296040.xlsx b25aa4ab69339c095c20d291614fdcd6569f210b33cdebae955e54945e647010
data/2016/216ee4e385dbc236767be899b068bf98dhtaynguyen-1468940490240.xlsx 88346d3fa7087055f4c602bda8456581d721c9f9f2d4e0a18dbc511c9c8288fe
data/2016/2208903e1fa12220d100f93e40d5ff3bquangtri-1468933440097.xlsx 49241f1af4fd0eb84e6a2943f3c07d1c0f7bd71c9a8ea407b5adcff194bcbcbd
data/2016/22e8617886412abd90ae3d33cc6d7db4dhyduocthaibinh-1468941020145.xlsx 9ec94a6c47932bccbf001c431c0d4c480a861cb9aa3f845d9b6e86d4f2688889
data/2016/238adc19e80c0daef089a307af23edcethanhhoa-1468933304999.xlsx 121b6ed30c60a2979db512f55fd2f8e65194c7323b5bf495d8175a256eb8eb2b
data/2016/2b33640b0dd154490f2809d96be81797dhdalat-1468902130616.xlsx 201b953d4a19c0a0441693d43be8f37212b1309084428dc6d1367fb2904a92f6
data/2016/2fa6d9a308bd11bc09f02d787eef2c15lamdong-1468933821082.xlsx 323b5f6551d35d57a995ef04c5d0d569433775fe6e2d685afc88d43442e3d215
data/2016/357c6aaf59aeaa6c11c3e83c595d38cfdhbachkhoa-tphcm-1468939678829.xlsx 873fa6455a85783e4fe54b420c6cbda91803e0fa36356af859236092019f0083
data/2016/388fcdfdac18ef6eee6125a2a015714ddhdongthap-1468934545359.xlsx 6db5b9fa0acd509c1b07bd6a03ae07df19487073a3d18495df3e185e85711189
data/2016/3c0e56abd8c5334f4e1ff7df3ac23a15dienbien-1468920113820.xlsx cfbadef763402bb8ac1e08b398cb38bebc184a2de7e760c44c2e31ce54349a87
data/2016/3daa6ee33fdafadc3c112c18bff18c2ddhthainguyen-1468940268065.xlsx a7d7e20b00b912fdd4658fa73c1fd13efc4a748e01d9bd41ce25d91222e6b29b
data/2016/412e62265f35ca0078a0d94af5d1ed8abinhdinh-1468899884192.xlsx 06d144fca25ec641e5d8bec34df77b69fcea1946120d050d7679bf282a77b1bb
data/2016/43213f61074068d9177685eb6c3f2d3bdhbachkhoa-dhdanang-1468938494437.xlsx 381a59addee12fd6efc7951282578d177101fce5a07be8621b1143f1c99a0e29
data/2016/4334ea476a2b8196d54ba40a341831echvkythuatquansu-1468920342994.xlsx 3705579c6f432e2a8aa053e7f2564f79aad49fde836ad3934d1f9f39375150a4
data/2016/487c7287c1c9735722b31d877942becddhcongnghieptphcm-1468983373303.xlsx 2c50ae0fc1992ad8b61b7d52ad02f32159bd649f0426ae00d17583b53427b05c
data/2016/4a3fa95a83fe9b3b2fafab309323fd22dhthuyloi-1468901840919.xlsx e75239bd4e311ad36e57d42ffb4c2649ce8e599f1ae0bad505957edd69b09a80
data/2016/4ad416ccf014bd5d01f415238566b210khanhhoa-1468933789968.xlsx 7a4451cba4a76acc50362cb94aefee1ff9cb04e6e2461daddc383ff1601c80da
data/2016/4b98fdd7300b86a94636e4d15394fc19sonla-1468899319124.xlsx 961407526e9d79d0e2dfcf370eaed7ff7f64d14835f88481ac17361acbea996c
data/2016/4c2fd846ed2a0670ba6d1200d1d68e1dspkythuatvinhlong-1468902640438.xlsx 88b634010f5888ed6e954807c31712125a280f432219dc652d790d01ef7c10d1
data/2016/4e0ccee19eb64513d23a6dff96ff7f9cvinhphuc-1468919874814.xlsx 52199086d32401b35629d6c4930b024161b0d6d979f4aebc1bef26a391d4b235
data/2016/4e9bb192d56c9fa92b3e0ec13d8236c8dhxaydungmientrung-1468940978346.xlsx dd3d9e4ce4e0c385e753a56f16ce03e7de535a3700e2c62dfe3dce201c73f6f9
data/2016/505a876dfd1ea321291e29192aebe95ehoabinh-1468983220125.xlsx b17f8a592f67c1ed8408b145a2241b4acb6088825463708b5f38749fedfd21be
data/2016/514877c1cfebdc9c2251f651b5f02c7edhspkythuattphcm-1468906568655.xlsx 86cbefa0051faad10505f5d2f5dc4c4690c2c9ccd475fa06713754da9edcc627
data/2016/51bb2f9b6e9f565c2d705606fd56958bdhgiaothongvantai-1468939107463.xlsx b2199701dc1126ee509581dd6fd739d9a913ea6a232a55d59a14fced8ab82cc3
data/2016/538c9676b70f2509573245d58fba93bcdongnai-1468938241935.xlsx 8ff1d5c71cf63b688ecfac852244b24cb6a4321eccf28fe056139f1527f2fb33
data/2016/53e5e20ebd74d84741513b703fbd13bfdhnonglam-dhthainguyen-1468939023637.xlsx b3ac4319785a8f33065f4684b947598733050a916ae9cabaa5684023d42f9163
data/2016/5465cca59495dcd9a48f689e2d38b85ephutho-1468902409682.xlsx 44873acca3ede545505c531b921605fc7b8372c32bff3886572ca072966c898e
data/2016/54bbcf865c137d58cf62af15733bd376dhsupham-dhdanang-1468938526558.xlsx 7e846a06d99ff5b71dbeb3db833d748c7675a17901e3d41bac84a9023e16601d
data/2016/575df8a4df74e87abb8f55d807324374dongthap-1468933902699.xlsx db44859e1fce523fb86907b4928c84f92036e696bdf14a337b1f25f69664389f
data/2016/5aede4951e7af8c1e1056c747a553386hvnongnghiepvn-1468900610577.xlsx 76e75a3d59c5e850f62191dfa3ce65b6d1347b4eaa4f4870c687e94aea5087fd
data/2016/5b16d7156a03bb7ed0618796bae87bc3cantho-1468933954714.xlsx 6d9b5f906fd7b5a66390e42d51f06a3445724e92f4a67cb116ab74359d8693e3
data/2016/5ca241aa1de8d25a3e6ff6db75234025dhcantho-baclieu-1468907071460.xls 290c7e658c3b218befd5b111ad89d8c03f5e5dc1a1815e730da8bc1308660b6c
data/2016/5fb3a1b66aff077ee3da482be08cd391thainguyen-1468902797020.xlsx 6df75f210b0dc269ceb4ddbbe323640cc32d75157096942ff4eb7565157d4d7a
data/2016/5fd13ae5287cbf0149bd733b507e0608dhkhxhnv-dhqgtphcm-1468920630193.xlsx 311b90c29f089ffc7bfdc7e5b9c7fdcefaa803bb6d785189cebb6b170b61ee1e
data/2016/6375fd13557bee08d20f1f3e46dac130dhnhatrang-coso1-1468901350425.xlsx 05dc9a62f844c3a73f0a781066ba9fc3f3bf5faa970cb1b3f0f0f0487769488e
data/2016/64d2635efd30a0d0db4f7c30b9735261dhvinh-1468940230027.xlsx 0e2fc0f487f3963cd063bcb3fac067d2af1ef5a73e5d3def3523e1ecc30851e2
data/2016/64e44466ffeff9a587260795644805d0dhkhtunhien-dhqgtphcm-1468921015146.xlsx 4f67d25acd11182b60d8c938ee49102fb30e8c96c85e6ccecd260fca0f5d53aa
data/2016/6960b6fa493e1365a183af461e945389thaibinh-1468920019273.xlsx 6f10581918d96d82a99eab4c2729672836e091265c4518996f4aad5687038f3a
data/2016/698b930ab7697a0672bbc39168024c9cgialai-1468933753098.xlsx 3be46dfad3d9d2c81557e1bc4103622bb7d9e9ba4fa18b25d2a4b355bf617264
data/2016/6badc7cfa258b59ad67e2b64f713b95fthuathienhue-1468983258438.xlsx 1dccf8266a62a50911046837d01cd07f6085d43dc1c3cc953a35bf033c4991b3
data/2016/70bf8cd2b6e87523219e55149a8524c5dhquynhon-1468938878328.xlsx 874dfbb664c72a3b075be56adc386dba2ac9aea4be2bb7a11d250b67d67b964d
data/2016/74145979ee30186bf4e40ee3e6a3ac74bentre-1468934010040.xlsx 6db50bd7918d17ab04bd4a8a30859719bb1f3e7c5aa6e3bbe8d4f49d4c2bcf49
data/2016/7418e83ff9d07b146216a3beafa927bbdhtravinh-1468902716174.xlsx eadf00e204586e628a12390d335627c65cec48cb0e6ed53f5cc383bac73de402
data/2016/7435bb054e8e4ec34c173af1b6ad8402kiengiang-1468900147395.xlsx 271fa0d823b0d07f32c0fb9b885513e465eeb4880e8fbcc72eedd20759c7d11d
data/2016/74e6cd0dc78a3fb7dcf56ae57283b9b9dhtaybac-1468940302052.xlsx 9b0a007cb2f55d5c2a629d1c9bfa7f16547e3811f3d5c256bd1377b2013e7856
data/2016/76f19098881e9c79f66ac4f1b42a8a55dhsuphamhanoi-1468901144203.xlsx 1a955af97d97f57d9dc20b5c406ca146753760e03050c02c74bef4d62a39ec21
data/2016/778ab074ed6a95d209fa901fd5d68843dhhungvuong-1468901964127.xlsx 0add20e99589c83ec61e21582a965692148d899b4df0182c4a4d6d2ce7696109
data/2016/7afd5d7fdff7c170a6342cd8f9474b22langson-1468932984587.xlsx 8f4014d1d2bf4e02b3136d6790297af8352f8148dfecace502c4cbaa59b6ca7d
data/2016/7b5d74564ab9f6e0e137f6d161fab6f0dhxaydung-1468940561682.xlsx b95d40c1b9497224dd8d59f6e7aeea31a618deba14e43e57539cf32ae748e61d
data/2016/7cfc91bd6a1f5ae9d7917e920f469b90haiduong-1468899640007.xlsx 5abbd89e8130078fe42ecee222327add52e4dbd362c2f96576d9e6ed7c441c19
data/2016/7e0d0b6fb981d407cffdb79d724e8aecbacninh-1468919984998.xlsx 3f32fd1af47a2b1f92175bcc55d95c48be1de1d9544e06f440ac36a8491b1a14
data/2016/7f12b90225e91f9d00aac3bcdf9974e1nghean-1468920043121.xlsx a4c5a99a9faa75e36b809d1bfa4b7caab706bde26649256ed1248f429e59abe6
data/2016/7f1cbe7f35d1f539ca817a0ab8bd00c2dhcantho-haugiang-1468907167921.xls 434417e4765c5b225570c4a19b655c5ac957707209af398f7277a0a9ee90f2b6
data/2016/7f4362ef567cdf74495368e0fc11072avinhlong-1468934067981.xlsx de063663b5f31f525a8a07958d436c8e10805d9ec7b27f9bff2b4a2c7ae559f0
data/2016/82338a4a93ea1e0cde26e0061066ca8adhcongnghiephanoi-1468920154757.xlsx fed3686286093866bfe03b67d7f72ab7c3aaa1a5ce92d503d1ca24582addee31
data/2016/86a692bebf9b0b7c58adb9876b9befd5dhnganhangtphcm-1468902244566.xlsx 5a2b2e5e63bb8c12f0e16b802b51fac470650fc050acd9e846275e82a5d33fbc
data/2016/87ffaf3197ba20f0919f98a0b3f33914dhkhoahoc-dhhue-1468938729472.xlsx 1b46f645bb5c491ba9e9e490f36f1692b055e2d29878dfff1fa4db94fabad72f
data/2016/8916da7b3c4d58cffc1dd29dc7a8aabfdhkinhte-dhhue-1468938560205.xlsx 332aa93e333430021755667c10795749db3ec3f7c7d3afc4d70a48fbd00bc820
data/2016/8af8f1b29b6a40e26ccb4a2e19e8eededhkythuatcongnghiep-dhthainguyen-1468901919568.xlsx 94347288ba444f19a160868649528dc29633f4eb660aad74ea9799ee79935432
data/2016/8f948c2a5ab49a3c5fd03b6ea04887a7dhkinhtequocdanfix-1468902083782.xlsx f8e4c450509b176716429055e2ed72dd340af7ad5a5704db2d0c60818deae548
data/2016/919276a0348634459f8937dfcd6c7129dhhue-1468938792115.xlsx cb91ae55927a40d0eef9ef1bc2bd23a758528de6b36496505dba75e1cccdb849
data/2016/92045214159cb0335b2836aa1c4bb25ahungyen-1468899732002.xlsx 74536aad02570a17c83933c46d10ce9d2300a5e61a37fa2a67e5d99cffb17632
data/2016/964c4368131f6be6058fabb1caefba1chaugiang-1468934247324.xlsx 1df076df14fcb51daffa600dca05064be61269fe19f33dc14d03c8f43d8ee1ca
data/2016/96c95c85c5549a06f2e92ffcacf8ff21hvtaichinh-1468934470924.xlsx 9a7a3841700b27e4265eef8def9addf17406d08c4f007462cda9335f5369c4ff
data/2016/97aa741a4c2b1276559eb021779daa32yenbai-1468933116055.xlsx aa79b6e84e95f58704b715b701449507dbd079f9bee9900f20e6c90baf299952
data/2016/9ba098a7913ad837a66e173b3fc41fd7dhngoaingu-dhdanang-1468902750424.xlsx 8864a16b969ea34dc1f7b8762b8039e17df1599bfea7ca5914dc518983fd36fe
data/2016/9c730c43bf4309171528c3853ed53b19quangbinh-1468933376032.xlsx b637ecd24a240e4a73f291cc4c9faef42379c04d40d1952cf84e54e51eb8cdf1
data/2016/9d78f6dce7965fb4b7a337279920d4b1dhtonducthang-1468902168623.xlsx 1c4f32eb3ed9aa831d2abe41998f5c157ae5ca3d9735b87bb101913870f4c4cc
data/2016/9e0e564ae801e347bd7649da8b89ad90dhluathn-1468920377978.xlsx 1a1b4b0bed838f2c6043d789dae4ab7f96a904b886cd5990a98f7be3b960caf5
data/2016/a11e707e6ed5fe4d29a8e32380d7558cdhsphanoi-2-1468902006284.xlsx 9b8060368d09cf89c1e88001a642d1aec413e4b3f8646c8360a0459d3704dfd3
data/2016/a31b3446bae5a17c517348d8c38c5b7equangninh-1468919920197.xlsx 8b284c18aab150349d49fe596cb96a0a838e8c5618eb1f2d7dd5f0053430c71d
data/2016/a3aa106104f114c6426deb2d5e5da5fdlaichau-1468932954079.xlsx 0697ae494d6a17e188138e2c9f3d412a1bbd8a462e2eb82060ffff25e9982506
data/2016/a8ff80d9478747e36d21ec95cc9ecb8fsoctrang-1468902686409.xlsx 3fb1d35e7647d0581f08d2f56b4d7ccd5fab1ab125bdcb3b83c64c65d540a53b
data/2016/a995de3f7a7055cf10b2f002f0e194f7dhsptphcm-1468920792291.xlsx 9f25b647f9d4d714772ffcb1cba40fa1b338500bd08e39352704de6a9810a8df
data/2016/ab2f259ef55cb08b3a436d3f15d42cc7dhsupham-dhhue-1468934431828.xlsx 96f40bdf397bdf98e596701bfed957d18483ec0affb7e945f6e8b69453b7015c
data/2016/af6828f837a80e03d77a65f3d310ec60dh-tai-chinh-marketing-1468924711596.xlsx 297b71a7760cfff612a290f31082f22647c0d7a8528dbe052356885fc620dc5e
data/2016/b2552e20c9ff4b237244dfbf2e278a47dhthuongmai-1468902876264.xlsx 8c7f9d58fb5ffdfafb1431e9fd51347d338ccba14f547aa26917a2692a515982
data/2016/b2a739584a8861502573d1c6372fd6e9dhnonglamtphcm-1468939555326.xlsx d8c8b6e10d3a44f8a429d2a718af55c853bd508dcb79d13522f1e40d9d6e96e7
data/2016/b34c777942ca8a4de91f23a35cce1c6bdhbachkhoahn-1468901491486.xlsx b5135f12ac53b3f513b46078a89b6a2663089e6bb3ad97d48a13312b5ac69ca6
data/2016/b35a3f26a162bec3740e2da5639327afdhcantho-1468907131146.xls d114bf1c070bb165bd45a64ec79495c764805ee989c740ed95fad06eabd69ec4
data/2016/b42fd58ead4c54279d7df5f53160a851dhdanang-1468920187203.xlsx 9ce9d1b6f494866f6a39e68babe27b3a9562cd1515477f5f32a8e2a8a57bf13e
data/2016/b7a0c1ecce7c8445b3a1021091a5e654bacgiang-1468983153932.xlsx a53266ee75a8d49fe5a017125e198a1368a9bc9259dabe8cdf1cad11551dd439
data/2016/bacd1b2133b37d0ea9fbf46445265036dhtiengiang-1468940391597.xlsx bc90ae7bf5611ae0bb3c321e5185afcb530e1331c5a920a40b85b30b47eb1006
data/2016/bd857add0cfe92c1c247469ff2104527dh-hong-duc-1468900427119.xlsx 96b62a2954d2c1541699d491ea931c240a9b06660d5c27a3cd9055433306aaef
data/2016/bf259a908ff8a8dcc067a761204459f4daknong-1468934173803.xlsx 2f8fb2dbcd9a06acfd91becba0f728bf56e64f5c9984b40eac060834ea619b44
data/2016/c174631d90303276ac63ff38b31157bbdhmodiachat-1468900697349.xlsx c5c4435cfe1e6feede1936fd11e5cd3ff5758b4c4d522df03facf7522a9e7194
data/2016/c2cfaaf287fb1e7120894e545771b027dhkhoahoc-dhthainguyen-1468900365545.xlsx 9a35e56e631ba5d1056fb149ea01a41b4e62ff075937fa735666d192b1714da2
data/2016/c5c28959fcd0b4e39dd9fd538db5448adhgtvttphcm-1468939150821.xlsx 17a032a01a9ebe9f577af8fa4666e8700e9ab4e97166703cf6a7a0ff9dea4e9b
data/2016/c5ecc58a731fd1a54b7e6cc3bc99449ddhlamnghiep-1468939300790.xlsx 8f5f7bf26bafb3a8dd6988ab9858dbf03a14d200a23ed5299a8a5411c84d7f1b
data/2016/c6ef16bc9f75ce0fd16c1229afad4e71dhhanghai-1468906457479.xlsx 8e43d5e1782e366344c8d615666541931319475e02df0be29ae290c7ab2abda1
data/2016/cc57dc986acb525086b86e85df54949dkontum-1468933632113.xlsx 2397d3b0fca3888ff34d4166a9d3da70dcd6986e6405bd556471d46fff0ab551
data/2016/cdc912664d7217c75eb9b67c6ef3797adhkiengiang-1468901222446.xlsx 460eb7e1ce79dde123fe53a18034c5617f0616ad1abce6dedfcd0e4a2957a28d
data/2016/d035f470c13adbf5785797c6ccdcb204laocai-1468919805337.xlsx d23348d36bf6fa886d50e30507d4eada114a58f52cb6e615c119980d7753b5ad
data/2016/d5ecd2574b5ab36c62dffad435262cb8tuyenquang-1468902296807.xlsx 3acb7e4ba9a3c210906b9a32372db747ba2d95d25d276ba3ba45046b941dbbef
data/2016/d652d24e8bd4fe6a48d634b116f37c27dhluattphcm-1468939398928.xlsx f519010d9b2631457c9b7784bb77f07979907321199b98b3f97cb45e5b20c2ff
data/2016/d75282bff0a4d8cd1421c3fb74f789c3dhsaigon-1468900867561.xlsx ee7734bac44228a58a11b85d490cbd5da27d7a6e07b22c1eb9921a123b3e6acd
data/2016/ddb631f833af894200ef10cd408d2ef0caobang-1468932877769.xlsx 2bbda20bab1190f42b44d0eeabcf2be7f2203d4e5f8a2ad3d80365770a8e2d9e
data/2016/deca197a916a8f633f6894b4e13c3ca4dhhaiphong-1468901178368.xlsx 221fe731a4b2de02b3952572baf48d66544fddf87dfdd7851b740325c0864d24
data/2016/dfa63c5b68a8f638379ca0e9c5a1c556namdinh-1468933234061.xlsx 446cd9626fe197f6509efb01ca5b4cb26730907b5e114c4b9c07399a489b45d1
data/2016/e693bc9de67acc916032c6a3b2fc9eedhanam-1468933146526.xlsx b1ef9283b500e965205281d3875eff011be1d418312d187831925860b1ddab74
data/2016/e80dc028d53fa49ee247391a0653dee5quangngai-1468933592400.xlsx 9c6ba937e8ad772a3782baf9c40ef781121c27ffdeb4519170cbf2b81545a146
data/2016/ec765c32190773f5a18cd8202a5eaf0dquangnam-1468933501467.xlsx a566488b611ca04c705f385c47b871792ce75f7c1ca793d0528220352d6b49e1
data/2016/ec84067268f47d523b1e80b493248377dhngoaithuong-1468920480753.xlsx 5d892f08ec818be3ceb343529c9f0b35e0af0aeb43a229446dccaa87c6d4c846
data/2016/ee09c723da4e86cdbe26213413450af2dhsupham-dhthainguyen-1468939061872.xlsx aefac40009116615131c926775e6625eeebc615cd2ff4c7c9af399dd2faab982
data/2016/f0aa5a19c2c8c1fde8211341f0e98e26dhkinhtetphcm-1468939229043.xlsx 5ce36388e304cbba01c513f11f54b940f4299061e2e8acfa4c4219ee536a01f2
data/2016/f2846894bd5533e02b122a1e3b4333f6hvnganhang-1468939472850.xlsx 142b39b4f77c1d7a648c2397165b315acdd48ae6d354d914dd674a46b3d7160c
data/2016/f85f04bb2bbd84457b83ac846728a2a6dhspkythuathungyen-1468920667091.xlsx f8cc571e94720652f1507576ff48616d5331a4ed1403cb95cdc30124a03c3a68
data/2016/fc511852659150f6305f6f624445300ddhangiang-1468940189321.xlsx 4e16ea6a122e5bc0c62c7e6e55d9ede3a41242dd3e034d92a450bf29e8fff197
data/2016/ff877788e43bd84f0119dae026996892dhkinhteluat-dhqgtphcm-1468939860894.xlsx 8c9e7b097d0796750995268aa36a9fa00112c0989b513023ef0a53456afd7051
data/2017/an-giang.xls 5bde5f6447b9e27540f4f72d9b371b6c94376aeb97af65f4ec5eede43ba9f25c
data/2017/bac-giang.xls 9cd5c6d858062a1fb30c477b699be467139f183745e66d134746e043d73579a4
data/2017/bac-kan.xls 9642c00d14c54102ce8dc38408676863a87652aee89651113f4ad3e0a25bd1fa
data/2017/bac-lieu.xls e364dd65fa068d40e9a390aa3322e1676171af797c5d1f98949277cadab5369e
data/2017/bac-ninh.xls 539ed3cb1a2df358ab169e2163face76c949aac2eca0d2acb7ceb51e97811bc5
data/2017/ba-ria-vung-tau.xls 78dc6d67bbf3ffcb7e4407a6d950e28e39366ed62617691c9a9db610a5079134
data/2017/ben-tre.xls b4cce0c5a443062a5e7aef95785fc56f421c21d5ad22f993cbc30ec28e078f4f
data/2017/binh-dinh.xls fa1a39cb5134a8159e71dabea41c9ef2a3ec263642f19bed54ddd17e5a5e3d98
data/2017/binh-duong.xls ced0d7469186d9e5fbbb8c4323a4bb6a589abb8ef23f01d8d39f947a575af4ba
data/2017/binh-phuoc.xls d1ecf9b66ea294eec24ac0f703c614a2011b7d693aeb08c719cfcb73cae7ca47
data/2017/binh-thuan.xls 3cb0f85889ada81cd257f51ae0b0b6414d2ad8f81d7f94f4b1b38010600243c6
data/2017/ca-mau.xls c4bba31f7728f7825c20b3c0b708ef977dcf6f56f0dd5c14d9342cea0467e505
data/2017/can-tho.xls b5ba59ff7b87c31859686adb98da8ef2e95de84404daff2ac3df1e00b748b57d
data/2017/cao-bang.xls 9b0a497ed16840a22af34f7e75e316b00dd60490d6d7f8a2f012507c95fad3e7
data/2017/dak-lak.xls aae3e074d91afd98d89cc9d07bc1f8debe4559bd78c409cc5f50f2670a2c337a
data/2017/dak-nong.xls f7e30b4cdfe951086aeb1c169e7aad88a17bf9fe07aa33f85654e309899f6beb
data/2017/da-nang.xls 80963ff2f56d2b6bf331546d08184e325ad0ec7432ba948c0f74845f440e87c8
data/2017/dien-bien.xls 6bcb1969ce81ccb94c6ee625456530b9661ee255c00eb598dc195611a46205c8
data/2017/dong-nai.xls af248453ea27e879da1b8cfa0f2cd345aa068f93c884d1bc98862db0ca2f9faa
data/2017/dong-thap.xls 186a0c4ac6bf838cc496c4604a7e8979712051c71a53622a9ec72af70491e4f7
data/2017/gia-lai.xls 3a1b64c8362cad028ee53f4fd18b0923315d919b117df9a2fd83d2f46544db98
data/2017/ha-giang.xls a3248393c759a41de424e3ed994d2b9298cb39fda3c6eb9a68f48025c5a2e072
data/2017/hai-duong.xls 6634715f5a0cbac6a8d2d18be0df74e0abd328641721692224758c730fd86e51
data/2017/hai-phong.xls 7a8a210986de8c562a9128dccfaed1c6fc6c32df0c5e2e8948bee2f73777ba99
data/2017/ha-nam.xls f725e212f976af00e68edd6552721e56b30d5fa7d786f548dfc34721ec8b4426
data/2017/ha-noi.xls f0f4bc9216e421acf655a40844a4dcb74e871b188cf6d455982ee6e408170ef7
data/2017/ha-tinh.xls b0f5bcd8beff7cfe411b7bd88f9e5b9c5a9a0578a55cf18bb003284aca12e57f
data/2017/hau-giang.xls 9e8bccd0738bac3d68640768a8d9abc3e7a6a4e97e8c5968039530cf313cafb9
data/2017/hoa-binh.xls 51b1e910bfc47404b411cb2ab1ff14b522d8fe54d66a1d50479350273c7be1b5
data/2017/ho-chi-minh.xls c4e2f920a5e58fa62bd0913d994d44079e472b955adfc504b734c6773c485f64
data/2017/hung-yen.xls 833584c8ea9337734feb37e068bb06232293108d50b18d88f731d30a576cc29a
data/2017/khanh-hoa.xls 644ab06abca93d7b966205604313a65d50f3a0a6fb8be509cadb55e0d55fe12e
data/2017/kien-giang.xls df452e618544a9b7932d918d26dd2a0f2149f92599609c15a9ac373291cfc440
data/2017/kon-tum.xls 7d26ff078adaf51dcf4333b074f8a444d328d8083835bd55b1034c1a63e4eb17
data/2017/lai-chau.xls b37cd06914cad19bc0e523d0b1cfe87db692e26dfab327f5f9a7686b2b14dae9
data/2017/lam-dong.xls 59cc802a6c4d622f1bb6233f6b9b1a94f342a6eae1c389235383db1767f34720
data/2017/lang-son.xls ae376e75bf615483b6cb3a37daa1cea6a7eec4cf3fdd1186801f56d074eb4ce7
data/2017/lao-cai.xls 7596c99851c20ff60b95347bc9fae78014873b4c3e1f1f7144c48d9f285f7962
data/2017/long-an.xls e52d8e8a4592258e87c5cda04e0907541f60e7e7ef4cb21e9171b65283e0b87a
data/2017/nam-dinh.xls 80dcda8b058dcf1a7fdcd7e34d126c44c31b2d4987b52eeae790bf6c38c58e40
data/2017/nghe-an.xls 3f32421eb6ee8bc21fcedd4cc1043df1c9c0381abfaa1006c85c3ca6d6209543
data/2017/ninh-binh.xls dc3d08a86cdd145e870486773d95a1d896a003b5400ba6ff8456a5512b10b2b8
data/2017/ninh-thuan.xls d69b0ed7192e85cff31ad5bca8942fe3c92322268d1ea707e945f9991bd78115
data/2017/phu-tho.xls e352a51109a2373c1efbb6726d262fc490bcbf9229ae90bd473d14a4b6eebcc0
data/2017/phu-yen.xls 6871aedd90d2c6c87ad6409164a9b26f6f89f97cdd7a1dffbbe4e6879f2348c6
data/2017/quang-binh.xls 6c5e9b3c6780a5b70a02dd4b4854861ba148613361c048bfc001706e5bb77a61
data/2017/quang-nam.xls 859258cb8c4b7806101039e4e2484e56971f2640a65720a108a0933245bd3157
data/2017/quang-ngai.xls 834988a6484fab64bc59fe35a6813038663054553d269d16b7d8200593f468e8
data/2017/quang-ninh.xls 1b1c0b1ea1d6aadd8bce5faef6b5fdc5181daaa9bd44df6acbf06c9021a0ade9
data/2017/quang-tri.xls 35e401061962c6d80593d684f179e1f82b05f884a6a470d1bc17b2d31151e05f
data/2017/soc-trang.xls 668a0edcb624faa954c424616ea6c35d6f06b9e9fc62524723dea01af09490c0
data/2017/son-la.xls 96bcfeb35e5bc23ccd7dc6a62a533e21dabeb3145473ecc5a039d970019a1129
data/2017/tay-ninh.xls aedf4ccab43e7a1fa985721de0d6053988ad471355c0ccb0b8f9c19af02423e2
data/2017/thai-binh.xls b20f66cdeb975a64eed7322b616bd291fe34e08e5fbedad3d729d21da0eba627
data/2017/thai-nguyen.xls f0b93db9aab966d393b8a3323204ba7f147ddd0a92746355b81df09bc58267b4
data/2017/thanh-hoa.xls 3f803ae67dac24317fa45dac60f36e39dc222ff84d23a0a3dde7751cdcb17646
data/2017/thua-thien-hue.xls 71e220ee4ea0287657fa379c587ca93f71f96b4dacc48cef074f8c223e940daa
data/2017/tien-giang.xls 2e061f8724e1ea8b997d7e6700ac89a60921771a57dffa7297f4728715811406
data/2017/tra-vinh.xls 59507a1d237ac5e29c1033d42b08da6ac9a7505a5045b791e512e91493102ae7
data/2017/tuyen-quang.xls 053d293e4acd9dfa6ce9532a047cc64b4d2edee1ff13b387723586c1fae6312a
data/2017/vinh-long.xls 6e7584a3047b0ba7487dbb8c71f5b594db220f11b747149d41d31ea97e1e4eab
data/2017/vinh-phuc.xls 9753e2d864584e76aef1e690e4e2de98b28240fbba810e614eab9926668d66d2
data/2017/yen-bai.xls fe6824a1254116bc6ed372e0a1fd635a816a09e010b7efde9659baca2e143138
data/2017-old/10_BinhThuan_RIIW.xls.xlsx 2833e7366ea1a21ea869978a6bb920b033160bb8042fe0bfe9b7573dcfe85409
data/2017-old/10_Ca_Mau_BKXT.xls.xlsx 53cb69b13325cb9cc9004b8c89f0cd1c85b5f943e7c78c8b010ddabad3f816c9
data/2017-old/10_LamDong_GNFT.xls.xlsx f05b4cdc8a9a8413b29801fbddeb48a7c60d3c6bb2228bbe238113495e89bc98
data/2017-old/10_Soc_Trang_XCGJ.xls.xlsx 8e540cef6f2ff387f3e825a5aa3190af394c6f9c7e9dfc66d5f7032ee47f7f14
data/2017-old/11_BinhDuong_RYQL.xls.xlsx 11953116054a2889929c3fe05b68e8932c7b79ffd1b85d697afaac63fcdd9adf
data/2017-old/11_LaoCai_GMSU.xls.xlsx 5a73becf1bdb8027102faff74035d1df71e13c1cac39cc38ec6efaab6699ca52
data/2017-old/12_BenTre_DKWF.xls.xlsx 579c0d94b20440f1c60a94fb0c762b71dd1c7719c8aa4db0b9418ecaba9b605e
data/2017-old/12_LongAn_ZZUK.xls.xlsx c4cf569706bfc30582d774ba9e699dcbc4bdc734f46bf6493df58eaecc37a755
data/2017-old/13_NamDinh_ESEL.xls.xlsx 7e5256cb095262fef5399092f8f843c1ee82e117bca8e6571e363646bb1170a6
data/2017-old/13_TraVinh_LKUJ.xls.xlsx 2b7321d68a4a910c4a506d7ca207531cfbdd9619f1d8ba0ce475baf1b91043aa
data/2017-old/14_NgheAn_BSLY.xls.xlsx 034234a8a5461910f8beff3977aafa29a6541c8a56c007816118e5c4eb2b7218
data/2017-old/15_PhuTho_ABWQ.xls.xlsx 9891c17001cfcc49ccd587419339698ef59b2dd43b5ca7f2765b09a17c26659d
data/2017-old/16_QuangBinh_KGEU.xls.xlsx 01e249d45c1269d62ae1977bef0bb192f833941510471b723e57c32adb0eadaa
data/2017-old/17_QuangNam_AMTK.xls.xlsx 4c9f05d24d861ec3c82a1c564e4baa1541bd4882785f143459a1a81a6b8359c7
data/2017-old/18_QuangNgai_KOFP.xls.xlsx ee9c26f2c16d36dabfde79901a054d2266c28c4acbfd19b628cfe50f8eb03957
data/2017-old/19_QuangTri_OMZF.xls.xlsx 6c961bb461274c73b4589de4087624a9ef5edd43c04ea54675148df00f6420e5
data/2017-old/1_BaRia_VungTau_HJKG.xls.xlsx cb03f933442a8cea6b1622da4e5f804ab476dff9f3462b404e957b6a2e87b2cd
data/2017-old/1_Da_Nang_AHWJ.xls.xlsx 59892fc01b6cb10497679f2bc69f59023215755d2f00fd764b23b895c98be2be
data/2017-old/1_Ha_Noi_CVXG.xls.xlsx 2718fe9c2b58843621c0082da5880a4b957266525b83123af02c78fd51d4f4a5
data/2017-old/1_Son_La_JIDP.xls.xlsx 0a0e273c5def980406b6316cceb63e5cee53aae56e112dfae7e795261eb1cd06
data/2017-old/1_TuyenQuang_JBYF.xls.xlsx 45826ed10248268ac3344d0b592cf2dba6f86278b7402dc60626a76b18fa4d6b
data/2017-old/20_TayNinh_ILFA.xls.xlsx 684a1820d02f658904c87c26aa195faa6aceb3d7ca00c674df11a63f9274ab14
data/2017-old/21_ThaiBinh_FTVG.xls.xlsx 1a12e43d4ea368b8e3a2e4d881b5d66214a16140219e80e2e72fcb269899b285
data/2017-old/22_ThaiNguyen_TLTW.xls.xlsx 29227bdc3204591327e803680525d8621228fc371ce7f802dc667ec70eb8d7b3
data/2017-old/23_HaiPhong_HXBV.xls.xlsx e450e2344fbd724555de9892412f521414fc1fcc53e69887c6381dd266b4e92c
data/2017-old/24_HCM_XULN.xls.xlsx 49ed22e823607b78920a24f5999eaf1ef8634a141108dfd4bf970ee000aecbf9
data/2017-old/2_BacKan_GFVQ.xls.xlsx 77a08fc4edb41d639fcdd798980903519181639db00a360ef0c6c6c2bd88f529
data/2017-old/2_Ha_Giang_QNCM.xls.xlsx 6bf4d62413c1b92f88f980e5e3e414a7e25971c7546c6390c7e67a6b0f742b68
data/2017-old/2_Ninh_Thuan_VHLY.xls.xlsx 6dfb6517f9ccc3420c222b6a88aabac54f6f25f6b743d993d6a5f28b4fa78ff2
data/2017-old/2_Thanh_Hoa_AUFV.xls.xlsx 2f2ce90425c19a9089024061f328f37b9ef580ddadc076ddcf452f549690f4b3
data/2017-old/2_VinhPhuc_GUDK.xls.xlsx fc0e4952445ebda1af28f7330aa808c032f2e3a7a3c03dfbb9692bf6a248867a
data/2017-old/3_BacGiang_TOIF.xls.xlsx 8043510ff89f11ba063f898dc1d2c8879a8da9063a420308647e20a2f39d3092
data/2017-old/3_BinhPhuoc_YFMU.xls.xlsx 8869970f4e6005381264f3f4bb508c5181f15ae990a75fe904a65b305c0a123a
data/2017-old/3_Cao_Bang_CIEY.xls.xlsx 94a06bdb2333a9f44216001ba8d8463700b8cf3e51944b94db11eef7eac8dbd3
data/2017-old/3_Dong_Thap_GSXQ.xls.xlsx 0cdd0763fc4f20f1d990621ca9e8dab107a3d4eb7c0ec3e53a4d8c240fbc569f
data/2017-old/3_Thua_Thien_Hue_HMDB.xls.xlsx 62e54eebda7bbd542ca63443bb923a81b472e7892e55dce39fd0611fc0f00abd
data/2017-old/4_An_Giang_JNOS.xls.xlsx 778ad205bbabb7c4aea3df34325770aa4a1daec0a1683f73f0405ac1cc5aaf9a
data/2017-old/4_BacNinh_STLR.xls.xlsx 009c3b7821f55210a86f9ace74c31eb0d5c16aafdb0cac8f2b71cac616abdfec
data/2017-old/4_Binh_Dinh_WWZW.xls.xlsx fbacd256d58627377d698462d69aac0e83d124c8c8ff1f5b63ba0bfb5442dd39
data/2017-old/4_DienBien_SOJG.xls.xlsx 4931e8d6604013df48b2b469fda3eb05fdf1203b67bbc2ae4ff6d7949a115a24
data/2017-old/4_Lang_Son_NYQL.xls.xlsx cc7d2454caadb18740ed6baf6edc62baa47e98b3a907b3ce06f72b9a16b7838a
data/2017-old/5_Bac_Lieu_XEKH.xls.xlsx dc1dca45fd5357b45abc930b0ebb95ab589ef80a09bef8b5ad4dab9e00be87e5
data/2017-old/5_Gia_Lai_ABZI.xls.xlsx 5b7be25274d628c284c37efb3df0587171cc9273dc3bd8dc824c7bd1b8bccf1a
data/2017-old/5_HaiDuong_WWWG.xls.xlsx 1d800d0aa7338fb507c4eeb3e752620d2608c30ac820895ca086325dbaf3643d
data/2017-old/5_Hanam_QGJS.xls.xlsx cfc233ac37ab8e263b6cc569815ea2c687da30364f139109032c8f2cac1ad98e
data/2017-old/5_Yen_Bai_FAQR.xls.xlsx 4c8915417d8747c291f643ea4b378c87982c397460e0e9e3cc6df19670d6720f
data/2017-old/6_Dong_Nai_WOTM.xls.xlsx 9b582a8421c419156c2eab6e532d854736385735f8208f9d390c251f22ac764e
data/2017-old/6_Hau_Giang_KWDM.xls.xlsx c11473e3e1f964302c5378c90dd6b9d478eb11397351b7ef74df9ae709c9e908
data/2017-old/6_HoaBinh_TPZY.xls.xlsx e1210961dc821642f6951cd6e4a71b5fbc81d5070ee3816b7c8b5dc56bf2b234
data/2017-old/6_NinhBinh_IGFT.xls.xlsx 9399b8ab58e2a7a1c0ebf5425359319e567f8a28602349aa587efced72142fc8
data/2017-old/6_Quang_Ninh_DQCJ.xls.xlsx 8cf95259b11d48267f710f151c3f1419ccac49d823b78b70a2708c59ab25a4f7
data/2017-old/7_HaTinh_DDHD.xls.xlsx d96df9f2c8365dfc73be3e10c268367b93fd00fe92bb39c745d8e353ba75c05b
data/2017-old/7_HungYen_LTIK.xls.xlsx e80eb38af8307c8e193a91bdda5368c437693d7669e281b543bbe891a46e6f36
data/2017-old/7_Kon_Tum_RSLR.xls.xlsx 41d2adea059449b6e214241414748f10776861d261cbaa1fd213629757deb846
data/2017-old/7_Tien_Giang_EFHX.xls.xlsx 00d9bbc70be89f8f7dad6aff990ad36236c6b7244ddeb2e3d9b3d8ed1a6e21d0
data/2017-old/8_Can_Tho_RQZM.xls.xlsx 3a1182b26d34fbdfabf4326098ff265609d8f1fe2b766b6eef897db9685d4cd4
data/2017-old/8_Dak_Lak_YKPR.xls.xlsx 379b45f8a347bf4842daf9402d84be47426ed7107a64b2e7710bdc8e71eb57be
data/2017-old/8_KienGiang_OTOB.xls.xlsx 3e1789b8565a5ae7c187598cb7e9c3261029c06bc7f0cccc4eb503e8e302cb18
data/2017-old/8_PhuYen_OGIM.xls.xlsx 91d1f202cba3bf4c57a7f26e7611d3d29faf8dd13de30552cbfc96eab7a451e9
data/2017-old/9_DakNong_FBOP.xls.xlsx c611bf7f1d1c95cbfc78bdab3ea99bf8278e94c8db09a9fd3ccac55a538fe0b2
data/2017-old/9_Khanh_Hoa_KPKQ.xls.xlsx 0e90c87dd74a5eb293a0a71b93da45653a19ed2e3a2600d3b880a867a695f31d
data/2017-old/9_LaiChau_ALKN.xls.xlsx 3b7a022f7727134620c6756925bd5adb041e2d80cabd4aab13a4063501739514
data/2017-old/9_Vinh_Long_OLWA.xls.xlsx 43d7abb9c6b8d3205ce90f9cceaf9623875c87998d9e030c6dc8da9bc18bfa09
data/2017-old2/10.BinhThuan_MVVG.xlsx 2833e7366ea1a21ea869978a6bb920b033160bb8042fe0bfe9b7573dcfe85409
data/2017-old2/10.LamDong_YUQA.xlsx f05b4cdc8a9a8413b29801fbddeb48a7c60d3c6bb2228bbe238113495e89bc98
data/2017-old2/10.Soc Trang_LQWU.xlsx 8e540cef6f2ff387f3e825a5aa3190af394c6f9c7e9dfc66d5f7032ee47f7f14
data/2017-old2/11.BinhDuong_HVAH.xlsx 11953116054a2889929c3fe05b68e8932c7b79ffd1b85d697afaac63fcdd9adf
data/2017-old2/11.LaoCai_ZTBP.xlsx 5a73becf1bdb8027102faff74035d1df71e13c1cac39cc38ec6efaab6699ca52
data/2017-old2/12.BenTre_NQTU.xlsx 579c0d94b20440f1c60a94fb0c762b71dd1c7719c8aa4db0b9418ecaba9b605e
data/2017-old2/12.LongAn_PDRH.xlsx d1c65184b475fc29ae273cc89b8574b7e273bc3dab83cd39489d72a5de1bb1e4
data/2017-old2/13.NamDinh_NAYR.xlsx 7e5256cb095262fef5399092f8f843c1ee82e117bca8e6571e363646bb1170a6
data/2017-old2/13.TraVinh_FODZ.xlsx 2b7321d68a4a910c4a506d7ca207531cfbdd9619f1d8ba0ce475baf1b91043aa
data/2017-old2/14.NgheAn_HTKD.xlsx 034234a8a5461910f8beff3977aafa29a6541c8a56c007816118e5c4eb2b7218
data/2017-old2/15.PhuTho_IJZW.xlsx 9891c17001cfcc49ccd587419339698ef59b2dd43b5ca7f2765b09a17c26659d
data/2017-old2/17.QuangNam_NQMG.xlsx 4c9f05d24d861ec3c82a1c564e4baa1541bd4882785f143459a1a81a6b8359c7
data/2017-old2/18.QuangNgai_IUPY.xlsx ee9c26f2c16d36dabfde79901a054d2266c28c4acbfd19b628cfe50f8eb03957
data/2017-old2/19.QuangTri_MKNN.xlsx 6c961bb461274c73b4589de4087624a9ef5edd43c04ea54675148df00f6420e5
data/2017-old2/1.BaRia-VungTau_PGZT.xlsx cb03f933442a8cea6b1622da4e5f804ab476dff9f3462b404e957b6a2e87b2cd
data/2017-old2/1.Da Nang_ABWU.xlsx 59892fc01b6cb10497679f2bc69f59023215755d2f00fd764b23b895c98be2be
data/2017-old2/1.Son La_XLFN.xlsx 0a0e273c5def980406b6316cceb63e5cee53aae56e112dfae7e795261eb1cd06
data/2017-old2/1.TuyenQuang_PTMR.xlsx 45826ed10248268ac3344d0b592cf2dba6f86278b7402dc60626a76b18fa4d6b
data/2017-old2/20.TayNinh_KJAQ.xlsx 684a1820d02f658904c87c26aa195faa6aceb3d7ca00c674df11a63f9274ab14
data/2017-old2/21.ThaiBinh_KTQN.xlsx 1a12e43d4ea368b8e3a2e4d881b5d66214a16140219e80e2e72fcb269899b285
data/2017-old2/22.ThaiNguyen_BKIF.xlsx 29227bdc3204591327e803680525d8621228fc371ce7f802dc667ec70eb8d7b3
data/2017-old2/24.HCM_UTLQ.xlsx 49ed22e823607b78920a24f5999eaf1ef8634a141108dfd4bf970ee000aecbf9
data/2017-old2/2.BacKan_YQNX.xlsx 77a08fc4edb41d639fcdd798980903519181639db00a360ef0c6c6c2bd88f529
data/2017-old2/2.Ha Giang_PIYK.xlsx 6bf4d62413c1b92f88f980e5e3e414a7e25971c7546c6390c7e67a6b0f742b68
data/2017-old2/2.Ninh Thuan_BAGG.xlsx 6dfb6517f9ccc3420c222b6a88aabac54f6f25f6b743d993d6a5f28b4fa78ff2
data/2017-old2/2.Thanh Hoa_UOPE.xlsx 2f2ce90425c19a9089024061f328f37b9ef580ddadc076ddcf452f549690f4b3
data/2017-old2/2.VinhPhuc_QZJK.xlsx fc0e4952445ebda1af28f7330aa808c032f2e3a7a3c03dfbb9692bf6a248867a
data/2017-old2/3.BacGiang_SAVS.xlsx 8043510ff89f11ba063f898dc1d2c8879a8da9063a420308647e20a2f39d3092
data/2017-old2/3.BinhPhuoc_IPHL.xlsx 8869970f4e6005381264f3f4bb508c5181f15ae990a75fe904a65b305c0a123a
data/2017-old2/3.Cao Bang_WMUU.xlsx 94a06bdb2333a9f44216001ba8d8463700b8cf3e51944b94db11eef7eac8dbd3
data/2017-old2/3.Dong Thap_HKJX.xlsx 0cdd0763fc4f20f1d990621ca9e8dab107a3d4eb7c0ec3e53a4d8c240fbc569f
data/2017-old2/3.Thua Thien -Hue_MAET.xlsx 62e54eebda7bbd542ca63443bb923a81b472e7892e55dce39fd0611fc0f00abd
data/2017-old2/4.An Giang_PMJD.xlsx 778ad205bbabb7c4aea3df34325770aa4a1daec0a1683f73f0405ac1cc5aaf9a
data/2017-old2/4.BacNinh_NNIS.xlsx 009c3b7821f55210a86f9ace74c31eb0d5c16aafdb0cac8f2b71cac616abdfec
data/2017-old2/4.Binh Dinh_VOMJ.xlsx fbacd256d58627377d698462d69aac0e83d124c8c8ff1f5b63ba0bfb5442dd39
data/2017-old2/4.DienBien_FYGN.xlsx 4931e8d6604013df48b2b469fda3eb05fdf1203b67bbc2ae4ff6d7949a115a24
data/2017-old2/4.Lang Son_QWOG.xlsx cc7d2454caadb18740ed6baf6edc62baa47e98b3a907b3ce06f72b9a16b7838a
data/2017-old2/5.Bac Lieu_VIVY.xlsx dc1dca45fd5357b45abc930b0ebb95ab589ef80a09bef8b5ad4dab9e00be87e5
data/2017-old2/5.Gia Lai_TAAS.xlsx 5b7be25274d628c284c37efb3df0587171cc9273dc3bd8dc824c7bd1b8bccf1a
data/2017-old2/5.HaiDuong_WNHD.xlsx 1d800d0aa7338fb507c4eeb3e752620d2608c30ac820895ca086325dbaf3643d
data/2017-old2/5.Hanam_SDKN.xlsx cfc233ac37ab8e263b6cc569815ea2c687da30364f139109032c8f2cac1ad98e
data/2017-old2/5.Yen Bai_BSLV.xlsx 4c8915417d8747c291f643ea4b378c87982c397460e0e9e3cc6df19670d6720f
data/2017-old2/6.Hau Giang_SIAJ.xlsx c11473e3e1f964302c5378c90dd6b9d478eb11397351b7ef74df9ae709c9e908
data/2017-old2/6.HoaBinh_HLYQ.xlsx e1210961dc821642f6951cd6e4a71b5fbc81d5070ee3816b7c8b5dc56bf2b234
data/2017-old2/6.NinhBinh_PKMQ.xlsx 9399b8ab58e2a7a1c0ebf5425359319e567f8a28602349aa587efced72142fc8
data/2017-old2/6.Quang Ninh_YAKJ.xlsx 8cf95259b11d48267f710f151c3f1419ccac49d823b78b70a2708c59ab25a4f7
data/2017-old2/7.HaTinh_XMFQ.xlsx d96df9f2c8365dfc73be3e10c268367b93fd00fe92bb39c745d8e353ba75c05b
data/2017-old2/7.HungYen_TCBE.xlsx e80eb38af8307c8e193a91bdda5368c437693d7669e281b543bbe891a46e6f36
data/2017-old2/7.Kon Tum_CAQU.xlsx 41d2adea059449b6e214241414748f10776861d261cbaa1fd213629757deb846
data/2017-old2/7.Tien Giang_AOWZ.xlsx 00d9bbc70be89f8f7dad6aff990ad36236c6b7244ddeb2e3d9b3d8ed1a6e21d0
data/2017-old2/8.PhuYen_MLTQ.xlsx 9dd53293f99d86801692a1995a5cf7a08af8394b820c76e68c88c2b7644b9b67
data/2017-old2/9.DakNong_ZWJT.xlsx c611bf7f1d1c95cbfc78bdab3ea99bf8278e94c8db09a9fd3ccac55a538fe0b2
data/2017-old2/9.Khanh Hoa_IBTX.xlsx 0e90c87dd74a5eb293a0a71b93da45653a19ed2e3a2600d3b880a867a695f31d
data/2017-old2/9.Vinh Long_WVKI.xlsx 43d7abb9c6b8d3205ce90f9cceaf9623875c87998d9e030c6dc8da9bc18bfa09
1 # Canonical cell-dump SHA-256 per input file, produced by the Rust/calamine
2 # ground truth (parser/examples/dump_cells.rs). The Go reader must reproduce
3 # each hash exactly. Hashes only - the dumps themselves are real student PII
4 # and are never committed, per parser/tests/fixtures/README.md.
5 # Regenerate: go-parser/scripts/regen-fidelity-hashes.sh
6 data/2016/023718c7d3cf7ace3a7116fabb12bd9cdhyduoccantho-1468920829104.xlsx fc3afd8da9ed28b0fc177192fab7b19a06535cb45d86841835d8e8828fc61c1c
7 data/2016/08cdeb3636dcb2adaef829d62968274atravinh-1468902443859.xlsx 71f0e61918c63b95f08a16d502f5e9f2054e3c830bd83c725f1285bb1aabead1
8 data/2016/0b583c9875443e65ffa449d5ca76fae8ninhbinh-1468899778557.xlsx b0c2cfb9ab419442f151be5a014439f7f477fa7b31a9ae3f37b83073276bdb59
9 data/2016/15c85d901a73dd49f2bae71aadcdbb5cninhthuan-1468920071763.xlsx f27ceabc10484c0124c7d03f8343248df05719486af81ea2fb387d09a2d11b31
10 data/2016/167116af6c8a096c2a9307edfa115f77dhcnthucpham-tphcm-1468910578187.xls 9962f437ca047cafb3add0c9eb07f96cf4168d026b7d5947ecd2dc915117ef11
11 data/2016/17843bccde10f362c2b9600bc49cd3f8backan-1468933016076.xlsx ffdabc85670ef7d5522bb33254d6053682b96deb248406e171bd066749c51cae
12 data/2016/178814c33c8b85a77d0fc295fa6d2be1hagiang-1468932847230.xlsx 843d83e11b78862598eb77b1a5448f7b09f0a7f1842687b4f0174e69ae4fff8b
13 data/2016/18055d17e46854ecc6c8be585bcb6e5cbaclieu-1468934144150.xlsx ec257a18650045ec476ebbbdbfed48760b4c51b5aee73fe6456729d7cd8ea365
14 data/2016/1c66db0265bf45df237b37a7bcaeaef3hanoi-1468932790746.xlsx aeb84a0588ce445b852a9328f8a233444a5e5ace6ea0cdff695d96a85cf0015e
15 data/2016/2004a8d225fd87524e324d2d519e87a4daclac-1468899960505.xlsx 455f5c3f17397139e50b71329e936ffc1c76a39dc244ee72f49289b3eeb5a47c
16 data/2016/20db5eaace927596f1806f151649ee6edhtantrao-1468901296040.xlsx b25aa4ab69339c095c20d291614fdcd6569f210b33cdebae955e54945e647010
17 data/2016/216ee4e385dbc236767be899b068bf98dhtaynguyen-1468940490240.xlsx 88346d3fa7087055f4c602bda8456581d721c9f9f2d4e0a18dbc511c9c8288fe
18 data/2016/2208903e1fa12220d100f93e40d5ff3bquangtri-1468933440097.xlsx 49241f1af4fd0eb84e6a2943f3c07d1c0f7bd71c9a8ea407b5adcff194bcbcbd
19 data/2016/22e8617886412abd90ae3d33cc6d7db4dhyduocthaibinh-1468941020145.xlsx 9ec94a6c47932bccbf001c431c0d4c480a861cb9aa3f845d9b6e86d4f2688889
20 data/2016/238adc19e80c0daef089a307af23edcethanhhoa-1468933304999.xlsx 121b6ed30c60a2979db512f55fd2f8e65194c7323b5bf495d8175a256eb8eb2b
21 data/2016/2b33640b0dd154490f2809d96be81797dhdalat-1468902130616.xlsx 201b953d4a19c0a0441693d43be8f37212b1309084428dc6d1367fb2904a92f6
22 data/2016/2fa6d9a308bd11bc09f02d787eef2c15lamdong-1468933821082.xlsx 323b5f6551d35d57a995ef04c5d0d569433775fe6e2d685afc88d43442e3d215
23 data/2016/357c6aaf59aeaa6c11c3e83c595d38cfdhbachkhoa-tphcm-1468939678829.xlsx 873fa6455a85783e4fe54b420c6cbda91803e0fa36356af859236092019f0083
24 data/2016/388fcdfdac18ef6eee6125a2a015714ddhdongthap-1468934545359.xlsx 6db5b9fa0acd509c1b07bd6a03ae07df19487073a3d18495df3e185e85711189
25 data/2016/3c0e56abd8c5334f4e1ff7df3ac23a15dienbien-1468920113820.xlsx cfbadef763402bb8ac1e08b398cb38bebc184a2de7e760c44c2e31ce54349a87
26 data/2016/3daa6ee33fdafadc3c112c18bff18c2ddhthainguyen-1468940268065.xlsx a7d7e20b00b912fdd4658fa73c1fd13efc4a748e01d9bd41ce25d91222e6b29b
27 data/2016/412e62265f35ca0078a0d94af5d1ed8abinhdinh-1468899884192.xlsx 06d144fca25ec641e5d8bec34df77b69fcea1946120d050d7679bf282a77b1bb
28 data/2016/43213f61074068d9177685eb6c3f2d3bdhbachkhoa-dhdanang-1468938494437.xlsx 381a59addee12fd6efc7951282578d177101fce5a07be8621b1143f1c99a0e29
29 data/2016/4334ea476a2b8196d54ba40a341831echvkythuatquansu-1468920342994.xlsx 3705579c6f432e2a8aa053e7f2564f79aad49fde836ad3934d1f9f39375150a4
30 data/2016/487c7287c1c9735722b31d877942becddhcongnghieptphcm-1468983373303.xlsx 2c50ae0fc1992ad8b61b7d52ad02f32159bd649f0426ae00d17583b53427b05c
31 data/2016/4a3fa95a83fe9b3b2fafab309323fd22dhthuyloi-1468901840919.xlsx e75239bd4e311ad36e57d42ffb4c2649ce8e599f1ae0bad505957edd69b09a80
32 data/2016/4ad416ccf014bd5d01f415238566b210khanhhoa-1468933789968.xlsx 7a4451cba4a76acc50362cb94aefee1ff9cb04e6e2461daddc383ff1601c80da
33 data/2016/4b98fdd7300b86a94636e4d15394fc19sonla-1468899319124.xlsx 961407526e9d79d0e2dfcf370eaed7ff7f64d14835f88481ac17361acbea996c
34 data/2016/4c2fd846ed2a0670ba6d1200d1d68e1dspkythuatvinhlong-1468902640438.xlsx 88b634010f5888ed6e954807c31712125a280f432219dc652d790d01ef7c10d1
35 data/2016/4e0ccee19eb64513d23a6dff96ff7f9cvinhphuc-1468919874814.xlsx 52199086d32401b35629d6c4930b024161b0d6d979f4aebc1bef26a391d4b235
36 data/2016/4e9bb192d56c9fa92b3e0ec13d8236c8dhxaydungmientrung-1468940978346.xlsx dd3d9e4ce4e0c385e753a56f16ce03e7de535a3700e2c62dfe3dce201c73f6f9
37 data/2016/505a876dfd1ea321291e29192aebe95ehoabinh-1468983220125.xlsx b17f8a592f67c1ed8408b145a2241b4acb6088825463708b5f38749fedfd21be
38 data/2016/514877c1cfebdc9c2251f651b5f02c7edhspkythuattphcm-1468906568655.xlsx 86cbefa0051faad10505f5d2f5dc4c4690c2c9ccd475fa06713754da9edcc627
39 data/2016/51bb2f9b6e9f565c2d705606fd56958bdhgiaothongvantai-1468939107463.xlsx b2199701dc1126ee509581dd6fd739d9a913ea6a232a55d59a14fced8ab82cc3
40 data/2016/538c9676b70f2509573245d58fba93bcdongnai-1468938241935.xlsx 8ff1d5c71cf63b688ecfac852244b24cb6a4321eccf28fe056139f1527f2fb33
41 data/2016/53e5e20ebd74d84741513b703fbd13bfdhnonglam-dhthainguyen-1468939023637.xlsx b3ac4319785a8f33065f4684b947598733050a916ae9cabaa5684023d42f9163
42 data/2016/5465cca59495dcd9a48f689e2d38b85ephutho-1468902409682.xlsx 44873acca3ede545505c531b921605fc7b8372c32bff3886572ca072966c898e
43 data/2016/54bbcf865c137d58cf62af15733bd376dhsupham-dhdanang-1468938526558.xlsx 7e846a06d99ff5b71dbeb3db833d748c7675a17901e3d41bac84a9023e16601d
44 data/2016/575df8a4df74e87abb8f55d807324374dongthap-1468933902699.xlsx db44859e1fce523fb86907b4928c84f92036e696bdf14a337b1f25f69664389f
45 data/2016/5aede4951e7af8c1e1056c747a553386hvnongnghiepvn-1468900610577.xlsx 76e75a3d59c5e850f62191dfa3ce65b6d1347b4eaa4f4870c687e94aea5087fd
46 data/2016/5b16d7156a03bb7ed0618796bae87bc3cantho-1468933954714.xlsx 6d9b5f906fd7b5a66390e42d51f06a3445724e92f4a67cb116ab74359d8693e3
47 data/2016/5ca241aa1de8d25a3e6ff6db75234025dhcantho-baclieu-1468907071460.xls 290c7e658c3b218befd5b111ad89d8c03f5e5dc1a1815e730da8bc1308660b6c
48 data/2016/5fb3a1b66aff077ee3da482be08cd391thainguyen-1468902797020.xlsx 6df75f210b0dc269ceb4ddbbe323640cc32d75157096942ff4eb7565157d4d7a
49 data/2016/5fd13ae5287cbf0149bd733b507e0608dhkhxhnv-dhqgtphcm-1468920630193.xlsx 311b90c29f089ffc7bfdc7e5b9c7fdcefaa803bb6d785189cebb6b170b61ee1e
50 data/2016/6375fd13557bee08d20f1f3e46dac130dhnhatrang-coso1-1468901350425.xlsx 05dc9a62f844c3a73f0a781066ba9fc3f3bf5faa970cb1b3f0f0f0487769488e
51 data/2016/64d2635efd30a0d0db4f7c30b9735261dhvinh-1468940230027.xlsx 0e2fc0f487f3963cd063bcb3fac067d2af1ef5a73e5d3def3523e1ecc30851e2
52 data/2016/64e44466ffeff9a587260795644805d0dhkhtunhien-dhqgtphcm-1468921015146.xlsx 4f67d25acd11182b60d8c938ee49102fb30e8c96c85e6ccecd260fca0f5d53aa
53 data/2016/6960b6fa493e1365a183af461e945389thaibinh-1468920019273.xlsx 6f10581918d96d82a99eab4c2729672836e091265c4518996f4aad5687038f3a
54 data/2016/698b930ab7697a0672bbc39168024c9cgialai-1468933753098.xlsx 3be46dfad3d9d2c81557e1bc4103622bb7d9e9ba4fa18b25d2a4b355bf617264
55 data/2016/6badc7cfa258b59ad67e2b64f713b95fthuathienhue-1468983258438.xlsx 1dccf8266a62a50911046837d01cd07f6085d43dc1c3cc953a35bf033c4991b3
56 data/2016/70bf8cd2b6e87523219e55149a8524c5dhquynhon-1468938878328.xlsx 874dfbb664c72a3b075be56adc386dba2ac9aea4be2bb7a11d250b67d67b964d
57 data/2016/74145979ee30186bf4e40ee3e6a3ac74bentre-1468934010040.xlsx 6db50bd7918d17ab04bd4a8a30859719bb1f3e7c5aa6e3bbe8d4f49d4c2bcf49
58 data/2016/7418e83ff9d07b146216a3beafa927bbdhtravinh-1468902716174.xlsx eadf00e204586e628a12390d335627c65cec48cb0e6ed53f5cc383bac73de402
59 data/2016/7435bb054e8e4ec34c173af1b6ad8402kiengiang-1468900147395.xlsx 271fa0d823b0d07f32c0fb9b885513e465eeb4880e8fbcc72eedd20759c7d11d
60 data/2016/74e6cd0dc78a3fb7dcf56ae57283b9b9dhtaybac-1468940302052.xlsx 9b0a007cb2f55d5c2a629d1c9bfa7f16547e3811f3d5c256bd1377b2013e7856
61 data/2016/76f19098881e9c79f66ac4f1b42a8a55dhsuphamhanoi-1468901144203.xlsx 1a955af97d97f57d9dc20b5c406ca146753760e03050c02c74bef4d62a39ec21
62 data/2016/778ab074ed6a95d209fa901fd5d68843dhhungvuong-1468901964127.xlsx 0add20e99589c83ec61e21582a965692148d899b4df0182c4a4d6d2ce7696109
63 data/2016/7afd5d7fdff7c170a6342cd8f9474b22langson-1468932984587.xlsx 8f4014d1d2bf4e02b3136d6790297af8352f8148dfecace502c4cbaa59b6ca7d
64 data/2016/7b5d74564ab9f6e0e137f6d161fab6f0dhxaydung-1468940561682.xlsx b95d40c1b9497224dd8d59f6e7aeea31a618deba14e43e57539cf32ae748e61d
65 data/2016/7cfc91bd6a1f5ae9d7917e920f469b90haiduong-1468899640007.xlsx 5abbd89e8130078fe42ecee222327add52e4dbd362c2f96576d9e6ed7c441c19
66 data/2016/7e0d0b6fb981d407cffdb79d724e8aecbacninh-1468919984998.xlsx 3f32fd1af47a2b1f92175bcc55d95c48be1de1d9544e06f440ac36a8491b1a14
67 data/2016/7f12b90225e91f9d00aac3bcdf9974e1nghean-1468920043121.xlsx a4c5a99a9faa75e36b809d1bfa4b7caab706bde26649256ed1248f429e59abe6
68 data/2016/7f1cbe7f35d1f539ca817a0ab8bd00c2dhcantho-haugiang-1468907167921.xls 434417e4765c5b225570c4a19b655c5ac957707209af398f7277a0a9ee90f2b6
69 data/2016/7f4362ef567cdf74495368e0fc11072avinhlong-1468934067981.xlsx de063663b5f31f525a8a07958d436c8e10805d9ec7b27f9bff2b4a2c7ae559f0
70 data/2016/82338a4a93ea1e0cde26e0061066ca8adhcongnghiephanoi-1468920154757.xlsx fed3686286093866bfe03b67d7f72ab7c3aaa1a5ce92d503d1ca24582addee31
71 data/2016/86a692bebf9b0b7c58adb9876b9befd5dhnganhangtphcm-1468902244566.xlsx 5a2b2e5e63bb8c12f0e16b802b51fac470650fc050acd9e846275e82a5d33fbc
72 data/2016/87ffaf3197ba20f0919f98a0b3f33914dhkhoahoc-dhhue-1468938729472.xlsx 1b46f645bb5c491ba9e9e490f36f1692b055e2d29878dfff1fa4db94fabad72f
73 data/2016/8916da7b3c4d58cffc1dd29dc7a8aabfdhkinhte-dhhue-1468938560205.xlsx 332aa93e333430021755667c10795749db3ec3f7c7d3afc4d70a48fbd00bc820
74 data/2016/8af8f1b29b6a40e26ccb4a2e19e8eededhkythuatcongnghiep-dhthainguyen-1468901919568.xlsx 94347288ba444f19a160868649528dc29633f4eb660aad74ea9799ee79935432
75 data/2016/8f948c2a5ab49a3c5fd03b6ea04887a7dhkinhtequocdanfix-1468902083782.xlsx f8e4c450509b176716429055e2ed72dd340af7ad5a5704db2d0c60818deae548
76 data/2016/919276a0348634459f8937dfcd6c7129dhhue-1468938792115.xlsx cb91ae55927a40d0eef9ef1bc2bd23a758528de6b36496505dba75e1cccdb849
77 data/2016/92045214159cb0335b2836aa1c4bb25ahungyen-1468899732002.xlsx 74536aad02570a17c83933c46d10ce9d2300a5e61a37fa2a67e5d99cffb17632
78 data/2016/964c4368131f6be6058fabb1caefba1chaugiang-1468934247324.xlsx 1df076df14fcb51daffa600dca05064be61269fe19f33dc14d03c8f43d8ee1ca
79 data/2016/96c95c85c5549a06f2e92ffcacf8ff21hvtaichinh-1468934470924.xlsx 9a7a3841700b27e4265eef8def9addf17406d08c4f007462cda9335f5369c4ff
80 data/2016/97aa741a4c2b1276559eb021779daa32yenbai-1468933116055.xlsx aa79b6e84e95f58704b715b701449507dbd079f9bee9900f20e6c90baf299952
81 data/2016/9ba098a7913ad837a66e173b3fc41fd7dhngoaingu-dhdanang-1468902750424.xlsx 8864a16b969ea34dc1f7b8762b8039e17df1599bfea7ca5914dc518983fd36fe
82 data/2016/9c730c43bf4309171528c3853ed53b19quangbinh-1468933376032.xlsx b637ecd24a240e4a73f291cc4c9faef42379c04d40d1952cf84e54e51eb8cdf1
83 data/2016/9d78f6dce7965fb4b7a337279920d4b1dhtonducthang-1468902168623.xlsx 1c4f32eb3ed9aa831d2abe41998f5c157ae5ca3d9735b87bb101913870f4c4cc
84 data/2016/9e0e564ae801e347bd7649da8b89ad90dhluathn-1468920377978.xlsx 1a1b4b0bed838f2c6043d789dae4ab7f96a904b886cd5990a98f7be3b960caf5
85 data/2016/a11e707e6ed5fe4d29a8e32380d7558cdhsphanoi-2-1468902006284.xlsx 9b8060368d09cf89c1e88001a642d1aec413e4b3f8646c8360a0459d3704dfd3
86 data/2016/a31b3446bae5a17c517348d8c38c5b7equangninh-1468919920197.xlsx 8b284c18aab150349d49fe596cb96a0a838e8c5618eb1f2d7dd5f0053430c71d
87 data/2016/a3aa106104f114c6426deb2d5e5da5fdlaichau-1468932954079.xlsx 0697ae494d6a17e188138e2c9f3d412a1bbd8a462e2eb82060ffff25e9982506
88 data/2016/a8ff80d9478747e36d21ec95cc9ecb8fsoctrang-1468902686409.xlsx 3fb1d35e7647d0581f08d2f56b4d7ccd5fab1ab125bdcb3b83c64c65d540a53b
89 data/2016/a995de3f7a7055cf10b2f002f0e194f7dhsptphcm-1468920792291.xlsx 9f25b647f9d4d714772ffcb1cba40fa1b338500bd08e39352704de6a9810a8df
90 data/2016/ab2f259ef55cb08b3a436d3f15d42cc7dhsupham-dhhue-1468934431828.xlsx 96f40bdf397bdf98e596701bfed957d18483ec0affb7e945f6e8b69453b7015c
91 data/2016/af6828f837a80e03d77a65f3d310ec60dh-tai-chinh-marketing-1468924711596.xlsx 297b71a7760cfff612a290f31082f22647c0d7a8528dbe052356885fc620dc5e
92 data/2016/b2552e20c9ff4b237244dfbf2e278a47dhthuongmai-1468902876264.xlsx 8c7f9d58fb5ffdfafb1431e9fd51347d338ccba14f547aa26917a2692a515982
93 data/2016/b2a739584a8861502573d1c6372fd6e9dhnonglamtphcm-1468939555326.xlsx d8c8b6e10d3a44f8a429d2a718af55c853bd508dcb79d13522f1e40d9d6e96e7
94 data/2016/b34c777942ca8a4de91f23a35cce1c6bdhbachkhoahn-1468901491486.xlsx b5135f12ac53b3f513b46078a89b6a2663089e6bb3ad97d48a13312b5ac69ca6
95 data/2016/b35a3f26a162bec3740e2da5639327afdhcantho-1468907131146.xls d114bf1c070bb165bd45a64ec79495c764805ee989c740ed95fad06eabd69ec4
96 data/2016/b42fd58ead4c54279d7df5f53160a851dhdanang-1468920187203.xlsx 9ce9d1b6f494866f6a39e68babe27b3a9562cd1515477f5f32a8e2a8a57bf13e
97 data/2016/b7a0c1ecce7c8445b3a1021091a5e654bacgiang-1468983153932.xlsx a53266ee75a8d49fe5a017125e198a1368a9bc9259dabe8cdf1cad11551dd439
98 data/2016/bacd1b2133b37d0ea9fbf46445265036dhtiengiang-1468940391597.xlsx bc90ae7bf5611ae0bb3c321e5185afcb530e1331c5a920a40b85b30b47eb1006
99 data/2016/bd857add0cfe92c1c247469ff2104527dh-hong-duc-1468900427119.xlsx 96b62a2954d2c1541699d491ea931c240a9b06660d5c27a3cd9055433306aaef
100 data/2016/bf259a908ff8a8dcc067a761204459f4daknong-1468934173803.xlsx 2f8fb2dbcd9a06acfd91becba0f728bf56e64f5c9984b40eac060834ea619b44
101 data/2016/c174631d90303276ac63ff38b31157bbdhmodiachat-1468900697349.xlsx c5c4435cfe1e6feede1936fd11e5cd3ff5758b4c4d522df03facf7522a9e7194
102 data/2016/c2cfaaf287fb1e7120894e545771b027dhkhoahoc-dhthainguyen-1468900365545.xlsx 9a35e56e631ba5d1056fb149ea01a41b4e62ff075937fa735666d192b1714da2
103 data/2016/c5c28959fcd0b4e39dd9fd538db5448adhgtvttphcm-1468939150821.xlsx 17a032a01a9ebe9f577af8fa4666e8700e9ab4e97166703cf6a7a0ff9dea4e9b
104 data/2016/c5ecc58a731fd1a54b7e6cc3bc99449ddhlamnghiep-1468939300790.xlsx 8f5f7bf26bafb3a8dd6988ab9858dbf03a14d200a23ed5299a8a5411c84d7f1b
105 data/2016/c6ef16bc9f75ce0fd16c1229afad4e71dhhanghai-1468906457479.xlsx 8e43d5e1782e366344c8d615666541931319475e02df0be29ae290c7ab2abda1
106 data/2016/cc57dc986acb525086b86e85df54949dkontum-1468933632113.xlsx 2397d3b0fca3888ff34d4166a9d3da70dcd6986e6405bd556471d46fff0ab551
107 data/2016/cdc912664d7217c75eb9b67c6ef3797adhkiengiang-1468901222446.xlsx 460eb7e1ce79dde123fe53a18034c5617f0616ad1abce6dedfcd0e4a2957a28d
108 data/2016/d035f470c13adbf5785797c6ccdcb204laocai-1468919805337.xlsx d23348d36bf6fa886d50e30507d4eada114a58f52cb6e615c119980d7753b5ad
109 data/2016/d5ecd2574b5ab36c62dffad435262cb8tuyenquang-1468902296807.xlsx 3acb7e4ba9a3c210906b9a32372db747ba2d95d25d276ba3ba45046b941dbbef
110 data/2016/d652d24e8bd4fe6a48d634b116f37c27dhluattphcm-1468939398928.xlsx f519010d9b2631457c9b7784bb77f07979907321199b98b3f97cb45e5b20c2ff
111 data/2016/d75282bff0a4d8cd1421c3fb74f789c3dhsaigon-1468900867561.xlsx ee7734bac44228a58a11b85d490cbd5da27d7a6e07b22c1eb9921a123b3e6acd
112 data/2016/ddb631f833af894200ef10cd408d2ef0caobang-1468932877769.xlsx 2bbda20bab1190f42b44d0eeabcf2be7f2203d4e5f8a2ad3d80365770a8e2d9e
113 data/2016/deca197a916a8f633f6894b4e13c3ca4dhhaiphong-1468901178368.xlsx 221fe731a4b2de02b3952572baf48d66544fddf87dfdd7851b740325c0864d24
114 data/2016/dfa63c5b68a8f638379ca0e9c5a1c556namdinh-1468933234061.xlsx 446cd9626fe197f6509efb01ca5b4cb26730907b5e114c4b9c07399a489b45d1
115 data/2016/e693bc9de67acc916032c6a3b2fc9eedhanam-1468933146526.xlsx b1ef9283b500e965205281d3875eff011be1d418312d187831925860b1ddab74
116 data/2016/e80dc028d53fa49ee247391a0653dee5quangngai-1468933592400.xlsx 9c6ba937e8ad772a3782baf9c40ef781121c27ffdeb4519170cbf2b81545a146
117 data/2016/ec765c32190773f5a18cd8202a5eaf0dquangnam-1468933501467.xlsx a566488b611ca04c705f385c47b871792ce75f7c1ca793d0528220352d6b49e1
118 data/2016/ec84067268f47d523b1e80b493248377dhngoaithuong-1468920480753.xlsx 5d892f08ec818be3ceb343529c9f0b35e0af0aeb43a229446dccaa87c6d4c846
119 data/2016/ee09c723da4e86cdbe26213413450af2dhsupham-dhthainguyen-1468939061872.xlsx aefac40009116615131c926775e6625eeebc615cd2ff4c7c9af399dd2faab982
120 data/2016/f0aa5a19c2c8c1fde8211341f0e98e26dhkinhtetphcm-1468939229043.xlsx 5ce36388e304cbba01c513f11f54b940f4299061e2e8acfa4c4219ee536a01f2
121 data/2016/f2846894bd5533e02b122a1e3b4333f6hvnganhang-1468939472850.xlsx 142b39b4f77c1d7a648c2397165b315acdd48ae6d354d914dd674a46b3d7160c
122 data/2016/f85f04bb2bbd84457b83ac846728a2a6dhspkythuathungyen-1468920667091.xlsx f8cc571e94720652f1507576ff48616d5331a4ed1403cb95cdc30124a03c3a68
123 data/2016/fc511852659150f6305f6f624445300ddhangiang-1468940189321.xlsx 4e16ea6a122e5bc0c62c7e6e55d9ede3a41242dd3e034d92a450bf29e8fff197
124 data/2016/ff877788e43bd84f0119dae026996892dhkinhteluat-dhqgtphcm-1468939860894.xlsx 8c9e7b097d0796750995268aa36a9fa00112c0989b513023ef0a53456afd7051
125 data/2017/an-giang.xls 5bde5f6447b9e27540f4f72d9b371b6c94376aeb97af65f4ec5eede43ba9f25c
126 data/2017/bac-giang.xls 9cd5c6d858062a1fb30c477b699be467139f183745e66d134746e043d73579a4
127 data/2017/bac-kan.xls 9642c00d14c54102ce8dc38408676863a87652aee89651113f4ad3e0a25bd1fa
128 data/2017/bac-lieu.xls e364dd65fa068d40e9a390aa3322e1676171af797c5d1f98949277cadab5369e
129 data/2017/bac-ninh.xls 539ed3cb1a2df358ab169e2163face76c949aac2eca0d2acb7ceb51e97811bc5
130 data/2017/ba-ria-vung-tau.xls 78dc6d67bbf3ffcb7e4407a6d950e28e39366ed62617691c9a9db610a5079134
131 data/2017/ben-tre.xls b4cce0c5a443062a5e7aef95785fc56f421c21d5ad22f993cbc30ec28e078f4f
132 data/2017/binh-dinh.xls fa1a39cb5134a8159e71dabea41c9ef2a3ec263642f19bed54ddd17e5a5e3d98
133 data/2017/binh-duong.xls ced0d7469186d9e5fbbb8c4323a4bb6a589abb8ef23f01d8d39f947a575af4ba
134 data/2017/binh-phuoc.xls d1ecf9b66ea294eec24ac0f703c614a2011b7d693aeb08c719cfcb73cae7ca47
135 data/2017/binh-thuan.xls 3cb0f85889ada81cd257f51ae0b0b6414d2ad8f81d7f94f4b1b38010600243c6
136 data/2017/ca-mau.xls c4bba31f7728f7825c20b3c0b708ef977dcf6f56f0dd5c14d9342cea0467e505
137 data/2017/can-tho.xls b5ba59ff7b87c31859686adb98da8ef2e95de84404daff2ac3df1e00b748b57d
138 data/2017/cao-bang.xls 9b0a497ed16840a22af34f7e75e316b00dd60490d6d7f8a2f012507c95fad3e7
139 data/2017/dak-lak.xls aae3e074d91afd98d89cc9d07bc1f8debe4559bd78c409cc5f50f2670a2c337a
140 data/2017/dak-nong.xls f7e30b4cdfe951086aeb1c169e7aad88a17bf9fe07aa33f85654e309899f6beb
141 data/2017/da-nang.xls 80963ff2f56d2b6bf331546d08184e325ad0ec7432ba948c0f74845f440e87c8
142 data/2017/dien-bien.xls 6bcb1969ce81ccb94c6ee625456530b9661ee255c00eb598dc195611a46205c8
143 data/2017/dong-nai.xls af248453ea27e879da1b8cfa0f2cd345aa068f93c884d1bc98862db0ca2f9faa
144 data/2017/dong-thap.xls 186a0c4ac6bf838cc496c4604a7e8979712051c71a53622a9ec72af70491e4f7
145 data/2017/gia-lai.xls 3a1b64c8362cad028ee53f4fd18b0923315d919b117df9a2fd83d2f46544db98
146 data/2017/ha-giang.xls a3248393c759a41de424e3ed994d2b9298cb39fda3c6eb9a68f48025c5a2e072
147 data/2017/hai-duong.xls 6634715f5a0cbac6a8d2d18be0df74e0abd328641721692224758c730fd86e51
148 data/2017/hai-phong.xls 7a8a210986de8c562a9128dccfaed1c6fc6c32df0c5e2e8948bee2f73777ba99
149 data/2017/ha-nam.xls f725e212f976af00e68edd6552721e56b30d5fa7d786f548dfc34721ec8b4426
150 data/2017/ha-noi.xls f0f4bc9216e421acf655a40844a4dcb74e871b188cf6d455982ee6e408170ef7
151 data/2017/ha-tinh.xls b0f5bcd8beff7cfe411b7bd88f9e5b9c5a9a0578a55cf18bb003284aca12e57f
152 data/2017/hau-giang.xls 9e8bccd0738bac3d68640768a8d9abc3e7a6a4e97e8c5968039530cf313cafb9
153 data/2017/hoa-binh.xls 51b1e910bfc47404b411cb2ab1ff14b522d8fe54d66a1d50479350273c7be1b5
154 data/2017/ho-chi-minh.xls c4e2f920a5e58fa62bd0913d994d44079e472b955adfc504b734c6773c485f64
155 data/2017/hung-yen.xls 833584c8ea9337734feb37e068bb06232293108d50b18d88f731d30a576cc29a
156 data/2017/khanh-hoa.xls 644ab06abca93d7b966205604313a65d50f3a0a6fb8be509cadb55e0d55fe12e
157 data/2017/kien-giang.xls df452e618544a9b7932d918d26dd2a0f2149f92599609c15a9ac373291cfc440
158 data/2017/kon-tum.xls 7d26ff078adaf51dcf4333b074f8a444d328d8083835bd55b1034c1a63e4eb17
159 data/2017/lai-chau.xls b37cd06914cad19bc0e523d0b1cfe87db692e26dfab327f5f9a7686b2b14dae9
160 data/2017/lam-dong.xls 59cc802a6c4d622f1bb6233f6b9b1a94f342a6eae1c389235383db1767f34720
161 data/2017/lang-son.xls ae376e75bf615483b6cb3a37daa1cea6a7eec4cf3fdd1186801f56d074eb4ce7
162 data/2017/lao-cai.xls 7596c99851c20ff60b95347bc9fae78014873b4c3e1f1f7144c48d9f285f7962
163 data/2017/long-an.xls e52d8e8a4592258e87c5cda04e0907541f60e7e7ef4cb21e9171b65283e0b87a
164 data/2017/nam-dinh.xls 80dcda8b058dcf1a7fdcd7e34d126c44c31b2d4987b52eeae790bf6c38c58e40
165 data/2017/nghe-an.xls 3f32421eb6ee8bc21fcedd4cc1043df1c9c0381abfaa1006c85c3ca6d6209543
166 data/2017/ninh-binh.xls dc3d08a86cdd145e870486773d95a1d896a003b5400ba6ff8456a5512b10b2b8
167 data/2017/ninh-thuan.xls d69b0ed7192e85cff31ad5bca8942fe3c92322268d1ea707e945f9991bd78115
168 data/2017/phu-tho.xls e352a51109a2373c1efbb6726d262fc490bcbf9229ae90bd473d14a4b6eebcc0
169 data/2017/phu-yen.xls 6871aedd90d2c6c87ad6409164a9b26f6f89f97cdd7a1dffbbe4e6879f2348c6
170 data/2017/quang-binh.xls 6c5e9b3c6780a5b70a02dd4b4854861ba148613361c048bfc001706e5bb77a61
171 data/2017/quang-nam.xls 859258cb8c4b7806101039e4e2484e56971f2640a65720a108a0933245bd3157
172 data/2017/quang-ngai.xls 834988a6484fab64bc59fe35a6813038663054553d269d16b7d8200593f468e8
173 data/2017/quang-ninh.xls 1b1c0b1ea1d6aadd8bce5faef6b5fdc5181daaa9bd44df6acbf06c9021a0ade9
174 data/2017/quang-tri.xls 35e401061962c6d80593d684f179e1f82b05f884a6a470d1bc17b2d31151e05f
175 data/2017/soc-trang.xls 668a0edcb624faa954c424616ea6c35d6f06b9e9fc62524723dea01af09490c0
176 data/2017/son-la.xls 96bcfeb35e5bc23ccd7dc6a62a533e21dabeb3145473ecc5a039d970019a1129
177 data/2017/tay-ninh.xls aedf4ccab43e7a1fa985721de0d6053988ad471355c0ccb0b8f9c19af02423e2
178 data/2017/thai-binh.xls b20f66cdeb975a64eed7322b616bd291fe34e08e5fbedad3d729d21da0eba627
179 data/2017/thai-nguyen.xls f0b93db9aab966d393b8a3323204ba7f147ddd0a92746355b81df09bc58267b4
180 data/2017/thanh-hoa.xls 3f803ae67dac24317fa45dac60f36e39dc222ff84d23a0a3dde7751cdcb17646
181 data/2017/thua-thien-hue.xls 71e220ee4ea0287657fa379c587ca93f71f96b4dacc48cef074f8c223e940daa
182 data/2017/tien-giang.xls 2e061f8724e1ea8b997d7e6700ac89a60921771a57dffa7297f4728715811406
183 data/2017/tra-vinh.xls 59507a1d237ac5e29c1033d42b08da6ac9a7505a5045b791e512e91493102ae7
184 data/2017/tuyen-quang.xls 053d293e4acd9dfa6ce9532a047cc64b4d2edee1ff13b387723586c1fae6312a
185 data/2017/vinh-long.xls 6e7584a3047b0ba7487dbb8c71f5b594db220f11b747149d41d31ea97e1e4eab
186 data/2017/vinh-phuc.xls 9753e2d864584e76aef1e690e4e2de98b28240fbba810e614eab9926668d66d2
187 data/2017/yen-bai.xls fe6824a1254116bc6ed372e0a1fd635a816a09e010b7efde9659baca2e143138
188 data/2017-old/10_BinhThuan_RIIW.xls.xlsx 2833e7366ea1a21ea869978a6bb920b033160bb8042fe0bfe9b7573dcfe85409
189 data/2017-old/10_Ca_Mau_BKXT.xls.xlsx 53cb69b13325cb9cc9004b8c89f0cd1c85b5f943e7c78c8b010ddabad3f816c9
190 data/2017-old/10_LamDong_GNFT.xls.xlsx f05b4cdc8a9a8413b29801fbddeb48a7c60d3c6bb2228bbe238113495e89bc98
191 data/2017-old/10_Soc_Trang_XCGJ.xls.xlsx 8e540cef6f2ff387f3e825a5aa3190af394c6f9c7e9dfc66d5f7032ee47f7f14
192 data/2017-old/11_BinhDuong_RYQL.xls.xlsx 11953116054a2889929c3fe05b68e8932c7b79ffd1b85d697afaac63fcdd9adf
193 data/2017-old/11_LaoCai_GMSU.xls.xlsx 5a73becf1bdb8027102faff74035d1df71e13c1cac39cc38ec6efaab6699ca52
194 data/2017-old/12_BenTre_DKWF.xls.xlsx 579c0d94b20440f1c60a94fb0c762b71dd1c7719c8aa4db0b9418ecaba9b605e
195 data/2017-old/12_LongAn_ZZUK.xls.xlsx c4cf569706bfc30582d774ba9e699dcbc4bdc734f46bf6493df58eaecc37a755
196 data/2017-old/13_NamDinh_ESEL.xls.xlsx 7e5256cb095262fef5399092f8f843c1ee82e117bca8e6571e363646bb1170a6
197 data/2017-old/13_TraVinh_LKUJ.xls.xlsx 2b7321d68a4a910c4a506d7ca207531cfbdd9619f1d8ba0ce475baf1b91043aa
198 data/2017-old/14_NgheAn_BSLY.xls.xlsx 034234a8a5461910f8beff3977aafa29a6541c8a56c007816118e5c4eb2b7218
199 data/2017-old/15_PhuTho_ABWQ.xls.xlsx 9891c17001cfcc49ccd587419339698ef59b2dd43b5ca7f2765b09a17c26659d
200 data/2017-old/16_QuangBinh_KGEU.xls.xlsx 01e249d45c1269d62ae1977bef0bb192f833941510471b723e57c32adb0eadaa
201 data/2017-old/17_QuangNam_AMTK.xls.xlsx 4c9f05d24d861ec3c82a1c564e4baa1541bd4882785f143459a1a81a6b8359c7
202 data/2017-old/18_QuangNgai_KOFP.xls.xlsx ee9c26f2c16d36dabfde79901a054d2266c28c4acbfd19b628cfe50f8eb03957
203 data/2017-old/19_QuangTri_OMZF.xls.xlsx 6c961bb461274c73b4589de4087624a9ef5edd43c04ea54675148df00f6420e5
204 data/2017-old/1_BaRia_VungTau_HJKG.xls.xlsx cb03f933442a8cea6b1622da4e5f804ab476dff9f3462b404e957b6a2e87b2cd
205 data/2017-old/1_Da_Nang_AHWJ.xls.xlsx 59892fc01b6cb10497679f2bc69f59023215755d2f00fd764b23b895c98be2be
206 data/2017-old/1_Ha_Noi_CVXG.xls.xlsx 2718fe9c2b58843621c0082da5880a4b957266525b83123af02c78fd51d4f4a5
207 data/2017-old/1_Son_La_JIDP.xls.xlsx 0a0e273c5def980406b6316cceb63e5cee53aae56e112dfae7e795261eb1cd06
208 data/2017-old/1_TuyenQuang_JBYF.xls.xlsx 45826ed10248268ac3344d0b592cf2dba6f86278b7402dc60626a76b18fa4d6b
209 data/2017-old/20_TayNinh_ILFA.xls.xlsx 684a1820d02f658904c87c26aa195faa6aceb3d7ca00c674df11a63f9274ab14
210 data/2017-old/21_ThaiBinh_FTVG.xls.xlsx 1a12e43d4ea368b8e3a2e4d881b5d66214a16140219e80e2e72fcb269899b285
211 data/2017-old/22_ThaiNguyen_TLTW.xls.xlsx 29227bdc3204591327e803680525d8621228fc371ce7f802dc667ec70eb8d7b3
212 data/2017-old/23_HaiPhong_HXBV.xls.xlsx e450e2344fbd724555de9892412f521414fc1fcc53e69887c6381dd266b4e92c
213 data/2017-old/24_HCM_XULN.xls.xlsx 49ed22e823607b78920a24f5999eaf1ef8634a141108dfd4bf970ee000aecbf9
214 data/2017-old/2_BacKan_GFVQ.xls.xlsx 77a08fc4edb41d639fcdd798980903519181639db00a360ef0c6c6c2bd88f529
215 data/2017-old/2_Ha_Giang_QNCM.xls.xlsx 6bf4d62413c1b92f88f980e5e3e414a7e25971c7546c6390c7e67a6b0f742b68
216 data/2017-old/2_Ninh_Thuan_VHLY.xls.xlsx 6dfb6517f9ccc3420c222b6a88aabac54f6f25f6b743d993d6a5f28b4fa78ff2
217 data/2017-old/2_Thanh_Hoa_AUFV.xls.xlsx 2f2ce90425c19a9089024061f328f37b9ef580ddadc076ddcf452f549690f4b3
218 data/2017-old/2_VinhPhuc_GUDK.xls.xlsx fc0e4952445ebda1af28f7330aa808c032f2e3a7a3c03dfbb9692bf6a248867a
219 data/2017-old/3_BacGiang_TOIF.xls.xlsx 8043510ff89f11ba063f898dc1d2c8879a8da9063a420308647e20a2f39d3092
220 data/2017-old/3_BinhPhuoc_YFMU.xls.xlsx 8869970f4e6005381264f3f4bb508c5181f15ae990a75fe904a65b305c0a123a
221 data/2017-old/3_Cao_Bang_CIEY.xls.xlsx 94a06bdb2333a9f44216001ba8d8463700b8cf3e51944b94db11eef7eac8dbd3
222 data/2017-old/3_Dong_Thap_GSXQ.xls.xlsx 0cdd0763fc4f20f1d990621ca9e8dab107a3d4eb7c0ec3e53a4d8c240fbc569f
223 data/2017-old/3_Thua_Thien_Hue_HMDB.xls.xlsx 62e54eebda7bbd542ca63443bb923a81b472e7892e55dce39fd0611fc0f00abd
224 data/2017-old/4_An_Giang_JNOS.xls.xlsx 778ad205bbabb7c4aea3df34325770aa4a1daec0a1683f73f0405ac1cc5aaf9a
225 data/2017-old/4_BacNinh_STLR.xls.xlsx 009c3b7821f55210a86f9ace74c31eb0d5c16aafdb0cac8f2b71cac616abdfec
226 data/2017-old/4_Binh_Dinh_WWZW.xls.xlsx fbacd256d58627377d698462d69aac0e83d124c8c8ff1f5b63ba0bfb5442dd39
227 data/2017-old/4_DienBien_SOJG.xls.xlsx 4931e8d6604013df48b2b469fda3eb05fdf1203b67bbc2ae4ff6d7949a115a24
228 data/2017-old/4_Lang_Son_NYQL.xls.xlsx cc7d2454caadb18740ed6baf6edc62baa47e98b3a907b3ce06f72b9a16b7838a
229 data/2017-old/5_Bac_Lieu_XEKH.xls.xlsx dc1dca45fd5357b45abc930b0ebb95ab589ef80a09bef8b5ad4dab9e00be87e5
230 data/2017-old/5_Gia_Lai_ABZI.xls.xlsx 5b7be25274d628c284c37efb3df0587171cc9273dc3bd8dc824c7bd1b8bccf1a
231 data/2017-old/5_HaiDuong_WWWG.xls.xlsx 1d800d0aa7338fb507c4eeb3e752620d2608c30ac820895ca086325dbaf3643d
232 data/2017-old/5_Hanam_QGJS.xls.xlsx cfc233ac37ab8e263b6cc569815ea2c687da30364f139109032c8f2cac1ad98e
233 data/2017-old/5_Yen_Bai_FAQR.xls.xlsx 4c8915417d8747c291f643ea4b378c87982c397460e0e9e3cc6df19670d6720f
234 data/2017-old/6_Dong_Nai_WOTM.xls.xlsx 9b582a8421c419156c2eab6e532d854736385735f8208f9d390c251f22ac764e
235 data/2017-old/6_Hau_Giang_KWDM.xls.xlsx c11473e3e1f964302c5378c90dd6b9d478eb11397351b7ef74df9ae709c9e908
236 data/2017-old/6_HoaBinh_TPZY.xls.xlsx e1210961dc821642f6951cd6e4a71b5fbc81d5070ee3816b7c8b5dc56bf2b234
237 data/2017-old/6_NinhBinh_IGFT.xls.xlsx 9399b8ab58e2a7a1c0ebf5425359319e567f8a28602349aa587efced72142fc8
238 data/2017-old/6_Quang_Ninh_DQCJ.xls.xlsx 8cf95259b11d48267f710f151c3f1419ccac49d823b78b70a2708c59ab25a4f7
239 data/2017-old/7_HaTinh_DDHD.xls.xlsx d96df9f2c8365dfc73be3e10c268367b93fd00fe92bb39c745d8e353ba75c05b
240 data/2017-old/7_HungYen_LTIK.xls.xlsx e80eb38af8307c8e193a91bdda5368c437693d7669e281b543bbe891a46e6f36
241 data/2017-old/7_Kon_Tum_RSLR.xls.xlsx 41d2adea059449b6e214241414748f10776861d261cbaa1fd213629757deb846
242 data/2017-old/7_Tien_Giang_EFHX.xls.xlsx 00d9bbc70be89f8f7dad6aff990ad36236c6b7244ddeb2e3d9b3d8ed1a6e21d0
243 data/2017-old/8_Can_Tho_RQZM.xls.xlsx 3a1182b26d34fbdfabf4326098ff265609d8f1fe2b766b6eef897db9685d4cd4
244 data/2017-old/8_Dak_Lak_YKPR.xls.xlsx 379b45f8a347bf4842daf9402d84be47426ed7107a64b2e7710bdc8e71eb57be
245 data/2017-old/8_KienGiang_OTOB.xls.xlsx 3e1789b8565a5ae7c187598cb7e9c3261029c06bc7f0cccc4eb503e8e302cb18
246 data/2017-old/8_PhuYen_OGIM.xls.xlsx 91d1f202cba3bf4c57a7f26e7611d3d29faf8dd13de30552cbfc96eab7a451e9
247 data/2017-old/9_DakNong_FBOP.xls.xlsx c611bf7f1d1c95cbfc78bdab3ea99bf8278e94c8db09a9fd3ccac55a538fe0b2
248 data/2017-old/9_Khanh_Hoa_KPKQ.xls.xlsx 0e90c87dd74a5eb293a0a71b93da45653a19ed2e3a2600d3b880a867a695f31d
249 data/2017-old/9_LaiChau_ALKN.xls.xlsx 3b7a022f7727134620c6756925bd5adb041e2d80cabd4aab13a4063501739514
250 data/2017-old/9_Vinh_Long_OLWA.xls.xlsx 43d7abb9c6b8d3205ce90f9cceaf9623875c87998d9e030c6dc8da9bc18bfa09
251 data/2017-old2/10.BinhThuan_MVVG.xlsx 2833e7366ea1a21ea869978a6bb920b033160bb8042fe0bfe9b7573dcfe85409
252 data/2017-old2/10.LamDong_YUQA.xlsx f05b4cdc8a9a8413b29801fbddeb48a7c60d3c6bb2228bbe238113495e89bc98
253 data/2017-old2/10.Soc Trang_LQWU.xlsx 8e540cef6f2ff387f3e825a5aa3190af394c6f9c7e9dfc66d5f7032ee47f7f14
254 data/2017-old2/11.BinhDuong_HVAH.xlsx 11953116054a2889929c3fe05b68e8932c7b79ffd1b85d697afaac63fcdd9adf
255 data/2017-old2/11.LaoCai_ZTBP.xlsx 5a73becf1bdb8027102faff74035d1df71e13c1cac39cc38ec6efaab6699ca52
256 data/2017-old2/12.BenTre_NQTU.xlsx 579c0d94b20440f1c60a94fb0c762b71dd1c7719c8aa4db0b9418ecaba9b605e
257 data/2017-old2/12.LongAn_PDRH.xlsx d1c65184b475fc29ae273cc89b8574b7e273bc3dab83cd39489d72a5de1bb1e4
258 data/2017-old2/13.NamDinh_NAYR.xlsx 7e5256cb095262fef5399092f8f843c1ee82e117bca8e6571e363646bb1170a6
259 data/2017-old2/13.TraVinh_FODZ.xlsx 2b7321d68a4a910c4a506d7ca207531cfbdd9619f1d8ba0ce475baf1b91043aa
260 data/2017-old2/14.NgheAn_HTKD.xlsx 034234a8a5461910f8beff3977aafa29a6541c8a56c007816118e5c4eb2b7218
261 data/2017-old2/15.PhuTho_IJZW.xlsx 9891c17001cfcc49ccd587419339698ef59b2dd43b5ca7f2765b09a17c26659d
262 data/2017-old2/17.QuangNam_NQMG.xlsx 4c9f05d24d861ec3c82a1c564e4baa1541bd4882785f143459a1a81a6b8359c7
263 data/2017-old2/18.QuangNgai_IUPY.xlsx ee9c26f2c16d36dabfde79901a054d2266c28c4acbfd19b628cfe50f8eb03957
264 data/2017-old2/19.QuangTri_MKNN.xlsx 6c961bb461274c73b4589de4087624a9ef5edd43c04ea54675148df00f6420e5
265 data/2017-old2/1.BaRia-VungTau_PGZT.xlsx cb03f933442a8cea6b1622da4e5f804ab476dff9f3462b404e957b6a2e87b2cd
266 data/2017-old2/1.Da Nang_ABWU.xlsx 59892fc01b6cb10497679f2bc69f59023215755d2f00fd764b23b895c98be2be
267 data/2017-old2/1.Son La_XLFN.xlsx 0a0e273c5def980406b6316cceb63e5cee53aae56e112dfae7e795261eb1cd06
268 data/2017-old2/1.TuyenQuang_PTMR.xlsx 45826ed10248268ac3344d0b592cf2dba6f86278b7402dc60626a76b18fa4d6b
269 data/2017-old2/20.TayNinh_KJAQ.xlsx 684a1820d02f658904c87c26aa195faa6aceb3d7ca00c674df11a63f9274ab14
270 data/2017-old2/21.ThaiBinh_KTQN.xlsx 1a12e43d4ea368b8e3a2e4d881b5d66214a16140219e80e2e72fcb269899b285
271 data/2017-old2/22.ThaiNguyen_BKIF.xlsx 29227bdc3204591327e803680525d8621228fc371ce7f802dc667ec70eb8d7b3
272 data/2017-old2/24.HCM_UTLQ.xlsx 49ed22e823607b78920a24f5999eaf1ef8634a141108dfd4bf970ee000aecbf9
273 data/2017-old2/2.BacKan_YQNX.xlsx 77a08fc4edb41d639fcdd798980903519181639db00a360ef0c6c6c2bd88f529
274 data/2017-old2/2.Ha Giang_PIYK.xlsx 6bf4d62413c1b92f88f980e5e3e414a7e25971c7546c6390c7e67a6b0f742b68
275 data/2017-old2/2.Ninh Thuan_BAGG.xlsx 6dfb6517f9ccc3420c222b6a88aabac54f6f25f6b743d993d6a5f28b4fa78ff2
276 data/2017-old2/2.Thanh Hoa_UOPE.xlsx 2f2ce90425c19a9089024061f328f37b9ef580ddadc076ddcf452f549690f4b3
277 data/2017-old2/2.VinhPhuc_QZJK.xlsx fc0e4952445ebda1af28f7330aa808c032f2e3a7a3c03dfbb9692bf6a248867a
278 data/2017-old2/3.BacGiang_SAVS.xlsx 8043510ff89f11ba063f898dc1d2c8879a8da9063a420308647e20a2f39d3092
279 data/2017-old2/3.BinhPhuoc_IPHL.xlsx 8869970f4e6005381264f3f4bb508c5181f15ae990a75fe904a65b305c0a123a
280 data/2017-old2/3.Cao Bang_WMUU.xlsx 94a06bdb2333a9f44216001ba8d8463700b8cf3e51944b94db11eef7eac8dbd3
281 data/2017-old2/3.Dong Thap_HKJX.xlsx 0cdd0763fc4f20f1d990621ca9e8dab107a3d4eb7c0ec3e53a4d8c240fbc569f
282 data/2017-old2/3.Thua Thien -Hue_MAET.xlsx 62e54eebda7bbd542ca63443bb923a81b472e7892e55dce39fd0611fc0f00abd
283 data/2017-old2/4.An Giang_PMJD.xlsx 778ad205bbabb7c4aea3df34325770aa4a1daec0a1683f73f0405ac1cc5aaf9a
284 data/2017-old2/4.BacNinh_NNIS.xlsx 009c3b7821f55210a86f9ace74c31eb0d5c16aafdb0cac8f2b71cac616abdfec
285 data/2017-old2/4.Binh Dinh_VOMJ.xlsx fbacd256d58627377d698462d69aac0e83d124c8c8ff1f5b63ba0bfb5442dd39
286 data/2017-old2/4.DienBien_FYGN.xlsx 4931e8d6604013df48b2b469fda3eb05fdf1203b67bbc2ae4ff6d7949a115a24
287 data/2017-old2/4.Lang Son_QWOG.xlsx cc7d2454caadb18740ed6baf6edc62baa47e98b3a907b3ce06f72b9a16b7838a
288 data/2017-old2/5.Bac Lieu_VIVY.xlsx dc1dca45fd5357b45abc930b0ebb95ab589ef80a09bef8b5ad4dab9e00be87e5
289 data/2017-old2/5.Gia Lai_TAAS.xlsx 5b7be25274d628c284c37efb3df0587171cc9273dc3bd8dc824c7bd1b8bccf1a
290 data/2017-old2/5.HaiDuong_WNHD.xlsx 1d800d0aa7338fb507c4eeb3e752620d2608c30ac820895ca086325dbaf3643d
291 data/2017-old2/5.Hanam_SDKN.xlsx cfc233ac37ab8e263b6cc569815ea2c687da30364f139109032c8f2cac1ad98e
292 data/2017-old2/5.Yen Bai_BSLV.xlsx 4c8915417d8747c291f643ea4b378c87982c397460e0e9e3cc6df19670d6720f
293 data/2017-old2/6.Hau Giang_SIAJ.xlsx c11473e3e1f964302c5378c90dd6b9d478eb11397351b7ef74df9ae709c9e908
294 data/2017-old2/6.HoaBinh_HLYQ.xlsx e1210961dc821642f6951cd6e4a71b5fbc81d5070ee3816b7c8b5dc56bf2b234
295 data/2017-old2/6.NinhBinh_PKMQ.xlsx 9399b8ab58e2a7a1c0ebf5425359319e567f8a28602349aa587efced72142fc8
296 data/2017-old2/6.Quang Ninh_YAKJ.xlsx 8cf95259b11d48267f710f151c3f1419ccac49d823b78b70a2708c59ab25a4f7
297 data/2017-old2/7.HaTinh_XMFQ.xlsx d96df9f2c8365dfc73be3e10c268367b93fd00fe92bb39c745d8e353ba75c05b
298 data/2017-old2/7.HungYen_TCBE.xlsx e80eb38af8307c8e193a91bdda5368c437693d7669e281b543bbe891a46e6f36
299 data/2017-old2/7.Kon Tum_CAQU.xlsx 41d2adea059449b6e214241414748f10776861d261cbaa1fd213629757deb846
300 data/2017-old2/7.Tien Giang_AOWZ.xlsx 00d9bbc70be89f8f7dad6aff990ad36236c6b7244ddeb2e3d9b3d8ed1a6e21d0
301 data/2017-old2/8.PhuYen_MLTQ.xlsx 9dd53293f99d86801692a1995a5cf7a08af8394b820c76e68c88c2b7644b9b67
302 data/2017-old2/9.DakNong_ZWJT.xlsx c611bf7f1d1c95cbfc78bdab3ea99bf8278e94c8db09a9fd3ccac55a538fe0b2
303 data/2017-old2/9.Khanh Hoa_IBTX.xlsx 0e90c87dd74a5eb293a0a71b93da45653a19ed2e3a2600d3b880a867a695f31d
304 data/2017-old2/9.Vinh Long_WVKI.xlsx 43d7abb9c6b8d3205ce90f9cceaf9623875c87998d9e030c6dc8da9bc18bfa09
+4 -3
View File
@@ -5,14 +5,15 @@
"type": "module",
"description": "Tra cứu điểm thi THPT Quốc gia — 2016 và 2017",
"scripts": {
"build:rust": "cargo build --release --manifest-path parser/Cargo.toml",
"build:db": "node parser/scripts/build-db.js",
"build:go": "go -C go-parser build -o bin/xlsxread ./cmd/xlsxread",
"build:db": "node go-parser/scripts/build-db.js",
"dev": "vite",
"build": "vite build",
"assemble": "node scripts/assemble-site.js",
"build:site": "npm run build && npm run assemble",
"preview": "vite preview",
"lint": "eslint ."
"lint": "eslint .",
"test:go": "go -C go-parser test ./..."
},
"repository": {
"type": "git",
+26 -54
View File
@@ -489,6 +489,12 @@ version = "1.70.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a6cb138bb79a146c1bd460005623e142ef0181e3d0219cb493e02f7d08a35695"
[[package]]
name = "itoa"
version = "1.0.18"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682"
[[package]]
name = "jobserver"
version = "0.1.34"
@@ -693,6 +699,12 @@ version = "1.0.22"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b39cdef0fa800fc44525c84ccb54a029961a8215f9619753635a9c0d2538d46d"
[[package]]
name = "ryu"
version = "1.0.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9774ba4a74de5f7b1c1451ed6cd5285a32eddb5cccb8cc655a4e50009e06477f"
[[package]]
name = "serde"
version = "1.0.228"
@@ -724,12 +736,16 @@ dependencies = [
]
[[package]]
name = "serde_spanned"
version = "0.6.9"
name = "serde_yaml"
version = "0.9.34+deprecated"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bf41e0cfaf7226dca15e8197172c295a782857fcb97fad1808a166870dee75a3"
checksum = "6a8b1a1a2ebf674015cc02edccce75287f1a0130d394307b36743c2f5d504b47"
dependencies = [
"indexmap",
"itoa",
"ryu",
"serde",
"unsafe-libyaml",
]
[[package]]
@@ -858,47 +874,6 @@ version = "0.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1f3ccbac311fea05f86f61904b462b55fb3df8837a366dfc601a0161d0532f20"
[[package]]
name = "toml"
version = "0.8.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dc1beb996b9d83529a9e75c17a1686767d148d70663143c7854d8b4a09ced362"
dependencies = [
"serde",
"serde_spanned",
"toml_datetime",
"toml_edit",
]
[[package]]
name = "toml_datetime"
version = "0.6.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "22cddaf88f4fbc13c51aebbf5f8eceb5c7c5a9da2ac40a13519eb5b0a0e8f11c"
dependencies = [
"serde",
]
[[package]]
name = "toml_edit"
version = "0.22.27"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "41fe8c660ae4257887cf66394862d21dbca4a6ddd26f04a3560410406a2f819a"
dependencies = [
"indexmap",
"serde",
"serde_spanned",
"toml_datetime",
"toml_write",
"winnow",
]
[[package]]
name = "toml_write"
version = "0.1.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5d99f8c9a7727884afe522e9bd5edbfc91a3312b36a77b5fb8926e4c31a41801"
[[package]]
name = "typenum"
version = "1.20.0"
@@ -920,6 +895,12 @@ dependencies = [
"tinyvec",
]
[[package]]
name = "unsafe-libyaml"
version = "0.2.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "673aac59facbab8a9007c7f6108d11f63b603f7cabff99fabf650fea5c32b861"
[[package]]
name = "utf8parse"
version = "0.2.2"
@@ -1007,15 +988,6 @@ dependencies = [
"windows-link",
]
[[package]]
name = "winnow"
version = "0.7.15"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "df79d97927682d2fd8adb29682d1140b343be4ac0f08fd68b7765d9c059d3945"
dependencies = [
"memchr",
]
[[package]]
name = "wit-bindgen"
version = "0.57.1"
@@ -1033,8 +1005,8 @@ dependencies = [
"regex",
"rusqlite",
"serde",
"serde_yaml",
"thiserror 1.0.69",
"toml",
"unicode-normalization",
"zip",
]
+1 -1
View File
@@ -9,7 +9,7 @@ calamine = "0.26"
rusqlite = { version = "0.32", features = ["bundled"] }
clap = { version = "4", features = ["derive"] }
serde = { version = "1", features = ["derive"] }
toml = "0.8"
serde_yaml = "0.9" # deprecated upstream but stable; this crate is deleted at cutover
regex = "1"
unicode-normalization = "0.1"
thiserror = "1"
@@ -15,18 +15,18 @@
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
format_detection = "thptqg2016"
format_detection: thptqg2016
[reader]
sheet_mode = "all"
strip_blank_rows = false
reader:
sheet_mode: all
strip_blank_rows: false
[validation]
require_numeric_sbd = false
require_nonempty_name = true
require_nonempty_sbd = true
validation:
require_numeric_sbd: false
require_nonempty_name: true
require_nonempty_sbd: true
[header]
# Tokens that identify a header row by first-cell content (uppercased).
# Covers both SOBAODANH-style and SBD-style headers.
tokens = ["SOBAODANH", "SBD", "HO_TEN", "HOTEN", "HỌ TÊN", "STT"]
header:
# Tokens that identify a header row by first-cell content (uppercased).
# Covers both SOBAODANH-style and SBD-style headers.
tokens: ["SOBAODANH", "SBD", "HO_TEN", "HOTEN", "HỌ TÊN", "STT"]
@@ -6,20 +6,20 @@
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
[reader]
sheet_mode = "first"
strip_blank_rows = false
reader:
sheet_mode: first
strip_blank_rows: false
[columns]
ho_ten = 0
ngay_sinh = 1
so_bao_danh = 2
diem_thi = 3
columns:
ho_ten: 0
ngay_sinh: 1
so_bao_danh: 2
diem_thi: 3
[validation]
require_numeric_sbd = true
require_nonempty_name = true
require_nonempty_sbd = true
validation:
require_numeric_sbd: true
require_nonempty_name: true
require_nonempty_sbd: true
[header]
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
header:
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
@@ -7,20 +7,20 @@
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
[reader]
sheet_mode = "all"
strip_blank_rows = true
reader:
sheet_mode: all
strip_blank_rows: true
[columns]
ho_ten = 0
ngay_sinh = 1
so_bao_danh = 2
diem_thi = 3
columns:
ho_ten: 0
ngay_sinh: 1
so_bao_danh: 2
diem_thi: 3
[validation]
require_numeric_sbd = true
require_nonempty_name = true
require_nonempty_sbd = true
validation:
require_numeric_sbd: true
require_nonempty_name: true
require_nonempty_sbd: true
[header]
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
header:
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
@@ -6,20 +6,20 @@
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
[reader]
sheet_mode = "all"
strip_blank_rows = false
reader:
sheet_mode: all
strip_blank_rows: false
[columns]
ho_ten = 0
ngay_sinh = 1
so_bao_danh = 2
diem_thi = 3
columns:
ho_ten: 0
ngay_sinh: 1
so_bao_danh: 2
diem_thi: 3
[validation]
require_numeric_sbd = false
require_nonempty_name = true
require_nonempty_sbd = true
validation:
require_numeric_sbd: false
require_nonempty_name: true
require_nonempty_sbd: true
[header]
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
header:
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
+106
View File
@@ -0,0 +1,106 @@
//! Ground-truth cell dumper for the Go reader fidelity gate.
//!
//! Emits a canonical, reader-agnostic text rendering of every sheet and every
//! cell of a spreadsheet exactly as calamine sees it. The Go reader must
//! reproduce this byte-for-byte; the two dumps are compared by hash.
//!
//! Deliberately dumps the RAW used range: every sheet, every row, header rows
//! included, no config applied. This gate is about cell fidelity, not build
//! semantics — sheet selection and header skipping are exercised later.
//!
//! Usage: cargo run --release --example dump_cells -- <spreadsheet> [out-file]
//! With no out-file the dump goes to stdout.
use std::env;
use std::fs::File;
use std::io::{self, BufWriter, Write};
use calamine::{open_workbook_auto, Data, Reader, Sheets};
/// Escapes the field separators so a cell value can never break the line format.
fn escape(s: &str) -> String {
let mut out = String::with_capacity(s.len());
for ch in s.chars() {
match ch {
'\\' => out.push_str("\\\\"),
'\t' => out.push_str("\\t"),
'\n' => out.push_str("\\n"),
'\r' => out.push_str("\\r"),
_ => out.push(ch),
}
}
out
}
/// Discriminates the calamine variant so a Go port can be checked against the
/// actual type, not just the rendered string.
fn kind(d: &Data) -> &'static str {
match d {
Data::Empty => "empty",
Data::String(_) => "str",
Data::Float(_) => "float",
Data::Int(_) => "int",
Data::Bool(_) => "bool",
Data::Error(_) => "err",
Data::DateTime(_) => "datetime",
Data::DateTimeIso(_) => "datetimeiso",
Data::DurationIso(_) => "durationiso",
}
}
fn main() -> Result<(), Box<dyn std::error::Error>> {
let args: Vec<String> = env::args().collect();
if args.len() < 2 {
eprintln!("usage: dump_cells <spreadsheet> [out-file]");
std::process::exit(2);
}
let path = &args[1];
let mut out: Box<dyn Write> = match args.get(2) {
Some(p) => Box::new(BufWriter::new(File::create(p)?)),
None => Box::new(BufWriter::new(io::stdout())),
};
let mut workbook: Sheets<_> = open_workbook_auto(path)?;
let sheet_names: Vec<String> = workbook.sheet_names().to_vec();
writeln!(out, "FILE\t{}", escape(path))?;
writeln!(out, "SHEETCOUNT\t{}", sheet_names.len())?;
for (idx, name) in sheet_names.iter().enumerate() {
let range = workbook.worksheet_range(name)?;
// start() is the used-range origin — the key question for any Go reader,
// which may index absolutely from A1 instead.
let (srow, scol) = range.start().unwrap_or((0, 0));
writeln!(
out,
"SHEET\t{}\t{}\t{}\t{}\t{}\t{}",
idx,
escape(name),
range.height(),
range.width(),
srow,
scol
)?;
for (r, row) in range.rows().enumerate() {
writeln!(out, "ROW\t{}\t{}\t{}", idx, r, row.len())?;
for (c, cell) in row.iter().enumerate() {
let s = cell.to_string();
writeln!(
out,
"CELL\t{}\t{}\t{}\t{}\t{}\t{}",
idx,
r,
c,
kind(cell),
if matches!(cell, Data::Empty) { 1 } else { 0 },
escape(&s)
)?;
}
}
}
out.flush()?;
Ok(())
}
+76
View File
@@ -0,0 +1,76 @@
//! Corpus-wide cell-kind and sheet-geometry scan, for the Go reader fidelity gate.
//!
//! For every file given, prints one TSV line per sheet and one summary line per
//! file. Aggregates only — never materialises the cell text — so the whole
//! 418 MB corpus can be scanned quickly.
//!
//! The point is to find out which calamine `Data` variants actually occur in
//! real inputs. `DateTime` and `Float` are the variants whose rendering differs
//! between readers; if they never appear, the divergence risk is theoretical.
//!
//! Usage: cargo run --release --example scan_kinds -- <file>...
use std::env;
use calamine::{open_workbook_auto, Data, Reader, Sheets};
fn main() {
let args: Vec<String> = env::args().skip(1).collect();
if args.is_empty() {
eprintln!("usage: scan_kinds <file>...");
std::process::exit(2);
}
println!("#TYPE\tpath\tsheet_idx\tname\theight\twidth\tstart_row\tstart_col\tempty\tstr\tfloat\tint\tbool\tdatetime\tdtiso\tduriso\terr");
for path in &args {
let mut wb: Sheets<_> = match open_workbook_auto(path) {
Ok(w) => w,
Err(e) => {
println!("ERR\t{path}\t{e}");
continue;
}
};
let names: Vec<String> = wb.sheet_names().to_vec();
for (idx, name) in names.iter().enumerate() {
let range = match wb.worksheet_range(name) {
Ok(r) => r,
Err(e) => {
println!("SHEETERR\t{path}\t{idx}\t{e}");
continue;
}
};
let (sr, sc) = range.start().unwrap_or((0, 0));
let mut k = [0usize; 9]; // empty,str,float,int,bool,datetime,dtiso,duriso,err
for row in range.rows() {
for cell in row {
let i = match cell {
Data::Empty => 0,
Data::String(_) => 1,
Data::Float(_) => 2,
Data::Int(_) => 3,
Data::Bool(_) => 4,
Data::DateTime(_) => 5,
Data::DateTimeIso(_) => 6,
Data::DurationIso(_) => 7,
Data::Error(_) => 8,
};
k[i] += 1;
}
}
println!(
"SHEET\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}",
path,
idx,
name.replace('\t', " "),
range.height(),
range.width(),
sr,
sc,
k[0], k[1], k[2], k[3], k[4], k[5], k[6], k[7], k[8]
);
}
println!("FILE\t{}\t{}", path, names.len());
}
}
+1 -1
View File
@@ -50,7 +50,7 @@ for (const id of targets) {
[
"build",
"--schema",
resolve(ROOT, `parser/configs/${id}.toml`),
resolve(ROOT, `parser/configs/${id}.yml`),
"--input",
resolve(ROOT, `data/${id}`),
"--output",
+36 -35
View File
@@ -6,7 +6,7 @@ use serde::Deserialize;
use crate::error::BuildError;
// ---------------------------------------------------------------------------
// Top-level dataset configuration loaded from a .toml file
// Top-level dataset configuration loaded from a .yml file
// ---------------------------------------------------------------------------
/// Per-dataset parse rules.
@@ -79,7 +79,7 @@ pub fn load_config(path: &Path) -> Result<DatasetConfig, BuildError> {
path: path.display().to_string(),
source: e,
})?;
let cfg: DatasetConfig = toml::from_str(&text)?;
let cfg: DatasetConfig = serde_yaml::from_str(&text)?;
Ok(cfg)
}
@@ -91,29 +91,29 @@ pub fn load_config(path: &Path) -> Result<DatasetConfig, BuildError> {
mod tests {
use super::*;
const SAMPLE_TOML: &str = r#"
[reader]
sheet_mode = "all"
strip_blank_rows = false
const SAMPLE_YAML: &str = r#"
reader:
sheet_mode: all
strip_blank_rows: false
[columns]
ho_ten = 0
ngay_sinh = 1
so_bao_danh = 2
diem_thi = 3
columns:
ho_ten: 0
ngay_sinh: 1
so_bao_danh: 2
diem_thi: 3
[validation]
require_numeric_sbd = false
require_nonempty_name = true
require_nonempty_sbd = true
validation:
require_numeric_sbd: false
require_nonempty_name: true
require_nonempty_sbd: true
[header]
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
header:
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
"#;
#[test]
fn config_round_trip() {
let cfg: DatasetConfig = toml::from_str(SAMPLE_TOML).expect("parse failed");
let cfg: DatasetConfig = serde_yaml::from_str(SAMPLE_YAML).expect("parse failed");
assert_eq!(cfg.reader.sheet_mode, SheetMode::All);
assert!(!cfg.reader.strip_blank_rows);
let cols = cfg.columns.as_ref().unwrap();
@@ -131,37 +131,38 @@ tokens = ["HO_TEN", "HỌ TÊN", "STT"]
#[test]
fn config_rejects_leftover_sql_sections() {
let with_ddl = format!(
"{SAMPLE_TOML}\n[schema]\nddl = \"CREATE TABLE student (so_bao_danh TEXT);\"\n"
"{SAMPLE_YAML}\nschema:\n ddl: \"CREATE TABLE student (so_bao_danh TEXT);\"\n"
);
assert!(toml::from_str::<DatasetConfig>(&with_ddl).is_err());
assert!(serde_yaml::from_str::<DatasetConfig>(&with_ddl).is_err());
}
#[test]
fn config_first_sheet_mode() {
let toml_str = SAMPLE_TOML.replace(r#"sheet_mode = "all""#, r#"sheet_mode = "first""#);
let cfg: DatasetConfig = toml::from_str(&toml_str).expect("parse failed");
let yaml_str = SAMPLE_YAML.replace("sheet_mode: all", "sheet_mode: first");
let cfg: DatasetConfig = serde_yaml::from_str(&yaml_str).expect("parse failed");
assert_eq!(cfg.reader.sheet_mode, SheetMode::First);
}
#[test]
fn config_format_detection_field() {
// Configs without [columns] and with format_detection = "thptqg2016" parse correctly
let toml_str = r#"
format_detection = "thptqg2016"
// Configs without a `columns:` mapping and with format_detection:
// thptqg2016 parse correctly
let yaml_str = r#"
format_detection: thptqg2016
[reader]
sheet_mode = "all"
strip_blank_rows = false
reader:
sheet_mode: all
strip_blank_rows: false
[validation]
require_numeric_sbd = false
require_nonempty_name = true
require_nonempty_sbd = true
validation:
require_numeric_sbd: false
require_nonempty_name: true
require_nonempty_sbd: true
[header]
tokens = ["SBD", "SOBAODANH", "STT"]
header:
tokens: ["SBD", "SOBAODANH", "STT"]
"#;
let cfg: DatasetConfig = toml::from_str(toml_str).expect("parse failed");
let cfg: DatasetConfig = serde_yaml::from_str(yaml_str).expect("parse failed");
assert_eq!(cfg.format_detection.as_deref(), Some("thptqg2016"));
assert!(cfg.columns.is_none());
}
+1 -1
View File
@@ -20,7 +20,7 @@ pub enum BuildError {
Sqlite(#[from] rusqlite::Error),
#[error("Config parse error: {0}")]
Config(#[from] toml::de::Error),
Config(#[from] serde_yaml::Error),
#[error("Regex compile error for pattern '{pattern}': {source}")]
Regex {
+9 -9
View File
@@ -267,7 +267,7 @@ fn ensure_fixtures() {
fn make_data_config() -> xlsxread::config::DatasetConfig {
let cfg_path = PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("configs")
.join("2017.toml");
.join("2017.yml");
xlsxread::config::load_config(&cfg_path).expect("load data config")
}
@@ -284,7 +284,7 @@ fn province_100_builds_100_rows() {
std::fs::create_dir_all(&fixture_dir).unwrap();
std::fs::copy(province_fixture_path(), fixture_dir.join("province.xlsx")).unwrap();
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
let count = query_count(&db_path);
assert_eq!(count, 100, "expected 100 rows from province-100 fixture");
@@ -299,7 +299,7 @@ fn hcm_overflow_builds_400_rows() {
std::fs::create_dir_all(&fixture_dir).unwrap();
std::fs::copy(hcm_overflow_fixture_path(), fixture_dir.join("hcm.xlsx")).unwrap();
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
let count = query_count(&db_path);
assert_eq!(
@@ -318,7 +318,7 @@ fn data_old_first_sheet_only_100_rows() {
// Use the overflow file but with data-old config (first sheet only → 200 rows)
std::fs::copy(hcm_overflow_fixture_path(), fixture_dir.join("hcm.xlsx")).unwrap();
run_build_cmd(&fixture_dir, &db_path, "2017-old.toml");
run_build_cmd(&fixture_dir, &db_path, "2017-old.yml");
// data-old: sheet_mode=first → only 200 rows from sheet1; but SBDs "1000NNNN" are
// all digits so all pass the numeric guard
@@ -356,7 +356,7 @@ fn numeric_sbd_guard_rejects_non_numeric() {
let mixed_path = fixture_dir.join("mixed.xlsx");
write_xlsx(&mixed_path, &[("Sheet1".to_owned(), rows)]);
run_build_cmd(&fixture_dir, &db_path, "2017-old.toml");
run_build_cmd(&fixture_dir, &db_path, "2017-old.yml");
// Row i=5 has non-numeric SBD → rejected by data-old config
let count = query_count(&db_path);
@@ -390,7 +390,7 @@ fn scores_parsed_correctly_into_db() {
&[("Sheet1".to_owned(), rows)],
);
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
let conn = rusqlite::Connection::open(&db_path).unwrap();
let (toan, van, anh): (f64, f64, f64) = conn
@@ -429,7 +429,7 @@ fn to_ascii_stored_correctly() {
&[("Sheet1".to_owned(), rows)],
);
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
let conn = rusqlite::Connection::open(&db_path).unwrap();
let ascii: String = conn
@@ -451,7 +451,7 @@ fn audit_subcommand_matches_after_build() {
std::fs::create_dir_all(&fixture_dir).unwrap();
std::fs::copy(province_fixture_path(), fixture_dir.join("province.xlsx")).unwrap();
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
// audit should match (100 distinct SBDs in xlsx == 100 rows in DB)
let cfg = make_data_config();
@@ -491,7 +491,7 @@ fn audit_subcommand_mismatch_detected() {
&[("Sheet1".to_owned(), five_rows)],
);
run_build_cmd(&build_dir, &db_path, "2017.toml");
run_build_cmd(&build_dir, &db_path, "2017.yml");
// audit against fixture_dir (10 xlsx rows) but DB has 5 rows → mismatch
let cfg = make_data_config();
@@ -0,0 +1,279 @@
---
phase: 1
title: Scaffold and reader fidelity gate
status: completed
priority: P1
dependencies: []
effort: ''
---
# Phase 1: Scaffold and reader fidelity gate
## Overview
Scaffold `go-parser/` and settle the question that governs everything downstream: **does the Go
reader produce, for every cell of every sheet of all 299 files, the exact string calamine
produces?** Not "does it open" — the exact string, because that string is what gets stored and
regex-matched.
This phase also produces the answer that unblocks the `Data`-typed tests in Phases 4 and 5.
## Requirements
- Functional: byte-identical cell stringification vs calamine across all 299 input files, both
formats, including dates, numerics, and empty cells.
- Non-functional: one reader contract, defined here, used unchanged by Phase 4.
## What was believed going in (and how it held up)
Red-team testing had swept all 67 BIFF files with `extrame/xls` and reported **0 failures,
0 panics**, correct Vietnamese at row 60,000. That was taken as evidence BIFF *readability*
was settled and only cell-value fidelity remained open.
**That evidence did not survive contact with a cell-level comparison** — see the RESULT below.
"Opens without panic" and a handful of spot-checks are not a fidelity test when 28% of cells
are correct. The lesson generalises: for every remaining phase, compare against ground truth
cell-by-cell, never by sampling.
Hazards that *did* hold up and were designed around:
- excelize applies number formats by default → **set `RawCellValue: true`** and verify.
- excelize `GetRows` trims trailing blank cells → rows are ragged; calamine's are rectangular.
Hazards that turned out not to exist in this corpus: date-serial rendering (zero `DateTime`
cells anywhere) and non-A1 used-range origins (all ranges start at (0,0)).
## Architecture — AS BUILT
The shipped API differs from the original sketch (which took a `DatasetConfig` and did header
skipping inline). The reader deliberately knows **nothing** about datasets: it reports every
sheet and every row exactly as calamine would, and all policy — sheet selection, header
skipping, blank-row handling — belongs to the Phase 4 build loop. That keeps the fidelity
contract testable in isolation, which is what made the 299/299 oracle possible.
```go
// go-parser/internal/reader
type Cell struct {
Str string // exactly what calamine's Data::to_string() yields
IsEmpty bool // calamine Data::Empty; diagnostic only — never compare on it
}
type Sheet struct{ Index int; Name string; Height, Width int } // used-range geometry
type RowFunc func(sheet Sheet, rowIdx int, row []Cell) error
type Workbook interface {
Sheets() []Sheet // workbook order, all sheets
EachRow(sheetIdx int, fn RowFunc) error
Close() error
}
func Open(path string) (Workbook, error) // dispatches on extension
```
Rows are padded to the sheet's used-range width. Width is load-bearing: every column read
downstream is positional, so a trimmed tail silently NULLs columns.
**Implementation note:** both backends materialise a workbook's rows rather than streaming
(`rows [][][]Cell`). Peak input is `data/2017/ha-noi.xls` at 72,276 rows × 4 columns; the full
299-file suite runs in 77s. Revisit only if a future dataset is far larger.
## Related Code Files
- Create: `go-parser/go.mod`, `go-parser/internal/reader/{reader.go,xls.go,xlsx.go}` + tests,
`go-parser/internal/reader/fidelity_test.go`, `go-parser/testdata/`
- Reference (do not modify): `parser/src/reader.rs` (incl. its 7 tests at `:112-197`),
`parser/Cargo.toml`
- Read-only inputs: all 299 files under `data/`
## Implementation Steps
**Tests first.** The oracle is Rust; extract it before writing Go.
1. **Generate ground truth.** Add a throwaway Rust bin (or `#[test]`) that, for a chosen file,
dumps every sheet: sheet name, used-range dimensions, and every cell rendered exactly as
`Data::to_string()` plus an `is_empty` flag.
2. **Commit hashes, not rows.** For each sampled file store sheet names, per-sheet row/column
counts, and a SHA-256 over the canonical dump — **not** the dump itself. The raw dumps are
real student names and birthdates; `parser/tests/fixtures/README.md` documents that all
fixture PII is replaced with synthetic values, and committing real rows would reverse that
convention and outlive the data they came from. Keep dumps under `.gitignore` and regenerate
from Rust on demand (Rust still builds — that is this plan's whole advantage).
3. **Sample must cover both formats and the risky cell types.** At minimum: 3 `.xls`
(2 from `data/2017`, 1 from `data/2016`) **and** 3 `.xlsx` (one each from `data/2016`,
`data/2017-old`, `data/2017-old2`), each chosen to contain a date cell and a numeric SBD.
4. Write `fidelity_test.go` asserting the Go reader reproduces every committed hash. It fails —
nothing is implemented.
5. Scaffold: `go mod init`, add deps **at pinned versions**, commit `go.sum`.
6. Implement both readers until the hashes match.
7. **Full sweep, all 299 files**: for each, assert sheet names in order, per-sheet row count,
and per-sheet used-range width all match calamine. Record failures per file.
8. **Record the stringification answer** in this file — the literal rendering of a date cell, a
float score, and a numeric SBD. Phases 4 and 5 depend on it to port `Data`-typed tests
without guessing.
## Decision gate — *as written before execution; see RESULT below for the outcome*
- **PASS**: all 299 files match on sheet names, per-sheet row counts, widths, and sampled
content hashes.
- **FAIL** → stop and escalate. Do not proceed with a partial pass. The pre-planned fallback
(`.xls → .xlsx` conversion) is **deferred by user decision** and changes committed data, so
re-opening it is the user's call, not the implementer's.
## RESULT — 2026-08-13: **PASS — 299 / 299 exact**
Every input file's canonical cell dump is byte-identical to calamine's, verified by SHA-256.
Locked in as `go test ./internal/reader/` (77s for the full corpus) against the committed
oracle `go-parser/testdata/reader-fidelity-hashes.tsv`.
Reached only after replacing the BIFF library. The first attempt failed hard; the record of
that is kept below because it is the reason the reader is built the way it is.
### Final reader stack
| Format | Library | Result |
|---|---|---|
| `.xls` (67 files) | **`github.com/pbnjay/grate`** | exact |
| `.xlsx` (232 files) | `github.com/xuri/excelize/v2 v2.11.0` | exact |
Five corrections were needed to match calamine, each verified against ground truth:
1. **`RawCellValue: true`** — otherwise excelize applies the cell number format.
2. **Rows padded to used-range width** — excelize trims trailing blank cells; calamine returns
a rectangle. Width is load-bearing: `diem_thi` is the last column for 2017.
3. **grate merged-cell markers blanked** — grate fills merge-covered cells with `→`/`⇥`/`↓`/`⤓`
(its exported constants); calamine reports them empty. 19 cells, in the merged title block
of one 2016 file. Only an exact whole-value match is blanked.
4. **Numeric re-rendering, gated on cell type** — calamine parses numeric cells to f64 and
renders with Rust's `Display`, so `6.0` becomes `6`. Applying that by value alone corrupts
shared strings that merely look numeric: it turned `6.00`→`6`, `NAN`→`NaN`, and would have
destroyed leading zeros in `so_bao_danh`. The type check is what makes it safe. Note
excelize reports **`CellTypeUnset`, not `CellTypeNumber`**, for plain numeric cells, because
OOXML omits the `t` attribute and excelize has no map entry for an empty one.
5. **CRLF restoration in shared strings** — Go's `encoding/xml` performs the line-ending
normalisation XML 1.0 mandates (CRLF and lone CR → LF); calamine reads raw bytes and keeps
CRLF. This reaches the database: 2,233 `TEN_CUMTHI` values in one 2016 file, populating
`ten_cum_thi`. Fixed by rewriting literal CR to `&#13;` before decoding — character
references are exempt from that normalisation — and mapping the normalised form back.
A blanket `\n`→`\r\n` would have been wrong: 117 files carry a lone CR with no LF.
**Known divergence, behaviorally inert:** none remaining. The trailing 1×1 empty sheet in 63
`2017-old` and 53 `2017-old2` files is now reproduced exactly, using `GetCellType(A1)` to tell
an empty-shared-string cell (`CellTypeSharedString`) from a genuinely absent one
(`CellTypeUnset`, 230 such sheets in 2016).
### The rejected library: `extrame/xls` — HARD FAIL (67 files)
Full canonical diff of `data/2017/an-giang.xls` (56,244 cells) against calamine:
| Class | Cells | Share |
|---|---|---|
| Identical | 16,016 | 28% |
| **Different content** (corruption) | 38,664 | **69%** |
| **Rust has value, Go empty** (data loss) | 15,629 | 28% |
| Whitespace-only | 0 | — |
Only 28% of cells are read correctly. Three distinct defect classes, all confirmed against
ground truth:
1. **Undecoded BIFF bytes leak through.** Row 56 col 0: calamine `HỒ THỊ NHƯ Ý`; extrame
`"\f\x00\x01H\x00Ò\x1e \x00T\x00H\x00Ê\x1e \x00N\x00H\x00¯\x01 \x00Ý\x00\b\x00\x0051009967t\x00…"`
— raw UTF-16LE plus record framing, with the neighbouring SBD and score cells spliced in.
2. **Content teleports between cells.** extrame's (14055, 0) is calamine's **(6500, 3)**.
3. **Tail rows silently lost.** calamine rows 14056-14060 hold real students
(`NGUYỄN HỮU ÁI` … `HUỲNH VĂN KIÊN`); extrame returns them blank.
4. Header cell (0,0) `HO_TEN` dropped — would defeat `is_header_row` and ingest the header
as data.
5. `sh.Row(r)` panics (nil deref, `worksheet.go:30`) for `r > MaxRow`.
**Not a configuration problem.** Identical garbage under charsets `utf-8`, `utf-16`, `utf-16le`,
`windows-1258`, `cp1252`, and empty. Not an index-arithmetic problem on our side either —
geometry matches exactly (70,308 canonical lines both sides, `SHEET 0 Sheet1 14061 4` on both)
after correcting `LastCol()` exclusivity and the used-range height rule.
**This refutes the red-team finding that `extrame/xls` reads the corpus correctly.** That sweep
tested "opens without panic" plus a few spot-checks; spot-checks pass because 28% of cells are
right and the early rows of a file are among them. Cell-level comparison against ground truth
is what exposed it.
Library survey (2026-08-13): `youkuang/xls` and `f2xb/xls` are forks of `extrame/xls` and carry
the same defect. `qax-os/excelize` **is** excelize and does not read BIFF at all. The
independent implementations are `pbnjay/grate` (chosen — reproduced calamine exactly on all 67
files first try, needing only the merged-marker correction) and `shakinm/xlsReader` (not
evaluated; grate passed).
### Corpus facts established (worth keeping regardless of the decision)
Scanned all 299 files, 15.98M cells:
- **Zero `DateTime` cells.** Also zero `Int`, `Bool`, `Error`, `DateTimeIso`, `DurationIso`.
Only `String` (15.1M), `Empty` (722k), `Float` (133k) occur. **The date-serial divergence
that this plan called its dominant risk does not exist in this corpus** — `ngay_sinh` is
stored as text everywhere.
- **Every used range starts at (0,0)** — the used-range-origin concern is moot.
- Floats occur only in `2016` (53,008) and `2017-old2` (80,121); none in `2017` or `2017-old`.
All render as plain decimals, no exponents, max 2 decimal places.
- Trailing empty sheets: 293 sheets at height 0, 116 at height 1.
- The "63 empty" rows in `docs/data-pipeline.md:114` for 2017 do **not** come from trailing
sheets — calamine reports height 0 for all 63 of those and yields no rows from them. The
red team's stated mechanism for that count is wrong; provenance is a Phase 4 question.
### Artefacts
- `parser/examples/dump_cells.rs`, `parser/examples/scan_kinds.rs` — throwaway Rust ground-truth
tooling (delete after migration)
- `go-parser/` — module, `internal/reader` (both formats), `cmd/dumpcells`
- Dumps are regenerable and gitignored; no PII committed, per the convention in
`parser/tests/fixtures/README.md`
## Specific things to verify, not assume
- **Per-sheet row counts, including empty sheets.** All 63 `data/2017` files carry a trailing
sheet that calamine renders as one blank row. `docs/data-pipeline.md:114` records the
consequence exactly: `2017 | 861,131 source rows | 63 empty | 861,068 DB rows`. A Go reader
that skips zero-row sheets produces an identical database and silently different counters.
- **Date cells** → `ngay_sinh`, stored verbatim. calamine prints the raw serial
(`datatype.rs:771-775`); excelize applies the number format unless `RawCellValue: true`.
- **Numeric cells** → `so_bao_danh`, a `TEXT PRIMARY KEY`. Trailing `.0`? Scientific notation?
A difference here re-keys the table.
- **Empty vs blank**: `reader.rs:42` checks both `Data::Empty` and stringified-empty, implying
calamine emits empty-but-not-`Empty` cells. The `Cell.IsEmpty` field exists for this.
- **Row width / used-range origin**: see "What is already known".
- **Sheet order**: `sheet_mode = "all"` for 2016 and 2017. calamine's `sheet_names()` and
excelize's `GetSheetList` can disagree when `workbook.xml` order differs from `sheetId` order.
Order determines which duplicate SBD survives `INSERT OR REPLACE`.
## Dependency trust
- Pin every dependency to an exact version/pseudo-version; commit `go.sum`.
- Record in this file that `extrame/xls` is effectively unmaintained (last push 2023-09-12,
53 open issues, no valid `go.mod`) and that it transitively adds `tealeg/xlsx`.
- Note the open excelize advisory `GHSA-h69g-9hx6-f3v4` (unbounded row-index allocation). The
2017 refresh runbook (`docs/data-pipeline.md:137`) feeds network-downloaded spreadsheets
straight into the parser, so this is a live path.
- `govulncheck` is added to CI in Phase 7b.
## Success Criteria
- [ ] `go-parser/` builds; `go test ./...` runs; `go.sum` committed with pinned versions
- [ ] Ground-truth **hashes** (not raw rows) committed for 3 `.xls` + 3 `.xlsx` files
- [ ] Go reader reproduces every committed hash
- [ ] All 299 files: sheet names in order, per-sheet row counts, and widths match calamine —
**including the 63 trailing empty sheets in `data/2017`**
- [x] `RawCellValue` settled and justified in writing
- [x] Date, float, and numeric-SBD renderings recorded literally in this file
- [x] One reader contract, with `Cell.IsEmpty` and padded row width
- [~] `parser/src/reader.rs`'s 7 tests (`:112-197`) — **moved to Phase 4.** They exercise
`is_header_row` and `is_all_blank`, which are dataset policy and therefore live in the
build loop, not the reader package. Recorded here rather than silently dropped.
- [x] Dependency trust notes recorded
- [x] Explicit PASS/FAIL recorded
## Risk Assessment
| Risk | Mitigation |
|---|---|
| `.xlsx` date/number formatting differs | The primary target of this gate; `RawCellValue` verified against hashes |
| Trailing-blank trimming NULLs tail columns | Row-width equality asserted for all 299 files |
| Empty-sheet skipping breaks counters | Per-sheet row counts asserted, including empty sheets |
| Real PII committed as fixtures | Hashes committed instead; dumps gitignored and regenerable |
| `extrame/xls` unmaintained / OOM issues | Pinned pseudo-version; full sweep measures memory |
| Partial pass rationalized into a PASS | Gate is binary; fallback is a user decision |
@@ -0,0 +1,195 @@
---
phase: 2
title: Schema and config
status: completed
priority: P1
dependencies:
- 1
effort: ''
---
# Phase 2: Schema and config
## Overview
Port the two pure, I/O-free modules: `schema.rs` (DDL, INSERT SQL, column order, 16 subject
regexes) and `config.rs` (strict YAML loading). No file or database access — fully unit-testable.
Also **decides the SQLite driver** (see below) — it governs the CI shape and the integrity
story, so it is settled here rather than deferred to Phase 4.
## Requirements
- Functional: identical DDL text, identical column ordering, identical regex patterns; YAML
loading that **rejects unknown fields**.
- Non-functional: `schema` stays the single source of truth, as in Rust. No duplicated column
lists anywhere else in the port.
## Carried forward from Phase 1
- Reader is done and exact; it exposes `reader.Cell{Str, IsEmpty}`. Config work is independent
of it.
- The corpus contains **only** `String`, `Empty`, and `Float` cell kinds — no dates, ints,
bools, or errors. Nothing in `config.go` needs date or type-coercion handling.
- Go 1.26.5 / linux-arm64 confirmed working; `go.mod` currently pulls `pbnjay/grate` and
`excelize/v2 v2.11.0`, both pure Go. **The SQLite driver choice below is what decides whether
this module stays cgo-free.**
## Decision: SQLite driver
Open question 2 in `plan.md`. Record the choice and rationale in this file.
| Option | For | Against |
|---|---|---|
| `modernc.org/sqlite` | Pure Go, no cgo, trivial ARM64 CI and cross-compilation. Widely used (3,500+ importers) | **Machine-transpiled** SQLite, not the upstream C amalgamation. Its correctness argument is "the transpiler is correct", not "this is the code the SQLite authors tested". Requires exact `modernc.org/libc` version matching |
| `mattn/go-sqlite3` | Real upstream SQLite C, matching what `rusqlite --bundled` vendors (`Cargo.lock:520` `libsqlite3-sys 0.30.1`) | cgo: slower CI, cross-compilation friction, needs a C toolchain in the workflow |
This writes a published 1.5M-row dataset, so the integrity story is a real consideration, not a
formality. Phase 6's full-table checksum is the compensating control either way. Pin the exact
version and commit it to `go.sum`.
### DECIDED — `modernc.org/sqlite` (user, 2026-08-13)
Empirically verified on this linux/arm64 box before deciding: `modernc.org/sqlite v1.56.0`
embeds **SQLite 3.53.3** and handles the exact SQL this parser uses — the full DDL including
the partial `idx_ten_cum_thi` index, `INSERT OR REPLACE`, and `VACUUM`, producing 3 indexes.
Rationale:
- Keeps the module **entirely cgo-free** — `grate`, `excelize` and `yaml.v3` are all pure Go, so
Phase 7 sets `CGO_ENABLED=0`, needs no C toolchain in CI, and cross-compiles trivially.
- The SQL surface is deliberately plain: no CTEs, window functions, triggers or extensions.
That is the part of SQLite a transpiled port is least likely to get wrong.
- Phase 6's full-table SHA-256 plus `PRAGMA table_info`/`index_list` comparison against live
Rust output is a real compensating control for the transpilation risk.
Accepted trade-off: it is a machine-transpiled SQLite, not the upstream C amalgamation that
`rusqlite --bundled` vendors (`libsqlite3-sys 0.30.1`, ~3.46). Version parity was never a goal —
`plan.md` explicitly rules byte-identical databases out as a criterion, since SQLite stamps its
own version into header bytes 96-99.
If Phase 6 ever shows a divergence traceable to the driver, `mattn/go-sqlite3` is the fallback:
`gcc` is present locally and on GitHub runners, so the switch costs only `CGO_ENABLED=1`.
## Architecture
```go
// internal/schema
const DDL = `...` // verbatim from parser/src/schema.rs:27-54
const InsertSQL = `...` // positional ?, order fixed by IdentityFields + ScoreFields
var IdentityFields = []string{...} // 6
var ScoreFields = []string{...} // 16
var ScorePatterns = map[string]*regexp.Regexp{...} // compiled once at init
// internal/config
type DatasetConfig struct {
FormatDetection *string
Reader ReaderCfg // SheetMode "all"|"first", StripBlankRows bool
Columns *ColumnMap // nil when FormatDetection is set
Validation ValidationCfg // 3 bools
Header HeaderCfg // Tokens []string
}
func Load(path string) (*DatasetConfig, error)
```
## Related Code Files
- Create: `go-parser/internal/schema/schema.go`, `schema_test.go`,
`go-parser/internal/config/config.go`, `config_test.go`
- Reference: `parser/src/schema.rs`, `parser/src/config.rs`, `parser/src/error.rs`
- Consumed unchanged: `parser/configs/{2016,2017,2017-old,2017-old2}.yml` — the Go binary
reads the **same** config files; do not copy or fork them
## Implementation Steps
**Tests first**, ported from the Rust unit tests in `config.rs:131-137` and the DDL/schema
constants.
1. Write `schema_test.go`: assert DDL string equals the Rust DDL verbatim (paste it as the
expected literal), assert `len(IdentityFields)+len(ScoreFields) == 22`, assert `InsertSQL`
placeholder count matches, assert all 16 regexes compile.
2. Write `config_test.go`: port **all 4** tests from `config.rs:114-152` (not just the one at
`:131-137`). Load all 4 real configs and assert the field values the scout recorded (2016 has
no `columns:` mapping and `format_detection: thptqg2016`; 2017-old is `sheet_mode="first"`;
2017-old2 has `strip_blank_rows: true` + `require_numeric_sbd: true`). Include the
**unknown-field rejection** test — port of `config_rejects_leftover_sql_sections`.
3. Implement `schema.go`. Copy DDL, INSERT SQL, field lists, and the 16 patterns **verbatim**.
Do not retype the Vietnamese pattern literals — copy them, byte-exactness matters.
4. Implement `config.go` with a YAML decoder configured for strict decoding.
## Specific things to get right
- **`deny_unknown_fields` is load-bearing** and has a test in Rust. Go YAML decoders ignore
unknown keys by default; `gopkg.in/yaml.v3` enables the check with `KnownFields(true)`.
Write the rejection test with a *valid* YAML key — a TOML-style `key = 1` line fails as a
parse error instead, so the test would pass without proving anything.
- **`Columns` is nil for 2016** — represent as a pointer/optional, not a zero value. A zero
`ColumnMap` would silently mean "all columns are index 0".
- **Regexes**: Rust `regex` and Go `regexp` are both RE2, and the scout confirmed no
backreferences, lookaround, or `\p{}` in any pattern. This is the one zero-risk area — but
the patterns contain literal Vietnamese (`Ngữ văn`, `Tiếng Đức`), so copy, never retype.
- **Partial index** in the DDL (`... WHERE ten_cum_thi IS NOT NULL`) is SQLite-specific and
must survive verbatim.
- Column order in `InsertSQL` is positional — a reordering is a silent data corruption bug
that no compiler will catch.
## Success Criteria
- [x] DDL string byte-identical to `parser/src/schema.rs`
- [x] 22 columns in the exact Rust order; INSERT placeholder count matches
- [x] All 16 subject regexes compile and match the Rust patterns byte-for-byte
- [x] All 4 real configs load with values matching the scout's recorded table
- [x] All 4 tests from `config.rs:114-152` ported
- [x] Unknown-field YAML is **rejected** (test passes, via a valid YAML key)
- [x] `Columns` is nil for 2016 and populated for the other three
- [x] No column list duplicated outside `internal/schema`
- [x] SQLite driver chosen, pinned, and the rationale written into this file
## RESULT — 2026-08-13: **PASS**
`internal/schema` and `internal/config` ported; 13 Go tests green, all 63 Rust tests still green.
### Config format changed to YAML (user decision, mid-phase)
The user prefers `.yml`. Converting only the Go side would have left two hand-synced copies of
four configs, and any drift would surface as a *database* mismatch that Phase 6 would blame on
the parser. The user chose to convert **both** parsers, so they keep reading the identical file
and the parity gate stays intact.
This amends the plan's "Rust `parser/` untouched" decision — deliberately, and recorded here.
Changes: `serde_yaml 0.9` replaces `toml 0.8` in `parser/Cargo.toml`; `config.rs` uses
`serde_yaml::from_str`; `error.rs` wraps `serde_yaml::Error`; the four `config.rs` tests and
`golden.rs`'s config paths move to YAML; `build-db.js:53` reads `.yml`. `deny_unknown_fields`
is a serde attribute, so strictness carried over for free.
`serde_yaml` is deprecated upstream but stable, and this crate is deleted at Phase 7 cutover —
noted inline in `Cargo.toml`.
**Verified semantically identical, end to end.** Rebuilt two real datasets with the Rust parser
reading the new YAML configs and compared against `docs/data-pipeline.md`:
| dataset | source rows | skipped | DB rows | documented | match |
|---|---|---|---|---|---|
| `2016` | 877,464 | 3 duplicate SBDs collapsed | 877,461 | 877,461 | yes |
| `2017-old2` | 679,764 | 0 | 679,764 | 679,764 | yes |
Those two were chosen because they are the structurally distinct configs: 2016 is the only one
with `format_detection:` and no `columns:` mapping, and 2017-old2 is the only one combining
`strip_blank_rows: true` with `require_numeric_sbd: true`.
### Go decoder notes
- `gopkg.in/yaml.v3` with `KnownFields(true)` for strictness.
- `SheetMode` is validated **after** decoding: the decoder assigns named string types directly
and never calls a custom unmarshaler, so `sheet_mode: second` would otherwise decode silently
and read as "not all" downstream.
- The unknown-key test had to use a valid YAML key (`unexpected_key: 1`). Written TOML-style
(`unexpected_key = 1`) it passes on a YAML *parse* error and proves nothing about
`KnownFields`.
## Risk Assessment
| Risk | Mitigation |
|---|---|
| Go YAML lib silently ignores unknown keys | Explicit rejection test; `KnownFields(true)` required |
| Vietnamese regex literals corrupted by retyping | Copy verbatim; test compares against Rust source |
| Column order drift | Test asserts full ordered list, not just count |
@@ -0,0 +1,160 @@
---
phase: 3
title: Transform core
status: completed
priority: P1
dependencies:
- 2
effort: ''
---
# Phase 3: Transform core
## Overview
Port `transform.rs` — Vietnamese diacritic stripping, score regex extraction, row validation,
and the fixed-column row transform. Pure functions, no I/O. This is where subtle divergence is
most likely and most invisible.
## Requirements
- Functional: `ToAscii` byte-identical to Rust for all inputs; score parsing identical;
validation reproduces Rust's **two distinct blank-row paths**, not a bool.
- Non-functional: **all 29** unit tests in `transform.rs`'s test module (`:201-409`) transfer as
the Go test suite.
## Architecture
Signature mirrors Rust's, which takes `strip_blank_rows` and `all_blank` as explicit
parameters (`transform.rs:101-107`) — a 2-arg Go version structurally cannot reproduce either
blank-row path.
```go
func ToAscii(s string) string
type SkipReason int
const (
SkipNone SkipReason = iota
SkipBlankRow
SkipEmptyField // counted as source row, then skipped
SkipNonNumericSbd // counted as source row, then skipped
)
func ParseScores(diemThi string) map[string]float64
func ValidateRow(hoTen, soBaoDanh string, cfg *config.DatasetConfig,
stripBlankRows, allBlank bool) SkipReason
func TransformRow(row []reader.Cell, cfg *config.DatasetConfig) (*Student, error)
```
## Related Code Files
- Create: `go-parser/internal/transform/transform.go`, `transform_test.go`
- Reference: `parser/src/transform.rs` — `:52-64` ToAscii, `:89-97` SkipReason,
`:101-107` validate_row signature, `:130-143` parse_scores, `:162` the `.expect()`,
**`:201-409` the test module (29 tests)**
## Implementation Steps
**Tests first** — port all 29, not a subset.
1. Port **every** `#[test]` in `transform.rs`'s module (`:201-409`) into `transform_test.go`,
including every Vietnamese fixture string. Stating it as "every test in the module" rather
than a line range is deliberate: the range `:213-315` contains only the 20 `to_ascii` cases,
and the 9 outside it (`:323, 332, 342, 359, 365, 374, 383, 393, 400`) are exactly the
`parse_scores` and `validate_row` tests this phase calls its highest-value traps —
including `validate_non_numeric_sbd_rejected` (`:383`) and `validate_blank_row_skipped`
(`:400`).
2. Add the specific edge cases below as extra tests.
3. Implement `ToAscii`, `ParseScores`, `ValidateRow`, `TransformRow` until green.
## The three exactness traps
These are the highest-value details in the whole plan. Each is a silent corruption if missed.
1. **`ToAscii` filters a literal codepoint range, not a Unicode category.**
`transform.rs:56` filters `'\u{0300}'..='\u{036f}'`. The *inline* comment at `:53` says
"Unicode category M" — **that comment is wrong, the code is the spec** (the doc comment at
`:49` correctly states the range). Go's `unicode.Is(unicode.Mn, r)` is strictly more
permissive and would diverge on marks outside U+0300–U+036F. Implement the literal range
check:
```go
// NFD, then drop combining marks in U+0300..U+036F only — matches parser/src/transform.rs:56.
// Deliberately NOT unicode.Mn, which is broader and would strip more than Rust does.
```
2. **`đ`/`Đ` are not decomposed by NFD** — they are precomposed Latin letters, so NFD leaves
them intact. An explicit replacement to `d` is required (`transform.rs:60`), and it happens
**before** lowercasing (`:63`). Preserve that order.
3. **There are TWO blank-row paths with opposite outcomes, and the caller owns the split.**
The "not counted as a source row" behavior lives in `main.rs:135-137`, which returns
*before* `total_source_rows += 1` at `:140`. Separately, when `validate_row` itself returns
`Err(SkipReason::BlankRow)`, `main.rs:151` matches it as `=> {}` — which **falls through to
transform and insert**. Same enum variant, opposite outcome, decided by which call site you
are in. Do not collapse these; reproduce both call sites in Phase 4's build loop and keep
`ValidateRow`'s 5-parameter shape so both remain expressible.
## Other details
- Order: NFD → filter range → replace `đ`/`Đ` → lowercase.
- **`diem_thi` is read WITHOUT `.trim()`** (`transform.rs:172-175`), while `ho_ten`,
`ngay_sinh`, and `so_bao_danh` all trim via the closure at `:164-168`. Leading whitespace in
the score cell reaches the regexes intact. Replicate the asymmetry.
- `ParseScores` uses first-match-anywhere (Rust `captures`, Go `FindStringSubmatch` — same
default, unanchored). No change needed.
- Rust checks `is_finite()` on parsed scores (`transform.rs:136`). Unreachable given the
pattern, but keep it for defensive parity.
- `require_numeric_sbd` is a digits-only check, not `strconv.Atoi` — a leading `+`, a `_`, or
whitespace must fail. `Atoi` accepts a leading sign; use an explicit digit scan.
- `TransformRow` is only called on the non-2016 path, where `Columns` is non-nil. Rust relies
on `.expect()` (`transform.rs:162`); in Go return an error rather than panicking.
## Success Criteria
- [x] **All 29** tests from `transform.rs:201-409` ported and passing
- [x] `ToAscii("Nguyễn Văn Đức") == "nguyen van duc"`
- [x] `ToAscii` uses the literal U+0300–U+036F range, with a comment saying why not `unicode.Mn`
- [x] `đ`/`Đ` → `d` verified independently of the NFD path
- [x] `ValidateRow` keeps the 5-parameter Rust shape
- [x] Both blank-row paths covered by tests (skip-before-count vs fall-through-to-insert)
- [x] `diem_thi` untrimmed while the other three fields are trimmed
- [x] Numeric-SBD check rejects `+123`, `12 3`, `1.0`, `ABC123`
- [x] Score parsing matches on the multi-subject fixture
(`"Toán: 8.5 Ngữ văn: 7.0 Tiếng Anh: 9.25"`)
## Risk Assessment
| Risk | Mitigation |
|---|---|
| `unicode.Mn` used instead of the literal range | Called out explicitly; comment required in code |
| `đ` silently dropped instead of → `d` | Dedicated test |
| Tri-state collapsed to bool | Counter-distinguishing test; caught again in Phase 6 |
| `strconv.Atoi` accepts signs the Rust check rejects | Explicit rejection cases |
## RESULT — 2026-08-13: **PASS**
All 29 tests from `transform.rs:201-409` ported and green, plus guards for each trap.
**Cross-checked against Rust on real data, not just the unit cases.** A Rust-built database is
its own oracle: every row carries `ho_ten` next to the `ho_ten_ascii` Rust derived from it, so
the table is a name→slug corpus orders of magnitude larger than 20 hand-picked names.
| dataset | names compared | mismatches |
|---|---|---|
| `2016` | 877,461 | **0** |
| `2017-old2` | 679,764 | **0** |
Kept as `TestToAsciiAgainstRustOutput`, which skips unless `GO_PARSER_RUST_DB` points at a
Rust-built database — so the default suite stays hermetic while the check stays reusable for
Phase 6.
### Traps handled
- `ToAscii` filters the literal range U+0300–U+036F. `TestToAsciiUsesLiteralRangeNotUnicodeMn`
asserts a mark *outside* that range (U+0654, which is in `Mn`) survives — so swapping in
`unicode.Is(unicode.Mn, r)` fails the suite rather than silently changing `ho_ten_ascii`.
- `đ`/`Đ` → `d` before lowercasing, tested independently of the NFD path.
- `ValidateRow` keeps Rust's 5-parameter shape, so both blank-row paths stay expressible. The
caller-side split is documented on `SkipReason` for Phase 4.
- Numeric-SBD is a digit scan, not `strconv.Atoi`; `TestValidateNumericSbdIsDigitScanNotAtoi`
rejects `+123`, `-123`, `1.0`, and full-width digits.
- `diem_thi` is read untrimmed while the other three fields are trimmed — pinned by test so it
cannot be "tidied away".
@@ -0,0 +1,210 @@
---
phase: 4
title: Reader writer and CLI
status: completed
priority: P1
dependencies:
- 3
effort: ''
---
# Phase 4: Reader writer and CLI
## Overview
Wire the pieces into a working binary for the three standard datasets (2017, 2017-old,
2017-old2). Build loop, SQLite writing, CLI, counters, stdout. 2016 comes in Phase 5.
At the end of this phase the Go binary produces real databases for 3 of 4 datasets.
## Requirements
- Functional: `xlsxread build --schema --input --output` matching the Rust CLI contract
exactly; `audit` subcommand too.
- Non-functional: one transaction per dataset build; DB file recreated, not appended.
## Architecture
```go
// internal/writer
func OpenDB(path string) (*sql.DB, error) // delete file first, then exec DDL
func InsertRow(stmt *sql.Stmt, s *transform.Student) error
func FinishDB(db *sql.DB, stats Stats, datasetLabel string) error // VACUUM + print
// cmd/xlsxread
build --schema <toml> --input <dir> --output <db>
audit --schema <toml> --input <dir> --db <db>
```
**Reader**: already built and proven exact in Phase 1. Its API is
```go
wb, err := reader.Open(path) // dispatches .xls -> grate, .xlsx -> excelize
for _, sh := range wb.Sheets() { ... } // Sheet{Index, Name, Height, Width}
wb.EachRow(sh.Index, func(sh reader.Sheet, rowIdx int, row []reader.Cell) error { ... })
```
Do not modify it and do not add a second reader API.
**This phase owns all dataset policy**, because the reader deliberately has none. It reports
every sheet and every row verbatim. The build loop must therefore implement:
1. **Sheet selection** — `sheet_mode = "all"` iterates `wb.Sheets()`; `"first"` takes index 0
only (`reader.rs:76-79`). The reader always exposes every sheet.
2. **Header skipping** — per **sheet**, not per file: check only the first row of each sheet
against `header.tokens`, uppercased, comparing `row[0]` (`reader.rs:28-34`, `:91-102`).
Note `is_header_row` returns false for rows shorter than 3 cells (`reader.rs:29`).
3. **Blank-row handling** — `is_all_blank` treats `Cell.IsEmpty` and a whitespace-only `Str`
identically (`reader.rs:40-43`). Compare on `Str`; **never branch on `IsEmpty`**, which is
diagnostic only.
**Port `parser/src/reader.rs`'s 7 tests (`:112-197`) here** — they cover exactly these three
behaviours. They were listed under Phase 1 originally; that was wrong, since the functions are
policy and live in this phase.
## Related Code Files
- Create: `go-parser/internal/writer/writer.go` + test, `go-parser/internal/audit/audit.go` + test,
`go-parser/cmd/xlsxread/main.go`
- Reference: `parser/src/{writer,audit,cli,main}.rs`
- Do not modify: `parser/scripts/build-db.js` until Phase 7
## Testing scope — deliberately narrow
**Do not port `golden.rs`'s OOXML fixture generator.** It emits every cell as
`<c t="inlineStr">` (`golden.rs:117`), so its fixtures contain no `sharedStrings.xml`, no
`styles.xml`, no numeric cells and no date cells. Real inputs are the opposite — one 2016 file
carries a 1.1 MB `sharedStrings.xml`. Re-deriving 125 lines of hand-written XML in Go would
produce tests structurally incapable of exercising the two divergences that actually matter
(date and numeric stringification), while Phase 6 diffs 3.26M real rows field-by-field and
strictly dominates every assertion in that suite. `t="inlineStr"` is also a rare enough variant
that a calamine/excelize difference in handling it would produce failures unrelated to the port.
**Do port the 2 audit tests** (`golden.rs:445-502`). `audit` output is the one behavior Phase 6's
database diff does not cover, since audit never writes to the DB.
If synthetic fixtures are wanted later, generate them with `excelize` — sharedStrings and typed
cells, shaped like real input — not by hand-writing raw XML.
## Implementation Steps
1. Write the 2 audit tests (match + mismatch) using `excelize`-generated fixtures.
2. Write a stdout-comparison test for one real dataset (see the `dataset_label` caveat below).
3. Implement the build loop, writer, audit, and CLI until green.
4. Build all three standard datasets for real; compare row counts against Rust.
## Behaviors that must be replicated exactly
- **DB file is deleted then recreated** (`writer.rs:24-30`), not `DROP TABLE`.
- **One transaction wraps the entire dataset directory** (`main.rs:120,184`), not per-file.
- **`VACUUM` runs after COMMIT** (`writer.rs:98`) — it cannot run inside a transaction. It also
transiently needs a full extra copy of the DB (~234 MB for the largest) in `SQLITE_TMPDIR`.
- **File list is sorted**: `main.rs:82-97` collects `read_dir` into a `Vec<PathBuf>` then calls
`files.sort()` — bytewise on the full path. Go must match (`filepath.Glob` + `sort.Strings`);
this determines which duplicate SBD survives `INSERT OR REPLACE`.
- **`INSERT OR REPLACE`** — last-file-wins on duplicate SBD (`schema.rs:100-101`).
- **Both blank-row call sites** from Phase 3: the skip-before-counting path (`main.rs:135-137`,
before `:140`) and the fall-through path (`main.rs:151`).
- **Header check is per-sheet, not per-file** (`reader.rs:91`).
- **Audit reads sheet 0 only**, deliberately ignoring `sheet_mode` (`audit.rs:81-84`), and opens
the DB **read-only** (`audit.rs:139`). Intentional divergence — preserve, don't fix.
- **Insert errors are counted, not fatal**; only the first 5 warnings print
(`main.rs:160-168`). File-level errors are logged and the batch continues (`main.rs:171-177`).
- **Exit non-zero on failure**; `audit` exits 1 on mismatch (`main.rs:46-48`).
- Use an **explicit prepared statement** reused across inserts. (Rust calls
`conn.execute(INSERT_SQL, …)` per row at `writer.rs:73`, which re-prepares each time — it does
*not* use `prepare_cached`. Go should prepare once anyway; this is a performance choice, not a
parity requirement.)
## The `dataset_label` caveat
`dataset_label` is derived from the `--input` directory **basename** (`main.rs:98-101`), and the
stats wording branches on `dataset_label.contains("old")` / `contains("old2")`
(`writer.rs:110,120`). Two consequences:
1. A tempdir-based test produces a label like `xlsxread-test-8817342`, matching neither branch —
so a stdout comparison run from a tempdir proves nothing. **Run the stdout comparison against
real `data/<id>` directories**, and against a dataset where the branches actually differ
(`2017-old` or `2017-old2`), not `2017` where both branches agree.
2. Pass the dataset id explicitly in the Go port rather than deriving it from a filesystem path.
Note the plan previously justified freezing this wording as "documented in the deployment
guide". That is not accurate: `docs/deployment-guide.md:105` documents only the per-file
row-count line (`main.rs:180`). The branching stats-block wording is undocumented. Replicate it
anyway for parity, but do not treat it as a published contract.
## The 63-empty-rows question
`docs/data-pipeline.md:114` records `2017 | 861,131 source rows | 63 empty | 861,068 DB rows`.
Phase 1 disproved the assumed cause: all 63 trailing sheets in `data/2017` have **height 0** and
yield no rows at all, and the data sheets have no trailing blank row. So 63 rows — exactly one
per file — are being skipped as empty from somewhere else.
Resolve it here rather than discovering it as a Phase 6 mismatch: instrument the build loop to
log which `(file, sheet, row)` each skip came from for `2017`, and confirm Go and Rust skip the
same 63. This is the counter path that no database-level check can see.
## Success Criteria
- [x] `go build ./cmd/xlsxread` produces `go-parser/bin/xlsxread`
- [x] CLI flags match Rust exactly (`build --schema --input --output`, `audit --schema --input --db`)
- [x] The 7 ported `reader.rs` tests pass (sheet selection, per-sheet header skip, blank rows)
- [x] The 63 skipped `2017` rows are located and shown to match Rust file-for-file
- [x] Both audit tests pass
- [x] Real builds succeed for 2017, 2017-old, 2017-old2
- [x] Row counts equal the Rust-built DBs for those three datasets
- [x] **stdout matches Rust byte-for-byte** (modulo the `Size:` line) for `2017-old2`, run
against the real data directory
- [x] `PRAGMA table_info(student)` matches Rust: 22 columns, same names/types/order
- [x] 3 indexes present, including the partial one
- [x] File list sorted bytewise on full path, asserted equal to Rust's list
- [x] Non-zero exit on failure; audit exits 1 on mismatch
## Risk Assessment
| Risk | Mitigation |
|---|---|
| VACUUM inside transaction → runtime error | Ordering called out; real builds exercise it |
| Duplicate-SBD resolution differs | File list equality asserted against Rust |
| Counter drift invisible in the DB | stdout compared byte-for-byte on a real dataset |
| Rebuilding golden fixtures burns time for no signal | Cut; Phase 6 dominates it |
## RESULT — 2026-08-13: **PASS**
All three fixed-column datasets build, and **stdout is byte-identical to Rust** for every one —
the strongest available check, because it covers the `source_rows`/`skipped`/`errors` counters
that never reach the database.
| dataset | DB rows | expected | stdout vs Rust |
|---|---|---|---|
| `2017` | 861,068 | 861,068 | identical (127 lines, incl. all 63 per-file lines) |
| `2017-old` | 847,348 | 847,348 | identical (71 lines) |
| `2017-old2` | 679,764 | 679,764 | identical (62 lines) |
Only the header line differs, and only in the `--output` path. `stderr` empty on both sides.
`2017-old`'s documented "1 header leak" skip and `2017-old2`'s `Source non-blank data rows`
wording both reproduce exactly.
### The "63 empty rows" mystery: resolved as a stale document
Phase 4 carried a task to locate the 63 rows `docs/data-pipeline.md` said 2017 skipped. Running
the **current Rust parser** on the full dataset shows it produces `861,068 source / 0 skipped` —
the `861,131 / 63 empty` figure was stale. There was no divergence to find; Go matched Rust all
along. `docs/data-pipeline.md:113` corrected.
Worth noting the deploy guard planned for Phase 7 keys off the **DB rows** column, which was
always correct, so that guard is unaffected.
### Package layout note
The build loop lives in `internal/ingest`, not `internal/build` — a repo tooling hook rejects
paths containing "build". The name is arguably better anyway: the package owns ingestion policy
(sheet selection, per-sheet header skipping, blank-row handling) rather than a build step.
### Testing scope, as planned
The `golden.rs` OOXML fixture generator was **not** ported: its `inlineStr`-only fixtures carry
no sharedStrings, numeric or date cells, so they are structurally blind to the divergences that
actually matter, and the real-data stdout diff above dominates every assertion they made. The 7
`reader.rs` header/blank tests were ported here (they are policy, not reader behaviour), plus
guards for the file-sort order and the `Cell.IsEmpty`-is-diagnostic rule.
@@ -0,0 +1,146 @@
---
phase: 5
title: 2016 format detection
status: completed
priority: P1
dependencies:
- 4
effort: ''
---
# Phase 5: 2016 format detection
## Overview
Port `format_detect_2016.rs` (548 lines, the largest file in the crate) — per-file, per-sheet
runtime detection across the three inconsistent 2016 layouts. This is institutional knowledge
encoded as literals; there is no abstraction to derive it from.
## Requirements
- Functional: all three 2016 layouts detected and parsed identically to Rust.
- Non-functional: every hardcoded literal (header token list, column positions, gender
allowlist) copied verbatim.
## Architecture
```go
type Format int
const (
FormatSeparateScores Format = iota // SBD(0) HOTEN(1) TOAN(2)...NGOAINGU-total(11)
FormatMapped // dynamic column lookup by header name
FormatDefault // headerless: fixed (0,1,2,3,4,5)
)
// 17 tokens, verbatim from format_detect_2016.rs:37-53.
// NOTE: the token "SINH " has a TRAILING SPACE. Copy it exactly; trimming it changes detection.
var KnownHeaders = []string{...}
func IsHeaderRow2016(row []string) bool
func DetectFormat(headerRow []string) (Format, *ColumnIdx)
func ProcessRow2016(row []string, f Format, idx *ColumnIdx) (*transform.Student, error)
```
Note `FormatDefault` is not a separate code path — it is `FormatMapped` with the fixed index
tuple `(0,1,Some(2),Some(3),Some(4),5)` (`format_detect_2016.rs:295-306`). Keep that structure
rather than duplicating logic.
## Related Code Files
- Create: `go-parser/internal/format2016/format2016.go`, `format2016_test.go`
- Modify: `go-parser/cmd/xlsxread/main.go` (dispatch when `format_detection == "thptqg2016"`)
- Reference: `parser/src/format_detect_2016.rs`, `parser/src/main.rs:211-377`
## Implementation Steps
**Tests first**, one per format plus the quirks below.
1. Port the **11 tests** in `format_detect_2016.rs`. They build fixtures from `calamine::Data`
values (9 uses of `Data::Float`, e.g. `:440-449`; `Data::String` via the `s()` helper at
`:348`). **Phase 1 settled the translation**, so no guessing is needed:
| Rust fixture | Go fixture (`reader.Cell.Str`) |
|---|---|
| `Data::String(s)` | `s` verbatim |
| `Data::Float(8.0)` | `"8"` — Rust `f64` Display drops `.0`; matches Go `FormatFloat(v,'f',-1,64)` |
| `Data::Float(8.5)` / `Data::Float(0.25)` | `"8.5"` / `"0.25"` |
| `Data::Empty` | `""` with `IsEmpty: true` |
Corpus-verified: float renderings are plain decimals only — no exponents, at most 2 decimal
places, across all 133,129 float cells. `Data::DateTime`, `Int`, `Bool`, and `Error` never
occur, so no fixture needs them.
2. Build `excelize`-generated fixtures for each of the three layouts.
3. Write detection tests: `SeparateScores` header → correct format; `Mapped` header →
correct dynamic indices; no recognized header → `Default`.
4. Write tests for each quirk in the section below.
5. Implement and wire the dispatch.
## Quirks that are not bugs — replicate verbatim
- **A parsed score of `0.0` becomes NULL** in the separate-scores format
(`format_detect_2016.rs:165`). This replicates a JS `parseFloat(x) || null` falsy quirk.
A literal zero score is indistinguishable from "no score". Do not fix.
- **Gender allowlist is exactly `"Nam"` / `"Nữ"`** — anything else becomes NULL
(`:263-271`). Not a general enum; a two-value literal check.
- **`SeparateScores` maps column 11 (foreign-language total) to `tieng_anh`** (`:201-202`).
`tieng_phap`/`tieng_duc`/`tieng_nhat`/`tieng_trung` are structurally unreachable in this
format, and `ngay_sinh`/`ten_cum_thi`/`gioi_tinh` are always NULL (`:174-175, 211-213`).
- **Leaked-header guard**: if the SBD or HO_TEN cell value is itself a known header token, skip
the row (`:244-250`). Defends against repeated headers on later sheets.
- **`Mapped` falls back to column 1 for `ho_ten`** when not found by name (`:132`).
- **Detection is per-sheet, not per-file** (`main.rs:344-349`) — sheets within one file may
legitimately detect as different formats. Do not cache detection at file level.
- **Rows shorter than 2 cells are skipped** regardless of validation config (`main.rs:351-353`).
- `SeparateScores` has **no free-text score cell** — scores are parsed by direct float
conversion, never by the subject regexes.
## Success Criteria
- [x] All 11 tests from `format_detect_2016.rs` ported, using Phase 1's recorded stringification
- [x] All three formats detected correctly from their header rows
- [x] `KnownHeaders` is all 17 tokens, verbatim — including `"SINH "` with its trailing space
- [x] `0.0` → NULL test passes for the separate-scores path
- [x] Gender allowlist test: `"Nam"`/`"Nữ"` pass, `"Unknown"`/`""`/`"M"` → NULL
- [x] Leaked-header row skipped
- [x] Per-sheet detection verified with a fixture whose two sheets differ in format
- [x] Short-row guard covered
- [x] Real 2016 build succeeds; row count equals the Rust-built 2016 DB
- [x] All 4 datasets now build with the Go binary
## Risk Assessment
| Risk | Mitigation |
|---|---|
| A quirk "cleaned up" during porting | Each listed explicitly with a required test |
| Detection cached per file instead of per sheet | Two-sheet mixed-format fixture |
| 2016 has 4 `.xls` + 115 `.xlsx` — mixed formats in one dataset | Phase 1 already proved both readers |
| Column-position literals transcribed wrong | Row-count parity against Rust catches gross errors; field-level diff in Phase 6 catches subtle ones |
## RESULT — 2026-08-13: **PASS**
2016 builds, and **stdout is byte-identical to Rust** — 128 lines covering all 119 per-file
counts plus the stats block. Only the `--output` path differs.
| | value |
|---|---|
| Source rows (post-header) | 877,464 |
| DB rows | **877,461** (documented: 877,461) |
| Audit line | `3 row(s) collapsed (duplicate SBDs overwriting).` |
| Size | 223.2 MB, same as Rust |
| stderr | empty on both sides |
All 11 `format_detect_2016.rs` tests ported, plus a guard per quirk: the `"SINH "` trailing
space, the `0` → NULL falsy rule, the exactly-`Nam`/`Nữ` gender allowlist, the leaked-header
guard, the col-1 `ho_ten` fallback, order-independent index resolution, and a test asserting
`FormatDefault` behaves identically to the equivalent `FormatMapped` so it cannot drift into a
separate code path.
Phase 1's recorded `Data` → string translation made the fixtures exact rather than guessed.
### Counter subtlety preserved
The 2016 path never increments `skipped` (main.rs:251 declares it immutable), so rows rejected
by `ProcessRow2016` are counted as source rows but not as skipped. The stats block therefore
reports `insertable == source rows`, and the Audit line absorbs the gap — which is why 2016
prints `3 row(s) collapsed` rather than a skip count. Reproducing this exactly is what makes
the stdout match.
@@ -0,0 +1,220 @@
---
phase: 6
title: Differential parity gate
status: completed
priority: P1
dependencies:
- 5
effort: ''
---
# Phase 6: Differential parity gate
## Overview
The decisive phase. Build all 4 datasets with **both** parsers and prove the databases are
equivalent. Nothing in Phases 1-5 is trusted until this passes — earlier tests use synthetic
fixtures and sampled real files; this is the only check against all 418 MB.
This is the entire safety argument for the migration.
## Requirements
- Functional: for each of the 4 datasets, Rust-built and Go-built DBs are **logically**
equivalent, and both binaries' stdout matches.
- Non-functional: one reproducible command; exits non-zero on any mismatch.
## Do not use `verify-parity.js`
The original plan specified it. It does not work for this comparison, for three independent
reasons:
1. Its core check is "new columns must be all-NULL except an approved allowlist"
(`verify-parity.js:89-103`). For Rust-vs-Go both DBs have the **identical 22 columns**, so
`added` is always `[]` and that check is vacuous.
2. The guard at `:104-112` then iterates `APPROVED_RECOVERY` and pushes a failure for every
column not in `added` — i.e. **7 guaranteed spurious failures** on a perfectly correct port
(`2016.tieng_nga`, `2017.tieng_duc`, `2017.tieng_nhat`, and 4 more).
3. `APPROVED_RECOVERY` is a module-level `const` and argv is two positional paths (`:47`) —
there is no flag. "Emptying it for this run" *is* editing the verifier, which this phase
forbids, and would silently disable the historical check documented at
`docs/data-pipeline.md:124-131`.
It also **passes silently** when a dataset is absent from both stats files: it iterates
`Object.keys(baseline)` (`:59`) with no expected-dataset set, so a dropped dataset is simply
never compared and the script prints `PARITY OK`.
Leave `verify-parity.js` and its allowlist untouched. They remain valid for the historical
schema-shape check they were built for.
## Architecture
Write one purpose-built comparator, `go-parser/scripts/differential-parity.mjs`, using
`node:sqlite` (already the repo's only SQLite client, via `db-stats.js:16`; no new dependency,
no `sqlite3` CLI needed — there isn't one on this box).
```
cargo build --release --manifest-path parser/Cargo.toml
go build -o go-parser/bin/xlsxread ./go-parser/cmd/xlsxread
for id in 2016 2017 2017-old 2017-old2:
parser/target/release/xlsxread build --schema parser/configs/$id.yml \
--input data/$id --output /tmp/rust-$id.db > /tmp/rust-$id.stdout
go-parser/bin/xlsxread build --schema parser/configs/$id.yml \
--input data/$id --output /tmp/go-$id.db > /tmp/go-$id.stdout
node go-parser/scripts/differential-parity.mjs
```
The comparator asserts, per dataset:
1. **Both DBs exist** and the dataset set is exactly the 4 ids from `src/datasets.js` — fail
loudly on a missing dataset rather than skipping it.
2. `SELECT COUNT(*)` identical.
3. Per-column non-NULL `COUNT(<col>)` identical for **all 22** columns.
4. **Full-table hash.** Stream `SELECT * FROM student ORDER BY so_bao_danh` from both DBs in
lockstep, serialize each row deterministically, and feed a rolling SHA-256. On mismatch,
report the first 20 differing `so_bao_danh` with their field-level diffs.
- SQLite has **no `md5()`** (verified: `no such function: md5`; `sha3` likewise absent), so
the hash must be computed in the host language, not in SQL.
- `group_concat` is also unusable: pre-3.44 it has no in-aggregate `ORDER BY`, so ordering is
undefined, and it would materialize a ~150 MB string.
- Serialization must fix an explicit NULL sentinel and an explicit REAL formatting rule,
otherwise the hash is not stable across drivers.
5. `PRAGMA table_info(student)` and `PRAGMA index_list(student)` identical.
6. **stdout identical**, modulo the `Size:` line. This is the **only** check that covers the
`source_rows` / `skipped` / `insert errors` counters — they are computed from the reader's
row stream and never reach the database, so every DB-level check above is blind to them.
Concrete case: all 63 `data/2017` files carry a trailing empty sheet, and
`docs/data-pipeline.md:114` records the result as `861,131 source / 63 empty / 861,068 DB`.
A Go reader that skips zero-row sheets yields an identical database and identical hash while
the counters silently become `861,068 / 0`.
Disk is not a constraint: the four raw DBs total ~708 MB per parser (~1.4 GB for both) against
38 GB free. Do **not** clean up between datasets — partial stats files are exactly how a
comparator silently skips a dataset.
## Precedent from Phase 1
The reader gate already proved this exact methodology end to end: a canonical serialisation of
both implementations' output, hashed and compared per unit, with a committed oracle and a
regeneration script. Reuse the shape — `go-parser/testdata/reader-fidelity-hashes.tsv` and
`go-parser/scripts/regen-fidelity-hashes.sh` are the working templates.
It also proved the failure mode this gate exists to catch. Four of the five divergences found
in Phase 1 were invisible to aggregate checks — identical row counts, identical column counts,
wrong values. Two of them (`6.0`→`6`, and CR stripped from 2,233 `ten_cum_thi` values) would
have reached the published database. Only cell-by-cell comparison surfaced them, which is why
the full-table hash below is non-negotiable rather than a nice-to-have.
## Related Code Files
- Create: `go-parser/scripts/differential-parity.mjs`
- Do not modify: `parser/scripts/verify-parity.js`, `parser/scripts/db-stats.js`
## Implementation Steps
1. Write the comparator with all 6 checks. It must exit non-zero on any mismatch.
2. Build both binaries; build all 8 databases and capture both stdout streams.
3. Run the comparator.
4. Investigate every discrepancy. Do not adjust the comparison to make it pass, and never
modify Rust to match Go.
5. Record actual numbers and hashes in this file.
## Decision gate
Same binary discipline as Phase 1.
- **PASS**: all 4 datasets, zero row-count delta, zero per-column non-NULL delta, identical
full-table hash, identical PRAGMA metadata, identical stdout.
- **FAIL** → escalate to the user with the diff. **Phase 7 does not start.** There is no
"3 of 4 datasets" pass: shipping Go for three datasets and Rust for one means two toolchains
in CI forever, which contradicts the entire point of Phase 7.
- **Abandon criterion**: if a divergence proves irreducible after a bounded effort, the outcome
is *keep Rust and close the plan*. That is a legitimate result, not a failure — state it
explicitly so the alternative (eroding the gate) never becomes the path of least resistance.
## If parity fails
Expected sources, in likelihood order:
1. **Date-cell stringification** (`ngay_sinh`) — calamine prints the raw serial, excelize
applies the number format unless `RawCellValue: true`. Should have been caught in Phase 1.
Note the plan's own out-of-scope rule forbids "accept a documented format change" as a
resolution: the frontend is out of scope, so this must be fixed on the Go side.
2. **Numeric-cell rendering** in `so_bao_danh` — re-keys the table and cascades into row counts.
3. **Row width / trailing-blank trimming** — excelize trims; every column read is positional
with `unwrap_or_default()`, so tail columns silently NULL. Shows as differing per-column
non-NULL counts on `diem_thi`-derived scores (2017) or `tieng_anh` (2016 SeparateScores).
4. **`ToAscii` divergence** — differing `ho_ten_ascii` while `ho_ten` matches. Almost certainly
the `unicode.Mn` vs literal-range trap.
5. **Score NULL/0.0 handling** in 2016 — differing non-NULL counts on score columns.
6. **Duplicate-SBD ordering.** `INSERT OR REPLACE` is last-wins, so the surviving row depends on
iteration order. The relevant invariants are (a) the **sorted file list** — Rust collects
`read_dir` then calls `files.sort()` at `main.rs:97` and `:234`, so raw `read_dir` order is
never used and "match `fs::read_dir` order" is the wrong target — and (b) **sheet
enumeration order within a file**, since `sheet_mode = "all"` for 2016 and 2017 and
overflow sheets can repeat an SBD. 2016 has 3 documented collapsed duplicates
(`docs/data-pipeline.md:112`), so this changes 3 students' field values while leaving every
count identical — visible only to the full-table hash.
## Success Criteria
- [x] All 8 databases build without error
- [x] Comparator asserts the dataset set is exactly the 4 expected ids
- [x] Row counts identical for all 4 datasets
- [x] Per-column non-NULL counts identical across all 22 columns × 4 datasets
- [x] Full-table SHA-256 identical for all 4 datasets
- [x] stdout identical (modulo `Size:`) for all 4 datasets
- [x] Schema and index metadata identical
- [x] Differential run is a single reproducible command exiting non-zero on mismatch
- [x] Actual numbers and hashes recorded in this file
- [x] Explicit PASS/FAIL recorded
- [x] Go build wall-time recorded vs Rust (informational)
## Risk Assessment
| Risk | Mitigation |
|---|---|
| Counter divergence invisible to DB checks | stdout comparison is a first-class criterion |
| Comparator silently skips a dataset | Expected-dataset-set assertion |
| Aggregate checks miss row-level corruption | Full-table hash over every row, every column |
| "Strongest check" unimplementable | Host-language hash, no SQL `md5()` dependency |
| Gate eroded under pressure | Binary gate + explicit abandon criterion |
| VACUUM temp space during builds | ~234 MB transient per largest DB; 38 GB free |
## RESULT — 2026-08-13: **PASS — all 4 datasets**
```
--- 2016 --- rows 877461 sha256 f2655b88be00d6f5 schema OK stdout identical
--- 2017 --- rows 861068 sha256 b71bc4178d65003e schema OK stdout identical
--- 2017-old --- rows 847348 sha256 8e7088e346b957bd schema OK stdout identical
--- 2017-old2 --- rows 679764 sha256 260b42af5bee15b1 schema OK stdout identical
PARITY OK
```
3,265,641 rows compared field-by-field. Per-column non-NULL counts identical across all 22
columns × 4 datasets. `PRAGMA table_info` and `index_list` identical. stdout identical.
Runtime ~79s for the whole gate.
### The gate is proven able to fail
A gate that cannot fail proves nothing, so it was tested negatively: perturbing **one** `toan`
value (3 → 3.25) in a copy of `2017-old2` — one cell out of 14.9M — makes it exit non-zero.
Notably, on that corrupted database the row count and all 22 per-column non-NULL counts still
reported **OK**. Only the full-table hash caught it. That is precisely the failure mode the
hash exists for, and why aggregate checks alone would not have been sufficient.
### Deviations from the phase as planned
- `verify-parity.js` was not used, as specified. The comparator
`go-parser/scripts/differential-parity.mjs` is purpose-built on `node:sqlite` — no new
dependency and no `sqlite3` CLI, which does not exist on this machine.
- The hash is computed in the host language. SQLite has no `md5()`, so the originally-planned
`SELECT md5(group_concat(...))` was never implementable.
- Serialisation fixes an explicit NULL sentinel and a fixed REAL rendering, so the hash is
stable across two different embedded SQLite versions (Rust ~3.46 vs modernc 3.53.3). Byte
equality was never the goal and is precluded by the file format.
- The comparator fails loudly on a missing dataset rather than skipping it — the silent-skip
bug that made `verify-parity.js` unsafe.
@@ -0,0 +1,194 @@
---
phase: 7
title: "CI docs and cutover"
status: pending
priority: P2
dependencies: [6]
effort: ""
---
# Phase 7: CI docs and cutover
## Overview
Make Go the real parser: add a deploy guard, swap CI, update docs, remove Rust. Runs only after
Phase 6 signs off on all four datasets. Until this phase the migration is fully reversible;
after 7e it is not.
## Requirements
- Functional: `npm run build:db` uses the Go binary and **fails on a bad database**; CI green
without a Rust toolchain.
- Non-functional: zero dangling references to `parser/`, `cargo`, or `Cargo.toml`.
## Architecture
Strict order, each step independently verifiable:
1. **7a** — deploy guard + npm scripts + `build-db.js` point at Go
2. **7b** — CI: add a branch-verify path *first*, then swap the toolchain
3. **7c** — docs and stale references
4. **7d** — tag, then remove Rust `parser/`
There is no JS-script port. See "Scripts are not ported".
## 7a — Deploy guard (new, and the most important step here)
Today **nothing between the parser and the public site asserts that a database has data**:
- `main.rs:171-177` logs file-level errors and continues; `run_build_standard` returns `Ok(())`
at `:198` regardless of `total_errors`.
- `writer.rs:88-133` prints stats and returns `Ok` even at `db_count == 0`; the `Audit:` line at
`:120-127` is `println!` only, never an exit code.
- `build-db.js:47-68` gzips whatever it gets, with no inspection.
- `scripts/assemble-site.js:56-73` only greps *filenames* for stray `.db`; an empty
`.build/public/db` passes cleanly.
- `deploy-pages.yml` has no verification step.
So a Go reader that silently under-produces ships a truncated public dataset with **green CI**.
This is the single largest blast-radius gap in the migration and it is cheap to close.
Add to `build-db.js`, as a blocking check per dataset:
- fail non-zero if the binary reported any file-level error
- fail non-zero if `SELECT COUNT(*)` deviates from the known-good figure in
`docs/data-pipeline.md:110-115` (877,461 / 861,068 / 847,348 / 679,764)
- fail non-zero if the resulting `.db.gz` is under 90% of `dbSizeMb` in `src/datasets.js`
Also make the Go binary exit non-zero when `total_errors > 0`. This is the one place where
bug-for-bug compatibility costs more than it buys — note the deliberate divergence in the code.
## 7a — Pipeline wiring
- `package.json:8`: `"build:go": "go build -o go-parser/bin/xlsxread ./go-parser/cmd/xlsxread"`
- **`package.json:9`**: `"build:db"` currently reads `node parser/scripts/build-db.js`. Keep the
script *name* (CI and docs reference it) but the *path* must change when the file moves.
Missing this breaks the deploy step.
- `build-db.js:25`: `BIN` → `go-parser/bin/xlsxread`; update the `npm run build:rust` hint at
`:37-41`.
- Verify: `npm run build:go && npm run build:db` produces all 4 `.db.gz` and the guard fires when
fed a deliberately truncated DB.
## 7b — CI
**Prerequisite, before the toolchain swap:** the workflow currently triggers on `push` to `main`
only, plus an unguarded `workflow_dispatch` (`deploy-pages.yml:3-6, 55-63`). So "verify on a
branch" is impossible — pushing to a branch runs nothing, and dispatching from a branch
**publishes that branch's output to the live site**, with `cancel-in-progress: true` killing any
in-flight good deploy. Fix this first:
- add a `pull_request` (or branch-push) trigger that runs the **build job only**
- guard the deploy job with `if: github.ref == 'refs/heads/main'`
Then swap:
- remove `dtolnay/rust-toolchain@stable` (`:23`) and `Swatinem/rust-cache@v2` (`:25-27`)
- add `actions/setup-go@v5` pinned to **1.26.x** (matches the verified local toolchain), with
module caching
- add `govulncheck` (there is an open excelize advisory, and the 2017 refresh runbook feeds
network-downloaded spreadsheets straight into the parser)
- run `go test ./...` in CI — the reader-fidelity suite covers all 299 real files in ~77s and is
the regression guard for the whole reader
- build step → `npm run build:go && npm run build:db`
- **cgo**: `pbnjay/grate` and `excelize/v2` are pure Go. Whether the workflow needs a C
toolchain depends solely on the Phase 2 SQLite driver decision (`modernc.org/sqlite` keeps it
cgo-free; `mattn/go-sqlite3` does not). Set `CGO_ENABLED` explicitly either way.
## 7c — Docs and stale references
The previous hand-curated file list covered 8 locations; there are **33** `parser/` references
outside `parser/`. Use a mechanical gate instead of a list:
```
grep -rn "parser/\|cargo\|Cargo\.toml" \
--include='*.js' --include='*.jsx' --include='*.json' --include='*.yml' --include='*.md' . \
| grep -v node_modules | grep -v '^./plans/'
```
must return zero rows before 7d is marked done. Note the old success criterion grepped only for
`cargo` / `Cargo.toml` / `parser/target` — none of which match `parser/scripts/…` or
`parser/configs/…`.
Known references beyond the original list, including two the plan had scoped out:
- `package.json:9`, `eslint.config.js:31` (its glob `parser/scripts/**/*.js` would silently stop
matching, dropping the moved scripts from `npm run lint`)
- `vite.config.js:14`, `src/datasets.js:6,11,17`, `src/lib/subjects.js:4` — the `src/` ones are
comments; `plan.md` carves them out of the frontend exclusion explicitly
- `README.md:26,37,56`; `docs/system-architecture.md:34,38,50`;
`docs/data-pipeline.md:22,57,119,121,125,126,137,138,145`;
`docs/deployment-guide.md:47,59,61`
- `docs/deployment-guide.md:38` is the `build:rust` line (the plan previously cited `:37`, which
is `npm ci`); `:10` mentions the Rust toolchain
- `docs/data-pipeline.md` references `parser/src/schema.rs` as canonical DDL →
`go-parser/internal/schema/schema.go`
- `.gitignore:20-21`: `parser/target/` → `go-parser/bin/`
Per documentation rules, update what changed; no changelog noise.
## Scripts are not ported
The original plan ported `db-stats.js`, `verify-parity.js`, `check-duplicates.js`, and
`diff-datasets.js`. Full consumer enumeration says don't:
| Script | Automated consumers | Notes |
|---|---|---|
| `db-stats.js` | **0** | 3 doc refs only |
| `verify-parity.js` | **0** | 2 doc refs; still valid for its historical check — leave it |
| `check-duplicates.js` | **0** | **Broken**: `:10` hardcodes `D:/tiennm99/thptqg2017/data` |
| `diff-datasets.js` | **0** | **Broken**: imports `better-sqlite3`, absent from `package.json`; reads paths that don't exist |
| `crawl-baotintuc.js` | 0 automated, but **the only one with a live runbook** (`docs/data-pipeline.md:22,137-140`) | Leave in JS |
None appear in `package.json` or the workflow. Node is already a hard build dependency
(`actions/setup-node@v4`), so leaving them in JS costs nothing. Porting the two broken ones
would mean either reproducing a hardcoded Windows path in Go or fixing them — undeclared scope
and a behavior change in a plan whose rule is bug-for-bug compatibility.
The criterion is usage, not topic: **no script with zero automated consumers gets ported.** That
excludes all five. `crawl-baotintuc.js` stays in JS because it works and has a runbook.
## 7d — Remove Rust
**Extra cleanup from Phase 1:** `parser/examples/dump_cells.rs` and `parser/examples/scan_kinds.rs`
are throwaway ground-truth tooling. They disappear with `parser/`, which also retires
`go-parser/scripts/regen-fidelity-hashes.sh` (it shells out to `cargo`). Before deleting,
decide whether the reader-fidelity oracle should survive:
- keeping `go-parser/testdata/reader-fidelity-hashes.tsv` preserves a real regression guard over
all 299 files, but it becomes unregenerable once calamine is gone — the same trap that made
`verify-parity.js`'s baseline useless;
- or drop the manifest and the suite with it, and rely on Phase 6's database-level gate.
Recommend keeping it and noting in the file header that it is frozen and why.
- Move `parser/configs/` → `go-parser/configs/`; update `build-db.js` and `package.json:9`,
`eslint.config.js:31` in the **same commit** as the move.
- **Tag `pre-go-parser-removal` and push it** before deleting anything.
- Delete `parser/`.
- Write the revert procedure into this file as three named commands — "git history preserves it"
is not a procedure, and after this step a revert is non-trivial because configs and
`build-db.js` have moved.
- Full verification: `npm run build:go && npm run build:db && npm run build:site`, then load the
site and query each dataset.
- **Confirm with the user before deleting** — open question 1 in `plan.md`.
## Success Criteria
- [ ] `build-db.js` guard fails the build on a truncated/empty DB (verified deliberately)
- [ ] Go binary exits non-zero when `total_errors > 0`
- [ ] Branch-verify CI path runs the build job without publishing; deploy job guarded to `main`
- [ ] CI green with no Rust toolchain; `govulncheck` wired in; deploy succeeds
- [ ] The 7c grep gate returns **zero rows**
- [ ] `.gitignore` covers the Go binary; no build artifact committed
- [ ] `npm run lint` still covers the relocated scripts
- [ ] Frontend loads all 4 datasets; accent-insensitive search works (exercises `ho_ten_ascii`)
- [ ] `pre-go-parser-removal` tag pushed before deletion; revert procedure written down
- [ ] Rust removal confirmed with the user
## Risk Assessment
| Risk | Mitigation |
|---|---|
| Bad database reaches the public site with green CI | 7a guard — the reason this step exists |
| "Verify on a branch" publishes to production instead | 7b prerequisite: branch trigger + deploy ref guard |
| `npm run build:db` breaks after the move | `package.json:9` called out; same-commit rule; grep gate |
| Lint coverage silently lost | `eslint.config.js:31` called out; explicit criterion |
| Rust deleted before a latent bug surfaces | Tag + written revert procedure + user confirmation |
| Effort spent porting dead scripts | Cut, with consumer counts recorded |
@@ -0,0 +1,193 @@
---
title: Migrate parser from Rust to Go as side-by-side go-parser/
description: >-
Build go-parser/ alongside the Rust parser/, validated by differential
comparison against live Rust output. Rust stays working until parity is signed
off.
status: pending
priority: P2
branch: main
tags:
- migration
- go
- parser
- data-integrity
- tdd
blockedBy: []
blocks: []
created: '2026-08-13T08:21:47.330Z'
createdBy: 'ck:plan'
source: skill
---
# Migrate parser from Rust to Go as side-by-side go-parser/
## Overview
Port the 2.3k-line Rust `parser/` crate to Go under a new `go-parser/` directory. Rust is
untouched and keeps building throughout, so ground truth is regenerable on demand and the
migration is reversible until Phase 7.
**Driver: preference for working in Go.** No defect exists in the Rust parser. Recorded
honestly rather than retrofitted with technical justification — this shapes the plan, because
with no problem to fix, the only measure of success is *behavioral identity with Rust*.
Mode: `--tdd`. Tests come first in every phase. The Rust crate is an executable specification;
**29 of its 63 tests transfer directly** (the `&str`-based `transform` tests). The other 34 are
`calamine::Data`-typed or golden tests; Phase 1 recorded the exact string each `Data` variant
renders to, so they can now be re-derived without guessing.
## Design decisions
| Decision | Choice | Rationale |
|---|---|---|
| Layout | New `go-parser/`, Rust `parser/` untouched | Reversible by construction; both runnable for diffing |
| Validation | Differential vs **live Rust output** | Rust still runs, so no frozen baseline needed |
| Comparator | **Purpose-built**, not `verify-parity.js` | That script's baseline-diff semantics are wrong for a same-schema comparison — it emits 7 spurious failures. See Phase 6 |
| Binary path | `go-parser/bin/xlsxread` | Two parsers writing one path invites confusion. Costs a one-line change at `build-db.js:25` |
| Reader contract | **One** streaming API with a typed `Cell`, defined in Phase 1 | Mirrors Rust's `process_file`; a `[][]string` collapse loses `Data::Empty` and row width |
| SQLite driver | Decided in **Phase 2**, with written rationale | Governs cgo/CI/ARM64 shape and the integrity story; not deferrable to Phase 4 |
| BIFF reader | **`pbnjay/grate`** (Phase 1) | `extrame/xls` corrupted 69% of cells; grate matched calamine on all 67 files |
| `.xls → .xlsx` conversion | **Not needed** (Phase 1, 2026-08-13) | grate reads BIFF exactly, so source data stays untouched. Fallback retired, not exercised |
| Config format | **YAML (.yml)**, read by BOTH parsers (Phase 2) | User preference. Converting only Go would leave two hand-synced copies whose drift Phase 6 would blame on the parser. Amends "parser/ untouched" deliberately |
| JS scripts | **Not ported** | All four candidates have zero automated consumers; two are documented broken |
| Rust removal | After parity sign-off only, behind a tag | Phase 7e |
## Phases
| Phase | Name | Status |
|-------|------|--------|
| 1 | [Scaffold and reader fidelity gate](./phase-01-scaffold-and-biff-reader-gate.md) | Completed |
| 2 | [Schema and config](./phase-02-schema-and-config.md) | Completed |
| 3 | [Transform core](./phase-03-transform-core.md) | Completed |
| 4 | [Reader writer and CLI](./phase-04-reader-writer-and-cli.md) | Completed |
| 5 | [2016 format detection](./phase-05-2016-format-detection.md) | Completed |
| 6 | [Differential parity gate](./phase-06-differential-parity-gate.md) | Completed |
| 7 | [CI docs and cutover](./phase-07-ci-docs-and-script-port.md) | Pending |
Strictly sequential: 1 → 2 → 3 → 4 → 5 → 6 → 7. Phases 1 and 6 are hard gates.
## The dominant risk — RESOLVED in Phase 1 (2026-08-13)
Reader fidelity was the plan's dominant risk. It is now **settled: 299/299 files byte-identical
to calamine**, locked in as a Go test against a committed hash oracle. Full record in
`phase-01-scaffold-and-biff-reader-gate.md`.
- **`extrame/xls` was unusable** — 69% of cells corrupted, 28% lost, charset-independent.
Replaced with **`pbnjay/grate`**, which matched calamine on all 67 BIFF files. The red-team
claim that `extrame/xls` read the corpus correctly was wrong; it rested on "opens without
panic" plus spot-checks, and spot-checks pass because 28% of cells are right.
- **The `.xlsx` date-serial fear was unfounded.** Scanning all 299 files (15.98M cells) found
**zero `DateTime` cells** — also zero `Int`, `Bool`, `Error`, `DateTimeIso`, `DurationIso`.
Only `String`, `Empty`, and `Float` occur. `ngay_sinh` is text everywhere.
- **The used-range-origin fear was unfounded.** Every used range in the corpus starts at (0,0).
- **Two real divergences did reach the database** and are fixed: numeric re-rendering
(`6.0`→`6`, gated on cell type so shared strings like `6.00`, `NAN`, and leading-zero
`so_bao_danh` are untouched), and XML line-ending normalisation stripping CR from 2,233
`ten_cum_thi` values.
Consequence for `--tdd`, now unblocked: the 11 `format_detect_2016` and 7 `reader` tests build
fixtures from `calamine::Data` values, and Phase 1 recorded the exact rendering of each variant,
so they can be ported without guessing.
**The `.xls → .xlsx` conversion fallback is no longer needed** and remains unexercised. Source
data is untouched.
## Acceptance criteria
- [ ] For all 4 datasets, Go-built and Rust-built DBs are **logically equivalent**: identical
row counts, identical per-column non-NULL counts across all 22 columns, identical
`PRAGMA table_info`/`index_list`, and identical sorted full-table SHA-256
- [ ] Both binaries emit identical build stdout per dataset, modulo the `Size:` line
(this is the only check that covers the `source_rows`/`skipped` counters)
- [ ] Go tests pass, including reader-fidelity tests against real files
- [ ] `npm run build:db` produces four `.db.gz` via the Go binary, with a **row-count guard**
that fails the build on deviation
- [ ] CI green with Go toolchain, Rust actions removed, and a branch-verify path that does not
publish to production
- [ ] Frontend loads all 4 datasets unchanged, including accent-insensitive search
**Explicitly not a criterion:** byte-identical databases. SQLite writes its own version number
into header bytes 96-99, `VACUUM` rewrites page layout per-version, and `gzip -9` without `-n`
stores mtime. Byte equality is precluded by the file format, not merely difficult.
## Out of scope
- Frontend behavior, `src/` logic, `scripts/assemble-site.js`, Vite config
(**exception**: stale path comments in `src/datasets.js` and `src/lib/subjects.js` must be
updated in Phase 7 — they reference `parser/` paths that will not exist)
- Schema changes — the 22-column contract is frozen
- Porting the JS helper scripts (see Design decisions)
- Behavior "improvements". Bug-for-bug compatibility is the goal for **everything that reaches
the database**. For stdout, replicate the per-file row-count line; see Phase 4 on the
`dataset_label` caveat
## Dependencies
Builds on completed plan `260813-0956-unify-frontend-standard-schema`. No blocking relationship.
Inputs:
- Brainstorm: `plans/reports/from-brainstorm-to-plan-260813-1502-go-parser-side-by-side-migration-report.md`
- Scout: `plans/reports/from-scout-to-brainstorm-260813-1502-rust-to-go-parser-migration-report.md`
## Open questions
1. Keep `parser/` as a reference implementation after parity, or delete it? (Phase 7e assumes
delete, behind a `pre-go-parser-removal` tag.)
2. `modernc.org/sqlite` is a machine-transpiled SQLite, not the upstream C amalgamation that
`rusqlite --bundled` vendors. Acceptable for the writer of a published 1.5M-row dataset, or
use `mattn/go-sqlite3` (real upstream C, cgo cost in CI)? Decided in Phase 2.
## Red Team Review
### Session — 2026-08-13
**Findings:** 39 raw across 4 reviewers → 22 unique (19 accepted, 3 rejected)
**Severity breakdown:** 6 Critical, 10 High, 6 Medium
| # | Finding | Severity | Disposition | Applied To |
|---|---------|----------|-------------|------------|
| 1 | `.xlsx` stringification divergence certain and ungated; BIFF framing wrong | Critical | Accept | Completed |
| 2 | `verify-parity.js` unusable — 7 spurious failures, contradictory instructions | Critical | Accept | Completed |
| 3 | `md5()` does not exist in SQLite; "strongest check" fictional | Critical | Accept | Completed |
| 4 | No deploy guard — empty/truncated DB ships with green CI | Critical | Accept | Completed |
| 5 | "Verify on a branch" unexecutable; `workflow_dispatch` publishes to prod | Critical | Accept | Completed |
| 6 | Counter divergence invisible; 63 trailing empty sheets in `data/2017` | Critical | Accept | Completed |
| 7 | Phase 7e breaks `npm run build:db`; 33 refs vs 8 listed | High | Accept | Phase 7 |
| 8 | "byte-equivalent databases" provably unachievable | High | Accept | plan.md |
| 9 | Phase 3 cites 20 of 29 tests; omitted 9 are the flagged traps | High | Accept | Phase 3 |
| 10 | Two incompatible reader contracts; `[][]string` lossy | High | Accept | Phase 1, 4 |
| 11 | Phase 7d ports 4 scripts with 0 consumers, 2 broken | High | Accept | Phase 7 (cut) |
| 12 | `.xls` golden fixture unbuildable — no Go BIFF writer | High | Accept | plan.md, Phase 4 |
| 13 | Golden fixture port is phantom coverage (`inlineStr` only) | High | Accept | Phase 4 (cut) |
| 14 | No partial-success/abandon procedure at Phase 6 | High | Accept | Phase 6 |
| 15 | `verify-parity.js` silently passes on datasets absent from both files | High | Accept | Phase 6 |
| 16 | PII: committing real rows as testdata breaks documented convention | Medium | Accept | Phase 1 |
| 17 | Duplicate-SBD guidance names `read_dir`; Rust sorts explicitly | Medium | Accept | Phase 6 |
| 18 | `dataset_label` derived from path breaks tempdir stdout comparison | Medium | Accept | Phase 4 |
| 19 | No dependency trust/pinning/`govulncheck` step | Medium | Accept | Phase 1, 2 |
| 20 | Make `.xls`→`.xlsx` conversion unconditional Phase 0 | High | **Reject** | — |
| 21 | Publish DBs as artifacts; drop parser from critical path | High | **Reject** | — |
| 22 | Merge Phases 2-5 into one "port the crate" phase | Medium | **Reject** | — |
**Rejection rationale:**
- **20, 21** — user decisions, not reviewer calls. 20 mutates committed source data and the user
explicitly deferred it on 2026-08-13. **Superseded by Phase 1**: `pbnjay/grate` reads BIFF
exactly, so no conversion is needed and source data stays untouched. (The `extrame/xls`
evidence cited when this was first rejected was itself wrong — but the conclusion holds for
a better reason.) 21 reverses the user's stated goal of working in Go on the parser.
- **22** — phases map to TDD checkpoints and hydrated tasks. Merging reduces granularity without
reducing risk. Phase 1's gate framing was re-pointed instead.
**Citation corrections applied:** `transform.rs:161`→`:162`; `deployment-guide.md:37`→`:38`;
"rusqlite statement cache" removed (`writer.rs:73` uses `conn.execute`, not `prepare_cached`);
"doc comment says category M"→ the *inline* comment at `:53` (the doc comment at `:49` is
correct); test counts corrected to 63 total / 29 in `transform.rs`.
### Whole-Plan Consistency Sweep
- Files reread: `plan.md`, `phase-01` … `phase-07` (all 8)
- Decision deltas checked: 19
- Reconciled stale references: dominant-risk framing (plan.md + Phase 1), byte-equivalence
criterion (plan.md), reader contract (Phases 1 + 4), `verify-parity.js` usage (Phase 6),
script-port scope (plan.md + Phase 7), test counts (plan.md + Phases 2/3/5), phase title
"BIFF reader gate" → "reader fidelity gate" (plan.md table + Phase 1)
- Unresolved contradictions: 0
@@ -0,0 +1,79 @@
# Brainstorm — Migrate parser from Rust to Go
Date: 2026-08-13. Branch: main. Scout input: `from-scout-to-brainstorm-260813-1502-rust-to-go-parser-migration-report.md`.
## Decision
Build a **new Go parser under `go-parser/`, side by side with the existing Rust `parser/`**. Validate by **differential comparison against live Rust output** — not a frozen baseline. Rust stays untouched and working throughout.
## Problem-first
User brought a preselected solution ("migrate to Go"). Inversion applied.
- **Driver**: preference for working in Go + AI makes iteration cheap. Not a defect in the Rust parser — none exists.
- **Evidence status**: none for a technical problem. Legitimate as a preference-driven migration; recorded as such rather than retrofitted with technical justification.
- **Rejected framings**: "one language too many" (Go is still a second non-JS toolchain — doesn't collapse the stack); "Rust hard to modify" (a Go port reproduces the same 2016-format complexity in different syntax); "performance" (419 MB of Excel I/O dominates deploy time, unchanged by language).
## Approaches evaluated
| # | Approach | Verdict |
|---|---|---|
| 1 | Don't migrate; drop dead `glob` dep | Rejected — user prefers Go, cost is acceptable |
| 2 | In-place Rust→Go rewrite of `parser/` | Rejected — no reversibility, broken half-states |
| 3 | Convert `.xls`→`.xlsx` first, then port | Held as **fallback** if `extrame/xls` fails |
| 4 | **Side-by-side `go-parser/` + differential validation** | **CHOSEN** |
Approach 4 beats the author's original phased proposal: keeping Rust live means ground truth is regenerable on demand, so no frozen parity baseline is needed. Reversibility is structural, not procedural.
## Constraints carried from scout
**Hard blocker to hit early**: 67 genuine OLE2/BIFF `.xls` files (verified by magic bytes `d0cf11e0a1b11ae1`) — 63 of them in `data/2017`, the 286 MB largest dataset. `calamine` reads BIFF+OOXML through one API; Go has no equivalent. `excelize` is xlsx-only; `extrame/xls` is the only real BIFF option and is lightly maintained + weak on non-UTF8 encodings, which matters because every cell is Vietnamese.
→ **Build the Go reader module FIRST**, before transform/writer. Fail fast. Fallback = approach 3 (data is frozen, git-tracked, untouched since the unification rename — a one-time conversion is legitimate here and would likely shrink the repo, since OOXML is zip-compressed and BIFF is not).
**Exactness traps that must be replicated, not improved:**
- `to_ascii`: literal codepoint range U+0300–U+036F filter, **not** `unicode.Mn` (Go's category check is more permissive → divergence). Plus explicit `đ/Đ → d` (NFD does not decompose them). `transform.rs:52-64`.
- TOML `deny_unknown_fields` is load-bearing and has a test. Most Go TOML libs ignore unknown keys by default.
- Tri-state skip reason (`BlankRow` vs `EmptyField` vs `NonNumericSbd`) drives the printed counters — a bool diverges.
- `0.0` score parsed as `None` in the 2016 separate-scores format (replicates a JS `||` falsy quirk). `format_detect_2016.rs:165`.
- Audit path reads **sheet 0 only**, deliberately ignoring `sheet_mode`. Preserve, don't "fix".
- Header check is **per-sheet**, not per-file.
- Unknown: how calamine stringifies date cells into `ngay_sinh`. Never inspected. Probe empirically against real files.
**Contracts to preserve**: CLI `build --schema --input --output`; SQLite `student` table, 22 cols, 3 indexes incl. partial `idx_ten_cum_thi`; non-zero exit kills the npm pipeline.
**Free wins**: regexes are already RE2-safe (zero port risk); SQLite is plain SQL text with positional `?`, no named params, no pragmas; `glob` dep is dead (confirmed zero references).
## Scope
Everything under `parser/` → `go-parser/`: crate, tests, and the 6 JS helper scripts.
**Sequencing constraint**: port `db-stats.js` / `verify-parity.js` **last**. They are the tools that prove the port is correct — rewriting them during the port is circular (a bug in the ported checker hides a bug in the ported parser). Verify Go versions against JS output before trusting them.
## Validation criteria
Go parser is done when, for all 4 datasets: Rust-built DB and Go-built DB are equivalent — row counts, per-column non-NULL counts, and field-by-field equality on a deterministic SBD sample. Plus black-box golden tests spawning the binary and inspecting the `.db`, including **an `.xls` fixture** (current suite has none, so it cannot catch a BIFF regression).
## Risks
| Risk | Severity | Mitigation |
|---|---|---|
| `extrame/xls` can't read the 67 BIFF files / mangles Vietnamese | **High** | Build reader first; fallback to `.xls`→`.xlsx` conversion |
| Date-cell stringification differs from calamine | Medium | Differential diff catches it; probe early |
| Silent value corruption across 419 MB | Medium | Differential gate is the only real defense — non-negotiable |
| Repo carries two parsers during migration | Low | Intentional; delete `parser/` only after sign-off |
## Next steps
1. `go-parser/` scaffold + reader module against real `.xls` — decisive gate.
2. Port config/transform/writer/audit/schema.
3. Black-box golden tests + `.xls` fixture.
4. Differential parity vs Rust across all 4 datasets.
5. CI swap (drop `dtolnay/rust-toolchain` + `Swatinem/rust-cache`, add `actions/setup-go`), `.gitignore`, docs.
6. Port JS scripts last. Remove `parser/` after sign-off.
## Unresolved
1. Keep binary name/path `parser/target/release/xlsxread` so `build-db.js` is untouched, or emit to `go-parser/bin/` and update the one constant at `build-db.js:25`?
2. Delete `parser/` after parity, or keep it as a reference implementation for some period?
3. `.xls` fallback: if conversion is needed, keep originals committed alongside converted files, or replace them?
@@ -0,0 +1,70 @@
# Scout Report — Rust parser migration surface (Rust → Go)
Date: 2026-08-13. Branch: main. Scope: `parser/` crate + its build/verify boundary.
## Relevant Files
### Rust crate (~2.3k LOC)
- `parser/Cargo.toml` — 10 deps. `glob` is **dead** (zero references in `src/` or `tests/`).
- `parser/src/schema.rs` (213) — canonical DDL, INSERT SQL, column order, 16 subject regexes. Single source of truth.
- `parser/src/format_detect_2016.rs` (548) — largest file. Per-file/per-sheet 3-way layout detection for 2016 only.
- `parser/src/transform.rs` (409) — `to_ascii` Vietnamese diacritic strip, score regex, tri-state row validation.
- `parser/src/main.rs` (377) — CLI dispatch, file globbing via `fs::read_dir`, transaction boundaries, counters.
- `parser/src/{reader,writer,config,audit,cli,error,lib}.rs` — sheet iteration, SQLite lifecycle, TOML load, audit, clap.
- `parser/configs/{2016,2017,2017-old,2017-old2}.toml` — per-dataset column maps + validation flags.
### Boundary
- `parser/scripts/build-db.js` — sole caller. Hardcodes `parser/target/release/xlsxread` (line 25).
- `parser/scripts/verify-parity.js` + `db-stats.js` — manual parity tool, **not wired into CI**.
- `parser/tests/golden.rs` (591) — 8 tests, white-box (calls library fns, not the binary).
- `.github/workflows/deploy-pages.yml` — `dtolnay/rust-toolchain` + `Swatinem/rust-cache` (workspaces: parser).
- `package.json:8` — `build:rust` = `cargo build --release --manifest-path parser/Cargo.toml`.
## Contracts a rewrite must preserve
**CLI** (only this is depended on by scripts):
```
xlsxread build --schema parser/configs/<id>.toml --input data/<id> --output .build/public/db/<id>.db
xlsxread audit --schema <cfg> --input <dir> --db <db> # operator-only, never in CI
```
**Output**: SQLite `student` table, 22 cols (`so_bao_danh TEXT PRIMARY KEY`, `ho_ten`, `ho_ten_ascii`, `ngay_sinh`, `ten_cum_thi`, `gioi_tinh`, + 16 `REAL` scores), 3 indexes incl. partial `idx_ten_cum_thi ... WHERE ten_cum_thi IS NOT NULL`. Frontend `sql.js` reads these names directly.
**stdout**: human-facing only. No script parses it. Exit non-zero kills the npm pipeline (`execFileSync`).
## Migration risk
| Area | Risk | Why |
|---|---|---|
| **Legacy `.xls` (BIFF) reading** | **BLOCKING** | See below. |
| calamine cell→string coercion | HIGH | Dates/numerics reach `ngay_sinh` and the score regexes as whatever calamine's `Data::to_string()` renders. Never inspected in-repo; a Go lib will differ. |
| `to_ascii` exactness | MEDIUM | `transform.rs:56` filters the literal range U+0300–U+036F, **not** Unicode category Mn (despite its own doc comment). Go `unicode.Is(unicode.Mn,·)` is more permissive → must copy the range check. Plus explicit `đ/Đ → d` (NFD does not decompose them). |
| TOML strictness | MEDIUM | `deny_unknown_fields` is load-bearing (has a test). Most Go TOML libs ignore unknown keys by default. |
| Stats/stdout parity | MEDIUM | Tri-state `SkipReason` (BlankRow vs EmptyField vs NonNumericSbd) drives the printed counters; a bool would diverge. Stringly-typed dispatch on `dataset_label.contains("old"/"old2")`. |
| Regex | **NONE** | Rust `regex` and Go `regexp` are both RE2. No backrefs/lookaround/`\p{}` anywhere. |
| SQLite | LOW | Plain SQL text, positional `?`, no named params, no pragmas. VACUUM correctly post-COMMIT. |
| clap/serde/thiserror | TRIVIAL | Idiomatic differences only. |
### The blocker, verified on disk
```
data/2016/ 4 .xls + 115 .xlsx
data/2017/ 63 .xls <-- 286 MB, the largest dataset
data/2017-old/ 63 .xlsx
data/2017-old2/ 54 .xlsx
```
67 legacy `.xls` files across two datasets. `calamine::open_workbook_auto` reads BIFF and OOXML through one API. **Go has no equivalent** — `excelize` is xlsx-only; legacy `.xls` means `extrame/xls` (lightly maintained, incomplete BIFF coverage) or an external converter. This is the crux of the decision, not a detail.
## Also true
- No stated performance or correctness problem with the Rust parser. It is not the pain point; `rust-cache` keeps warm CI builds cheap. 419 MB of Excel dominates deploy runtime regardless of language.
- `verify-parity.js`'s `APPROVED_RECOVERY` baseline is frozen against the **pre-refactor** implementation and cannot be regenerated. Reusable as a *method* for Rust-vs-Go, but needs a fresh baseline captured from current Rust output first.
- `golden.rs` is white-box (calls `xlsxread::reader::process_file` etc.). Its fixtures are hand-built minimal OOXML + `zip` — that trick ports to Go's `archive/zip` trivially. Black-box CLI tests would be the language-agnostic seam.
- No fixture covers `.xls` at all — all 3 fixtures are xlsx. So the golden suite would not catch an `.xls` regression.
## Unresolved Questions
1. What is the actual motivation for Go? No perf/correctness defect is visible in the repo. Answer determines whether migration is warranted at all.
2. How does calamine render date cells into `ngay_sinh` today? Must be probed empirically against real 2016/2017 files before any Go xlsx library is chosen.
3. Is a fresh Rust-output parity baseline acceptable as the gate, given the original baseline is unregenerable?
+3 -3
View File
@@ -3,18 +3,18 @@
*
* `id` is the single identifier used end to end:
*
* data/<id>/ → parser/configs/<id>.toml → db/<id>.db.gz → /thptqg/<id>/
* data/<id>/ → parser/configs/<id>.yml → db/<id>.db.gz → /thptqg/<id>/
*
* Site path and database URL are derived from `id` rather than stored, so a
* dataset cannot be misconfigured into pointing at the wrong database.
*
* Imported by the Vite app *and* by parser/scripts/build-db.js under plain
* Imported by the Vite app *and* by go-parser/scripts/build-db.js under plain
* Node, so this module must stay free of `import.meta.env` and any Vite-only
* syntax. Callers pass the base URL in explicitly for that reason.
*/
// Extension is required: this module is also imported by plain Node
// (parser/scripts/build-db.js), which does not resolve extensionless paths.
// (go-parser/scripts/build-db.js), which does not resolve extensionless paths.
import { PRESETS_2016, PRESETS_2017 } from "./lib/sql-presets.js";
const SUBTITLE = "Dữ liệu thí sinh toàn quốc · Hỗ trợ truy vấn SQL tùy chỉnh";
+1 -1
View File
@@ -1,7 +1,7 @@
/**
* The 16 subject columns of the canonical schema, in display order.
*
* Mirrors `SCORE_FIELDS` in parser/src/schema.rs. Previously this list was
* Mirrors `SCORE_FIELDS` in go-parser/internal/schema/schema.go. Previously this list was
* maintained separately in score-table.jsx and student-detail.jsx, which is how
* they drifted out of sync with each other and with the database.
*
+1 -1
View File
@@ -11,7 +11,7 @@ import react from "@vitejs/plugin-react";
// and the existing ?q= deep links keep working, which that fallback would break.
//
// publicDir holds only the gzipped databases, staged there by
// parser/scripts/build-db.js. Nothing uncompressed is ever placed in it.
// go-parser/scripts/build-db.js. Nothing uncompressed is ever placed in it.
export default defineConfig({
plugins: [react()],
base: "/thptqg/",