mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
refactor(parser): reimplement the parser in Go alongside the Rust crate
Adds go-parser/, a Go reimplementation of the xlsxread parser, verified byte-for-byte against the Rust original before any cutover. Reader fidelity is exact across all 299 input files: the canonical cell dump of every sheet matches calamine's, locked in as a test against a committed hash oracle. Reaching that required replacing extrame/xls, which corrupted 69% of cells and dropped a further 28% on the BIFF corpus, with pbnjay/grate; correcting excelize's number-format application and trailing-cell trimming; restoring carriage returns that XML line-ending normalisation strips from 2,233 ten_cum_thi values; and gating numeric re-rendering on cell type so shared strings that merely look numeric keep their leading zeros. The differential gate compares both parsers over all four datasets: 3,265,641 rows with identical full-table SHA-256, identical per-column non-NULL counts, identical schema metadata and identical stdout. Config moves from TOML to YAML for both parsers, so they keep reading the same files and the gate stays meaningful. Verified by rebuilding 2016 and 2017-old2 with Rust under the new configs and matching the recorded counts. build-db.js now refuses to publish a database whose row count does not match the known figure, closing a path where an under-producing parser could ship a truncated public dataset with green CI. The deploy workflow gains a pull_request trigger and guards deploy to main, so branch verification can no longer publish to production.
This commit is contained in:
1 parent
8902232747
commit
0eb174721d
60 files changed
+6050
-189
No files matched your search
@@ -3,6 +3,11 @@ name: Deploy to GitHub Pages
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
# Pull requests run the build job only. Before this, the workflow triggered on
|
||||
# main-push and workflow_dispatch alone, so "verify on a branch first" was not
|
||||
# actually possible: pushing to a branch ran nothing, and dispatching from one
|
||||
# published that branch straight to the live site.
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
@@ -17,14 +22,18 @@ concurrency:
|
||||
jobs:
|
||||
build:
|
||||
runs-on: ubuntu-latest
|
||||
env:
|
||||
# The parser is pure Go — grate, excelize, yaml.v3 and modernc.org/sqlite
|
||||
# are all cgo-free — so no C toolchain is needed. Set explicitly rather
|
||||
# than relying on the default.
|
||||
CGO_ENABLED: '0'
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- uses: dtolnay/rust-toolchain@stable
|
||||
|
||||
- uses: Swatinem/rust-cache@v2
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
workspaces: parser
|
||||
go-version: '1.26'
|
||||
cache-dependency-path: go-parser/go.sum
|
||||
|
||||
- uses: actions/setup-node@v4
|
||||
with:
|
||||
@@ -34,12 +43,26 @@ jobs:
|
||||
|
||||
- run: npm ci
|
||||
|
||||
# The reader-fidelity suite compares all 299 real input files against a
|
||||
# committed hash oracle, so it is the regression guard for the whole
|
||||
# reader. Runs before anything is built.
|
||||
- name: Test parser
|
||||
run: npm run test:go
|
||||
|
||||
# excelize carries an open advisory, and the 2017 refresh runbook feeds
|
||||
# network-downloaded spreadsheets straight into the parser.
|
||||
- name: Vulnerability scan
|
||||
working-directory: go-parser
|
||||
run: |
|
||||
go install golang.org/x/vuln/cmd/govulncheck@latest
|
||||
"$(go env GOPATH)/bin/govulncheck" ./...
|
||||
|
||||
# One parser binary builds every dataset; build-db.js reads the dataset
|
||||
# list from src/datasets.js and gzips each database in place, leaving no
|
||||
# uncompressed file behind.
|
||||
# list from src/datasets.js, verifies each database against its known row
|
||||
# count, then gzips it in place leaving no uncompressed file behind.
|
||||
- name: Build databases
|
||||
run: |
|
||||
npm run build:rust
|
||||
npm run build:go
|
||||
npm run build:db
|
||||
|
||||
# One Vite build produces every page. scripts/assemble-site.js copies the
|
||||
@@ -53,6 +76,10 @@ jobs:
|
||||
path: _site
|
||||
|
||||
deploy:
|
||||
# Guarded to main. Without this, a workflow_dispatch from any branch would
|
||||
# publish that branch's output to the live site, and concurrency
|
||||
# cancel-in-progress would kill an in-flight good deploy on the way.
|
||||
if: github.ref == 'refs/heads/main'
|
||||
needs: build
|
||||
runs-on: ubuntu-latest
|
||||
environment:
|
||||
|
||||
@@ -22,3 +22,7 @@ parser/target/
|
||||
|
||||
### Assembled Pages artifact ###
|
||||
_site/
|
||||
|
||||
# Go parser build output and regenerable ground-truth dumps
|
||||
go-parser/bin/
|
||||
go-parser/testdata/dumps/
|
||||
@@ -25,7 +25,7 @@ index.html + src/ the frontend — one app serving all four datasets and the
|
||||
data/<id>/ raw Excel files, one directory per dataset
|
||||
parser/ the Rust parser
|
||||
src/schema.rs canonical 22-column table: DDL, INSERT, subject regexes
|
||||
configs/<id>.toml per-dataset parse rules only, no SQL
|
||||
configs/<id>.yml per-dataset parse rules only, no SQL
|
||||
scripts/ database build, crawler, parity verification
|
||||
scripts/ site assembly
|
||||
docs/ architecture, data pipeline, deployment
|
||||
@@ -34,14 +34,14 @@ docs/ architecture, data pipeline, deployment
|
||||
The dataset id is one identifier end to end:
|
||||
|
||||
```
|
||||
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
|
||||
data/2017-old/ → parser/configs/2017-old.yml → db/2017-old.db.gz → /thptqg/2017-old/
|
||||
```
|
||||
|
||||
## Build
|
||||
|
||||
```bash
|
||||
npm ci
|
||||
npm run build:rust # compile the parser
|
||||
npm run build:go # compile the parser
|
||||
npm run build:db # build + gzip all four databases (add an id for just one)
|
||||
npm run build:site # one Vite build, then assemble into _site/
|
||||
npx serve _site
|
||||
@@ -53,7 +53,7 @@ Pushing to `main` runs the same steps in
|
||||
## Adding a dataset
|
||||
|
||||
1. Put the Excel files in `data/<id>/`
|
||||
2. Add `parser/configs/<id>.toml` — sheet mode, column indices, validation
|
||||
2. Add `parser/configs/<id>.yml` — sheet mode, column indices, validation
|
||||
guards. No SQL; the schema is canonical.
|
||||
3. Add an entry to `DATASETS` in `src/datasets.js`
|
||||
|
||||
|
||||
@@ -4,8 +4,8 @@ From raw Excel files to a compressed SQLite file the browser can load.
|
||||
|
||||
One Rust binary (`parser/`) builds every dataset. What differs per dataset is
|
||||
parse rules only — sheet strategy, column layout, validation guards — declared
|
||||
in `parser/configs/<id>.toml`. The table shape, the INSERT and the subject
|
||||
regexes are canonical and live in `parser/src/schema.rs`.
|
||||
in `parser/configs/<id>.yml`. The table shape, the INSERT and the subject
|
||||
regexes are canonical and live in `go-parser/internal/schema/schema.go`.
|
||||
|
||||
## Sources
|
||||
|
||||
@@ -54,7 +54,7 @@ which is why only 2016 populates those columns.
|
||||
|
||||
## Score text parsing
|
||||
|
||||
`SCORE_PATTERNS` in `parser/src/schema.rs` defines one regex per subject, and
|
||||
`SCORE_PATTERNS` in `go-parser/internal/schema/schema.go` defines one regex per subject, and
|
||||
**all 16 run against every dataset**. A subject a given exam year did not offer
|
||||
simply never matches and stays NULL.
|
||||
|
||||
@@ -110,7 +110,7 @@ silently drops 13,720 students** (Hanoi +7,275, HCM +6,445). That is what
|
||||
| id | Source rows | Skipped | DB rows |
|
||||
| --- | --- | --- | --- |
|
||||
| `2016` | 877,464 | 3 duplicate SBDs collapsed | **877,461** |
|
||||
| `2017` | 861,131 | 63 empty | **861,068** |
|
||||
| `2017` | 861,068 | 0 | **861,068** |
|
||||
| `2017-old` | 847,349 | 1 header leak | **847,348** |
|
||||
| `2017-old2` | 679,764 | 0 | **679,764** |
|
||||
|
||||
@@ -135,7 +135,7 @@ regenerated — the two old crates no longer exist. Both scripts use the built-i
|
||||
```bash
|
||||
rm data/2017/*.xls
|
||||
node parser/scripts/crawl-baotintuc.js
|
||||
node parser/scripts/build-db.js 2017
|
||||
node go-parser/scripts/build-db.js 2017
|
||||
```
|
||||
|
||||
Then re-run the parity check above and confirm the row count still matches.
|
||||
|
||||
@@ -7,8 +7,8 @@ One-time setup: **Settings → Pages → Source: GitHub Actions**.
|
||||
|
||||
## What the workflow does
|
||||
|
||||
1. Checkout, Rust toolchain, Node 24, `npm ci`
|
||||
2. `npm run build:rust` — one parser binary
|
||||
1. Checkout, Go toolchain, Node 24, `npm ci`
|
||||
2. `npm run build:go` — one parser binary
|
||||
3. `npm run build:db` — builds and gzips all four databases into
|
||||
`.build/public/db/`
|
||||
4. `npm run build:site` — one Vite build, then `scripts/assemble-site.js`
|
||||
@@ -35,7 +35,7 @@ string intact.
|
||||
|
||||
```bash
|
||||
npm ci
|
||||
npm run build:rust
|
||||
npm run build:go
|
||||
npm run build:db # all four; pass an id to build just one
|
||||
npm run build:site # vite build + assemble into _site/
|
||||
npx serve _site
|
||||
@@ -44,7 +44,7 @@ npx serve _site
|
||||
To rebuild a single dataset:
|
||||
|
||||
```bash
|
||||
node parser/scripts/build-db.js 2017-old
|
||||
node go-parser/scripts/build-db.js 2017-old
|
||||
```
|
||||
|
||||
## Base path
|
||||
@@ -56,9 +56,9 @@ up as a blank page with 404s on `/assets/...`.
|
||||
## Adding a dataset
|
||||
|
||||
1. Put the Excel files in `data/<id>/`
|
||||
2. Add `parser/configs/<id>.toml` with the parse rules — sheet mode, column
|
||||
2. Add `parser/configs/<id>.yml` with the parse rules — sheet mode, column
|
||||
indices, SBD validation, header tokens, blank-row stripping. No SQL: the
|
||||
schema is canonical and lives in `parser/src/schema.rs`
|
||||
schema is canonical and lives in `go-parser/internal/schema/schema.go`
|
||||
3. Add an entry to `DATASETS` in `src/datasets.js`
|
||||
|
||||
Nothing else. The build script, the site assembly and the router all read that
|
||||
|
||||
@@ -31,11 +31,11 @@ data/<id>/*.xls(x)
|
||||
One identifier ties the whole pipeline together:
|
||||
|
||||
```
|
||||
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
|
||||
data/2017-old/ → parser/configs/2017-old.yml → db/2017-old.db.gz → /thptqg/2017-old/
|
||||
```
|
||||
|
||||
`src/datasets.js` declares the four ids once. The frontend, the database build
|
||||
(`parser/scripts/build-db.js`) and the site assembly all import that list, so
|
||||
(`go-parser/scripts/build-db.js`) and the site assembly all import that list, so
|
||||
adding a dataset means adding one entry and one config file.
|
||||
|
||||
| id | Exam | Rows | Source |
|
||||
@@ -47,7 +47,7 @@ adding a dataset means adding one entry and one config file.
|
||||
|
||||
## Canonical schema
|
||||
|
||||
Defined once in `parser/src/schema.rs` — DDL, INSERT, column order and the 16
|
||||
Defined once in `go-parser/internal/schema/schema.go` — DDL, INSERT, column order and the 16
|
||||
subject regexes. The four TOML configs carry no SQL at all, only per-dataset
|
||||
parse rules. Config parsing uses `deny_unknown_fields`, so a leftover `[schema]`
|
||||
block fails loudly instead of looking effective while `schema.rs` drives the
|
||||
|
||||
+1
-1
@@ -28,7 +28,7 @@ export default defineConfig([
|
||||
},
|
||||
{
|
||||
// Node-executed files (Vite config, parser tooling) run with Node globals.
|
||||
files: ['vite.config.js', 'scripts/**/*.js', 'parser/scripts/**/*.js'],
|
||||
files: ['vite.config.js', 'scripts/**/*.js', 'go-parser/scripts/**/*.js', 'parser/scripts/**/*.js'],
|
||||
languageOptions: {
|
||||
globals: { ...globals.node },
|
||||
},
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
// Command dumpcells emits the canonical cell rendering of a spreadsheet, for
|
||||
// comparison against the Rust/calamine ground truth produced by
|
||||
// parser/examples/dump_cells.rs.
|
||||
//
|
||||
// The canonical stream carries geometry and rendered cell values only. The
|
||||
// calamine Data variant is deliberately excluded: Data::Empty and
|
||||
// Data::String("") both render "" and both count as blank everywhere
|
||||
// downstream, so the distinction cannot affect the database.
|
||||
//
|
||||
// Usage: dumpcells <spreadsheet> [out-file]
|
||||
package main
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
)
|
||||
|
||||
// escape mirrors the Rust dumper so field separators can never break the format.
|
||||
func escape(s string) string {
|
||||
var b strings.Builder
|
||||
b.Grow(len(s))
|
||||
for _, ch := range s {
|
||||
switch ch {
|
||||
case '\\':
|
||||
b.WriteString(`\\`)
|
||||
case '\t':
|
||||
b.WriteString(`\t`)
|
||||
case '\n':
|
||||
b.WriteString(`\n`)
|
||||
case '\r':
|
||||
b.WriteString(`\r`)
|
||||
default:
|
||||
b.WriteRune(ch)
|
||||
}
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
func main() {
|
||||
if len(os.Args) < 2 {
|
||||
fmt.Fprintln(os.Stderr, "usage: dumpcells <spreadsheet> [out-file]")
|
||||
os.Exit(2)
|
||||
}
|
||||
path := os.Args[1]
|
||||
|
||||
out := os.Stdout
|
||||
if len(os.Args) > 2 {
|
||||
f, err := os.Create(os.Args[2])
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "create: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
defer f.Close()
|
||||
out = f
|
||||
}
|
||||
w := bufio.NewWriterSize(out, 1<<20)
|
||||
defer w.Flush()
|
||||
|
||||
wb, err := reader.Open(path)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "open: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
defer wb.Close()
|
||||
|
||||
sheets := wb.Sheets()
|
||||
fmt.Fprintf(w, "FILE\t%s\n", escape(path))
|
||||
fmt.Fprintf(w, "SHEETCOUNT\t%d\n", len(sheets))
|
||||
|
||||
for _, sh := range sheets {
|
||||
fmt.Fprintf(w, "SHEET\t%d\t%s\t%d\t%d\n", sh.Index, escape(sh.Name), sh.Height, sh.Width)
|
||||
err := wb.EachRow(sh.Index, func(s reader.Sheet, rowIdx int, row []reader.Cell) error {
|
||||
fmt.Fprintf(w, "ROW\t%d\t%d\t%d\n", s.Index, rowIdx, len(row))
|
||||
for c, cell := range row {
|
||||
fmt.Fprintf(w, "CELL\t%d\t%d\t%d\t%s\n", s.Index, rowIdx, c, escape(cell.Str))
|
||||
}
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "rows: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,115 @@
|
||||
// Command xlsxread reads .xls/.xlsx files and builds SQLite databases for the
|
||||
// thptqg datasets.
|
||||
//
|
||||
// The CLI contract is fixed by parser/scripts/build-db.js and must match the
|
||||
// Rust binary exactly:
|
||||
//
|
||||
// xlsxread build --schema <config.yml> --input <dir> --output <db>
|
||||
// xlsxread audit --schema <config.yml> --input <dir> --db <db>
|
||||
//
|
||||
// Implemented with the standard flag package rather than a CLI framework: two
|
||||
// subcommands with three flags each do not justify a dependency.
|
||||
package main
|
||||
|
||||
import (
|
||||
"flag"
|
||||
"fmt"
|
||||
"os"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/audit"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/config"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/ingest"
|
||||
)
|
||||
|
||||
func usage() {
|
||||
fmt.Fprint(os.Stderr, `xlsxread — read .xls/.xlsx files and build SQLite databases for thptqg datasets
|
||||
|
||||
Usage:
|
||||
xlsxread build --schema <config.yml> --input <dir> --output <db>
|
||||
xlsxread audit --schema <config.yml> --input <dir> --db <db>
|
||||
`)
|
||||
}
|
||||
|
||||
func main() {
|
||||
if len(os.Args) < 2 {
|
||||
usage()
|
||||
os.Exit(2)
|
||||
}
|
||||
|
||||
switch os.Args[1] {
|
||||
case "build":
|
||||
runBuild(os.Args[2:])
|
||||
case "audit":
|
||||
runAudit(os.Args[2:])
|
||||
case "-h", "--help", "help":
|
||||
usage()
|
||||
default:
|
||||
fmt.Fprintf(os.Stderr, "unknown subcommand %q\n\n", os.Args[1])
|
||||
usage()
|
||||
os.Exit(2)
|
||||
}
|
||||
}
|
||||
|
||||
func runBuild(args []string) {
|
||||
fs := flag.NewFlagSet("build", flag.ExitOnError)
|
||||
schemaPath := fs.String("schema", "", "path to the dataset YAML config file")
|
||||
inputDir := fs.String("input", "", "directory containing the .xls / .xlsx source files")
|
||||
outputPath := fs.String("output", "", "output SQLite database path")
|
||||
fs.Parse(args)
|
||||
|
||||
if *schemaPath == "" || *inputDir == "" || *outputPath == "" {
|
||||
fmt.Fprintln(os.Stderr, "build requires --schema, --input and --output")
|
||||
os.Exit(2)
|
||||
}
|
||||
|
||||
cfg, err := config.Load(*schemaPath)
|
||||
if err != nil {
|
||||
fatalf("Failed to load config: %v", err)
|
||||
}
|
||||
|
||||
// The 2016 dataset selects its column layout per file at runtime; every other
|
||||
// dataset uses the fixed columns: mapping (main.rs:63-67).
|
||||
if cfg.FormatDetection != nil && *cfg.FormatDetection == "thptqg2016" {
|
||||
if err := ingest.Detect2016(cfg, *inputDir, *outputPath); err != nil {
|
||||
fatalf("%v", err)
|
||||
}
|
||||
return
|
||||
}
|
||||
if err := ingest.Standard(cfg, *inputDir, *outputPath); err != nil {
|
||||
fatalf("%v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func runAudit(args []string) {
|
||||
fs := flag.NewFlagSet("audit", flag.ExitOnError)
|
||||
schemaPath := fs.String("schema", "", "path to the dataset YAML config file")
|
||||
inputDir := fs.String("input", "", "directory containing the .xlsx source files")
|
||||
dbPath := fs.String("db", "", "SQLite database to compare against")
|
||||
fs.Parse(args)
|
||||
|
||||
if *schemaPath == "" || *inputDir == "" || *dbPath == "" {
|
||||
fmt.Fprintln(os.Stderr, "audit requires --schema, --input and --db")
|
||||
os.Exit(2)
|
||||
}
|
||||
|
||||
cfg, err := config.Load(*schemaPath)
|
||||
if err != nil {
|
||||
fatalf("Failed to load config: %v", err)
|
||||
}
|
||||
res, err := audit.Run(*inputDir, *dbPath, cfg)
|
||||
if err != nil {
|
||||
fatalf("%v", err)
|
||||
}
|
||||
audit.PrintReport(res)
|
||||
|
||||
// A mismatch is a non-zero exit even though the audit itself succeeded
|
||||
// (main.rs:46-48) — CI treats it as a failure signal.
|
||||
if !res.Matched {
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
func fatalf(format string, a ...any) {
|
||||
fmt.Fprintf(os.Stderr, "Error: "+format+"\n", a...)
|
||||
os.Exit(1)
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
module github.com/tiennm99/thptqg/go-parser
|
||||
|
||||
go 1.26.5
|
||||
|
||||
require (
|
||||
github.com/pbnjay/grate v0.0.0-20231006022435-3f8e65d74a14
|
||||
github.com/xuri/excelize/v2 v2.11.0
|
||||
golang.org/x/text v0.41.0
|
||||
gopkg.in/yaml.v3 v3.0.1
|
||||
modernc.org/sqlite v1.56.0
|
||||
)
|
||||
|
||||
require (
|
||||
github.com/dustin/go-humanize v1.0.1 // indirect
|
||||
github.com/google/uuid v1.6.0 // indirect
|
||||
github.com/mattn/go-isatty v0.0.24 // indirect
|
||||
github.com/ncruces/go-strftime v1.0.0 // indirect
|
||||
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
|
||||
github.com/richardlehane/mscfb v1.0.7 // indirect
|
||||
github.com/richardlehane/msoleps v1.0.6 // indirect
|
||||
github.com/tiendc/go-deepcopy v1.7.2 // indirect
|
||||
github.com/xuri/efp v0.0.1 // indirect
|
||||
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9 // indirect
|
||||
golang.org/x/crypto v0.53.0 // indirect
|
||||
golang.org/x/net v0.56.0 // indirect
|
||||
golang.org/x/sys v0.47.0 // indirect
|
||||
modernc.org/libc v1.74.4 // indirect
|
||||
modernc.org/mathutil v1.7.1 // indirect
|
||||
modernc.org/memory v1.11.0 // indirect
|
||||
)
|
||||
@@ -0,0 +1,82 @@
|
||||
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
|
||||
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
||||
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
|
||||
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
|
||||
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3 h1:LMLX+LgTNWpfvCBdFebv6EsYotImrt/Ppc5cXIriCSo=
|
||||
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3/go.mod h1:jl5iWTm0/hd5PjEYEOuwAJ57L/CibdZfrqZ5XA5GrCk=
|
||||
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
|
||||
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
|
||||
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
|
||||
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
|
||||
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
|
||||
github.com/mattn/go-isatty v0.0.24/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
|
||||
github.com/ncruces/go-strftime v1.0.0 h1:HMFp8mLCTPp341M/ZnA4qaf7ZlsbTc+miZjCLOFAw7w=
|
||||
github.com/ncruces/go-strftime v1.0.0/go.mod h1:Fwc5htZGVVkseilnfgOVb9mKy6w1naJmn9CehxcKcls=
|
||||
github.com/pbnjay/grate v0.0.0-20231006022435-3f8e65d74a14 h1:ZfXdW7GIVZT3Z9oejLJ+GHrrQv/ezU2Bwqn0BF37s4g=
|
||||
github.com/pbnjay/grate v0.0.0-20231006022435-3f8e65d74a14/go.mod h1:VaZEKQrYbYr2untVA/EFNdC6hM7GyARRNM+k4+5CmA0=
|
||||
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
|
||||
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
|
||||
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE=
|
||||
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo=
|
||||
github.com/richardlehane/mscfb v1.0.7 h1:oeoiM0WE79vHwE8RpIYYvIAc8ajTH2mb6UZm55/+EB0=
|
||||
github.com/richardlehane/mscfb v1.0.7/go.mod h1:pe0+IUIc0AHh0+teNzBlJCtSyZdFOGgV4ZK9bsoV+Jo=
|
||||
github.com/richardlehane/msoleps v1.0.6 h1:9BvkpjvD+iUBalUY4esMwv6uBkfOip/Lzvd93jvR9gg=
|
||||
github.com/richardlehane/msoleps v1.0.6/go.mod h1:BWev5JBpU9Ko2WAgmZEuiz4/u3ZYTKbjLycmwiWUfWg=
|
||||
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
|
||||
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
|
||||
github.com/tiendc/go-deepcopy v1.7.2 h1:Ut2yYR7W9tWjTQitganoIue4UGxZwCcJy3orjrrIj44=
|
||||
github.com/tiendc/go-deepcopy v1.7.2/go.mod h1:4bKjNC2r7boYOkD2IOuZpYjmlDdzjbpTRyCx+goBCJQ=
|
||||
github.com/xuri/efp v0.0.1 h1:fws5Rv3myXyYni8uwj2qKjVaRP30PdjeYe2Y6FDsCL8=
|
||||
github.com/xuri/efp v0.0.1/go.mod h1:ybY/Jr0T0GTCnYjKqmdwxyxn2BQf2RcQIIvex5QldPI=
|
||||
github.com/xuri/excelize/v2 v2.11.0 h1:HxaEFl6sRN2+8J5a8HaKq+0M4FsjBGMnWWtjOCPSG88=
|
||||
github.com/xuri/excelize/v2 v2.11.0/go.mod h1:jxFLbzaIwGQ5ufFNvYfUOHqXhfPaNmP14KWfmNz2Uak=
|
||||
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9 h1:+C0TIdyyYmzadGaL/HBLbf3WdLgC29pgyhTjAT/0nuE=
|
||||
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9/go.mod h1:WwHg+CVyzlv/TX9xqBFXEZAuxOPxn2k1GNHwG41IIUQ=
|
||||
golang.org/x/crypto v0.53.0 h1:QZ4Muo8THX6CizN2vPPd5fBGHyogrdK9fG4wLPFUsto=
|
||||
golang.org/x/crypto v0.53.0/go.mod h1:DNLU434OwVakk9PzuwV8w62mAJpRJL3vsgcfp4Qnsio=
|
||||
golang.org/x/image v0.38.0 h1:5l+q+Y9JDC7mBOMjo4/aPhMDcxEptsX+Tt3GgRQRPuE=
|
||||
golang.org/x/image v0.38.0/go.mod h1:/3f6vaXC+6CEanU4KJxbcUZyEePbyKbaLoDOe4ehFYY=
|
||||
golang.org/x/mod v0.38.0 h1:MECBjubtXD7yj4HrhIUcywNaGeNVUdfVnxmPajOk4yk=
|
||||
golang.org/x/mod v0.38.0/go.mod h1:V6Xz0pq8TQ3dGqVQ1FVHuelZpAL0uNhSkk9ogYP3c40=
|
||||
golang.org/x/net v0.56.0 h1:Rw8j/hFzGvJUZwNBXnAtf5sVDVt+65SK2C7IxCxZt5o=
|
||||
golang.org/x/net v0.56.0/go.mod h1:D3Ku6r+V6JROoZK144D2XfMHFcMq/0zSfLelVTCFKec=
|
||||
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
|
||||
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
|
||||
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
|
||||
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
|
||||
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
|
||||
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
|
||||
golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE=
|
||||
golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk=
|
||||
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM=
|
||||
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
|
||||
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
|
||||
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
|
||||
modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI=
|
||||
modernc.org/cc/v4 v4.29.1/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
|
||||
modernc.org/ccgo/v4 v4.34.6 h1:sBgfIwyN0TQ9C5hwIeuqyeAKyMWnbvj2fvpF4L11uzU=
|
||||
modernc.org/ccgo/v4 v4.34.6/go.mod h1:SZ8YcN9NG7XVsQYdm6jYBvi8PQP1qi+kqB6OhjqI3Fk=
|
||||
modernc.org/fileutil v1.4.0 h1:j6ZzNTftVS054gi281TyLjHPp6CPHr2KCxEXjEbD6SM=
|
||||
modernc.org/fileutil v1.4.0/go.mod h1:EqdKFDxiByqxLk8ozOxObDSfcVOv/54xDs/DUHdvCUU=
|
||||
modernc.org/gc/v2 v2.6.5 h1:nyqdV8q46KvTpZlsw66kWqwXRHdjIlJOhG6kxiV/9xI=
|
||||
modernc.org/gc/v2 v2.6.5/go.mod h1:YgIahr1ypgfe7chRuJi2gD7DBQiKSLMPgBQe9oIiito=
|
||||
modernc.org/gc/v3 v3.1.4 h1:2g65LGVSmFQrXeITAw97x7hCRvZFcyE1uDP+7Vng7JI=
|
||||
modernc.org/gc/v3 v3.1.4/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
|
||||
modernc.org/goabi0 v0.2.0 h1:HvEowk7LxcPd0eq6mVOAEMai46V+i7Jrj13t4AzuNks=
|
||||
modernc.org/goabi0 v0.2.0/go.mod h1:CEFRnnJhKvWT1c1JTI3Avm+tgOWbkOu5oPA8eH8LnMI=
|
||||
modernc.org/libc v1.74.4 h1:fX1Omw4o2/1C2iRkkIsrQTasJQldLhRmuPreXLoWs9k=
|
||||
modernc.org/libc v1.74.4/go.mod h1:eeQAS9W3sZeKYMFubydxJpII9ybHWshk+7or7bLG9co=
|
||||
modernc.org/mathutil v1.7.1 h1:GCZVGXdaN8gTqB1Mf/usp1Y/hSqgI2vAGGP4jZMCxOU=
|
||||
modernc.org/mathutil v1.7.1/go.mod h1:4p5IwJITfppl0G4sUEDtCr4DthTaT47/N3aT6MhfgJg=
|
||||
modernc.org/memory v1.11.0 h1:o4QC8aMQzmcwCK3t3Ux/ZHmwFPzE6hf2Y5LbkRs+hbI=
|
||||
modernc.org/memory v1.11.0/go.mod h1:/JP4VbVC+K5sU2wZi9bHoq2MAkCnrt2r98UGeSK7Mjw=
|
||||
modernc.org/opt v0.2.0 h1:tGyef5ApycA7FSEOMraay9SaTk5zmbx7Tu+cJs4QKZg=
|
||||
modernc.org/opt v0.2.0/go.mod h1:03fq9lsNfvkYSfxrfUhZCWPk1lm4cq4N+Bh//bEtgns=
|
||||
modernc.org/sortutil v1.2.1 h1:+xyoGf15mM3NMlPDnFqrteY07klSFxLElE2PVuWIJ7w=
|
||||
modernc.org/sortutil v1.2.1/go.mod h1:7ZI3a3REbai7gzCLcotuw9AC4VZVpYMjDzETGsSMqJE=
|
||||
modernc.org/sqlite v1.56.0 h1:/D8e2RfFqoy/Zc6PuC76U28zFwmI/sYx1Kjm4yEn9e0=
|
||||
modernc.org/sqlite v1.56.0/go.mod h1:yCJ2cmAaIkHQ25oXWrF8H4O1lIfPYPR26yCEDj2P3pQ=
|
||||
modernc.org/strutil v1.2.1 h1:UneZBkQA+DX2Rp35KcM69cSsNES9ly8mQWD71HKlOA0=
|
||||
modernc.org/strutil v1.2.1/go.mod h1:EHkiggD70koQxjVdSBM3JKM7k6L0FbGE5eymy9i3B9A=
|
||||
modernc.org/token v1.1.0 h1:Xl7Ap9dKaEs5kLoOQeQmPWevfnk/DM5qcLcYlA8ys6Y=
|
||||
modernc.org/token v1.1.0/go.mod h1:UGzOrNV1mAFSEB63lOFHIpNRUVMvYTc6yu1SMY/XTDM=
|
||||
@@ -0,0 +1,148 @@
|
||||
// Package audit compares distinct SBDs in the source spreadsheets against the
|
||||
// row count in a built database — a port of parser/src/audit.rs.
|
||||
//
|
||||
// Two deliberate divergences from the build path are preserved, both inherited
|
||||
// from audit-row-counts.js:
|
||||
//
|
||||
// - only .xlsx files are considered (audit.rs:53-59)
|
||||
// - only sheet 0 is read, regardless of sheet_mode (audit.rs:81-84)
|
||||
//
|
||||
// These are intentional, not oversights. Do not "fix" them.
|
||||
package audit
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/config"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/ingest"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/sqlitedb"
|
||||
)
|
||||
|
||||
// Result carries the audit counters.
|
||||
type Result struct {
|
||||
TotalDataRows uint64
|
||||
BothEmpty uint64
|
||||
EmptyName uint64
|
||||
EmptySbd uint64
|
||||
DistinctSbds int
|
||||
DBCount int64
|
||||
Matched bool
|
||||
}
|
||||
|
||||
// Run collects distinct SBDs from the .xlsx files in inputDir and compares the
|
||||
// total against the student row count in dbPath.
|
||||
func Run(inputDir, dbPath string, cfg *config.DatasetConfig) (*Result, error) {
|
||||
entries, err := os.ReadDir(inputDir)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read input dir %s: %w", inputDir, err)
|
||||
}
|
||||
var files []string
|
||||
for _, e := range entries {
|
||||
if e.IsDir() {
|
||||
continue
|
||||
}
|
||||
if strings.EqualFold(filepath.Ext(e.Name()), ".xlsx") {
|
||||
files = append(files, filepath.Join(inputDir, e.Name()))
|
||||
}
|
||||
}
|
||||
sort.Strings(files)
|
||||
|
||||
res := &Result{}
|
||||
seen := make(map[string]struct{})
|
||||
|
||||
// The audit uses positional defaults for format-detection configs, mirroring
|
||||
// audit-row-counts.js's fixed column assumption (audit.rs:104-111).
|
||||
hoTenCol, sbdCol := 1, 0
|
||||
if cfg.Columns != nil {
|
||||
hoTenCol, sbdCol = cfg.Columns.HoTen, cfg.Columns.SoBaoDanh
|
||||
}
|
||||
|
||||
for _, file := range files {
|
||||
wb, err := reader.Open(file)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sheets := wb.Sheets()
|
||||
if len(sheets) == 0 {
|
||||
wb.Close()
|
||||
continue
|
||||
}
|
||||
|
||||
firstRow := true
|
||||
err = wb.EachRow(sheets[0].Index, func(_ reader.Sheet, _ int, row []reader.Cell) error {
|
||||
if firstRow {
|
||||
firstRow = false
|
||||
if ingest.IsHeaderRow(row, cfg.Header.Tokens) {
|
||||
return nil
|
||||
}
|
||||
}
|
||||
res.TotalDataRows++
|
||||
|
||||
hoTen := cellAt(row, hoTenCol)
|
||||
sbd := cellAt(row, sbdCol)
|
||||
|
||||
if hoTen == "" && sbd == "" {
|
||||
res.BothEmpty++
|
||||
return nil
|
||||
}
|
||||
if hoTen == "" {
|
||||
res.EmptyName++
|
||||
}
|
||||
if sbd == "" {
|
||||
res.EmptySbd++
|
||||
}
|
||||
if sbd != "" {
|
||||
seen[sbd] = struct{}{}
|
||||
}
|
||||
return nil
|
||||
})
|
||||
wb.Close()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
|
||||
// Opened read-only: the audit must never mutate the database it inspects
|
||||
// (audit.rs:139).
|
||||
db, err := sql.Open(sqlitedb.DriverName, "file:"+dbPath+"?mode=ro")
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("open db %s: %w", dbPath, err)
|
||||
}
|
||||
defer db.Close()
|
||||
if err := db.QueryRow("SELECT COUNT(*) FROM student").Scan(&res.DBCount); err != nil {
|
||||
return nil, fmt.Errorf("count rows: %w", err)
|
||||
}
|
||||
|
||||
res.DistinctSbds = len(seen)
|
||||
res.Matched = int64(res.DistinctSbds) == res.DBCount
|
||||
return res, nil
|
||||
}
|
||||
|
||||
// PrintReport mirrors audit-row-counts.js:54-62 exactly.
|
||||
func PrintReport(r *Result) {
|
||||
fmt.Println("=== Source vs DB ===")
|
||||
fmt.Printf("Source: total data rows across all files: %d\n", r.TotalDataRows)
|
||||
fmt.Printf("Source: rows with empty name AND sbd (skipped): %d\n", r.BothEmpty)
|
||||
fmt.Printf("Source: rows with missing name only: %d\n", r.EmptyName)
|
||||
fmt.Printf("Source: rows with missing sbd only: %d\n", r.EmptySbd)
|
||||
fmt.Printf("Source: distinct SBDs: %d\n", r.DistinctSbds)
|
||||
fmt.Printf("DB: row count: %d\n", r.DBCount)
|
||||
if r.Matched {
|
||||
fmt.Println("Match: YES — all unique SBDs accounted for")
|
||||
} else {
|
||||
fmt.Printf("Match: NO — gap of %d\n", int64(r.DistinctSbds)-r.DBCount)
|
||||
}
|
||||
}
|
||||
|
||||
func cellAt(row []reader.Cell, idx int) string {
|
||||
if idx < 0 || idx >= len(row) {
|
||||
return ""
|
||||
}
|
||||
return strings.TrimSpace(row[idx].Str)
|
||||
}
|
||||
@@ -0,0 +1,113 @@
|
||||
// Package config loads per-dataset parse rules — a port of parser/src/config.rs.
|
||||
//
|
||||
// The config deliberately carries no SQL. The table shape, the INSERT and the
|
||||
// subject regexes are identical for every dataset and live in internal/schema;
|
||||
// keeping them here meant four copies of the same DDL, which is how the 2016 and
|
||||
// 2017 schemas drifted apart.
|
||||
package config
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"fmt"
|
||||
"os"
|
||||
|
||||
"gopkg.in/yaml.v3"
|
||||
)
|
||||
|
||||
// SheetMode selects which sheets of a workbook are read.
|
||||
type SheetMode string
|
||||
|
||||
const (
|
||||
// SheetModeAll iterates every sheet, which is what recovers the Hanoi and
|
||||
// HCM rows that overflow past Excel's 65,536-row cap into a second sheet.
|
||||
SheetModeAll SheetMode = "all"
|
||||
// SheetModeFirst reads sheet 0 only.
|
||||
SheetModeFirst SheetMode = "first"
|
||||
)
|
||||
|
||||
// valid reports whether m is one of the two values the Rust enum accepts.
|
||||
//
|
||||
// Checked after decoding rather than via a custom unmarshaler: the decoder
|
||||
// assigns named string types directly, so a typo would otherwise decode
|
||||
// silently and be read as "not all" downstream.
|
||||
func (m SheetMode) valid() bool {
|
||||
return m == SheetModeAll || m == SheetModeFirst
|
||||
}
|
||||
|
||||
// DatasetConfig is the per-dataset parse rule set.
|
||||
type DatasetConfig struct {
|
||||
Reader ReaderCfg `yaml:"reader"`
|
||||
// Columns holds fixed column indices. Nil when FormatDetection handles
|
||||
// per-file mapping — a pointer rather than a value because a zero ColumnMap
|
||||
// would silently mean "every column is index 0".
|
||||
Columns *ColumnMap `yaml:"columns"`
|
||||
Validation ValidationCfg `yaml:"validation"`
|
||||
Header HeaderCfg `yaml:"header"`
|
||||
// FormatDetection, when set to "thptqg2016", enables per-file format
|
||||
// auto-detection: each file's header row is inspected at runtime to choose
|
||||
// the column layout (separate-scores / mapped / default-positional).
|
||||
FormatDetection *string `yaml:"format_detection"`
|
||||
}
|
||||
|
||||
type ReaderCfg struct {
|
||||
SheetMode SheetMode `yaml:"sheet_mode"`
|
||||
// StripBlankRows skips rows where every cell is empty before counting them
|
||||
// as source rows (a 2017-old2 quirk).
|
||||
StripBlankRows bool `yaml:"strip_blank_rows"`
|
||||
}
|
||||
|
||||
// ColumnMap holds zero-indexed column positions in the source row. Used by the
|
||||
// 2017 configs; 2016 uses runtime format detection instead.
|
||||
type ColumnMap struct {
|
||||
HoTen int `yaml:"ho_ten"`
|
||||
NgaySinh int `yaml:"ngay_sinh"`
|
||||
SoBaoDanh int `yaml:"so_bao_danh"`
|
||||
DiemThi int `yaml:"diem_thi"`
|
||||
}
|
||||
|
||||
type ValidationCfg struct {
|
||||
// RequireNumericSbd mirrors build-database-old.js / -old2.js, which require
|
||||
// so_bao_danh to match ^\d+$.
|
||||
RequireNumericSbd bool `yaml:"require_numeric_sbd"`
|
||||
RequireNonemptyName bool `yaml:"require_nonempty_name"`
|
||||
RequireNonemptySbd bool `yaml:"require_nonempty_sbd"`
|
||||
}
|
||||
|
||||
type HeaderCfg struct {
|
||||
// Tokens are matched against the uppercased first cell to detect a header row.
|
||||
Tokens []string `yaml:"tokens"`
|
||||
}
|
||||
|
||||
// Parse decodes a config, rejecting any key the struct does not declare.
|
||||
//
|
||||
// Strictness is load-bearing and has its own test. Rust gets it from serde's
|
||||
// deny_unknown_fields; yaml.v3 needs KnownFields(true) explicitly, and Go YAML
|
||||
// decoders ignore unknown keys by default. Without it a leftover `schema:`
|
||||
// mapping would look effective while internal/schema actually drove the build —
|
||||
// the drift that produced two divergent schemas before the unification.
|
||||
func Parse(src []byte) (*DatasetConfig, error) {
|
||||
var cfg DatasetConfig
|
||||
dec := yaml.NewDecoder(bytes.NewReader(src))
|
||||
dec.KnownFields(true)
|
||||
if err := dec.Decode(&cfg); err != nil {
|
||||
return nil, fmt.Errorf("decode config: %w", err)
|
||||
}
|
||||
if !cfg.Reader.SheetMode.valid() {
|
||||
return nil, fmt.Errorf("invalid sheet_mode %q (want %q or %q)",
|
||||
cfg.Reader.SheetMode, SheetModeAll, SheetModeFirst)
|
||||
}
|
||||
return &cfg, nil
|
||||
}
|
||||
|
||||
// Load reads and parses a config file.
|
||||
func Load(path string) (*DatasetConfig, error) {
|
||||
src, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read config %s: %w", path, err)
|
||||
}
|
||||
cfg, err := Parse(src)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", path, err)
|
||||
}
|
||||
return cfg, nil
|
||||
}
|
||||
@@ -0,0 +1,227 @@
|
||||
package config
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// sampleYAML mirrors SAMPLE_YAML in parser/src/config.rs.
|
||||
const sampleYAML = `
|
||||
reader:
|
||||
sheet_mode: all
|
||||
strip_blank_rows: false
|
||||
|
||||
columns:
|
||||
ho_ten: 0
|
||||
ngay_sinh: 1
|
||||
so_bao_danh: 2
|
||||
diem_thi: 3
|
||||
|
||||
validation:
|
||||
require_numeric_sbd: false
|
||||
require_nonempty_name: true
|
||||
require_nonempty_sbd: true
|
||||
|
||||
header:
|
||||
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
`
|
||||
|
||||
// TestConfigRoundTrip ports config_round_trip (config.rs:115).
|
||||
func TestConfigRoundTrip(t *testing.T) {
|
||||
cfg, err := Parse([]byte(sampleYAML))
|
||||
if err != nil {
|
||||
t.Fatalf("parse failed: %v", err)
|
||||
}
|
||||
if cfg.Reader.SheetMode != SheetModeAll {
|
||||
t.Errorf("sheet_mode = %q, want all", cfg.Reader.SheetMode)
|
||||
}
|
||||
if cfg.Reader.StripBlankRows {
|
||||
t.Error("strip_blank_rows should be false")
|
||||
}
|
||||
if cfg.Columns == nil {
|
||||
t.Fatal("columns should be present")
|
||||
}
|
||||
if cfg.Columns.HoTen != 0 {
|
||||
t.Errorf("ho_ten = %d, want 0", cfg.Columns.HoTen)
|
||||
}
|
||||
if cfg.Columns.DiemThi != 3 {
|
||||
t.Errorf("diem_thi = %d, want 3", cfg.Columns.DiemThi)
|
||||
}
|
||||
if cfg.Validation.RequireNumericSbd {
|
||||
t.Error("require_numeric_sbd should be false")
|
||||
}
|
||||
if !cfg.Validation.RequireNonemptyName {
|
||||
t.Error("require_nonempty_name should be true")
|
||||
}
|
||||
if len(cfg.Header.Tokens) != 3 {
|
||||
t.Errorf("tokens = %d, want 3", len(cfg.Header.Tokens))
|
||||
}
|
||||
if cfg.FormatDetection != nil {
|
||||
t.Errorf("format_detection = %v, want nil", *cfg.FormatDetection)
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigRejectsLeftoverSQLSections ports config_rejects_leftover_sql_sections
|
||||
// (config.rs:132). This is the load-bearing one: Rust gets the behaviour free
|
||||
// from serde's deny_unknown_fields, whereas Go YAML decoders ignore unknown keys
|
||||
// unless KnownFields(true) is set. Without it a stale schema: mapping would look effective while
|
||||
// internal/schema silently drove the build — exactly the drift that produced two
|
||||
// divergent schemas before the unification.
|
||||
func TestConfigRejectsLeftoverSQLSections(t *testing.T) {
|
||||
withDDL := sampleYAML + "\nschema:\n ddl: \"CREATE TABLE student (so_bao_danh TEXT);\"\n"
|
||||
if _, err := Parse([]byte(withDDL)); err == nil {
|
||||
t.Fatal("config with a leftover schema: mapping must be rejected")
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigRejectsUnknownScalarKey guards the same property for a stray scalar,
|
||||
// not just a stray mapping.
|
||||
func TestConfigRejectsUnknownScalarKey(t *testing.T) {
|
||||
withKey := sampleYAML + "\nunexpected_key: 1\n"
|
||||
if _, err := Parse([]byte(withKey)); err == nil {
|
||||
t.Fatal("config with an unknown scalar key must be rejected")
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigFirstSheetMode ports config_first_sheet_mode (config.rs:140).
|
||||
func TestConfigFirstSheetMode(t *testing.T) {
|
||||
src := []byte(replaceAll(sampleYAML, "sheet_mode: all", "sheet_mode: first"))
|
||||
cfg, err := Parse(src)
|
||||
if err != nil {
|
||||
t.Fatalf("parse failed: %v", err)
|
||||
}
|
||||
if cfg.Reader.SheetMode != SheetModeFirst {
|
||||
t.Errorf("sheet_mode = %q, want first", cfg.Reader.SheetMode)
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigRejectsUnknownSheetMode: the Rust enum accepts only "all"/"first",
|
||||
// so anything else must fail rather than defaulting.
|
||||
func TestConfigRejectsUnknownSheetMode(t *testing.T) {
|
||||
src := []byte(replaceAll(sampleYAML, "sheet_mode: all", "sheet_mode: second"))
|
||||
if _, err := Parse(src); err == nil {
|
||||
t.Fatal("unknown sheet_mode must be rejected")
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigFormatDetectionField ports config_format_detection_field
|
||||
// (config.rs:147): a 2016-style config has no columns: mapping at all.
|
||||
func TestConfigFormatDetectionField(t *testing.T) {
|
||||
const src = `
|
||||
format_detection: thptqg2016
|
||||
|
||||
reader:
|
||||
sheet_mode: all
|
||||
strip_blank_rows: false
|
||||
|
||||
validation:
|
||||
require_numeric_sbd: false
|
||||
require_nonempty_name: true
|
||||
require_nonempty_sbd: true
|
||||
|
||||
header:
|
||||
tokens: ["SBD", "SOBAODANH", "STT"]
|
||||
`
|
||||
cfg, err := Parse([]byte(src))
|
||||
if err != nil {
|
||||
t.Fatalf("parse failed: %v", err)
|
||||
}
|
||||
if cfg.FormatDetection == nil || *cfg.FormatDetection != "thptqg2016" {
|
||||
t.Errorf("format_detection = %v, want thptqg2016", cfg.FormatDetection)
|
||||
}
|
||||
if cfg.Columns != nil {
|
||||
t.Error("columns must be nil when format_detection drives the layout")
|
||||
}
|
||||
}
|
||||
|
||||
// TestLoadRealConfigs loads the four shipped configs — the Go binary reads the
|
||||
// same files as Rust, never a fork — and asserts the per-dataset differences
|
||||
// recorded during scouting.
|
||||
func TestLoadRealConfigs(t *testing.T) {
|
||||
root := repoRoot(t)
|
||||
want := map[string]struct {
|
||||
sheetMode SheetMode
|
||||
stripBlank bool
|
||||
numericSbd bool
|
||||
hasColumns bool
|
||||
formatDet string
|
||||
tokenCount int
|
||||
}{
|
||||
"2016": {SheetModeAll, false, false, false, "thptqg2016", 6},
|
||||
"2017": {SheetModeAll, false, false, true, "", 3},
|
||||
"2017-old": {SheetModeFirst, false, true, true, "", 3},
|
||||
"2017-old2": {SheetModeAll, true, true, true, "", 3},
|
||||
}
|
||||
for id, w := range want {
|
||||
t.Run(id, func(t *testing.T) {
|
||||
cfg, err := Load(filepath.Join(root, "parser", "configs", id+".yml"))
|
||||
if err != nil {
|
||||
t.Fatalf("load: %v", err)
|
||||
}
|
||||
if cfg.Reader.SheetMode != w.sheetMode {
|
||||
t.Errorf("sheet_mode = %q, want %q", cfg.Reader.SheetMode, w.sheetMode)
|
||||
}
|
||||
if cfg.Reader.StripBlankRows != w.stripBlank {
|
||||
t.Errorf("strip_blank_rows = %v, want %v", cfg.Reader.StripBlankRows, w.stripBlank)
|
||||
}
|
||||
if cfg.Validation.RequireNumericSbd != w.numericSbd {
|
||||
t.Errorf("require_numeric_sbd = %v, want %v", cfg.Validation.RequireNumericSbd, w.numericSbd)
|
||||
}
|
||||
if (cfg.Columns != nil) != w.hasColumns {
|
||||
t.Errorf("columns present = %v, want %v", cfg.Columns != nil, w.hasColumns)
|
||||
}
|
||||
got := ""
|
||||
if cfg.FormatDetection != nil {
|
||||
got = *cfg.FormatDetection
|
||||
}
|
||||
if got != w.formatDet {
|
||||
t.Errorf("format_detection = %q, want %q", got, w.formatDet)
|
||||
}
|
||||
if len(cfg.Header.Tokens) != w.tokenCount {
|
||||
t.Errorf("tokens = %d, want %d", len(cfg.Header.Tokens), w.tokenCount)
|
||||
}
|
||||
// Every dataset requires a non-empty name and SBD.
|
||||
if !cfg.Validation.RequireNonemptyName || !cfg.Validation.RequireNonemptySbd {
|
||||
t.Error("both non-empty validations should be true for every dataset")
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func replaceAll(s, old, new string) string {
|
||||
out := ""
|
||||
for {
|
||||
i := indexOf(s, old)
|
||||
if i < 0 {
|
||||
return out + s
|
||||
}
|
||||
out += s[:i] + new
|
||||
s = s[i+len(old):]
|
||||
}
|
||||
}
|
||||
|
||||
func indexOf(s, sub string) int {
|
||||
for i := 0; i+len(sub) <= len(s); i++ {
|
||||
if s[i:i+len(sub)] == sub {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
|
||||
func repoRoot(t *testing.T) string {
|
||||
t.Helper()
|
||||
dir, err := os.Getwd()
|
||||
if err != nil {
|
||||
t.Fatalf("getwd: %v", err)
|
||||
}
|
||||
for i := 0; i < 6; i++ {
|
||||
if _, err := os.Stat(filepath.Join(dir, "parser", "configs")); err == nil {
|
||||
return dir
|
||||
}
|
||||
dir = filepath.Dir(dir)
|
||||
}
|
||||
t.Fatal("could not locate repo root")
|
||||
return ""
|
||||
}
|
||||
@@ -0,0 +1,382 @@
|
||||
package ingest
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/config"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/transform"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/writer"
|
||||
)
|
||||
|
||||
// The 2016 dataset's 119 files were produced by inconsistent tooling and use
|
||||
// three different column layouts, chosen per sheet at runtime. This is a port of
|
||||
// parser/src/format_detect_2016.rs — institutional knowledge encoded as
|
||||
// literals, with no abstraction to derive it from, so everything here is copied
|
||||
// verbatim rather than rationalised.
|
||||
|
||||
// KnownHeaders are the upper-cased first-cell values that identify a header row
|
||||
// (format_detect_2016.rs:36-54, mirroring KNOWN_HEADERS in build-database.js).
|
||||
//
|
||||
// NOTE "SINH " carries a TRAILING SPACE, exactly as in the Rust source. Trimming
|
||||
// it would change which rows are recognised as headers.
|
||||
var KnownHeaders = []string{
|
||||
"SOBAODANH",
|
||||
"SBD",
|
||||
"HO_TEN",
|
||||
"HOTEN",
|
||||
"HỌ TÊN",
|
||||
"NGAY_SINH",
|
||||
"TEN_CUMTHI",
|
||||
"GIOI_TINH",
|
||||
"DIEM_THI",
|
||||
"STT",
|
||||
"TOAN",
|
||||
"VAN",
|
||||
"LY",
|
||||
"HOA",
|
||||
"SINH ",
|
||||
"SU",
|
||||
"DIA",
|
||||
}
|
||||
|
||||
func isKnownHeader(s string) bool {
|
||||
for _, h := range KnownHeaders {
|
||||
if h == s {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// IsHeaderRow2016 reports whether row[0] is a known header token
|
||||
// (format_detect_2016.rs:58-64).
|
||||
//
|
||||
// The guard here is len < 2, not the len < 3 used by the 2017 header check.
|
||||
func IsHeaderRow2016(row []reader.Cell) bool {
|
||||
if len(row) < 2 {
|
||||
return false
|
||||
}
|
||||
return isKnownHeader(strings.ToUpper(strings.TrimSpace(row[0].Str)))
|
||||
}
|
||||
|
||||
// FormatKind is one of the three 2016 layouts.
|
||||
type FormatKind int
|
||||
|
||||
const (
|
||||
// FormatSeparateScores has one column per subject rather than a free-text
|
||||
// DIEM_THI cell — the dhhanghai-style files.
|
||||
FormatSeparateScores FormatKind = iota
|
||||
// FormatMapped resolves column indices from the header by name.
|
||||
FormatMapped
|
||||
// FormatDefault is the positional 6-column layout used when no header is
|
||||
// recognised. It is FormatMapped with fixed indices, not a separate path.
|
||||
FormatDefault
|
||||
)
|
||||
|
||||
// Format is a detected layout plus, for the mapped case, its column indices.
|
||||
type Format struct {
|
||||
Kind FormatKind
|
||||
Sbd int
|
||||
HoTen int
|
||||
NgaySinh *int
|
||||
TenCumThi *int
|
||||
GioiTinh *int
|
||||
DiemThi int
|
||||
}
|
||||
|
||||
// defaultFormat is FormatMapped with the fixed positional indices
|
||||
// (format_detect_2016.rs:295-306).
|
||||
func defaultFormat() Format {
|
||||
two, three, four := 2, 3, 4
|
||||
return Format{
|
||||
Kind: FormatDefault, Sbd: 0, HoTen: 1,
|
||||
NgaySinh: &two, TenCumThi: &three, GioiTinh: &four, DiemThi: 5,
|
||||
}
|
||||
}
|
||||
|
||||
// DetectFormat inspects a header row and decides which layout applies
|
||||
// (format_detect_2016.rs:95-145).
|
||||
func DetectFormat(headerRow []reader.Cell) Format {
|
||||
cols := make([]string, len(headerRow))
|
||||
for i, c := range headerRow {
|
||||
cols[i] = strings.ToUpper(strings.TrimSpace(c.Str))
|
||||
}
|
||||
|
||||
// Format 1: SBD in col 0 AND TOAN in col 2.
|
||||
if len(cols) > 2 && cols[0] == "SBD" && cols[2] == "TOAN" {
|
||||
return Format{Kind: FormatSeparateScores}
|
||||
}
|
||||
|
||||
// Format 2: resolve indices by header name, order-independent.
|
||||
var sbdIdx, hoTenIdx, ngaySinhIdx, tenCumThiIdx, gioiTinhIdx, diemThiIdx *int
|
||||
for i := range cols {
|
||||
idx := i
|
||||
switch cols[i] {
|
||||
case "SOBAODANH", "SBD":
|
||||
sbdIdx = &idx
|
||||
case "HO_TEN", "HOTEN", "HỌ TÊN":
|
||||
hoTenIdx = &idx
|
||||
case "NGAY_SINH":
|
||||
ngaySinhIdx = &idx
|
||||
case "TEN_CUMTHI":
|
||||
tenCumThiIdx = &idx
|
||||
case "GIOI_TINH":
|
||||
gioiTinhIdx = &idx
|
||||
case "DIEM_THI":
|
||||
diemThiIdx = &idx
|
||||
}
|
||||
}
|
||||
|
||||
if sbdIdx != nil && diemThiIdx != nil {
|
||||
hoTen := 1 // fallback: col 1, present in all known files
|
||||
if hoTenIdx != nil {
|
||||
hoTen = *hoTenIdx
|
||||
}
|
||||
return Format{
|
||||
Kind: FormatMapped, Sbd: *sbdIdx, HoTen: hoTen,
|
||||
NgaySinh: ngaySinhIdx, TenCumThi: tenCumThiIdx, GioiTinh: gioiTinhIdx,
|
||||
DiemThi: *diemThiIdx,
|
||||
}
|
||||
}
|
||||
|
||||
// Unrecognised header, or none at all.
|
||||
return defaultFormat()
|
||||
}
|
||||
|
||||
// parseFloatCell parses a per-subject score cell.
|
||||
//
|
||||
// A parsed 0.0 becomes "no score" — this replicates JavaScript's
|
||||
// `parseFloat(row[N]) || null`, where 0 is falsy (format_detect_2016.rs:165).
|
||||
// It means a genuine zero is indistinguishable from a blank. Not obviously
|
||||
// correct, but it is the shipped behaviour and the published data depends on it.
|
||||
func parseFloatCell(row []reader.Cell, idx int) (float64, bool) {
|
||||
s := cellAt(row, idx)
|
||||
if s == "" {
|
||||
return 0, false
|
||||
}
|
||||
v, err := strconv.ParseFloat(s, 64)
|
||||
if err != nil || v == 0.0 {
|
||||
return 0, false
|
||||
}
|
||||
return v, true
|
||||
}
|
||||
|
||||
// processSeparateScoresRow handles the fixed 12-column layout
|
||||
// (format_detect_2016.rs:176-216):
|
||||
//
|
||||
// 0=SBD 1=HOTEN 2=TOAN 3=VAN 4=LY 5=HOA 6=SINH 7=SU 8=DIA
|
||||
// 9=NGOAINGUTN 10=NGOAINGUTL 11=NGOAINGU(total -> tieng_anh)
|
||||
//
|
||||
// tieng_phap / tieng_duc / tieng_nhat / tieng_trung are structurally unreachable
|
||||
// in this format, and ngay_sinh / ten_cum_thi / gioi_tinh are always nil.
|
||||
// Scores are read as floats directly; this layout has no free-text score cell,
|
||||
// so the subject regexes never run.
|
||||
func processSeparateScoresRow(row []reader.Cell) *transform.ParsedRow {
|
||||
sbd := cellAt(row, 0)
|
||||
hoTen := cellAt(row, 1)
|
||||
if sbd == "" || hoTen == "" {
|
||||
return nil
|
||||
}
|
||||
|
||||
scores := make(map[string]float64)
|
||||
for field, idx := range map[string]int{
|
||||
"toan": 2, "ngu_van": 3, "vat_ly": 4, "hoa_hoc": 5,
|
||||
"sinh_hoc": 6, "lich_su": 7, "dia_ly": 8,
|
||||
"tieng_anh": 11, // NGOAINGU total
|
||||
} {
|
||||
if v, ok := parseFloatCell(row, idx); ok {
|
||||
scores[field] = v
|
||||
}
|
||||
}
|
||||
|
||||
return &transform.ParsedRow{
|
||||
SoBaoDanh: sbd,
|
||||
HoTen: hoTen,
|
||||
HoTenAscii: transform.ToAscii(hoTen),
|
||||
Scores: scores,
|
||||
}
|
||||
}
|
||||
|
||||
// processMappedRow handles header-derived column indices
|
||||
// (format_detect_2016.rs:226-289).
|
||||
func processMappedRow(row []reader.Cell, f Format) *transform.ParsedRow {
|
||||
sbd := cellAt(row, f.Sbd)
|
||||
hoTen := cellAt(row, f.HoTen)
|
||||
if sbd == "" || hoTen == "" {
|
||||
return nil
|
||||
}
|
||||
|
||||
// Leaked-header guard: a row whose SBD or name cell is itself a header token
|
||||
// is a repeated header, not data (format_detect_2016.rs:244-250).
|
||||
if isKnownHeader(strings.ToUpper(sbd)) || isKnownHeader(strings.ToUpper(hoTen)) {
|
||||
return nil
|
||||
}
|
||||
|
||||
optional := func(idx *int) *string {
|
||||
if idx == nil {
|
||||
return nil
|
||||
}
|
||||
if s := cellAt(row, *idx); s != "" {
|
||||
return &s
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Gender is a two-value allowlist, not a general enum: anything other than
|
||||
// exactly "Nam" or "Nữ" becomes nil (format_detect_2016.rs:263-271).
|
||||
var gioiTinh *string
|
||||
if f.GioiTinh != nil {
|
||||
if s := cellAt(row, *f.GioiTinh); s == "Nam" || s == "Nữ" {
|
||||
gioiTinh = &s
|
||||
}
|
||||
}
|
||||
|
||||
// Read untrimmed, matching format_detect_2016.rs:273-276.
|
||||
diemThi := ""
|
||||
if f.DiemThi >= 0 && f.DiemThi < len(row) {
|
||||
diemThi = row[f.DiemThi].Str
|
||||
}
|
||||
|
||||
return &transform.ParsedRow{
|
||||
SoBaoDanh: sbd,
|
||||
HoTen: hoTen,
|
||||
HoTenAscii: transform.ToAscii(hoTen),
|
||||
NgaySinh: optional(f.NgaySinh),
|
||||
TenCumThi: optional(f.TenCumThi),
|
||||
GioiTinh: gioiTinh,
|
||||
Scores: transform.ParseScores(diemThi),
|
||||
}
|
||||
}
|
||||
|
||||
// ProcessRow2016 dispatches a data row through the detected layout. A nil return
|
||||
// means the row is empty or invalid and should be skipped.
|
||||
func ProcessRow2016(row []reader.Cell, f Format) *transform.ParsedRow {
|
||||
if f.Kind == FormatSeparateScores {
|
||||
return processSeparateScoresRow(row)
|
||||
}
|
||||
// FormatMapped and FormatDefault share one implementation; Default is just a
|
||||
// fixed index tuple.
|
||||
return processMappedRow(row, f)
|
||||
}
|
||||
|
||||
// Detect2016 ingests the 2016 dataset — the port of run_build_2016 and
|
||||
// process_file_2016 (main.rs:211-377).
|
||||
func Detect2016(cfg *config.DatasetConfig, inputDir, outputPath string) error {
|
||||
files, err := InputFiles(inputDir)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
label := DatasetLabel(inputDir)
|
||||
fmt.Printf("[build:2016] %s/ → %s (%d files)\n", label, outputPath, len(files))
|
||||
|
||||
db, err := writer.OpenDB(outputPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer db.Close()
|
||||
|
||||
tx, err := db.Begin()
|
||||
if err != nil {
|
||||
return fmt.Errorf("begin: %w", err)
|
||||
}
|
||||
ins, err := writer.Prepare(tx)
|
||||
if err != nil {
|
||||
tx.Rollback()
|
||||
return err
|
||||
}
|
||||
|
||||
var st writer.Stats // Skipped stays 0 on this path, matching main.rs:251
|
||||
for _, file := range files {
|
||||
base := filepath.Base(file)
|
||||
fileRows, err := processFile2016(file, cfg, ins, base, &st)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, " [error] %s: %v\n", base, err)
|
||||
st.Errors++
|
||||
continue
|
||||
}
|
||||
fmt.Printf(" %s: %d rows\n", base, fileRows)
|
||||
}
|
||||
|
||||
if err := ins.Close(); err != nil {
|
||||
tx.Rollback()
|
||||
return err
|
||||
}
|
||||
if err := tx.Commit(); err != nil {
|
||||
return fmt.Errorf("commit: %w", err)
|
||||
}
|
||||
return writer.Finish(db, outputPath, st, label, false)
|
||||
}
|
||||
|
||||
// processFile2016 detects the layout per SHEET and processes that sheet's rows.
|
||||
//
|
||||
// Detection is per sheet, not per file (main.rs:344-349): within one workbook,
|
||||
// sheet 2 may legitimately detect differently from sheet 1, which matters for
|
||||
// the province files that overflow past Excel's row cap.
|
||||
func processFile2016(path string, cfg *config.DatasetConfig, ins *writer.Inserter, base string, st *writer.Stats) (uint64, error) {
|
||||
wb, err := reader.Open(path)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
defer wb.Close()
|
||||
|
||||
sheets := wb.Sheets()
|
||||
if len(sheets) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
if cfg.Reader.SheetMode == config.SheetModeFirst {
|
||||
sheets = sheets[:1]
|
||||
}
|
||||
|
||||
var fileRows uint64
|
||||
for _, sh := range sheets {
|
||||
var rows [][]reader.Cell
|
||||
if err := wb.EachRow(sh.Index, func(_ reader.Sheet, _ int, row []reader.Cell) error {
|
||||
rows = append(rows, row)
|
||||
return nil
|
||||
}); err != nil {
|
||||
return fileRows, err
|
||||
}
|
||||
if len(rows) == 0 {
|
||||
continue
|
||||
}
|
||||
|
||||
f := defaultFormat()
|
||||
startIdx := 0
|
||||
if IsHeaderRow2016(rows[0]) {
|
||||
f = DetectFormat(rows[0])
|
||||
startIdx = 1
|
||||
}
|
||||
|
||||
for _, row := range rows[startIdx:] {
|
||||
// Rows shorter than 2 cells are dropped BEFORE the counter
|
||||
// (main.rs:351-353), so they never appear in the source total.
|
||||
if len(row) < 2 {
|
||||
continue
|
||||
}
|
||||
st.SourceRows++
|
||||
|
||||
parsed := ProcessRow2016(row, f)
|
||||
if parsed == nil {
|
||||
// Empty or invalid. Note this is NOT counted as skipped — the
|
||||
// 2016 path leaves that counter at zero (main.rs:251), so the
|
||||
// stats block reports insertable == source rows and the Audit
|
||||
// line absorbs the difference.
|
||||
continue
|
||||
}
|
||||
if err := ins.Insert(parsed); err != nil {
|
||||
st.Errors++
|
||||
if st.Errors <= 5 {
|
||||
fmt.Fprintf(os.Stderr, " [warn] %s: %v\n", base, err)
|
||||
}
|
||||
continue
|
||||
}
|
||||
fileRows++
|
||||
}
|
||||
}
|
||||
return fileRows, nil
|
||||
}
|
||||
@@ -0,0 +1,252 @@
|
||||
package ingest
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
)
|
||||
|
||||
// Ports the 11 tests in parser/src/format_detect_2016.rs, plus a guard per quirk.
|
||||
//
|
||||
// Phase 1 settled the Data -> string translation these fixtures need, so no
|
||||
// guessing: Data::String(s) is s verbatim, Data::Float(8.0) renders "8" (Rust's
|
||||
// f64 Display drops the .0, matching Go's FormatFloat(v,'f',-1,64)),
|
||||
// Data::Float(8.5) renders "8.5", and Data::Empty is "" with IsEmpty true.
|
||||
|
||||
// --- header detection ---
|
||||
|
||||
func TestIsHeaderRow2016(t *testing.T) {
|
||||
if !IsHeaderRow2016(cells("SBD", "HOTEN", "TOAN")) {
|
||||
t.Error("SBD header not detected")
|
||||
}
|
||||
if !IsHeaderRow2016(cells("sobaodanh", "x")) {
|
||||
t.Error("lowercase SOBAODANH not detected")
|
||||
}
|
||||
if IsHeaderRow2016(cells("Nguyen Van A", "01/01/2000")) {
|
||||
t.Error("data row wrongly detected as header")
|
||||
}
|
||||
// The 2016 guard is len < 2, unlike the 2017 check's len < 3.
|
||||
if IsHeaderRow2016(cells("SBD")) {
|
||||
t.Error("single-cell row must not be a header")
|
||||
}
|
||||
if !IsHeaderRow2016(cells("SBD", "x")) {
|
||||
t.Error("two-cell row with a header token must be a header")
|
||||
}
|
||||
}
|
||||
|
||||
// TestKnownHeadersHasTrailingSpaceToken pins the "SINH " literal. Trimming it
|
||||
// would silently change which rows count as headers.
|
||||
func TestKnownHeadersHasTrailingSpaceToken(t *testing.T) {
|
||||
var found bool
|
||||
for _, h := range KnownHeaders {
|
||||
if h == "SINH " {
|
||||
found = true
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
t.Error(`KnownHeaders must contain "SINH " WITH its trailing space`)
|
||||
}
|
||||
if len(KnownHeaders) != 17 {
|
||||
t.Errorf("KnownHeaders has %d tokens, want 17", len(KnownHeaders))
|
||||
}
|
||||
}
|
||||
|
||||
// --- format detection ---
|
||||
|
||||
func TestDetectFormatSeparateScores(t *testing.T) {
|
||||
f := DetectFormat(cells("SBD", "HOTEN", "TOAN", "VAN"))
|
||||
if f.Kind != FormatSeparateScores {
|
||||
t.Errorf("Kind = %v, want FormatSeparateScores", f.Kind)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDetectFormatMapped(t *testing.T) {
|
||||
f := DetectFormat(cells("STT", "SOBAODANH", "HO_TEN", "NGAY_SINH", "TEN_CUMTHI", "GIOI_TINH", "DIEM_THI"))
|
||||
if f.Kind != FormatMapped {
|
||||
t.Fatalf("Kind = %v, want FormatMapped", f.Kind)
|
||||
}
|
||||
if f.Sbd != 1 || f.HoTen != 2 || f.DiemThi != 6 {
|
||||
t.Errorf("indices sbd=%d ho_ten=%d diem_thi=%d", f.Sbd, f.HoTen, f.DiemThi)
|
||||
}
|
||||
if f.NgaySinh == nil || *f.NgaySinh != 3 || f.TenCumThi == nil || *f.TenCumThi != 4 || f.GioiTinh == nil || *f.GioiTinh != 5 {
|
||||
t.Error("optional indices not resolved")
|
||||
}
|
||||
}
|
||||
|
||||
// TestDetectFormatMappedIsOrderIndependent: indices are resolved by name.
|
||||
func TestDetectFormatMappedIsOrderIndependent(t *testing.T) {
|
||||
f := DetectFormat(cells("DIEM_THI", "HO_TEN", "SBD"))
|
||||
if f.Kind != FormatMapped || f.Sbd != 2 || f.DiemThi != 0 || f.HoTen != 1 {
|
||||
t.Errorf("got %+v", f)
|
||||
}
|
||||
}
|
||||
|
||||
// TestDetectFormatMappedHoTenFallback ports the col-1 fallback
|
||||
// (format_detect_2016.rs:132).
|
||||
func TestDetectFormatMappedHoTenFallback(t *testing.T) {
|
||||
f := DetectFormat(cells("SBD", "SOMETHING", "DIEM_THI"))
|
||||
if f.Kind != FormatMapped {
|
||||
t.Fatalf("Kind = %v, want FormatMapped", f.Kind)
|
||||
}
|
||||
if f.HoTen != 1 {
|
||||
t.Errorf("ho_ten = %d, want fallback 1", f.HoTen)
|
||||
}
|
||||
}
|
||||
|
||||
// TestDetectFormatDefault: a header lacking SBD or DIEM_THI falls back to the
|
||||
// positional layout.
|
||||
func TestDetectFormatDefault(t *testing.T) {
|
||||
f := DetectFormat(cells("A", "B", "C"))
|
||||
if f.Kind != FormatDefault {
|
||||
t.Errorf("Kind = %v, want FormatDefault", f.Kind)
|
||||
}
|
||||
if f.Sbd != 0 || f.HoTen != 1 || f.DiemThi != 5 {
|
||||
t.Errorf("default indices wrong: %+v", f)
|
||||
}
|
||||
}
|
||||
|
||||
// --- row processing ---
|
||||
|
||||
func TestProcessSeparateScoresRow(t *testing.T) {
|
||||
// 0=SBD 1=HOTEN 2=TOAN 3=VAN 4=LY 5=HOA 6=SINH 7=SU 8=DIA 9,10=NN 11=NN total
|
||||
row := cells("1000", "Nguyễn Văn Đức", "8", "7.5", "", "6", "", "5", "", "", "", "9.25")
|
||||
got := ProcessRow2016(row, Format{Kind: FormatSeparateScores})
|
||||
if got == nil {
|
||||
t.Fatal("row rejected")
|
||||
}
|
||||
if got.SoBaoDanh != "1000" || got.HoTenAscii != "nguyen van duc" {
|
||||
t.Errorf("sbd=%q ascii=%q", got.SoBaoDanh, got.HoTenAscii)
|
||||
}
|
||||
want := map[string]float64{"toan": 8, "ngu_van": 7.5, "hoa_hoc": 6, "lich_su": 5, "tieng_anh": 9.25}
|
||||
if len(got.Scores) != len(want) {
|
||||
t.Errorf("scores = %v, want %v", got.Scores, want)
|
||||
}
|
||||
for k, v := range want {
|
||||
if got.Scores[k] != v {
|
||||
t.Errorf("%s = %v, want %v", k, got.Scores[k], v)
|
||||
}
|
||||
}
|
||||
// These columns are structurally unreachable in this layout.
|
||||
for _, absent := range []string{"tieng_phap", "tieng_duc", "tieng_nhat", "tieng_trung", "khtn", "khxh", "gdcd"} {
|
||||
if _, ok := got.Scores[absent]; ok {
|
||||
t.Errorf("%s must be unreachable in separate-scores", absent)
|
||||
}
|
||||
}
|
||||
if got.NgaySinh != nil || got.TenCumThi != nil || got.GioiTinh != nil {
|
||||
t.Error("ngay_sinh/ten_cum_thi/gioi_tinh are always nil in separate-scores")
|
||||
}
|
||||
}
|
||||
|
||||
// TestZeroScoreBecomesNull pins the JS falsy quirk: parseFloat(x) || null means
|
||||
// a literal 0 is indistinguishable from "no score" (format_detect_2016.rs:165).
|
||||
func TestZeroScoreBecomesNull(t *testing.T) {
|
||||
row := cells("1000", "A", "0", "0.0", "1", "", "", "", "", "", "", "")
|
||||
got := ProcessRow2016(row, Format{Kind: FormatSeparateScores})
|
||||
if got == nil {
|
||||
t.Fatal("row rejected")
|
||||
}
|
||||
if _, ok := got.Scores["toan"]; ok {
|
||||
t.Error(`a "0" score must become NULL, not 0`)
|
||||
}
|
||||
if _, ok := got.Scores["ngu_van"]; ok {
|
||||
t.Error(`a "0.0" score must become NULL, not 0`)
|
||||
}
|
||||
if got.Scores["vat_ly"] != 1 {
|
||||
t.Error("a non-zero score must survive")
|
||||
}
|
||||
}
|
||||
|
||||
// TestGenderAllowlist ports format_detect_2016.rs:263-271 — exactly two values.
|
||||
func TestGenderAllowlist(t *testing.T) {
|
||||
f := defaultFormat()
|
||||
for in, want := range map[string]string{"Nam": "Nam", "Nữ": "Nữ"} {
|
||||
row := cells("1", "A", "", "", in, "")
|
||||
got := ProcessRow2016(row, f)
|
||||
if got == nil || got.GioiTinh == nil || *got.GioiTinh != want {
|
||||
t.Errorf("gioi_tinh for %q not preserved", in)
|
||||
}
|
||||
}
|
||||
for _, in := range []string{"nam", "NAM", "Unknown", "M", "F", "nữ", ""} {
|
||||
row := cells("1", "A", "", "", in, "")
|
||||
got := ProcessRow2016(row, f)
|
||||
if got == nil {
|
||||
t.Fatalf("row rejected for %q", in)
|
||||
}
|
||||
if got.GioiTinh != nil {
|
||||
t.Errorf("gioi_tinh for %q = %q, want nil", in, *got.GioiTinh)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestLeakedHeaderRowSkipped ports format_detect_2016.rs:244-250.
|
||||
func TestLeakedHeaderRowSkipped(t *testing.T) {
|
||||
f := defaultFormat()
|
||||
if ProcessRow2016(cells("SBD", "HO_TEN", "", "", "", ""), f) != nil {
|
||||
t.Error("a repeated header row must be skipped")
|
||||
}
|
||||
if ProcessRow2016(cells("123", "SOBAODANH", "", "", "", ""), f) != nil {
|
||||
t.Error("a row whose name cell is a header token must be skipped")
|
||||
}
|
||||
if ProcessRow2016(cells("123", "Nguyen Van A", "", "", "", ""), f) == nil {
|
||||
t.Error("a genuine data row must not be skipped")
|
||||
}
|
||||
}
|
||||
|
||||
// TestMappedRowEmptyFieldsRejected: either identity field empty rejects the row.
|
||||
func TestMappedRowEmptyFieldsRejected(t *testing.T) {
|
||||
f := defaultFormat()
|
||||
if ProcessRow2016(cells("", "A", "", "", "", ""), f) != nil {
|
||||
t.Error("empty sbd must reject")
|
||||
}
|
||||
if ProcessRow2016(cells("1", "", "", "", "", ""), f) != nil {
|
||||
t.Error("empty ho_ten must reject")
|
||||
}
|
||||
}
|
||||
|
||||
// TestMappedRowPopulates2016OnlyColumns: this is the only dataset that fills
|
||||
// ten_cum_thi and gioi_tinh.
|
||||
func TestMappedRowPopulates2016OnlyColumns(t *testing.T) {
|
||||
row := cells("123", "Lê Văn Long", "01/01/1998", "Cụm thi số 1", "Nam", "Toán: 7.5 Tiếng Đức: 6")
|
||||
got := ProcessRow2016(row, defaultFormat())
|
||||
if got == nil {
|
||||
t.Fatal("row rejected")
|
||||
}
|
||||
if got.NgaySinh == nil || *got.NgaySinh != "01/01/1998" {
|
||||
t.Errorf("ngay_sinh = %v", got.NgaySinh)
|
||||
}
|
||||
if got.TenCumThi == nil || *got.TenCumThi != "Cụm thi số 1" {
|
||||
t.Errorf("ten_cum_thi = %v", got.TenCumThi)
|
||||
}
|
||||
if got.Scores["toan"] != 7.5 || got.Scores["tieng_duc"] != 6 {
|
||||
t.Errorf("scores = %v", got.Scores)
|
||||
}
|
||||
}
|
||||
|
||||
// TestDefaultIsMappedWithFixedIndices: FormatDefault must not be a separate code
|
||||
// path (format_detect_2016.rs:295-306).
|
||||
func TestDefaultIsMappedWithFixedIndices(t *testing.T) {
|
||||
row := cells("123", "A", "01/01/2000", "Cluster", "Nữ", "Toán: 5")
|
||||
viaDefault := ProcessRow2016(row, defaultFormat())
|
||||
two, three, four := 2, 3, 4
|
||||
viaMapped := ProcessRow2016(row, Format{
|
||||
Kind: FormatMapped, Sbd: 0, HoTen: 1,
|
||||
NgaySinh: &two, TenCumThi: &three, GioiTinh: &four, DiemThi: 5,
|
||||
})
|
||||
if viaDefault == nil || viaMapped == nil {
|
||||
t.Fatal("row rejected")
|
||||
}
|
||||
if viaDefault.SoBaoDanh != viaMapped.SoBaoDanh ||
|
||||
*viaDefault.TenCumThi != *viaMapped.TenCumThi ||
|
||||
*viaDefault.GioiTinh != *viaMapped.GioiTinh ||
|
||||
viaDefault.Scores["toan"] != viaMapped.Scores["toan"] {
|
||||
t.Error("FormatDefault must behave identically to the equivalent FormatMapped")
|
||||
}
|
||||
}
|
||||
|
||||
// TestShortRowGuardIsSeparateFromValidation documents that rows under 2 cells
|
||||
// are dropped by the loop before the counter (main.rs:351-353), not here.
|
||||
func TestShortRowGuardIsSeparateFromValidation(t *testing.T) {
|
||||
if got := ProcessRow2016([]reader.Cell{{Str: "123"}}, defaultFormat()); got != nil {
|
||||
t.Error("a 1-cell row has no name and must be rejected")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,243 @@
|
||||
// Package ingest owns all dataset policy: which sheets to read, which rows are
|
||||
// headers, which are blank, and how rows are counted.
|
||||
//
|
||||
// The reader deliberately has none of this — it reports every sheet and every
|
||||
// row exactly as calamine would, which is what made its fidelity independently
|
||||
// testable. This package ports the build loop in parser/src/main.rs plus the
|
||||
// header/blank helpers in parser/src/reader.rs.
|
||||
package ingest
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/config"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/transform"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/writer"
|
||||
)
|
||||
|
||||
// IsHeaderRow reports whether row is a header, by matching its uppercased first
|
||||
// cell against the configured tokens (reader.rs:28-34).
|
||||
//
|
||||
// Rows shorter than 3 cells are never headers (reader.rs:29) — a 1- or 2-cell
|
||||
// row is a stray fragment, not a real header.
|
||||
func IsHeaderRow(row []reader.Cell, tokens []string) bool {
|
||||
if len(row) < 3 {
|
||||
return false
|
||||
}
|
||||
first := strings.ToUpper(strings.TrimSpace(row[0].Str))
|
||||
for _, t := range tokens {
|
||||
if strings.ToUpper(t) == first {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// IsAllBlank reports whether every cell is empty or whitespace-only
|
||||
// (reader.rs:40-43).
|
||||
//
|
||||
// Compares on Str only. Cell.IsEmpty is diagnostic: calamine distinguishes
|
||||
// Data::Empty from an empty string cell, but both render "" and both count as
|
||||
// blank here, so branching on the flag would invent a distinction the Rust
|
||||
// original never acts on.
|
||||
func IsAllBlank(row []reader.Cell) bool {
|
||||
for _, c := range row {
|
||||
if strings.TrimSpace(c.Str) != "" {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// InputFiles lists a dataset directory's spreadsheets, sorted.
|
||||
//
|
||||
// The sort is load-bearing, not cosmetic: INSERT OR REPLACE is last-wins, so
|
||||
// file order decides which row survives a duplicate SBD. Rust collects read_dir
|
||||
// then calls files.sort() (main.rs:82-97) — a bytewise sort on the full path.
|
||||
func InputFiles(dir string) ([]string, error) {
|
||||
entries, err := os.ReadDir(dir)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("cannot read input dir %s: %w", dir, err)
|
||||
}
|
||||
var out []string
|
||||
for _, e := range entries {
|
||||
if e.IsDir() {
|
||||
continue
|
||||
}
|
||||
switch strings.ToLower(filepath.Ext(e.Name())) {
|
||||
case ".xls", ".xlsx":
|
||||
out = append(out, filepath.Join(dir, e.Name()))
|
||||
}
|
||||
}
|
||||
sort.Strings(out)
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// DatasetLabel derives the stats-wording key from the input directory basename,
|
||||
// matching main.rs:98-101.
|
||||
func DatasetLabel(inputDir string) string {
|
||||
base := filepath.Base(strings.TrimRight(inputDir, string(filepath.Separator)))
|
||||
if base == "" || base == "." || base == string(filepath.Separator) {
|
||||
return "data"
|
||||
}
|
||||
return base
|
||||
}
|
||||
|
||||
// RowFn consumes one data row of one sheet, after header skipping.
|
||||
type RowFn func(sheetIdx int, row []reader.Cell)
|
||||
|
||||
// ProcessFile applies sheet selection and per-sheet header skipping, invoking fn
|
||||
// for every remaining row — the port of reader.rs:54-105.
|
||||
//
|
||||
// The header check is per SHEET, not per file: first_row resets inside the sheet
|
||||
// loop (reader.rs:91), so a workbook whose second sheet repeats the header has
|
||||
// it skipped there too.
|
||||
func ProcessFile(path string, cfg *config.DatasetConfig, fn RowFn) error {
|
||||
wb, err := reader.Open(path)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer wb.Close()
|
||||
|
||||
sheets := wb.Sheets()
|
||||
if len(sheets) == 0 {
|
||||
return fmt.Errorf("no sheets in %s", path)
|
||||
}
|
||||
if cfg.Reader.SheetMode == config.SheetModeFirst {
|
||||
sheets = sheets[:1]
|
||||
}
|
||||
|
||||
for _, sh := range sheets {
|
||||
firstRow := true
|
||||
err := wb.EachRow(sh.Index, func(s reader.Sheet, _ int, row []reader.Cell) error {
|
||||
if firstRow {
|
||||
firstRow = false
|
||||
if IsHeaderRow(row, cfg.Header.Tokens) {
|
||||
return nil
|
||||
}
|
||||
}
|
||||
fn(s.Index, row)
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Standard runs the fixed-column path for the 2017-family datasets — the port of
|
||||
// run_build_standard (main.rs:74-199).
|
||||
func Standard(cfg *config.DatasetConfig, inputDir, outputPath string) error {
|
||||
files, err := InputFiles(inputDir)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
label := DatasetLabel(inputDir)
|
||||
fmt.Printf("[build] %s/ → %s (%d files)\n", label, outputPath, len(files))
|
||||
|
||||
db, err := writer.OpenDB(outputPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer db.Close()
|
||||
|
||||
isOld2 := strings.Contains(label, "old2")
|
||||
stripBlank := cfg.Reader.StripBlankRows
|
||||
|
||||
// One transaction spans the whole dataset directory (main.rs:120,184).
|
||||
tx, err := db.Begin()
|
||||
if err != nil {
|
||||
return fmt.Errorf("begin: %w", err)
|
||||
}
|
||||
ins, err := writer.Prepare(tx)
|
||||
if err != nil {
|
||||
tx.Rollback()
|
||||
return err
|
||||
}
|
||||
|
||||
var st writer.Stats
|
||||
for _, file := range files {
|
||||
base := filepath.Base(file)
|
||||
var fileRows, fileSkipped, fileErrors uint64
|
||||
|
||||
procErr := ProcessFile(file, cfg, func(_ int, row []reader.Cell) {
|
||||
allBlank := IsAllBlank(row)
|
||||
// 2017-old2: blank rows drop out BEFORE the source-row counter
|
||||
// (main.rs:134-137, ahead of the increment at :140).
|
||||
if stripBlank && allBlank {
|
||||
return
|
||||
}
|
||||
st.SourceRows++
|
||||
|
||||
hoTen, soBaoDanh := "", ""
|
||||
if cols := cfg.Columns; cols != nil {
|
||||
hoTen = cellAt(row, cols.HoTen)
|
||||
soBaoDanh = cellAt(row, cols.SoBaoDanh)
|
||||
}
|
||||
|
||||
switch transform.ValidateRow(hoTen, soBaoDanh, &cfg.Validation, stripBlank, allBlank) {
|
||||
case transform.SkipBlankRow:
|
||||
// Falls through to transform and insert, matching main.rs:150.
|
||||
// Unreachable here: BlankRow requires stripBlank && allBlank,
|
||||
// which returned above. Kept so the two call sites with opposite
|
||||
// outcomes stay visibly distinct.
|
||||
case transform.SkipNone:
|
||||
// proceed
|
||||
default:
|
||||
fileSkipped++
|
||||
return
|
||||
}
|
||||
|
||||
parsed, err := transform.TransformRow(row, cfg)
|
||||
if err != nil {
|
||||
fileErrors++
|
||||
return
|
||||
}
|
||||
if err := ins.Insert(parsed); err != nil {
|
||||
fileErrors++
|
||||
// Only the first five insert warnings print (main.rs:164).
|
||||
if st.Errors+fileErrors <= 5 {
|
||||
fmt.Fprintf(os.Stderr, " [warn] %s: %v\n", base, err)
|
||||
}
|
||||
return
|
||||
}
|
||||
fileRows++
|
||||
})
|
||||
if procErr != nil {
|
||||
// A file that cannot be read is logged and counted, never fatal
|
||||
// (main.rs:171-177) — one corrupt file must not abandon the batch.
|
||||
fmt.Fprintf(os.Stderr, " [error] %s: %v\n", base, procErr)
|
||||
fileErrors++
|
||||
}
|
||||
|
||||
st.Skipped += fileSkipped
|
||||
st.Errors += fileErrors
|
||||
fmt.Printf(" %s: %d rows\n", base, fileRows)
|
||||
}
|
||||
|
||||
if err := ins.Close(); err != nil {
|
||||
tx.Rollback()
|
||||
return err
|
||||
}
|
||||
if err := tx.Commit(); err != nil {
|
||||
return fmt.Errorf("commit: %w", err)
|
||||
}
|
||||
|
||||
// VACUUM only after COMMIT — SQLite refuses it inside a transaction.
|
||||
return writer.Finish(db, outputPath, st, label, isOld2)
|
||||
}
|
||||
|
||||
// cellAt returns the trimmed cell at idx, or "" when out of range — the
|
||||
// unwrap_or_default() behaviour of transform.rs:163-167.
|
||||
func cellAt(row []reader.Cell, idx int) string {
|
||||
if idx < 0 || idx >= len(row) {
|
||||
return ""
|
||||
}
|
||||
return strings.TrimSpace(row[idx].Str)
|
||||
}
|
||||
@@ -0,0 +1,131 @@
|
||||
package ingest
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
)
|
||||
|
||||
// Ports the 7 tests in parser/src/reader.rs:116-197. They were listed under
|
||||
// Phase 1 originally, which was wrong: they exercise header and blank-row
|
||||
// policy, which lives here rather than in the reader package.
|
||||
|
||||
func cells(vals ...string) []reader.Cell {
|
||||
out := make([]reader.Cell, len(vals))
|
||||
for i, v := range vals {
|
||||
out[i] = reader.Cell{Str: v, IsEmpty: v == ""}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
var stdTokens = []string{"HO_TEN", "HỌ TÊN", "STT"}
|
||||
|
||||
// header_detects_ho_ten (reader.rs:128)
|
||||
func TestHeaderDetectsHoTen(t *testing.T) {
|
||||
if !IsHeaderRow(cells("HO_TEN", "NGAY_SINH", "SBD"), stdTokens) {
|
||||
t.Error("HO_TEN header not detected")
|
||||
}
|
||||
}
|
||||
|
||||
// header_detects_stt (reader.rs:139)
|
||||
func TestHeaderDetectsStt(t *testing.T) {
|
||||
if !IsHeaderRow(cells("STT", "B", "C"), stdTokens) {
|
||||
t.Error("STT header not detected")
|
||||
}
|
||||
}
|
||||
|
||||
// header_detects_ho_ten_unicode (reader.rs:150)
|
||||
func TestHeaderDetectsHoTenUnicode(t *testing.T) {
|
||||
if !IsHeaderRow(cells("HỌ TÊN", "B", "C"), stdTokens) {
|
||||
t.Error("HỌ TÊN header not detected")
|
||||
}
|
||||
}
|
||||
|
||||
// header_rejects_data_row (reader.rs:161)
|
||||
func TestHeaderRejectsDataRow(t *testing.T) {
|
||||
if IsHeaderRow(cells("Nguyen Van A", "01/01/2000", "12345678"), stdTokens) {
|
||||
t.Error("data row wrongly detected as header")
|
||||
}
|
||||
}
|
||||
|
||||
// header_rejects_short_row (reader.rs:172) — the <3 cell guard.
|
||||
func TestHeaderRejectsShortRow(t *testing.T) {
|
||||
if IsHeaderRow(cells("HO_TEN", ""), stdTokens) {
|
||||
t.Error("a 2-cell row must never be a header, even with a matching token")
|
||||
}
|
||||
}
|
||||
|
||||
// header_case_insensitive (reader.rs:179)
|
||||
func TestHeaderCaseInsensitive(t *testing.T) {
|
||||
if !IsHeaderRow(cells("ho_ten", "B", "C"), stdTokens) {
|
||||
t.Error("lowercase header not detected")
|
||||
}
|
||||
}
|
||||
|
||||
// blank_row_detection (reader.rs:190) — note the third cell is an empty *string*
|
||||
// cell, not Data::Empty, and must still count as blank.
|
||||
func TestBlankRowDetection(t *testing.T) {
|
||||
blank := []reader.Cell{{IsEmpty: true}, {IsEmpty: true}, {Str: ""}}
|
||||
if !IsAllBlank(blank) {
|
||||
t.Error("row of empty cells should be blank")
|
||||
}
|
||||
if IsAllBlank(cells("Nguyen", "", "")) {
|
||||
t.Error("row with content should not be blank")
|
||||
}
|
||||
}
|
||||
|
||||
// TestIsAllBlankIgnoresIsEmptyFlag pins the rule that blankness is decided by
|
||||
// the rendered string, never by Cell.IsEmpty — calamine emits empty-but-not-Empty
|
||||
// cells, and treating the flag as authoritative would diverge.
|
||||
func TestIsAllBlankIgnoresIsEmptyFlag(t *testing.T) {
|
||||
// Content present but IsEmpty wrongly set: still not blank.
|
||||
if IsAllBlank([]reader.Cell{{Str: "x", IsEmpty: true}}) {
|
||||
t.Error("a cell with content must not be blank regardless of IsEmpty")
|
||||
}
|
||||
// Whitespace only: blank, matching Rust's trim().is_empty().
|
||||
if !IsAllBlank([]reader.Cell{{Str: " "}, {Str: "\t"}}) {
|
||||
t.Error("whitespace-only cells should be blank")
|
||||
}
|
||||
}
|
||||
|
||||
// TestDatasetLabel covers the basename derivation that drives the stats wording.
|
||||
func TestDatasetLabel(t *testing.T) {
|
||||
for in, want := range map[string]string{
|
||||
"data/2017": "2017",
|
||||
"data/2017-old2": "2017-old2",
|
||||
"data/2017-old2/": "2017-old2",
|
||||
"/abs/path/2016": "2016",
|
||||
} {
|
||||
if got := DatasetLabel(in); got != want {
|
||||
t.Errorf("DatasetLabel(%q) = %q, want %q", in, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestInputFilesSortedAndFiltered: the sort decides which duplicate SBD survives
|
||||
// INSERT OR REPLACE, so it is behaviour, not presentation.
|
||||
func TestInputFilesSortedAndFiltered(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
for _, name := range []string{"b.xlsx", "a.xls", "c.XLSX", "notes.txt", "d.csv"} {
|
||||
if err := writeEmpty(filepath.Join(dir, name)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
got, err := InputFiles(dir)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
want := []string{"a.xls", "b.xlsx", "c.XLSX"}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("got %d files %v, want %d", len(got), got, len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if filepath.Base(got[i]) != want[i] {
|
||||
t.Errorf("file %d = %q, want %q", i, filepath.Base(got[i]), want[i])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func writeEmpty(path string) error { return os.WriteFile(path, nil, 0o644) }
|
||||
@@ -0,0 +1,128 @@
|
||||
package reader_test
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
)
|
||||
|
||||
// TestReaderFidelity asserts the Go reader reproduces calamine byte-for-byte on
|
||||
// every real input file.
|
||||
//
|
||||
// The oracle is a committed SHA-256 per file over a canonical cell dump. The
|
||||
// dumps themselves are real student names and birthdates, so only the hashes are
|
||||
// committed — regenerate the dumps from the Rust side on demand
|
||||
// (parser/examples/dump_cells.rs), which is possible because parser/ still
|
||||
// builds.
|
||||
//
|
||||
// The canonical form carries geometry and rendered cell values. The calamine
|
||||
// Data variant is excluded on purpose: Data::Empty and Data::String("") both
|
||||
// render "" and both count as blank in is_all_blank and transform, so the
|
||||
// distinction cannot reach the database.
|
||||
// Runs by default so CI and `go test ./...` keep the full guarantee; skipped
|
||||
// under -short, which is how to iterate without paying ~77s to re-read 418 MB.
|
||||
func TestReaderFidelity(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("-short: skipping the 299-file corpus sweep")
|
||||
}
|
||||
root := repoRoot(t)
|
||||
manifest := filepath.Join(root, "go-parser", "testdata", "reader-fidelity-hashes.tsv")
|
||||
|
||||
f, err := os.Open(manifest)
|
||||
if err != nil {
|
||||
t.Fatalf("open manifest: %v", err)
|
||||
}
|
||||
defer f.Close()
|
||||
|
||||
var checked int
|
||||
sc := bufio.NewScanner(f)
|
||||
for sc.Scan() {
|
||||
line := sc.Text()
|
||||
if line == "" || strings.HasPrefix(line, "#") {
|
||||
continue
|
||||
}
|
||||
rel, want, ok := strings.Cut(line, "\t")
|
||||
if !ok {
|
||||
t.Fatalf("malformed manifest line: %q", line)
|
||||
}
|
||||
path := filepath.Join(root, rel)
|
||||
if _, err := os.Stat(path); err != nil {
|
||||
t.Skipf("input data not present (%s); skipping fidelity suite", rel)
|
||||
}
|
||||
|
||||
checked++
|
||||
t.Run(rel, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
got, err := canonicalHash(path)
|
||||
if err != nil {
|
||||
t.Fatalf("hash %s: %v", rel, err)
|
||||
}
|
||||
if got != want {
|
||||
t.Errorf("cell dump diverges from calamine\n want %s\n got %s", want, got)
|
||||
}
|
||||
})
|
||||
}
|
||||
if err := sc.Err(); err != nil {
|
||||
t.Fatalf("read manifest: %v", err)
|
||||
}
|
||||
if checked == 0 {
|
||||
t.Fatal("manifest contained no entries")
|
||||
}
|
||||
}
|
||||
|
||||
// canonicalHash renders one file in the canonical form and hashes it. Kept
|
||||
// byte-identical to the awk canonicalisation used to build the manifest.
|
||||
func canonicalHash(path string) (string, error) {
|
||||
wb, err := reader.Open(path)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
defer wb.Close()
|
||||
|
||||
h := sha256.New()
|
||||
sheets := wb.Sheets()
|
||||
fmt.Fprintf(h, "SHEETCOUNT\t%d\n", len(sheets))
|
||||
for _, sh := range sheets {
|
||||
fmt.Fprintf(h, "SHEET\t%d\t%s\t%d\t%d\n", sh.Index, escape(sh.Name), sh.Height, sh.Width)
|
||||
err := wb.EachRow(sh.Index, func(s reader.Sheet, rowIdx int, row []reader.Cell) error {
|
||||
fmt.Fprintf(h, "ROW\t%d\t%d\t%d\n", s.Index, rowIdx, len(row))
|
||||
for c, cell := range row {
|
||||
fmt.Fprintf(h, "CELL\t%d\t%d\t%d\t%s\n", s.Index, rowIdx, c, escape(cell.Str))
|
||||
}
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
}
|
||||
return hex.EncodeToString(h.Sum(nil)), nil
|
||||
}
|
||||
|
||||
func escape(s string) string {
|
||||
return strings.NewReplacer("\\", `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`).Replace(s)
|
||||
}
|
||||
|
||||
// repoRoot walks up from the test's working directory to the directory holding
|
||||
// the data/ corpus.
|
||||
func repoRoot(t *testing.T) string {
|
||||
t.Helper()
|
||||
dir, err := os.Getwd()
|
||||
if err != nil {
|
||||
t.Fatalf("getwd: %v", err)
|
||||
}
|
||||
for i := 0; i < 6; i++ {
|
||||
if _, err := os.Stat(filepath.Join(dir, "data")); err == nil {
|
||||
return dir
|
||||
}
|
||||
dir = filepath.Dir(dir)
|
||||
}
|
||||
t.Fatal("could not locate repo root (no data/ directory found)")
|
||||
return ""
|
||||
}
|
||||
@@ -0,0 +1,79 @@
|
||||
// Package reader wraps the two spreadsheet libraries behind one streaming
|
||||
// contract, mirroring parser/src/reader.rs (which wraps calamine's
|
||||
// open_workbook_auto for both formats).
|
||||
//
|
||||
// The contract is deliberately row-streaming rather than whole-workbook
|
||||
// materialising: data/2017/ha-noi.xls alone holds 72k rows across two sheets,
|
||||
// and the Rust original keeps at most one sheet range live at a time.
|
||||
package reader
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// Cell is one spreadsheet cell rendered the way calamine's Data::to_string()
|
||||
// renders it.
|
||||
//
|
||||
// IsEmpty tracks calamine's Data::Empty variant separately from a string cell
|
||||
// that happens to be empty. Rust distinguishes them at reader.rs:42, but both
|
||||
// render "" and both count as blank in is_all_blank and in transform, so the
|
||||
// flag is diagnostic only — never compare on it.
|
||||
type Cell struct {
|
||||
Str string
|
||||
IsEmpty bool
|
||||
}
|
||||
|
||||
// Sheet carries a sheet's identity and geometry. Height and Width describe the
|
||||
// used range, matching calamine's Range::height()/width(); every used range in
|
||||
// the corpus starts at (0,0), verified across all 299 files.
|
||||
type Sheet struct {
|
||||
Index int
|
||||
Name string
|
||||
Height int
|
||||
Width int
|
||||
}
|
||||
|
||||
// RowFunc receives each row of each sheet. Rows are padded to the sheet's used
|
||||
// width — width is load-bearing because every column read downstream is
|
||||
// positional with an unwrap_or_default() equivalent, so a short row silently
|
||||
// NULLs its tail columns.
|
||||
type RowFunc func(sheet Sheet, rowIdx int, row []Cell) error
|
||||
|
||||
// Workbook is one opened spreadsheet.
|
||||
type Workbook interface {
|
||||
// Sheets returns sheet identity and geometry in workbook order.
|
||||
Sheets() []Sheet
|
||||
// EachRow streams every row of the given sheet in order.
|
||||
EachRow(sheetIdx int, fn RowFunc) error
|
||||
Close() error
|
||||
}
|
||||
|
||||
// Open dispatches on file extension, mirroring calamine's open_workbook_auto.
|
||||
func Open(path string) (Workbook, error) {
|
||||
switch strings.ToLower(filepath.Ext(path)) {
|
||||
case ".xls":
|
||||
return openXLS(path)
|
||||
case ".xlsx", ".xlsm":
|
||||
return openXLSX(path)
|
||||
default:
|
||||
return nil, fmt.Errorf("unsupported extension: %s", path)
|
||||
}
|
||||
}
|
||||
|
||||
// padRow extends row to width with empty cells, and truncates if longer.
|
||||
func padRow(row []Cell, width int) []Cell {
|
||||
if len(row) == width {
|
||||
return row
|
||||
}
|
||||
if len(row) > width {
|
||||
return row[:width]
|
||||
}
|
||||
out := make([]Cell, width)
|
||||
copy(out, row)
|
||||
for i := len(row); i < width; i++ {
|
||||
out[i] = Cell{IsEmpty: true}
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,119 @@
|
||||
package reader
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
|
||||
"github.com/pbnjay/grate"
|
||||
// Registers the BIFF backend with grate.Open.
|
||||
_ "github.com/pbnjay/grate/xls"
|
||||
)
|
||||
|
||||
// xlsWorkbook reads legacy BIFF through pbnjay/grate.
|
||||
//
|
||||
// grate replaced extrame/xls, which was measured against calamine ground truth
|
||||
// and found to corrupt 69% of cells and drop a further 28%: undecoded UTF-16LE
|
||||
// and BIFF record framing leaked into cell values, content moved between rows
|
||||
// and columns, and tail rows came back blank. That was charset-independent and
|
||||
// unfixable from the outside. grate reproduces calamine exactly on the same
|
||||
// files.
|
||||
//
|
||||
// The one normalisation grate needs is trailing blank rows: it yields rows past
|
||||
// the end of calamine's used range (one for a populated sheet, two for an empty
|
||||
// one), so trailing all-blank rows are trimmed. Note this is the opposite of
|
||||
// the xlsx path, where excelize already trims and calamine keeps a 1x1 empty
|
||||
// range — in both cases the rule is "match calamine's used range".
|
||||
type xlsWorkbook struct {
|
||||
sheets []Sheet
|
||||
rows [][][]Cell
|
||||
}
|
||||
|
||||
func openXLS(path string) (Workbook, error) {
|
||||
wb, err := grate.Open(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("grate open %s: %w", path, err)
|
||||
}
|
||||
defer wb.Close()
|
||||
|
||||
names, err := wb.List()
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("grate list %s: %w", path, err)
|
||||
}
|
||||
|
||||
out := &xlsWorkbook{}
|
||||
for idx, name := range names {
|
||||
sh, err := wb.Get(name)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("grate get %s/%s: %w", path, name, err)
|
||||
}
|
||||
|
||||
var raw [][]string
|
||||
for sh.Next() {
|
||||
row := sh.Strings()
|
||||
cp := make([]string, len(row))
|
||||
for i, v := range row {
|
||||
cp[i] = demergeMarker(v)
|
||||
}
|
||||
raw = append(raw, cp)
|
||||
}
|
||||
|
||||
// Trim to calamine's used range.
|
||||
height := len(raw)
|
||||
for height > 0 && rowAllBlank(raw[height-1]) {
|
||||
height--
|
||||
}
|
||||
raw = raw[:height]
|
||||
|
||||
width := 0
|
||||
for _, r := range raw {
|
||||
if len(r) > width {
|
||||
width = len(r)
|
||||
}
|
||||
}
|
||||
|
||||
cells := make([][]Cell, len(raw))
|
||||
for i, r := range raw {
|
||||
row := make([]Cell, len(r))
|
||||
for j, v := range r {
|
||||
row[j] = Cell{Str: v, IsEmpty: v == ""}
|
||||
}
|
||||
cells[i] = padRow(row, width)
|
||||
}
|
||||
|
||||
out.sheets = append(out.sheets, Sheet{Index: idx, Name: name, Height: height, Width: width})
|
||||
out.rows = append(out.rows, cells)
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// demergeMarker blanks grate's merged-cell continuation markers.
|
||||
//
|
||||
// grate fills the cells covered by a merge with sentinel runes; calamine
|
||||
// reports them as empty. In this corpus they occur only in the merged title
|
||||
// block of the 2016 spreadsheets (19 cells in rows 0-2 of one file). Only an
|
||||
// exact whole-value match is blanked, so a real cell that merely contains an
|
||||
// arrow is untouched.
|
||||
func demergeMarker(v string) string {
|
||||
switch v {
|
||||
case grate.ContinueColumnMerged, grate.EndColumnMerged,
|
||||
grate.ContinueRowMerged, grate.EndRowMerged:
|
||||
return ""
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
func (w *xlsWorkbook) Sheets() []Sheet { return w.sheets }
|
||||
|
||||
func (w *xlsWorkbook) EachRow(sheetIdx int, fn RowFunc) error {
|
||||
if sheetIdx < 0 || sheetIdx >= len(w.sheets) {
|
||||
return fmt.Errorf("sheet index %d out of range", sheetIdx)
|
||||
}
|
||||
sh := w.sheets[sheetIdx]
|
||||
for i, row := range w.rows[sheetIdx] {
|
||||
if err := fn(sh, i, row); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (w *xlsWorkbook) Close() error { return nil }
|
||||
@@ -0,0 +1,105 @@
|
||||
package reader
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
|
||||
"github.com/xuri/excelize/v2"
|
||||
)
|
||||
|
||||
// xlsxWorkbook reads OOXML through excelize.
|
||||
//
|
||||
// Two excelize behaviours must be corrected to match calamine:
|
||||
//
|
||||
// 1. GetRows applies the cell number format by default, while calamine renders
|
||||
// the underlying value. RawCellValue: true disables that.
|
||||
// 2. GetRows trims trailing blank cells, so rows are ragged; calamine returns a
|
||||
// rectangular used range. Rows are padded back out to the sheet width.
|
||||
type xlsxWorkbook struct {
|
||||
f *excelize.File
|
||||
sheets []Sheet
|
||||
rows [][][]Cell // [sheetIdx][rowIdx][colIdx]
|
||||
}
|
||||
|
||||
func openXLSX(path string) (Workbook, error) {
|
||||
f, err := excelize.OpenFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("excelize open %s: %w", path, err)
|
||||
}
|
||||
|
||||
wb := &xlsxWorkbook{f: f}
|
||||
crFixups := buildCRFixups(path)
|
||||
for idx, name := range f.GetSheetList() {
|
||||
raw, err := f.GetRows(name, excelize.Options{RawCellValue: true})
|
||||
if err != nil {
|
||||
f.Close()
|
||||
return nil, fmt.Errorf("excelize GetRows %s/%s: %w", path, name, err)
|
||||
}
|
||||
|
||||
// Do NOT trim trailing blank rows: calamine's used range keeps them, and
|
||||
// excelize's GetRows already drops trailing fully-empty rows itself.
|
||||
//
|
||||
// One correction is needed. 63 sheets in 2017-old and 53 in 2017-old2
|
||||
// hold a single empty shared-string cell at A1; calamine reports those
|
||||
// as a 1x1 range, while GetRows returns nothing. A genuinely empty sheet
|
||||
// (230 of them in 2016) is height 0 on both sides. GetCellType tells the
|
||||
// two apart: the empty-shared-string cell exists in the XML and types as
|
||||
// CellTypeSharedString, an absent cell types as CellTypeUnset.
|
||||
if len(raw) == 0 {
|
||||
if t, terr := f.GetCellType(name, "A1"); terr == nil && t != excelize.CellTypeUnset {
|
||||
raw = [][]string{{""}}
|
||||
}
|
||||
}
|
||||
height := len(raw)
|
||||
|
||||
width := 0
|
||||
for _, r := range raw {
|
||||
if len(r) > width {
|
||||
width = len(r)
|
||||
}
|
||||
}
|
||||
|
||||
cells := make([][]Cell, len(raw))
|
||||
for i, r := range raw {
|
||||
row := make([]Cell, len(r))
|
||||
for j, v := range r {
|
||||
if fixed, ok := crFixups[v]; ok {
|
||||
v = fixed
|
||||
} else {
|
||||
v = normalizeNumeric(f, name, j, i, v)
|
||||
}
|
||||
row[j] = Cell{Str: v, IsEmpty: v == ""}
|
||||
}
|
||||
cells[i] = padRow(row, width)
|
||||
}
|
||||
|
||||
wb.sheets = append(wb.sheets, Sheet{Index: idx, Name: name, Height: height, Width: width})
|
||||
wb.rows = append(wb.rows, cells)
|
||||
}
|
||||
return wb, nil
|
||||
}
|
||||
|
||||
func rowAllBlank(r []string) bool {
|
||||
for _, v := range r {
|
||||
if v != "" {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
func (w *xlsxWorkbook) Sheets() []Sheet { return w.sheets }
|
||||
|
||||
func (w *xlsxWorkbook) EachRow(sheetIdx int, fn RowFunc) error {
|
||||
if sheetIdx < 0 || sheetIdx >= len(w.sheets) {
|
||||
return fmt.Errorf("sheet index %d out of range", sheetIdx)
|
||||
}
|
||||
sh := w.sheets[sheetIdx]
|
||||
for i, row := range w.rows[sheetIdx] {
|
||||
if err := fn(sh, i, row); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (w *xlsxWorkbook) Close() error { return w.f.Close() }
|
||||
@@ -0,0 +1,155 @@
|
||||
package reader
|
||||
|
||||
import (
|
||||
"archive/zip"
|
||||
"bytes"
|
||||
"encoding/xml"
|
||||
"io"
|
||||
"strconv"
|
||||
|
||||
"github.com/xuri/excelize/v2"
|
||||
)
|
||||
|
||||
// normalizeNumeric reproduces calamine's rendering of a numeric cell.
|
||||
//
|
||||
// calamine parses a numeric cell to f64 and renders it with Rust's f64 Display,
|
||||
// so the stored literal "6.0" becomes "6". RawCellValue hands back the literal.
|
||||
//
|
||||
// The cell type must be consulted, not guessed: "01063476", "6.00" and "NAN" are
|
||||
// all shared strings that survive ParseFloat, and renumbering them would drop a
|
||||
// leading zero, drop a trailing zero, or recase NaN. GetCellType is only called
|
||||
// when re-rendering would actually change the text, which keeps it off the hot
|
||||
// path for the ~99% of cells that are already canonical or plainly non-numeric.
|
||||
func normalizeNumeric(f *excelize.File, sheet string, col, row int, v string) string {
|
||||
if v == "" {
|
||||
return v
|
||||
}
|
||||
fv, err := strconv.ParseFloat(v, 64)
|
||||
if err != nil {
|
||||
return v
|
||||
}
|
||||
out := strconv.FormatFloat(fv, 'f', -1, 64)
|
||||
if out == v {
|
||||
return v
|
||||
}
|
||||
axis, err := excelize.CoordinatesToCellName(col+1, row+1)
|
||||
if err != nil {
|
||||
return v
|
||||
}
|
||||
// OOXML omits the t attribute on numeric cells, and excelize has no map
|
||||
// entry for an empty t, so a plain number reports CellTypeUnset rather than
|
||||
// CellTypeNumber. Unset is only reachable here for a cell that exists and
|
||||
// parsed as a float — an absent cell is "" and returned above — so both
|
||||
// values mean "numeric". Shared strings report CellTypeSharedString and are
|
||||
// left alone, which is what protects "01063476", "6.00" and "NAN".
|
||||
if t, err := f.GetCellType(sheet, axis); err == nil &&
|
||||
(t == excelize.CellTypeNumber || t == excelize.CellTypeUnset) {
|
||||
return out
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
// buildCRFixups maps a shared string's line-ending-normalised form back to its
|
||||
// raw form, for the strings that contain a carriage return.
|
||||
//
|
||||
// Go's encoding/xml performs the line-ending normalisation the XML 1.0 spec
|
||||
// mandates (CRLF and lone CR both become LF), so excelize returns "a\nb" where
|
||||
// calamine — which reads the raw bytes — returns "a\r\nb". That difference
|
||||
// reaches the database: in one 2016 file it affects 2,233 TEN_CUMTHI values,
|
||||
// which populate the ten_cum_thi column.
|
||||
//
|
||||
// The trick is that character references are exempt from that normalisation, so
|
||||
// rewriting literal CR bytes to before decoding round-trips them intact.
|
||||
//
|
||||
// Returns nil when the file has no CR at all, which is the common case.
|
||||
func buildCRFixups(path string) map[string]string {
|
||||
zr, err := zip.OpenReader(path)
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
defer zr.Close()
|
||||
|
||||
var entry *zip.File
|
||||
for _, zf := range zr.File {
|
||||
if zf.Name == "xl/sharedStrings.xml" {
|
||||
entry = zf
|
||||
break
|
||||
}
|
||||
}
|
||||
if entry == nil {
|
||||
return nil
|
||||
}
|
||||
rc, err := entry.Open()
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
defer rc.Close()
|
||||
data, err := io.ReadAll(rc)
|
||||
if err != nil || !bytes.ContainsRune(data, '\r') {
|
||||
return nil
|
||||
}
|
||||
data = bytes.ReplaceAll(data, []byte{'\r'}, []byte(" "))
|
||||
|
||||
fixups := make(map[string]string)
|
||||
dec := xml.NewDecoder(bytes.NewReader(data))
|
||||
var cur bytes.Buffer
|
||||
inSI, inT := false, false
|
||||
for {
|
||||
tok, err := dec.Token()
|
||||
if err != nil {
|
||||
break
|
||||
}
|
||||
switch t := tok.(type) {
|
||||
case xml.StartElement:
|
||||
switch t.Name.Local {
|
||||
case "si":
|
||||
inSI, cur = true, bytes.Buffer{}
|
||||
case "t":
|
||||
inT = true
|
||||
}
|
||||
case xml.CharData:
|
||||
if inSI && inT {
|
||||
cur.Write(t)
|
||||
}
|
||||
case xml.EndElement:
|
||||
switch t.Name.Local {
|
||||
case "t":
|
||||
inT = false
|
||||
case "si":
|
||||
if inSI {
|
||||
raw := cur.String()
|
||||
if norm := normalizeEOL(raw); norm != raw {
|
||||
fixups[norm] = raw
|
||||
}
|
||||
}
|
||||
inSI = false
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(fixups) == 0 {
|
||||
return nil
|
||||
}
|
||||
return fixups
|
||||
}
|
||||
|
||||
// normalizeEOL applies XML 1.0 line-ending normalisation: CRLF and lone CR
|
||||
// both collapse to LF. This is what encoding/xml does to the raw bytes, so it
|
||||
// reproduces the text excelize hands back.
|
||||
func normalizeEOL(s string) string {
|
||||
if !bytes.ContainsRune([]byte(s), '\r') {
|
||||
return s
|
||||
}
|
||||
var b bytes.Buffer
|
||||
b.Grow(len(s))
|
||||
for i := 0; i < len(s); i++ {
|
||||
if s[i] == '\r' {
|
||||
if i+1 < len(s) && s[i+1] == '\n' {
|
||||
continue // CRLF: the LF is emitted on the next pass
|
||||
}
|
||||
b.WriteByte('\n') // lone CR
|
||||
continue
|
||||
}
|
||||
b.WriteByte(s[i])
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
@@ -0,0 +1,144 @@
|
||||
// Package schema is the single source of truth for the SQL shape of every
|
||||
// dataset — a direct port of parser/src/schema.rs.
|
||||
//
|
||||
// All four datasets (2016, 2017, 2017-old, 2017-old2) write into the same
|
||||
// 22-column student table. Columns a dataset has no data for bind NULL.
|
||||
//
|
||||
// Column provenance:
|
||||
//
|
||||
// ten_cum_thi, gioi_tinh, tieng_duc, tieng_nhat -> 2016 only
|
||||
// khtn, khxh, gdcd, tieng_nga -> 2017 datasets only
|
||||
// everything else -> both
|
||||
//
|
||||
// Before this consolidation the DDL, the INSERT and the subject regexes were
|
||||
// duplicated across four TOML configs, which is how the 2016 and 2017 schemas
|
||||
// drifted apart. The configs now carry only per-dataset parse rules.
|
||||
package schema
|
||||
|
||||
import "regexp"
|
||||
|
||||
// DDL is executed verbatim after the output database is (re)created.
|
||||
//
|
||||
// idx_ten_cum_thi is partial, so it holds zero entries on the three 2017
|
||||
// datasets — where the column is always NULL — while staying useful for the
|
||||
// 2016 cluster-grouping queries. Partial indexes are SQLite-specific.
|
||||
//
|
||||
// Byte-identical to parser/src/schema.rs:26-54; TestDDLMatchesRust enforces it.
|
||||
const DDL = `
|
||||
CREATE TABLE student (
|
||||
so_bao_danh TEXT PRIMARY KEY,
|
||||
ho_ten TEXT NOT NULL,
|
||||
ho_ten_ascii TEXT NOT NULL,
|
||||
ngay_sinh TEXT,
|
||||
ten_cum_thi TEXT,
|
||||
gioi_tinh TEXT,
|
||||
toan REAL,
|
||||
ngu_van REAL,
|
||||
vat_ly REAL,
|
||||
hoa_hoc REAL,
|
||||
sinh_hoc REAL,
|
||||
khtn REAL,
|
||||
lich_su REAL,
|
||||
dia_ly REAL,
|
||||
gdcd REAL,
|
||||
khxh REAL,
|
||||
tieng_anh REAL,
|
||||
tieng_phap REAL,
|
||||
tieng_nga REAL,
|
||||
tieng_duc REAL,
|
||||
tieng_nhat REAL,
|
||||
tieng_trung REAL
|
||||
);
|
||||
CREATE INDEX idx_ho_ten ON student(ho_ten);
|
||||
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
|
||||
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
|
||||
`
|
||||
|
||||
// IdentityFields are the identity columns, in INSERT parameter order.
|
||||
var IdentityFields = []string{
|
||||
"so_bao_danh",
|
||||
"ho_ten",
|
||||
"ho_ten_ascii",
|
||||
"ngay_sinh",
|
||||
"ten_cum_thi",
|
||||
"gioi_tinh",
|
||||
}
|
||||
|
||||
// ScoreFields are the subject columns, in INSERT parameter order. Bound NULL
|
||||
// when a row has no score for that subject.
|
||||
var ScoreFields = []string{
|
||||
"toan",
|
||||
"ngu_van",
|
||||
"vat_ly",
|
||||
"hoa_hoc",
|
||||
"sinh_hoc",
|
||||
"khtn",
|
||||
"lich_su",
|
||||
"dia_ly",
|
||||
"gdcd",
|
||||
"khxh",
|
||||
"tieng_anh",
|
||||
"tieng_phap",
|
||||
"tieng_nga",
|
||||
"tieng_duc",
|
||||
"tieng_nhat",
|
||||
"tieng_trung",
|
||||
}
|
||||
|
||||
// ParamCount is the total bound parameters per row.
|
||||
const ParamCount = 22
|
||||
|
||||
// InsertSQL is a positional INSERT matching IdentityFields then ScoreFields.
|
||||
//
|
||||
// OR REPLACE is a behavioural contract, not an optimisation: a repeated SBD
|
||||
// overwrites the earlier row rather than aborting the transaction, so the last
|
||||
// file to supply a duplicate wins.
|
||||
const InsertSQL = `
|
||||
INSERT OR REPLACE INTO student
|
||||
(so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh, ten_cum_thi, gioi_tinh,
|
||||
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn,
|
||||
lich_su, dia_ly, gdcd, khxh,
|
||||
tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung)
|
||||
VALUES
|
||||
(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
`
|
||||
|
||||
// scorePatternSources holds the regex per subject, applied to the DIEM_THI cell
|
||||
// text. Copied verbatim from parser/src/schema.rs:126-143 — the literals contain
|
||||
// Vietnamese text and must never be retyped.
|
||||
//
|
||||
// Every pattern runs against every dataset. A subject absent from a given exam
|
||||
// year simply never matches and stays NULL: 2016 files contain no "KHTN:" or
|
||||
// "Tiếng Nga:" tokens, and 2017 files contain no "Tiếng Đức:" or "Tiếng Nhật:".
|
||||
//
|
||||
// Go's regexp and Rust's regex crate are both RE2, and these patterns use no
|
||||
// backreferences, lookaround or Unicode classes, so they port with zero risk.
|
||||
var scorePatternSources = map[string]string{
|
||||
"toan": `Toán:\s*(\d+(?:\.\d+)?)`,
|
||||
"ngu_van": `Ngữ văn:\s*(\d+(?:\.\d+)?)`,
|
||||
"vat_ly": `Vật lí:\s*(\d+(?:\.\d+)?)`,
|
||||
"hoa_hoc": `Hóa học:\s*(\d+(?:\.\d+)?)`,
|
||||
"sinh_hoc": `Sinh học:\s*(\d+(?:\.\d+)?)`,
|
||||
"khtn": `KHTN:\s*(\d+(?:\.\d+)?)`,
|
||||
"lich_su": `Lịch sử:\s*(\d+(?:\.\d+)?)`,
|
||||
"dia_ly": `Địa lí:\s*(\d+(?:\.\d+)?)`,
|
||||
"gdcd": `GDCD:\s*(\d+(?:\.\d+)?)`,
|
||||
"khxh": `KHXH:\s*(\d+(?:\.\d+)?)`,
|
||||
"tieng_anh": `Tiếng Anh:\s*(\d+(?:\.\d+)?)`,
|
||||
"tieng_phap": `Tiếng Pháp:\s*(\d+(?:\.\d+)?)`,
|
||||
"tieng_nga": `Tiếng Nga:\s*(\d+(?:\.\d+)?)`,
|
||||
"tieng_duc": `Tiếng Đức:\s*(\d+(?:\.\d+)?)`,
|
||||
"tieng_nhat": `Tiếng Nhật:\s*(\d+(?:\.\d+)?)`,
|
||||
"tieng_trung": `Tiếng Trung:\s*(\d+(?:\.\d+)?)`,
|
||||
}
|
||||
|
||||
// ScorePatterns holds the compiled subject regexes, compiled once at init.
|
||||
// Rust compiles them once per run in CompiledPatterns::new; a package-level map
|
||||
// is the equivalent for a single-threaded CLI.
|
||||
var ScorePatterns = func() map[string]*regexp.Regexp {
|
||||
out := make(map[string]*regexp.Regexp, len(scorePatternSources))
|
||||
for field, src := range scorePatternSources {
|
||||
out[field] = regexp.MustCompile(src)
|
||||
}
|
||||
return out
|
||||
}()
|
||||
@@ -0,0 +1,150 @@
|
||||
package schema
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Ports the four tests in parser/src/schema.rs:149-213. Their purpose is to stop
|
||||
// the DDL, the INSERT column list and the field-order constants from drifting
|
||||
// apart — a drift that silently lands values in the wrong columns.
|
||||
|
||||
// TestInsertMatchesFieldOrder ports insert_matches_field_order (schema.rs:156).
|
||||
func TestInsertMatchesFieldOrder(t *testing.T) {
|
||||
if ParamCount != 22 {
|
||||
t.Errorf("ParamCount = %d, want 22", ParamCount)
|
||||
}
|
||||
if got := strings.Count(InsertSQL, "?"); got != ParamCount {
|
||||
t.Errorf("INSERT placeholders = %d, want %d", got, ParamCount)
|
||||
}
|
||||
|
||||
open := strings.Index(InsertSQL, "(")
|
||||
closeIdx := strings.Index(InsertSQL, ")")
|
||||
if open < 0 || closeIdx < 0 {
|
||||
t.Fatal("INSERT must contain a column list")
|
||||
}
|
||||
var listed []string
|
||||
for _, c := range strings.Split(InsertSQL[open+1:closeIdx], ",") {
|
||||
if c = strings.TrimSpace(c); c != "" {
|
||||
listed = append(listed, c)
|
||||
}
|
||||
}
|
||||
|
||||
want := append(append([]string{}, IdentityFields...), ScoreFields...)
|
||||
if len(listed) != len(want) {
|
||||
t.Fatalf("INSERT lists %d columns, want %d", len(listed), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if listed[i] != want[i] {
|
||||
t.Errorf("column %d: INSERT has %q, field order has %q", i, listed[i], want[i])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestScorePatternsCoverScoreFields ports score_patterns_cover_score_fields
|
||||
// (schema.rs:183).
|
||||
func TestScorePatternsCoverScoreFields(t *testing.T) {
|
||||
if len(ScorePatterns) != len(ScoreFields) {
|
||||
t.Fatalf("%d patterns for %d score columns", len(ScorePatterns), len(ScoreFields))
|
||||
}
|
||||
inFields := make(map[string]bool, len(ScoreFields))
|
||||
for _, f := range ScoreFields {
|
||||
inFields[f] = true
|
||||
}
|
||||
for field := range ScorePatterns {
|
||||
if !inFields[field] {
|
||||
t.Errorf("pattern %q has no column", field)
|
||||
}
|
||||
}
|
||||
for _, field := range ScoreFields {
|
||||
if _, ok := ScorePatterns[field]; !ok {
|
||||
t.Errorf("column %q has no pattern", field)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestDDLColumnsMatchInsert ports ddl_columns_match_insert (schema.rs:201).
|
||||
func TestDDLColumnsMatchInsert(t *testing.T) {
|
||||
for _, field := range append(append([]string{}, IdentityFields...), ScoreFields...) {
|
||||
if !strings.Contains(DDL, field) {
|
||||
t.Errorf("DDL missing column %q", field)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestScorePatternsCompile ports score_patterns_compile (schema.rs:208).
|
||||
// Compilation happens in the package initialiser, so reaching this point already
|
||||
// proves it; the explicit checks guard against an empty or partial table.
|
||||
func TestScorePatternsCompile(t *testing.T) {
|
||||
for field, re := range ScorePatterns {
|
||||
if re == nil {
|
||||
t.Errorf("pattern %q is nil", field)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestDDLMatchesRust asserts the DDL is byte-identical to parser/src/schema.rs.
|
||||
// Anything less and the two parsers can produce structurally different databases
|
||||
// while every row-level check still passes.
|
||||
func TestDDLMatchesRust(t *testing.T) {
|
||||
const want = `
|
||||
CREATE TABLE student (
|
||||
so_bao_danh TEXT PRIMARY KEY,
|
||||
ho_ten TEXT NOT NULL,
|
||||
ho_ten_ascii TEXT NOT NULL,
|
||||
ngay_sinh TEXT,
|
||||
ten_cum_thi TEXT,
|
||||
gioi_tinh TEXT,
|
||||
toan REAL,
|
||||
ngu_van REAL,
|
||||
vat_ly REAL,
|
||||
hoa_hoc REAL,
|
||||
sinh_hoc REAL,
|
||||
khtn REAL,
|
||||
lich_su REAL,
|
||||
dia_ly REAL,
|
||||
gdcd REAL,
|
||||
khxh REAL,
|
||||
tieng_anh REAL,
|
||||
tieng_phap REAL,
|
||||
tieng_nga REAL,
|
||||
tieng_duc REAL,
|
||||
tieng_nhat REAL,
|
||||
tieng_trung REAL
|
||||
);
|
||||
CREATE INDEX idx_ho_ten ON student(ho_ten);
|
||||
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
|
||||
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
|
||||
`
|
||||
if DDL != want {
|
||||
t.Errorf("DDL diverges from parser/src/schema.rs:26-54\n--- got ---\n%s\n--- want ---\n%s", DDL, want)
|
||||
}
|
||||
}
|
||||
|
||||
// TestScorePatternsMatchScores exercises each pattern against the shape the
|
||||
// DIEM_THI cell actually carries, including the wide runs of spaces seen in the
|
||||
// real corpus.
|
||||
func TestScorePatternsMatchScores(t *testing.T) {
|
||||
const cell = "Toán: 8.50 Ngữ văn: 7.00 Tiếng Đức: 9 KHXH: 5.58 "
|
||||
cases := map[string]string{
|
||||
"toan": "8.50",
|
||||
"ngu_van": "7.00",
|
||||
"tieng_duc": "9",
|
||||
"khxh": "5.58",
|
||||
"tieng_nhat": "", // absent from the cell -> no match
|
||||
}
|
||||
for field, want := range cases {
|
||||
re, ok := ScorePatterns[field]
|
||||
if !ok {
|
||||
t.Fatalf("no pattern for %q", field)
|
||||
}
|
||||
m := re.FindStringSubmatch(cell)
|
||||
got := ""
|
||||
if m != nil {
|
||||
got = m[1]
|
||||
}
|
||||
if got != want {
|
||||
t.Errorf("%s: matched %q, want %q", field, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
// Package sqlitedb registers the SQLite driver the parser writes with and names
|
||||
// it in one place.
|
||||
//
|
||||
// modernc.org/sqlite is a pure-Go SQLite, chosen so the whole module stays
|
||||
// cgo-free — grate, excelize and yaml.v3 are pure Go too, so CI can compile with
|
||||
// CGO_ENABLED=0 and no C toolchain.
|
||||
//
|
||||
// It is a machine-transpiled SQLite rather than the upstream C amalgamation that
|
||||
// Rust's rusqlite --bundled vendors (libsqlite3-sys 0.30.1). Two things make that
|
||||
// acceptable: this parser uses only plain SQL — no CTEs, window functions,
|
||||
// triggers or extensions — and the differential gate compares a full-table
|
||||
// SHA-256 plus PRAGMA table_info/index_list against live Rust output.
|
||||
//
|
||||
// Verified on linux/arm64 with v1.56.0 (SQLite 3.53.3): the full DDL including
|
||||
// the partial idx_ten_cum_thi index, INSERT OR REPLACE, and VACUUM.
|
||||
//
|
||||
// Fallback if the differential gate ever implicates the driver: mattn/go-sqlite3
|
||||
// is upstream C at the cost of cgo.
|
||||
package sqlitedb
|
||||
|
||||
// Registers "sqlite" with database/sql.
|
||||
import _ "modernc.org/sqlite"
|
||||
|
||||
// DriverName is the database/sql driver name to pass to sql.Open.
|
||||
const DriverName = "sqlite"
|
||||
@@ -0,0 +1,65 @@
|
||||
package transform_test
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"os"
|
||||
"testing"
|
||||
|
||||
_ "github.com/tiennm99/thptqg/go-parser/internal/sqlitedb"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/transform"
|
||||
)
|
||||
|
||||
// TestToAsciiAgainstRustOutput cross-checks ToAscii against Rust on real data.
|
||||
//
|
||||
// A Rust-built database is its own oracle: every row carries ho_ten alongside
|
||||
// the ho_ten_ascii that Rust derived from it, so the whole table is a
|
||||
// name -> expected-slug corpus far broader than the 20 hand-picked unit cases.
|
||||
//
|
||||
// Point it at a Rust-built database:
|
||||
//
|
||||
// GO_PARSER_RUST_DB=/tmp/rust-2016.db go test ./internal/transform/
|
||||
//
|
||||
// Skips when unset, so the default suite stays hermetic.
|
||||
func TestToAsciiAgainstRustOutput(t *testing.T) {
|
||||
path := os.Getenv("GO_PARSER_RUST_DB")
|
||||
if path == "" {
|
||||
t.Skip("GO_PARSER_RUST_DB not set; skipping cross-check against Rust output")
|
||||
}
|
||||
if _, err := os.Stat(path); err != nil {
|
||||
t.Skipf("GO_PARSER_RUST_DB=%s not readable: %v", path, err)
|
||||
}
|
||||
|
||||
db, err := sql.Open("sqlite", "file:"+path+"?mode=ro")
|
||||
if err != nil {
|
||||
t.Fatalf("open %s: %v", path, err)
|
||||
}
|
||||
defer db.Close()
|
||||
|
||||
rows, err := db.Query("SELECT ho_ten, ho_ten_ascii FROM student")
|
||||
if err != nil {
|
||||
t.Fatalf("query: %v", err)
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
var checked, bad int
|
||||
for rows.Next() {
|
||||
var name, rustAscii string
|
||||
if err := rows.Scan(&name, &rustAscii); err != nil {
|
||||
t.Fatalf("scan: %v", err)
|
||||
}
|
||||
checked++
|
||||
if got := transform.ToAscii(name); got != rustAscii {
|
||||
bad++
|
||||
if bad <= 5 {
|
||||
t.Errorf("ToAscii(%q)\n rust = %q\n go = %q", name, rustAscii, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
t.Fatalf("iterate: %v", err)
|
||||
}
|
||||
if checked == 0 {
|
||||
t.Fatal("database contained no rows")
|
||||
}
|
||||
t.Logf("compared %d real names, %d mismatches", checked, bad)
|
||||
}
|
||||
@@ -0,0 +1,223 @@
|
||||
// Package transform performs row transformation: ASCII normalisation, score
|
||||
// regex parsing, and validation — a port of parser/src/transform.rs.
|
||||
//
|
||||
// ToAscii replicates build-lib.js toAscii exactly:
|
||||
//
|
||||
// str.normalize("NFD").replace(/[̀-ͯ]/g,"").replace(/đ/gi,"d").toLowerCase()
|
||||
package transform
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"math"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"golang.org/x/text/unicode/norm"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/config"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/schema"
|
||||
)
|
||||
|
||||
// ToAscii normalises a Vietnamese name to an ASCII slug.
|
||||
//
|
||||
// 1. NFD decompose (splits base + combining diacritics)
|
||||
// 2. Drop combining marks in U+0300..U+036F
|
||||
// 3. Replace đ/Đ with d (NFD does not decompose them)
|
||||
// 4. Lowercase
|
||||
//
|
||||
// Step 2 filters a LITERAL CODEPOINT RANGE, not a Unicode category. The inline
|
||||
// comment at transform.rs:53 says "Unicode category M", but the code at :56 is
|
||||
// the specification and it checks '\u{0300}'..='\u{036f}'. unicode.Is(unicode.Mn, r)
|
||||
// is strictly broader and would strip marks Rust keeps, silently changing
|
||||
// ho_ten_ascii — the column the site's accent-insensitive search runs on.
|
||||
//
|
||||
// Step order matters and mirrors transform.rs:54-63: the đ/Đ replacement happens
|
||||
// before lowercasing.
|
||||
func ToAscii(s string) string {
|
||||
if s == "" {
|
||||
return ""
|
||||
}
|
||||
decomposed := norm.NFD.String(s)
|
||||
|
||||
var b strings.Builder
|
||||
b.Grow(len(decomposed))
|
||||
for _, r := range decomposed {
|
||||
if r >= 0x0300 && r <= 0x036F {
|
||||
continue // combining mark, in the range Rust drops
|
||||
}
|
||||
switch r {
|
||||
case 'đ', 'Đ':
|
||||
b.WriteByte('d')
|
||||
default:
|
||||
b.WriteRune(r)
|
||||
}
|
||||
}
|
||||
return strings.ToLower(b.String())
|
||||
}
|
||||
|
||||
// ParsedRow is one row ready for insertion.
|
||||
type ParsedRow struct {
|
||||
SoBaoDanh string
|
||||
HoTen string
|
||||
HoTenAscii string
|
||||
NgaySinh *string
|
||||
// TenCumThi is 2016 only: examination cluster name (TEN_CUMTHI column).
|
||||
TenCumThi *string
|
||||
// GioiTinh is 2016 only: gender, normalised to "Nam"/"Nữ" or nil.
|
||||
GioiTinh *string
|
||||
// Scores maps subject field -> value. Absent subjects are simply missing and
|
||||
// bind NULL.
|
||||
Scores map[string]float64
|
||||
}
|
||||
|
||||
// SkipReason says why a row was skipped, or SkipNone when it passed.
|
||||
//
|
||||
// The distinction is load-bearing for the printed counters, and the two
|
||||
// non-blank reasons are counted as source rows while BlankRow is not — but note
|
||||
// that split lives in the CALLER, not here. parser/src/main.rs has two call
|
||||
// sites with opposite outcomes for BlankRow: at :135-137 it returns before the
|
||||
// counter at :140, while at :151 it matches Err(BlankRow) => {} and falls
|
||||
// through to transform and insert. The build loop must reproduce both.
|
||||
type SkipReason int
|
||||
|
||||
const (
|
||||
SkipNone SkipReason = iota
|
||||
// SkipBlankRow: row is fully blank (2017-old2 only, checked before the
|
||||
// source-row counter).
|
||||
SkipBlankRow
|
||||
// SkipEmptyField: so_bao_danh or ho_ten empty/missing.
|
||||
SkipEmptyField
|
||||
// SkipNonNumericSbd: so_bao_danh contains non-digit characters
|
||||
// (2017-old / 2017-old2 guard).
|
||||
SkipNonNumericSbd
|
||||
)
|
||||
|
||||
func (s SkipReason) String() string {
|
||||
switch s {
|
||||
case SkipNone:
|
||||
return "none"
|
||||
case SkipBlankRow:
|
||||
return "blank_row"
|
||||
case SkipEmptyField:
|
||||
return "empty_field"
|
||||
case SkipNonNumericSbd:
|
||||
return "non_numeric_sbd"
|
||||
}
|
||||
return "unknown"
|
||||
}
|
||||
|
||||
// ValidateRow checks a row against the dataset's validation rules.
|
||||
//
|
||||
// Signature mirrors transform.rs:101-107, taking stripBlankRows and allBlank
|
||||
// explicitly; a shorter signature could not express both blank-row paths.
|
||||
func ValidateRow(hoTen, soBaoDanh string, cfg *config.ValidationCfg, stripBlankRows, allBlank bool) SkipReason {
|
||||
// 2017-old2: skip fully blank rows BEFORE counting source rows.
|
||||
if stripBlankRows && allBlank {
|
||||
return SkipBlankRow
|
||||
}
|
||||
if cfg.RequireNonemptySbd && soBaoDanh == "" {
|
||||
return SkipEmptyField
|
||||
}
|
||||
if cfg.RequireNonemptyName && hoTen == "" {
|
||||
return SkipEmptyField
|
||||
}
|
||||
if cfg.RequireNumericSbd && !allASCIIDigits(soBaoDanh) {
|
||||
return SkipNonNumericSbd
|
||||
}
|
||||
return SkipNone
|
||||
}
|
||||
|
||||
// allASCIIDigits mirrors Rust's chars().all(|c| c.is_ascii_digit()).
|
||||
//
|
||||
// Deliberately not strconv.Atoi: Atoi accepts a leading sign, so "+123" would
|
||||
// pass a check Rust rejects. Empty input returns true, matching Rust's all() on
|
||||
// an empty iterator — the empty case is caught earlier by RequireNonemptySbd.
|
||||
func allASCIIDigits(s string) bool {
|
||||
for i := 0; i < len(s); i++ {
|
||||
if s[i] < '0' || s[i] > '9' {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// ParseScores extracts subject scores from a DIEM_THI cell.
|
||||
//
|
||||
// Every one of the 16 patterns runs against every dataset; a subject absent from
|
||||
// a given exam year never matches and stays NULL. Matching is unanchored
|
||||
// first-match, like Rust's Regex::captures.
|
||||
func ParseScores(diemThi string) map[string]float64 {
|
||||
out := make(map[string]float64)
|
||||
if diemThi == "" {
|
||||
return out
|
||||
}
|
||||
for field, re := range schema.ScorePatterns {
|
||||
m := re.FindStringSubmatch(diemThi)
|
||||
if m == nil {
|
||||
continue
|
||||
}
|
||||
v, err := strconv.ParseFloat(m[1], 64)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
// Unreachable given the pattern shape, but kept for parity with
|
||||
// transform.rs:136 (is_finite).
|
||||
if math.IsInf(v, 0) || math.IsNaN(v) {
|
||||
continue
|
||||
}
|
||||
out[field] = v
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// ErrNoColumns is returned when the fixed-column path is used on a config that
|
||||
// has no columns: mapping. Rust panics here via .expect() (transform.rs:162);
|
||||
// returning an error is the Go-idiomatic equivalent and is unreachable in
|
||||
// practice, since only the non-2016 path calls this.
|
||||
var ErrNoColumns = errors.New("transform: config has no columns mapping")
|
||||
|
||||
// TransformRow extracts one row into a ParsedRow using fixed column indices,
|
||||
// the 2017-family path. 2016 uses runtime format detection instead.
|
||||
func TransformRow(raw []reader.Cell, cfg *config.DatasetConfig) (*ParsedRow, error) {
|
||||
cols := cfg.Columns
|
||||
if cols == nil {
|
||||
return nil, ErrNoColumns
|
||||
}
|
||||
|
||||
// Trimmed accessor, mirroring the closure at transform.rs:163-167. Out-of-range
|
||||
// indices yield "" rather than an error, matching unwrap_or_default().
|
||||
get := func(idx int) string {
|
||||
if idx < 0 || idx >= len(raw) {
|
||||
return ""
|
||||
}
|
||||
return strings.TrimSpace(raw[idx].Str)
|
||||
}
|
||||
|
||||
hoTen := get(cols.HoTen)
|
||||
ngaySinh := get(cols.NgaySinh)
|
||||
soBaoDanh := get(cols.SoBaoDanh)
|
||||
|
||||
// diem_thi is read WITHOUT trimming (transform.rs:172-175), unlike the three
|
||||
// fields above. Harmless because the score patterns are unanchored, but it is
|
||||
// the shipped behaviour — do not "tidy" it.
|
||||
diemThi := ""
|
||||
if cols.DiemThi >= 0 && cols.DiemThi < len(raw) {
|
||||
diemThi = raw[cols.DiemThi].Str
|
||||
}
|
||||
|
||||
var ngaySinhOpt *string
|
||||
if ngaySinh != "" {
|
||||
ngaySinhOpt = &ngaySinh
|
||||
}
|
||||
|
||||
return &ParsedRow{
|
||||
SoBaoDanh: soBaoDanh,
|
||||
HoTen: hoTen,
|
||||
HoTenAscii: ToAscii(hoTen),
|
||||
NgaySinh: ngaySinhOpt,
|
||||
TenCumThi: nil,
|
||||
GioiTinh: nil,
|
||||
Scores: ParseScores(diemThi),
|
||||
}, nil
|
||||
}
|
||||
@@ -0,0 +1,283 @@
|
||||
package transform
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/config"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/reader"
|
||||
)
|
||||
|
||||
// Ports every test in parser/src/transform.rs's test module (:201-409) — 29 in
|
||||
// total, not the 20 in the :213-315 range, which covers only ToAscii. The nine
|
||||
// outside that range are the ParseScores and ValidateRow cases, i.e. exactly the
|
||||
// behaviours this package's traps concern.
|
||||
|
||||
// --- ToAscii: the 20 cases at transform.rs:213-315 ---
|
||||
|
||||
func TestToAscii(t *testing.T) {
|
||||
cases := []struct{ name, in, want string }{
|
||||
{"plain_latin", "Nguyen Van A", "nguyen van a"},
|
||||
{"nguyen_thi_hoa", "Nguyễn Thị Hoa", "nguyen thi hoa"},
|
||||
{"tran_van_duc", "Trần Văn Đức", "tran van duc"},
|
||||
{"le_thi_my_duyen", "Lê Thị Mỹ Duyên", "le thi my duyen"},
|
||||
{"pham_thi_lan", "Phạm Thị Lan", "pham thi lan"},
|
||||
{"bui_thi_thu", "Bùi Thị Thu", "bui thi thu"},
|
||||
{"hoang_van_truong", "Hoàng Văn Trường", "hoang van truong"},
|
||||
{"do_thi_ngan", "Đỗ Thị Ngân", "do thi ngan"},
|
||||
{"nguyen_van_khanh", "Nguyễn Văn Khánh", "nguyen van khanh"},
|
||||
{"trinh_thi_bich_ngoc", "Trịnh Thị Bích Ngọc", "trinh thi bich ngoc"},
|
||||
{"vu_thi_dieu", "Vũ Thị Diệu", "vu thi dieu"},
|
||||
{"nguyen_thi_tuong_vi", "Nguyễn Thị Tường Vi", "nguyen thi tuong vi"},
|
||||
{"lowercase_d_stroke", "đặng thị hằng", "dang thi hang"},
|
||||
{"uppercase_d_stroke", "ĐẶNG THỊ HẰNG", "dang thi hang"},
|
||||
{"mixed_case", "NGUYỄN VĂN AN", "nguyen van an"},
|
||||
{"tran_thi_kim_anh", "Trần Thị Kim Anh", "tran thi kim anh"},
|
||||
{"nguyen_thi_phuong_thao", "Nguyễn Thị Phương Thảo", "nguyen thi phuong thao"},
|
||||
{"le_van_long", "Lê Văn Long", "le van long"},
|
||||
{"vo_thi_xuan_mai", "Võ Thị Xuân Mai", "vo thi xuan mai"},
|
||||
{"empty_string", "", ""},
|
||||
}
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
if got := ToAscii(c.in); got != c.want {
|
||||
t.Errorf("ToAscii(%q) = %q, want %q", c.in, got, c.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestToAsciiUsesLiteralRangeNotUnicodeMn guards the highest-value trap in this
|
||||
// package. transform.rs:56 filters the literal range U+0300..U+036F; the inline
|
||||
// comment at :53 calls it "Unicode category M", but the code is the spec.
|
||||
// unicode.Mn is strictly broader, so using it would strip marks Rust keeps.
|
||||
// U+0654 (ARABIC HAMZA ABOVE) is in Mn but outside the range: Rust keeps it.
|
||||
func TestToAsciiUsesLiteralRangeNotUnicodeMn(t *testing.T) {
|
||||
const in = "aٔb"
|
||||
if got := ToAscii(in); got != in {
|
||||
t.Errorf("ToAscii(%q) = %q — a combining mark outside U+0300..U+036F must survive; "+
|
||||
"stripping it means unicode.Mn was used instead of the literal range", in, got)
|
||||
}
|
||||
// And a mark inside the range must be stripped.
|
||||
if got := ToAscii("áb"); got != "ab" {
|
||||
t.Errorf("ToAscii(\"a\\u0301b\") = %q, want \"ab\"", got)
|
||||
}
|
||||
}
|
||||
|
||||
// TestToAsciiDStrokeIndependentOfNFD proves the đ/Đ replacement is a separate
|
||||
// step: NFD does not decompose them, so relying on the mark filter alone loses
|
||||
// the letter entirely.
|
||||
func TestToAsciiDStrokeIndependentOfNFD(t *testing.T) {
|
||||
for _, c := range []struct{ in, want string }{
|
||||
{"đ", "d"}, {"Đ", "d"}, {"đĐ", "dd"},
|
||||
} {
|
||||
if got := ToAscii(c.in); got != c.want {
|
||||
t.Errorf("ToAscii(%q) = %q, want %q", c.in, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// --- ParseScores: transform.rs:323, :332, :342 ---
|
||||
|
||||
func TestParseScoresSingle(t *testing.T) {
|
||||
s := ParseScores("Toán: 8.5")
|
||||
if v, ok := s["toan"]; !ok || v != 8.5 {
|
||||
t.Errorf("toan = %v (present=%v), want 8.5", v, ok)
|
||||
}
|
||||
if _, ok := s["ngu_van"]; ok {
|
||||
t.Error("ngu_van should be absent")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseScoresMultiple(t *testing.T) {
|
||||
s := ParseScores("Toán: 7.25 Ngữ văn: 6.0 Vật lí: 9")
|
||||
for field, want := range map[string]float64{"toan": 7.25, "ngu_van": 6.0, "vat_ly": 9.0} {
|
||||
if v, ok := s[field]; !ok || v != want {
|
||||
t.Errorf("%s = %v (present=%v), want %v", field, v, ok, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseScoresEmptyCell(t *testing.T) {
|
||||
if s := ParseScores(""); len(s) != 0 {
|
||||
t.Errorf("ParseScores(\"\") = %v, want empty", s)
|
||||
}
|
||||
}
|
||||
|
||||
// TestParseScoresRealCellShape uses the wide space runs seen in the corpus.
|
||||
func TestParseScoresRealCellShape(t *testing.T) {
|
||||
const cell = "Toán: 4.60 Ngữ văn: 5.50 Lịch sử: 4.50 "
|
||||
s := ParseScores(cell)
|
||||
if len(s) != 3 {
|
||||
t.Fatalf("matched %d subjects, want 3: %v", len(s), s)
|
||||
}
|
||||
if s["toan"] != 4.60 || s["ngu_van"] != 5.50 || s["lich_su"] != 4.50 {
|
||||
t.Errorf("got %v", s)
|
||||
}
|
||||
}
|
||||
|
||||
// --- ValidateRow: transform.rs:359, :365, :374, :383, :393, :400 ---
|
||||
|
||||
func defaultValidation() *config.ValidationCfg {
|
||||
return &config.ValidationCfg{
|
||||
RequireNumericSbd: false,
|
||||
RequireNonemptyName: true,
|
||||
RequireNonemptySbd: true,
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateOK(t *testing.T) {
|
||||
if r := ValidateRow("Nguyen Van A", "12345678", defaultValidation(), false, false); r != SkipNone {
|
||||
t.Errorf("got %v, want SkipNone", r)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateEmptySbd(t *testing.T) {
|
||||
if r := ValidateRow("Nguyen Van A", "", defaultValidation(), false, false); r != SkipEmptyField {
|
||||
t.Errorf("got %v, want SkipEmptyField", r)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateEmptyName(t *testing.T) {
|
||||
if r := ValidateRow("", "12345678", defaultValidation(), false, false); r != SkipEmptyField {
|
||||
t.Errorf("got %v, want SkipEmptyField", r)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateNonNumericSbdRejected(t *testing.T) {
|
||||
v := defaultValidation()
|
||||
v.RequireNumericSbd = true
|
||||
if r := ValidateRow("Nguyen Van A", "12AB5678", v, false, false); r != SkipNonNumericSbd {
|
||||
t.Errorf("got %v, want SkipNonNumericSbd", r)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateNumericSbdAccepted(t *testing.T) {
|
||||
v := defaultValidation()
|
||||
v.RequireNumericSbd = true
|
||||
if r := ValidateRow("Nguyen Van A", "12345678", v, false, false); r != SkipNone {
|
||||
t.Errorf("got %v, want SkipNone", r)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateBlankRowSkipped(t *testing.T) {
|
||||
if r := ValidateRow("", "", defaultValidation(), true, true); r != SkipBlankRow {
|
||||
t.Errorf("got %v, want SkipBlankRow", r)
|
||||
}
|
||||
}
|
||||
|
||||
// TestValidateNumericSbdIsDigitScanNotAtoi: Rust uses chars().all(is_ascii_digit),
|
||||
// which strconv.Atoi does not reproduce — Atoi accepts a leading sign, and would
|
||||
// wrongly admit "+123".
|
||||
func TestValidateNumericSbdIsDigitScanNotAtoi(t *testing.T) {
|
||||
v := defaultValidation()
|
||||
v.RequireNumericSbd = true
|
||||
for _, sbd := range []string{"+123", "-123", "12 3", "1.0", "ABC123", "123"} {
|
||||
if r := ValidateRow("Nguyen Van A", sbd, v, false, false); r != SkipNonNumericSbd {
|
||||
t.Errorf("ValidateRow(sbd=%q) = %v, want SkipNonNumericSbd", sbd, r)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestValidateBlankRowOnlyWhenStripEnabled: with strip_blank_rows false, an
|
||||
// all-blank row falls through to the empty-field checks instead (transform.rs:109).
|
||||
func TestValidateBlankRowOnlyWhenStripEnabled(t *testing.T) {
|
||||
if r := ValidateRow("", "", defaultValidation(), false, true); r != SkipEmptyField {
|
||||
t.Errorf("got %v, want SkipEmptyField when strip_blank_rows is off", r)
|
||||
}
|
||||
}
|
||||
|
||||
// --- TransformRow ---
|
||||
|
||||
func fixedColumnCfg() *config.DatasetConfig {
|
||||
return &config.DatasetConfig{
|
||||
Columns: &config.ColumnMap{HoTen: 0, NgaySinh: 1, SoBaoDanh: 2, DiemThi: 3},
|
||||
Validation: *defaultValidation(),
|
||||
}
|
||||
}
|
||||
|
||||
func cells(vals ...string) []reader.Cell {
|
||||
out := make([]reader.Cell, len(vals))
|
||||
for i, v := range vals {
|
||||
out[i] = reader.Cell{Str: v, IsEmpty: v == ""}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func TestTransformRow(t *testing.T) {
|
||||
row := cells("Nguyễn Văn Đức", "04/04/1999", "51002167", "Toán: 8.5 Ngữ văn: 7")
|
||||
got, err := TransformRow(row, fixedColumnCfg())
|
||||
if err != nil {
|
||||
t.Fatalf("TransformRow: %v", err)
|
||||
}
|
||||
if got.HoTen != "Nguyễn Văn Đức" || got.HoTenAscii != "nguyen van duc" {
|
||||
t.Errorf("ho_ten=%q ascii=%q", got.HoTen, got.HoTenAscii)
|
||||
}
|
||||
if got.SoBaoDanh != "51002167" {
|
||||
t.Errorf("so_bao_danh = %q", got.SoBaoDanh)
|
||||
}
|
||||
if got.NgaySinh == nil || *got.NgaySinh != "04/04/1999" {
|
||||
t.Errorf("ngay_sinh = %v", got.NgaySinh)
|
||||
}
|
||||
// 2016-only columns are never populated on the fixed-column path.
|
||||
if got.TenCumThi != nil || got.GioiTinh != nil {
|
||||
t.Error("ten_cum_thi and gioi_tinh must stay nil on the 2017 path")
|
||||
}
|
||||
if got.Scores["toan"] != 8.5 || got.Scores["ngu_van"] != 7 {
|
||||
t.Errorf("scores = %v", got.Scores)
|
||||
}
|
||||
}
|
||||
|
||||
// TestTransformRowEmptyNgaySinhBecomesNil ports transform.rs:179-183.
|
||||
func TestTransformRowEmptyNgaySinhBecomesNil(t *testing.T) {
|
||||
got, err := TransformRow(cells("A", "", "1", ""), fixedColumnCfg())
|
||||
if err != nil {
|
||||
t.Fatalf("TransformRow: %v", err)
|
||||
}
|
||||
if got.NgaySinh != nil {
|
||||
t.Errorf("empty ngay_sinh should be nil, got %q", *got.NgaySinh)
|
||||
}
|
||||
}
|
||||
|
||||
// TestTransformRowShortRowYieldsEmptyFields ports the unwrap_or_default()
|
||||
// behaviour at transform.rs:163-176: a row shorter than the configured indices
|
||||
// yields empty strings rather than an error.
|
||||
func TestTransformRowShortRowYieldsEmptyFields(t *testing.T) {
|
||||
got, err := TransformRow(cells("OnlyName"), fixedColumnCfg())
|
||||
if err != nil {
|
||||
t.Fatalf("TransformRow: %v", err)
|
||||
}
|
||||
if got.HoTen != "OnlyName" || got.SoBaoDanh != "" || got.NgaySinh != nil {
|
||||
t.Errorf("got ho_ten=%q sbd=%q ngay_sinh=%v", got.HoTen, got.SoBaoDanh, got.NgaySinh)
|
||||
}
|
||||
}
|
||||
|
||||
// TestTransformRowDiemThiIsNotTrimmed pins an asymmetry that is easy to
|
||||
// "tidy away": ho_ten, ngay_sinh and so_bao_danh are trimmed through the closure
|
||||
// at transform.rs:164-168, but diem_thi is read raw at :172-175.
|
||||
func TestTransformRowDiemThiIsNotTrimmed(t *testing.T) {
|
||||
row := cells(" A ", " 01/01/2000 ", " 123 ", " Toán: 5 ")
|
||||
got, err := TransformRow(row, fixedColumnCfg())
|
||||
if err != nil {
|
||||
t.Fatalf("TransformRow: %v", err)
|
||||
}
|
||||
if got.HoTen != "A" || got.SoBaoDanh != "123" {
|
||||
t.Errorf("trimmed fields wrong: ho_ten=%q sbd=%q", got.HoTen, got.SoBaoDanh)
|
||||
}
|
||||
if got.NgaySinh == nil || *got.NgaySinh != "01/01/2000" {
|
||||
t.Errorf("ngay_sinh = %v, want trimmed", got.NgaySinh)
|
||||
}
|
||||
// Untrimmed diem_thi still parses — the regexes are unanchored.
|
||||
if got.Scores["toan"] != 5 {
|
||||
t.Errorf("scores = %v", got.Scores)
|
||||
}
|
||||
}
|
||||
|
||||
// TestTransformRowRequiresColumns: the fixed-column path is only reachable when
|
||||
// the config has a columns: mapping. Rust panics via .expect() (transform.rs:162);
|
||||
// Go returns an error instead.
|
||||
func TestTransformRowRequiresColumns(t *testing.T) {
|
||||
cfg := &config.DatasetConfig{Validation: *defaultValidation()}
|
||||
if _, err := TransformRow(cells("A", "B", "C", "D"), cfg); err == nil {
|
||||
t.Fatal("TransformRow without a columns: mapping must return an error")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,177 @@
|
||||
// Package writer handles SQLite output: DDL setup, INSERT OR REPLACE, VACUUM
|
||||
// and the stats block — a port of parser/src/writer.rs.
|
||||
//
|
||||
// Every dataset writes the same canonical table (internal/schema), so there is
|
||||
// exactly one insert path. Columns a dataset carries no data for bind NULL.
|
||||
//
|
||||
// The stats lines are reproduced verbatim because they are the operator-facing
|
||||
// output of the build, and docs/deployment-guide.md points at the per-file row
|
||||
// counts for troubleshooting.
|
||||
package writer
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/schema"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/sqlitedb"
|
||||
"github.com/tiennm99/thptqg/go-parser/internal/transform"
|
||||
)
|
||||
|
||||
// OpenDB deletes any existing database at dbPath, recreates it, and executes the
|
||||
// canonical DDL.
|
||||
//
|
||||
// Deleting the file rather than issuing DROP TABLE mirrors build-lib.js:54 via
|
||||
// writer.rs:24-30. A consequence worth knowing: a concurrent reader sees the file
|
||||
// vanish mid-rebuild rather than a transactional swap.
|
||||
func OpenDB(dbPath string) (*sql.DB, error) {
|
||||
if _, err := os.Stat(dbPath); err == nil {
|
||||
if err := os.Remove(dbPath); err != nil {
|
||||
return nil, fmt.Errorf("remove existing db %s: %w", dbPath, err)
|
||||
}
|
||||
}
|
||||
if parent := filepath.Dir(dbPath); parent != "" && parent != "." {
|
||||
if err := os.MkdirAll(parent, 0o755); err != nil {
|
||||
return nil, fmt.Errorf("create %s: %w", parent, err)
|
||||
}
|
||||
}
|
||||
|
||||
db, err := sql.Open(sqlitedb.DriverName, dbPath)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("open db %s: %w", dbPath, err)
|
||||
}
|
||||
if _, err := db.Exec(schema.DDL); err != nil {
|
||||
db.Close()
|
||||
return nil, fmt.Errorf("execute DDL: %w", err)
|
||||
}
|
||||
return db, nil
|
||||
}
|
||||
|
||||
// Inserter wraps a prepared INSERT statement.
|
||||
//
|
||||
// Rust calls conn.execute(INSERT_SQL, ...) per row (writer.rs:73), re-preparing
|
||||
// each time. Preparing once is a performance choice, not a parity requirement —
|
||||
// the SQL and its bindings are identical either way.
|
||||
type Inserter struct{ stmt *sql.Stmt }
|
||||
|
||||
// Prepare compiles the canonical INSERT against tx.
|
||||
func Prepare(tx *sql.Tx) (*Inserter, error) {
|
||||
stmt, err := tx.Prepare(schema.InsertSQL)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("prepare insert: %w", err)
|
||||
}
|
||||
return &Inserter{stmt: stmt}, nil
|
||||
}
|
||||
|
||||
// Close releases the prepared statement.
|
||||
func (i *Inserter) Close() error { return i.stmt.Close() }
|
||||
|
||||
// Insert binds one parsed row and executes the INSERT.
|
||||
//
|
||||
// Parameter order is IdentityFields then ScoreFields. Subjects absent from
|
||||
// row.Scores — and the two identity columns only the 2016 layouts populate —
|
||||
// bind NULL.
|
||||
func (i *Inserter) Insert(row *transform.ParsedRow) error {
|
||||
args := make([]any, 0, schema.ParamCount)
|
||||
args = append(args,
|
||||
row.SoBaoDanh,
|
||||
row.HoTen,
|
||||
row.HoTenAscii,
|
||||
nullableString(row.NgaySinh),
|
||||
nullableString(row.TenCumThi),
|
||||
nullableString(row.GioiTinh),
|
||||
)
|
||||
for _, field := range schema.ScoreFields {
|
||||
if v, ok := row.Scores[field]; ok {
|
||||
args = append(args, v)
|
||||
} else {
|
||||
args = append(args, nil)
|
||||
}
|
||||
}
|
||||
if len(args) != schema.ParamCount {
|
||||
return fmt.Errorf("built %d params, want %d", len(args), schema.ParamCount)
|
||||
}
|
||||
if _, err := i.stmt.Exec(args...); err != nil {
|
||||
return err
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func nullableString(s *string) any {
|
||||
if s == nil {
|
||||
return nil
|
||||
}
|
||||
return *s
|
||||
}
|
||||
|
||||
// Stats carries the counters the build loop accumulates.
|
||||
type Stats struct {
|
||||
SourceRows uint64
|
||||
Skipped uint64
|
||||
Errors uint64
|
||||
}
|
||||
|
||||
// Finish runs VACUUM and prints the stats block.
|
||||
//
|
||||
// VACUUM must run AFTER the transaction commits — SQLite refuses it inside one.
|
||||
//
|
||||
// The wording branches on datasetLabel, which Rust derives from the input
|
||||
// directory's basename (main.rs:98-101). That makes the output depend on a
|
||||
// filesystem path rather than on config; it is reproduced here for parity, and
|
||||
// the caller passes the label explicitly so tests are not at the mercy of a
|
||||
// temp-directory name.
|
||||
func Finish(db *sql.DB, dbPath string, st Stats, datasetLabel string, isOld2 bool) error {
|
||||
if _, err := db.Exec("VACUUM"); err != nil {
|
||||
return fmt.Errorf("vacuum: %w", err)
|
||||
}
|
||||
|
||||
var dbCount int64
|
||||
if err := db.QueryRow("SELECT COUNT(*) FROM student").Scan(&dbCount); err != nil {
|
||||
return fmt.Errorf("count rows: %w", err)
|
||||
}
|
||||
|
||||
insertable := st.SourceRows - st.Skipped
|
||||
|
||||
fmt.Println()
|
||||
if isOld2 {
|
||||
fmt.Printf("Source non-blank data rows: %d\n", st.SourceRows)
|
||||
fmt.Printf(" skipped (empty/non-numeric SBD): %d\n", st.Skipped)
|
||||
} else {
|
||||
fmt.Printf("Source data rows (post-header): %d\n", st.SourceRows)
|
||||
if containsOld(datasetLabel) {
|
||||
fmt.Printf(" skipped (empty/non-numeric SBD): %d\n", st.Skipped)
|
||||
} else {
|
||||
fmt.Printf(" skipped (empty/invalid): %d\n", st.Skipped)
|
||||
}
|
||||
}
|
||||
fmt.Printf(" insertable: %d\n", insertable)
|
||||
fmt.Printf(" insert errors: %d\n", st.Errors)
|
||||
fmt.Printf("DB rows (distinct SBD): %d\n", dbCount)
|
||||
|
||||
if !containsOld(datasetLabel) && st.Errors == 0 {
|
||||
gap := int64(insertable) - dbCount
|
||||
if gap == 0 {
|
||||
fmt.Println("Audit: OK — every source row made it in.")
|
||||
} else {
|
||||
fmt.Printf("Audit: %d row(s) collapsed (duplicate SBDs overwriting).\n", gap)
|
||||
}
|
||||
}
|
||||
|
||||
var size int64
|
||||
if fi, err := os.Stat(dbPath); err == nil {
|
||||
size = fi.Size()
|
||||
}
|
||||
fmt.Printf("Size: %.1f MB\n", float64(size)/1024.0/1024.0)
|
||||
return nil
|
||||
}
|
||||
|
||||
func containsOld(label string) bool {
|
||||
for i := 0; i+3 <= len(label); i++ {
|
||||
if label[i:i+3] == "old" {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,126 @@
|
||||
#!/usr/bin/env node
|
||||
/**
|
||||
* Build the SQLite database for one or all datasets, verify it, then gzip it.
|
||||
*
|
||||
* The dataset list comes from src/datasets.js so it is written in exactly one
|
||||
* place.
|
||||
*
|
||||
* Output goes to .build/public/db/ — the directory Vite copies as its publicDir.
|
||||
* Only the .gz survives: shipping a 100+ MB uncompressed database is made
|
||||
* structurally impossible rather than left to a cleanup step.
|
||||
*
|
||||
* VERIFICATION IS THE POINT OF THIS SCRIPT, not an extra.
|
||||
*
|
||||
* Until now nothing between the parser and the public site asserted that a
|
||||
* database actually had data in it. The parser logs a file-level failure and
|
||||
* continues, returns success regardless, and finishes cleanly even at zero rows;
|
||||
* this script gzipped whatever it got; and scripts/assemble-site.js only greps
|
||||
* *filenames* for stray .db files. So a reader that silently under-produced
|
||||
* would publish a truncated dataset with green CI and no red signal anywhere.
|
||||
*
|
||||
* The guard below closes that: a build whose row count does not match the known
|
||||
* figure, or whose artifact is implausibly small, fails the pipeline.
|
||||
*
|
||||
* Usage:
|
||||
* node go-parser/scripts/build-db.js # all four datasets
|
||||
* node go-parser/scripts/build-db.js 2017-old # just one
|
||||
*/
|
||||
|
||||
import { execFileSync } from "node:child_process";
|
||||
import { mkdirSync, rmSync, existsSync, statSync } from "node:fs";
|
||||
import { dirname, resolve } from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
import { DatabaseSync } from "node:sqlite";
|
||||
|
||||
import { DATASET_IDS, DATASETS } from "../../src/datasets.js";
|
||||
|
||||
const ROOT = resolve(dirname(fileURLToPath(import.meta.url)), "../..");
|
||||
const BIN = resolve(ROOT, "go-parser/bin/xlsxread");
|
||||
const OUT_DIR = resolve(ROOT, ".build/public/db");
|
||||
|
||||
/**
|
||||
* Known-good row counts, from docs/data-pipeline.md.
|
||||
*
|
||||
* The inputs are frozen historical exam results, so these are exact, not
|
||||
* approximate. A deviation of even one row means something changed that nobody
|
||||
* intended — treat it as a build failure, not a warning.
|
||||
*/
|
||||
const EXPECTED_ROWS = {
|
||||
"2016": 877461,
|
||||
"2017": 861068,
|
||||
"2017-old": 847348,
|
||||
"2017-old2": 679764,
|
||||
};
|
||||
|
||||
/** A gzipped database far below its usual size means a truncated build. */
|
||||
const MIN_SIZE_RATIO = 0.9;
|
||||
|
||||
const requested = process.argv.slice(2);
|
||||
const unknown = requested.filter((id) => !DATASET_IDS.includes(id));
|
||||
if (unknown.length) {
|
||||
console.error(`unknown dataset(s): ${unknown.join(", ")}`);
|
||||
console.error(`known: ${DATASET_IDS.join(", ")}`);
|
||||
process.exit(2);
|
||||
}
|
||||
const targets = requested.length ? requested : DATASET_IDS;
|
||||
|
||||
if (!existsSync(BIN)) {
|
||||
console.error(`parser binary not found at ${BIN}`);
|
||||
console.error("run: npm run build:go");
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
mkdirSync(OUT_DIR, { recursive: true });
|
||||
|
||||
for (const id of targets) {
|
||||
const db = resolve(OUT_DIR, `${id}.db`);
|
||||
|
||||
// execFileSync throws on a non-zero exit, so a parser failure aborts the run.
|
||||
execFileSync(
|
||||
BIN,
|
||||
[
|
||||
"build",
|
||||
"--schema",
|
||||
resolve(ROOT, `parser/configs/${id}.yml`),
|
||||
"--input",
|
||||
resolve(ROOT, `data/${id}`),
|
||||
"--output",
|
||||
db,
|
||||
],
|
||||
{ stdio: "inherit" },
|
||||
);
|
||||
|
||||
// --- guard: the database must contain what it is supposed to contain ---
|
||||
const expected = EXPECTED_ROWS[id];
|
||||
if (expected === undefined) {
|
||||
console.error(`no expected row count recorded for ${id}; add one to EXPECTED_ROWS`);
|
||||
process.exit(1);
|
||||
}
|
||||
const conn = new DatabaseSync(db, { readOnly: true });
|
||||
const actual = conn.prepare("SELECT COUNT(*) c FROM student").get().c;
|
||||
conn.close();
|
||||
if (actual !== expected) {
|
||||
console.error(`\n${id}: row count ${actual}, expected ${expected}`);
|
||||
console.error("Refusing to publish — the build did not reproduce the known dataset.");
|
||||
process.exit(1);
|
||||
}
|
||||
console.log(` ✓ ${id}: ${actual} rows (matches expected)`);
|
||||
|
||||
// -9 without -k: the raw .db must not reach the published artifact.
|
||||
rmSync(`${db}.gz`, { force: true });
|
||||
execFileSync("gzip", ["-9", db], { stdio: "inherit" });
|
||||
|
||||
const gz = `${db}.gz`;
|
||||
const sizeMb = statSync(gz).size / 1024 / 1024;
|
||||
const nominal = DATASETS.find((d) => d.id === id)?.dbSizeMb;
|
||||
if (nominal && sizeMb < nominal * MIN_SIZE_RATIO) {
|
||||
console.error(
|
||||
`\n${id}: ${sizeMb.toFixed(1)} MB is below ${(nominal * MIN_SIZE_RATIO).toFixed(1)} MB ` +
|
||||
`(${MIN_SIZE_RATIO * 100}% of the expected ${nominal} MB)`,
|
||||
);
|
||||
console.error("Refusing to publish — the artifact looks truncated.");
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
console.log(` → db/${id}.db.gz (${sizeMb.toFixed(1)} MB)\n`);
|
||||
}
|
||||
Binary file not shown.
Executable
+34
@@ -0,0 +1,34 @@
|
||||
#!/usr/bin/env bash
|
||||
# Regenerates the reader-fidelity oracle from the Rust/calamine ground truth.
|
||||
#
|
||||
# Emits SHA-256 per input file over the canonical cell dump. Only hashes are
|
||||
# committed: the dumps are real student PII, and parser/tests/fixtures/README.md
|
||||
# establishes that fixtures in this repo carry synthetic data only.
|
||||
#
|
||||
# Requires the Rust parser to still build. Run from the repo root.
|
||||
set -euo pipefail
|
||||
|
||||
OUT=go-parser/testdata/reader-fidelity-hashes.tsv
|
||||
CANON='BEGIN{OFS="\t"}
|
||||
$1=="FILE"{next}
|
||||
$1=="SHEETCOUNT"{print;next}
|
||||
$1=="SHEET"{print $1,$2,$3,$4,$5;next}
|
||||
$1=="ROW"{print;next}
|
||||
$1=="CELL"{print $1,$2,$3,$4,$7;next}'
|
||||
|
||||
{
|
||||
echo "# Canonical cell-dump SHA-256 per input file, produced by the Rust/calamine"
|
||||
echo "# ground truth (parser/examples/dump_cells.rs). The Go reader must reproduce"
|
||||
echo "# each hash exactly. Hashes only - the dumps themselves are real student PII"
|
||||
echo "# and are never committed, per parser/tests/fixtures/README.md."
|
||||
echo "# Regenerate: go-parser/scripts/regen-fidelity-hashes.sh"
|
||||
} > "$OUT"
|
||||
|
||||
for f in data/2016/* data/2017/* data/2017-old/* data/2017-old2/*; do
|
||||
h=$(cargo run --release --quiet --manifest-path parser/Cargo.toml \
|
||||
--example dump_cells -- "$f" /dev/stdout \
|
||||
| awk -F'\t' "$CANON" | sha256sum | cut -d' ' -f1)
|
||||
printf '%s\t%s\n' "$f" "$h" >> "$OUT"
|
||||
done
|
||||
|
||||
echo "wrote $(grep -cv '^#' "$OUT") hashes to $OUT"
|
||||
+304
@@ -0,0 +1,304 @@
|
||||
# Canonical cell-dump SHA-256 per input file, produced by the Rust/calamine
|
||||
# ground truth (parser/examples/dump_cells.rs). The Go reader must reproduce
|
||||
# each hash exactly. Hashes only - the dumps themselves are real student PII
|
||||
# and are never committed, per parser/tests/fixtures/README.md.
|
||||
# Regenerate: go-parser/scripts/regen-fidelity-hashes.sh
|
||||
data/2016/023718c7d3cf7ace3a7116fabb12bd9cdhyduoccantho-1468920829104.xlsx fc3afd8da9ed28b0fc177192fab7b19a06535cb45d86841835d8e8828fc61c1c
|
||||
data/2016/08cdeb3636dcb2adaef829d62968274atravinh-1468902443859.xlsx 71f0e61918c63b95f08a16d502f5e9f2054e3c830bd83c725f1285bb1aabead1
|
||||
data/2016/0b583c9875443e65ffa449d5ca76fae8ninhbinh-1468899778557.xlsx b0c2cfb9ab419442f151be5a014439f7f477fa7b31a9ae3f37b83073276bdb59
|
||||
data/2016/15c85d901a73dd49f2bae71aadcdbb5cninhthuan-1468920071763.xlsx f27ceabc10484c0124c7d03f8343248df05719486af81ea2fb387d09a2d11b31
|
||||
data/2016/167116af6c8a096c2a9307edfa115f77dhcnthucpham-tphcm-1468910578187.xls 9962f437ca047cafb3add0c9eb07f96cf4168d026b7d5947ecd2dc915117ef11
|
||||
data/2016/17843bccde10f362c2b9600bc49cd3f8backan-1468933016076.xlsx ffdabc85670ef7d5522bb33254d6053682b96deb248406e171bd066749c51cae
|
||||
data/2016/178814c33c8b85a77d0fc295fa6d2be1hagiang-1468932847230.xlsx 843d83e11b78862598eb77b1a5448f7b09f0a7f1842687b4f0174e69ae4fff8b
|
||||
data/2016/18055d17e46854ecc6c8be585bcb6e5cbaclieu-1468934144150.xlsx ec257a18650045ec476ebbbdbfed48760b4c51b5aee73fe6456729d7cd8ea365
|
||||
data/2016/1c66db0265bf45df237b37a7bcaeaef3hanoi-1468932790746.xlsx aeb84a0588ce445b852a9328f8a233444a5e5ace6ea0cdff695d96a85cf0015e
|
||||
data/2016/2004a8d225fd87524e324d2d519e87a4daclac-1468899960505.xlsx 455f5c3f17397139e50b71329e936ffc1c76a39dc244ee72f49289b3eeb5a47c
|
||||
data/2016/20db5eaace927596f1806f151649ee6edhtantrao-1468901296040.xlsx b25aa4ab69339c095c20d291614fdcd6569f210b33cdebae955e54945e647010
|
||||
data/2016/216ee4e385dbc236767be899b068bf98dhtaynguyen-1468940490240.xlsx 88346d3fa7087055f4c602bda8456581d721c9f9f2d4e0a18dbc511c9c8288fe
|
||||
data/2016/2208903e1fa12220d100f93e40d5ff3bquangtri-1468933440097.xlsx 49241f1af4fd0eb84e6a2943f3c07d1c0f7bd71c9a8ea407b5adcff194bcbcbd
|
||||
data/2016/22e8617886412abd90ae3d33cc6d7db4dhyduocthaibinh-1468941020145.xlsx 9ec94a6c47932bccbf001c431c0d4c480a861cb9aa3f845d9b6e86d4f2688889
|
||||
data/2016/238adc19e80c0daef089a307af23edcethanhhoa-1468933304999.xlsx 121b6ed30c60a2979db512f55fd2f8e65194c7323b5bf495d8175a256eb8eb2b
|
||||
data/2016/2b33640b0dd154490f2809d96be81797dhdalat-1468902130616.xlsx 201b953d4a19c0a0441693d43be8f37212b1309084428dc6d1367fb2904a92f6
|
||||
data/2016/2fa6d9a308bd11bc09f02d787eef2c15lamdong-1468933821082.xlsx 323b5f6551d35d57a995ef04c5d0d569433775fe6e2d685afc88d43442e3d215
|
||||
data/2016/357c6aaf59aeaa6c11c3e83c595d38cfdhbachkhoa-tphcm-1468939678829.xlsx 873fa6455a85783e4fe54b420c6cbda91803e0fa36356af859236092019f0083
|
||||
data/2016/388fcdfdac18ef6eee6125a2a015714ddhdongthap-1468934545359.xlsx 6db5b9fa0acd509c1b07bd6a03ae07df19487073a3d18495df3e185e85711189
|
||||
data/2016/3c0e56abd8c5334f4e1ff7df3ac23a15dienbien-1468920113820.xlsx cfbadef763402bb8ac1e08b398cb38bebc184a2de7e760c44c2e31ce54349a87
|
||||
data/2016/3daa6ee33fdafadc3c112c18bff18c2ddhthainguyen-1468940268065.xlsx a7d7e20b00b912fdd4658fa73c1fd13efc4a748e01d9bd41ce25d91222e6b29b
|
||||
data/2016/412e62265f35ca0078a0d94af5d1ed8abinhdinh-1468899884192.xlsx 06d144fca25ec641e5d8bec34df77b69fcea1946120d050d7679bf282a77b1bb
|
||||
data/2016/43213f61074068d9177685eb6c3f2d3bdhbachkhoa-dhdanang-1468938494437.xlsx 381a59addee12fd6efc7951282578d177101fce5a07be8621b1143f1c99a0e29
|
||||
data/2016/4334ea476a2b8196d54ba40a341831echvkythuatquansu-1468920342994.xlsx 3705579c6f432e2a8aa053e7f2564f79aad49fde836ad3934d1f9f39375150a4
|
||||
data/2016/487c7287c1c9735722b31d877942becddhcongnghieptphcm-1468983373303.xlsx 2c50ae0fc1992ad8b61b7d52ad02f32159bd649f0426ae00d17583b53427b05c
|
||||
data/2016/4a3fa95a83fe9b3b2fafab309323fd22dhthuyloi-1468901840919.xlsx e75239bd4e311ad36e57d42ffb4c2649ce8e599f1ae0bad505957edd69b09a80
|
||||
data/2016/4ad416ccf014bd5d01f415238566b210khanhhoa-1468933789968.xlsx 7a4451cba4a76acc50362cb94aefee1ff9cb04e6e2461daddc383ff1601c80da
|
||||
data/2016/4b98fdd7300b86a94636e4d15394fc19sonla-1468899319124.xlsx 961407526e9d79d0e2dfcf370eaed7ff7f64d14835f88481ac17361acbea996c
|
||||
data/2016/4c2fd846ed2a0670ba6d1200d1d68e1dspkythuatvinhlong-1468902640438.xlsx 88b634010f5888ed6e954807c31712125a280f432219dc652d790d01ef7c10d1
|
||||
data/2016/4e0ccee19eb64513d23a6dff96ff7f9cvinhphuc-1468919874814.xlsx 52199086d32401b35629d6c4930b024161b0d6d979f4aebc1bef26a391d4b235
|
||||
data/2016/4e9bb192d56c9fa92b3e0ec13d8236c8dhxaydungmientrung-1468940978346.xlsx dd3d9e4ce4e0c385e753a56f16ce03e7de535a3700e2c62dfe3dce201c73f6f9
|
||||
data/2016/505a876dfd1ea321291e29192aebe95ehoabinh-1468983220125.xlsx b17f8a592f67c1ed8408b145a2241b4acb6088825463708b5f38749fedfd21be
|
||||
data/2016/514877c1cfebdc9c2251f651b5f02c7edhspkythuattphcm-1468906568655.xlsx 86cbefa0051faad10505f5d2f5dc4c4690c2c9ccd475fa06713754da9edcc627
|
||||
data/2016/51bb2f9b6e9f565c2d705606fd56958bdhgiaothongvantai-1468939107463.xlsx b2199701dc1126ee509581dd6fd739d9a913ea6a232a55d59a14fced8ab82cc3
|
||||
data/2016/538c9676b70f2509573245d58fba93bcdongnai-1468938241935.xlsx 8ff1d5c71cf63b688ecfac852244b24cb6a4321eccf28fe056139f1527f2fb33
|
||||
data/2016/53e5e20ebd74d84741513b703fbd13bfdhnonglam-dhthainguyen-1468939023637.xlsx b3ac4319785a8f33065f4684b947598733050a916ae9cabaa5684023d42f9163
|
||||
data/2016/5465cca59495dcd9a48f689e2d38b85ephutho-1468902409682.xlsx 44873acca3ede545505c531b921605fc7b8372c32bff3886572ca072966c898e
|
||||
data/2016/54bbcf865c137d58cf62af15733bd376dhsupham-dhdanang-1468938526558.xlsx 7e846a06d99ff5b71dbeb3db833d748c7675a17901e3d41bac84a9023e16601d
|
||||
data/2016/575df8a4df74e87abb8f55d807324374dongthap-1468933902699.xlsx db44859e1fce523fb86907b4928c84f92036e696bdf14a337b1f25f69664389f
|
||||
data/2016/5aede4951e7af8c1e1056c747a553386hvnongnghiepvn-1468900610577.xlsx 76e75a3d59c5e850f62191dfa3ce65b6d1347b4eaa4f4870c687e94aea5087fd
|
||||
data/2016/5b16d7156a03bb7ed0618796bae87bc3cantho-1468933954714.xlsx 6d9b5f906fd7b5a66390e42d51f06a3445724e92f4a67cb116ab74359d8693e3
|
||||
data/2016/5ca241aa1de8d25a3e6ff6db75234025dhcantho-baclieu-1468907071460.xls 290c7e658c3b218befd5b111ad89d8c03f5e5dc1a1815e730da8bc1308660b6c
|
||||
data/2016/5fb3a1b66aff077ee3da482be08cd391thainguyen-1468902797020.xlsx 6df75f210b0dc269ceb4ddbbe323640cc32d75157096942ff4eb7565157d4d7a
|
||||
data/2016/5fd13ae5287cbf0149bd733b507e0608dhkhxhnv-dhqgtphcm-1468920630193.xlsx 311b90c29f089ffc7bfdc7e5b9c7fdcefaa803bb6d785189cebb6b170b61ee1e
|
||||
data/2016/6375fd13557bee08d20f1f3e46dac130dhnhatrang-coso1-1468901350425.xlsx 05dc9a62f844c3a73f0a781066ba9fc3f3bf5faa970cb1b3f0f0f0487769488e
|
||||
data/2016/64d2635efd30a0d0db4f7c30b9735261dhvinh-1468940230027.xlsx 0e2fc0f487f3963cd063bcb3fac067d2af1ef5a73e5d3def3523e1ecc30851e2
|
||||
data/2016/64e44466ffeff9a587260795644805d0dhkhtunhien-dhqgtphcm-1468921015146.xlsx 4f67d25acd11182b60d8c938ee49102fb30e8c96c85e6ccecd260fca0f5d53aa
|
||||
data/2016/6960b6fa493e1365a183af461e945389thaibinh-1468920019273.xlsx 6f10581918d96d82a99eab4c2729672836e091265c4518996f4aad5687038f3a
|
||||
data/2016/698b930ab7697a0672bbc39168024c9cgialai-1468933753098.xlsx 3be46dfad3d9d2c81557e1bc4103622bb7d9e9ba4fa18b25d2a4b355bf617264
|
||||
data/2016/6badc7cfa258b59ad67e2b64f713b95fthuathienhue-1468983258438.xlsx 1dccf8266a62a50911046837d01cd07f6085d43dc1c3cc953a35bf033c4991b3
|
||||
data/2016/70bf8cd2b6e87523219e55149a8524c5dhquynhon-1468938878328.xlsx 874dfbb664c72a3b075be56adc386dba2ac9aea4be2bb7a11d250b67d67b964d
|
||||
data/2016/74145979ee30186bf4e40ee3e6a3ac74bentre-1468934010040.xlsx 6db50bd7918d17ab04bd4a8a30859719bb1f3e7c5aa6e3bbe8d4f49d4c2bcf49
|
||||
data/2016/7418e83ff9d07b146216a3beafa927bbdhtravinh-1468902716174.xlsx eadf00e204586e628a12390d335627c65cec48cb0e6ed53f5cc383bac73de402
|
||||
data/2016/7435bb054e8e4ec34c173af1b6ad8402kiengiang-1468900147395.xlsx 271fa0d823b0d07f32c0fb9b885513e465eeb4880e8fbcc72eedd20759c7d11d
|
||||
data/2016/74e6cd0dc78a3fb7dcf56ae57283b9b9dhtaybac-1468940302052.xlsx 9b0a007cb2f55d5c2a629d1c9bfa7f16547e3811f3d5c256bd1377b2013e7856
|
||||
data/2016/76f19098881e9c79f66ac4f1b42a8a55dhsuphamhanoi-1468901144203.xlsx 1a955af97d97f57d9dc20b5c406ca146753760e03050c02c74bef4d62a39ec21
|
||||
data/2016/778ab074ed6a95d209fa901fd5d68843dhhungvuong-1468901964127.xlsx 0add20e99589c83ec61e21582a965692148d899b4df0182c4a4d6d2ce7696109
|
||||
data/2016/7afd5d7fdff7c170a6342cd8f9474b22langson-1468932984587.xlsx 8f4014d1d2bf4e02b3136d6790297af8352f8148dfecace502c4cbaa59b6ca7d
|
||||
data/2016/7b5d74564ab9f6e0e137f6d161fab6f0dhxaydung-1468940561682.xlsx b95d40c1b9497224dd8d59f6e7aeea31a618deba14e43e57539cf32ae748e61d
|
||||
data/2016/7cfc91bd6a1f5ae9d7917e920f469b90haiduong-1468899640007.xlsx 5abbd89e8130078fe42ecee222327add52e4dbd362c2f96576d9e6ed7c441c19
|
||||
data/2016/7e0d0b6fb981d407cffdb79d724e8aecbacninh-1468919984998.xlsx 3f32fd1af47a2b1f92175bcc55d95c48be1de1d9544e06f440ac36a8491b1a14
|
||||
data/2016/7f12b90225e91f9d00aac3bcdf9974e1nghean-1468920043121.xlsx a4c5a99a9faa75e36b809d1bfa4b7caab706bde26649256ed1248f429e59abe6
|
||||
data/2016/7f1cbe7f35d1f539ca817a0ab8bd00c2dhcantho-haugiang-1468907167921.xls 434417e4765c5b225570c4a19b655c5ac957707209af398f7277a0a9ee90f2b6
|
||||
data/2016/7f4362ef567cdf74495368e0fc11072avinhlong-1468934067981.xlsx de063663b5f31f525a8a07958d436c8e10805d9ec7b27f9bff2b4a2c7ae559f0
|
||||
data/2016/82338a4a93ea1e0cde26e0061066ca8adhcongnghiephanoi-1468920154757.xlsx fed3686286093866bfe03b67d7f72ab7c3aaa1a5ce92d503d1ca24582addee31
|
||||
data/2016/86a692bebf9b0b7c58adb9876b9befd5dhnganhangtphcm-1468902244566.xlsx 5a2b2e5e63bb8c12f0e16b802b51fac470650fc050acd9e846275e82a5d33fbc
|
||||
data/2016/87ffaf3197ba20f0919f98a0b3f33914dhkhoahoc-dhhue-1468938729472.xlsx 1b46f645bb5c491ba9e9e490f36f1692b055e2d29878dfff1fa4db94fabad72f
|
||||
data/2016/8916da7b3c4d58cffc1dd29dc7a8aabfdhkinhte-dhhue-1468938560205.xlsx 332aa93e333430021755667c10795749db3ec3f7c7d3afc4d70a48fbd00bc820
|
||||
data/2016/8af8f1b29b6a40e26ccb4a2e19e8eededhkythuatcongnghiep-dhthainguyen-1468901919568.xlsx 94347288ba444f19a160868649528dc29633f4eb660aad74ea9799ee79935432
|
||||
data/2016/8f948c2a5ab49a3c5fd03b6ea04887a7dhkinhtequocdanfix-1468902083782.xlsx f8e4c450509b176716429055e2ed72dd340af7ad5a5704db2d0c60818deae548
|
||||
data/2016/919276a0348634459f8937dfcd6c7129dhhue-1468938792115.xlsx cb91ae55927a40d0eef9ef1bc2bd23a758528de6b36496505dba75e1cccdb849
|
||||
data/2016/92045214159cb0335b2836aa1c4bb25ahungyen-1468899732002.xlsx 74536aad02570a17c83933c46d10ce9d2300a5e61a37fa2a67e5d99cffb17632
|
||||
data/2016/964c4368131f6be6058fabb1caefba1chaugiang-1468934247324.xlsx 1df076df14fcb51daffa600dca05064be61269fe19f33dc14d03c8f43d8ee1ca
|
||||
data/2016/96c95c85c5549a06f2e92ffcacf8ff21hvtaichinh-1468934470924.xlsx 9a7a3841700b27e4265eef8def9addf17406d08c4f007462cda9335f5369c4ff
|
||||
data/2016/97aa741a4c2b1276559eb021779daa32yenbai-1468933116055.xlsx aa79b6e84e95f58704b715b701449507dbd079f9bee9900f20e6c90baf299952
|
||||
data/2016/9ba098a7913ad837a66e173b3fc41fd7dhngoaingu-dhdanang-1468902750424.xlsx 8864a16b969ea34dc1f7b8762b8039e17df1599bfea7ca5914dc518983fd36fe
|
||||
data/2016/9c730c43bf4309171528c3853ed53b19quangbinh-1468933376032.xlsx b637ecd24a240e4a73f291cc4c9faef42379c04d40d1952cf84e54e51eb8cdf1
|
||||
data/2016/9d78f6dce7965fb4b7a337279920d4b1dhtonducthang-1468902168623.xlsx 1c4f32eb3ed9aa831d2abe41998f5c157ae5ca3d9735b87bb101913870f4c4cc
|
||||
data/2016/9e0e564ae801e347bd7649da8b89ad90dhluathn-1468920377978.xlsx 1a1b4b0bed838f2c6043d789dae4ab7f96a904b886cd5990a98f7be3b960caf5
|
||||
data/2016/a11e707e6ed5fe4d29a8e32380d7558cdhsphanoi-2-1468902006284.xlsx 9b8060368d09cf89c1e88001a642d1aec413e4b3f8646c8360a0459d3704dfd3
|
||||
data/2016/a31b3446bae5a17c517348d8c38c5b7equangninh-1468919920197.xlsx 8b284c18aab150349d49fe596cb96a0a838e8c5618eb1f2d7dd5f0053430c71d
|
||||
data/2016/a3aa106104f114c6426deb2d5e5da5fdlaichau-1468932954079.xlsx 0697ae494d6a17e188138e2c9f3d412a1bbd8a462e2eb82060ffff25e9982506
|
||||
data/2016/a8ff80d9478747e36d21ec95cc9ecb8fsoctrang-1468902686409.xlsx 3fb1d35e7647d0581f08d2f56b4d7ccd5fab1ab125bdcb3b83c64c65d540a53b
|
||||
data/2016/a995de3f7a7055cf10b2f002f0e194f7dhsptphcm-1468920792291.xlsx 9f25b647f9d4d714772ffcb1cba40fa1b338500bd08e39352704de6a9810a8df
|
||||
data/2016/ab2f259ef55cb08b3a436d3f15d42cc7dhsupham-dhhue-1468934431828.xlsx 96f40bdf397bdf98e596701bfed957d18483ec0affb7e945f6e8b69453b7015c
|
||||
data/2016/af6828f837a80e03d77a65f3d310ec60dh-tai-chinh-marketing-1468924711596.xlsx 297b71a7760cfff612a290f31082f22647c0d7a8528dbe052356885fc620dc5e
|
||||
data/2016/b2552e20c9ff4b237244dfbf2e278a47dhthuongmai-1468902876264.xlsx 8c7f9d58fb5ffdfafb1431e9fd51347d338ccba14f547aa26917a2692a515982
|
||||
data/2016/b2a739584a8861502573d1c6372fd6e9dhnonglamtphcm-1468939555326.xlsx d8c8b6e10d3a44f8a429d2a718af55c853bd508dcb79d13522f1e40d9d6e96e7
|
||||
data/2016/b34c777942ca8a4de91f23a35cce1c6bdhbachkhoahn-1468901491486.xlsx b5135f12ac53b3f513b46078a89b6a2663089e6bb3ad97d48a13312b5ac69ca6
|
||||
data/2016/b35a3f26a162bec3740e2da5639327afdhcantho-1468907131146.xls d114bf1c070bb165bd45a64ec79495c764805ee989c740ed95fad06eabd69ec4
|
||||
data/2016/b42fd58ead4c54279d7df5f53160a851dhdanang-1468920187203.xlsx 9ce9d1b6f494866f6a39e68babe27b3a9562cd1515477f5f32a8e2a8a57bf13e
|
||||
data/2016/b7a0c1ecce7c8445b3a1021091a5e654bacgiang-1468983153932.xlsx a53266ee75a8d49fe5a017125e198a1368a9bc9259dabe8cdf1cad11551dd439
|
||||
data/2016/bacd1b2133b37d0ea9fbf46445265036dhtiengiang-1468940391597.xlsx bc90ae7bf5611ae0bb3c321e5185afcb530e1331c5a920a40b85b30b47eb1006
|
||||
data/2016/bd857add0cfe92c1c247469ff2104527dh-hong-duc-1468900427119.xlsx 96b62a2954d2c1541699d491ea931c240a9b06660d5c27a3cd9055433306aaef
|
||||
data/2016/bf259a908ff8a8dcc067a761204459f4daknong-1468934173803.xlsx 2f8fb2dbcd9a06acfd91becba0f728bf56e64f5c9984b40eac060834ea619b44
|
||||
data/2016/c174631d90303276ac63ff38b31157bbdhmodiachat-1468900697349.xlsx c5c4435cfe1e6feede1936fd11e5cd3ff5758b4c4d522df03facf7522a9e7194
|
||||
data/2016/c2cfaaf287fb1e7120894e545771b027dhkhoahoc-dhthainguyen-1468900365545.xlsx 9a35e56e631ba5d1056fb149ea01a41b4e62ff075937fa735666d192b1714da2
|
||||
data/2016/c5c28959fcd0b4e39dd9fd538db5448adhgtvttphcm-1468939150821.xlsx 17a032a01a9ebe9f577af8fa4666e8700e9ab4e97166703cf6a7a0ff9dea4e9b
|
||||
data/2016/c5ecc58a731fd1a54b7e6cc3bc99449ddhlamnghiep-1468939300790.xlsx 8f5f7bf26bafb3a8dd6988ab9858dbf03a14d200a23ed5299a8a5411c84d7f1b
|
||||
data/2016/c6ef16bc9f75ce0fd16c1229afad4e71dhhanghai-1468906457479.xlsx 8e43d5e1782e366344c8d615666541931319475e02df0be29ae290c7ab2abda1
|
||||
data/2016/cc57dc986acb525086b86e85df54949dkontum-1468933632113.xlsx 2397d3b0fca3888ff34d4166a9d3da70dcd6986e6405bd556471d46fff0ab551
|
||||
data/2016/cdc912664d7217c75eb9b67c6ef3797adhkiengiang-1468901222446.xlsx 460eb7e1ce79dde123fe53a18034c5617f0616ad1abce6dedfcd0e4a2957a28d
|
||||
data/2016/d035f470c13adbf5785797c6ccdcb204laocai-1468919805337.xlsx d23348d36bf6fa886d50e30507d4eada114a58f52cb6e615c119980d7753b5ad
|
||||
data/2016/d5ecd2574b5ab36c62dffad435262cb8tuyenquang-1468902296807.xlsx 3acb7e4ba9a3c210906b9a32372db747ba2d95d25d276ba3ba45046b941dbbef
|
||||
data/2016/d652d24e8bd4fe6a48d634b116f37c27dhluattphcm-1468939398928.xlsx f519010d9b2631457c9b7784bb77f07979907321199b98b3f97cb45e5b20c2ff
|
||||
data/2016/d75282bff0a4d8cd1421c3fb74f789c3dhsaigon-1468900867561.xlsx ee7734bac44228a58a11b85d490cbd5da27d7a6e07b22c1eb9921a123b3e6acd
|
||||
data/2016/ddb631f833af894200ef10cd408d2ef0caobang-1468932877769.xlsx 2bbda20bab1190f42b44d0eeabcf2be7f2203d4e5f8a2ad3d80365770a8e2d9e
|
||||
data/2016/deca197a916a8f633f6894b4e13c3ca4dhhaiphong-1468901178368.xlsx 221fe731a4b2de02b3952572baf48d66544fddf87dfdd7851b740325c0864d24
|
||||
data/2016/dfa63c5b68a8f638379ca0e9c5a1c556namdinh-1468933234061.xlsx 446cd9626fe197f6509efb01ca5b4cb26730907b5e114c4b9c07399a489b45d1
|
||||
data/2016/e693bc9de67acc916032c6a3b2fc9eedhanam-1468933146526.xlsx b1ef9283b500e965205281d3875eff011be1d418312d187831925860b1ddab74
|
||||
data/2016/e80dc028d53fa49ee247391a0653dee5quangngai-1468933592400.xlsx 9c6ba937e8ad772a3782baf9c40ef781121c27ffdeb4519170cbf2b81545a146
|
||||
data/2016/ec765c32190773f5a18cd8202a5eaf0dquangnam-1468933501467.xlsx a566488b611ca04c705f385c47b871792ce75f7c1ca793d0528220352d6b49e1
|
||||
data/2016/ec84067268f47d523b1e80b493248377dhngoaithuong-1468920480753.xlsx 5d892f08ec818be3ceb343529c9f0b35e0af0aeb43a229446dccaa87c6d4c846
|
||||
data/2016/ee09c723da4e86cdbe26213413450af2dhsupham-dhthainguyen-1468939061872.xlsx aefac40009116615131c926775e6625eeebc615cd2ff4c7c9af399dd2faab982
|
||||
data/2016/f0aa5a19c2c8c1fde8211341f0e98e26dhkinhtetphcm-1468939229043.xlsx 5ce36388e304cbba01c513f11f54b940f4299061e2e8acfa4c4219ee536a01f2
|
||||
data/2016/f2846894bd5533e02b122a1e3b4333f6hvnganhang-1468939472850.xlsx 142b39b4f77c1d7a648c2397165b315acdd48ae6d354d914dd674a46b3d7160c
|
||||
data/2016/f85f04bb2bbd84457b83ac846728a2a6dhspkythuathungyen-1468920667091.xlsx f8cc571e94720652f1507576ff48616d5331a4ed1403cb95cdc30124a03c3a68
|
||||
data/2016/fc511852659150f6305f6f624445300ddhangiang-1468940189321.xlsx 4e16ea6a122e5bc0c62c7e6e55d9ede3a41242dd3e034d92a450bf29e8fff197
|
||||
data/2016/ff877788e43bd84f0119dae026996892dhkinhteluat-dhqgtphcm-1468939860894.xlsx 8c9e7b097d0796750995268aa36a9fa00112c0989b513023ef0a53456afd7051
|
||||
data/2017/an-giang.xls 5bde5f6447b9e27540f4f72d9b371b6c94376aeb97af65f4ec5eede43ba9f25c
|
||||
data/2017/bac-giang.xls 9cd5c6d858062a1fb30c477b699be467139f183745e66d134746e043d73579a4
|
||||
data/2017/bac-kan.xls 9642c00d14c54102ce8dc38408676863a87652aee89651113f4ad3e0a25bd1fa
|
||||
data/2017/bac-lieu.xls e364dd65fa068d40e9a390aa3322e1676171af797c5d1f98949277cadab5369e
|
||||
data/2017/bac-ninh.xls 539ed3cb1a2df358ab169e2163face76c949aac2eca0d2acb7ceb51e97811bc5
|
||||
data/2017/ba-ria-vung-tau.xls 78dc6d67bbf3ffcb7e4407a6d950e28e39366ed62617691c9a9db610a5079134
|
||||
data/2017/ben-tre.xls b4cce0c5a443062a5e7aef95785fc56f421c21d5ad22f993cbc30ec28e078f4f
|
||||
data/2017/binh-dinh.xls fa1a39cb5134a8159e71dabea41c9ef2a3ec263642f19bed54ddd17e5a5e3d98
|
||||
data/2017/binh-duong.xls ced0d7469186d9e5fbbb8c4323a4bb6a589abb8ef23f01d8d39f947a575af4ba
|
||||
data/2017/binh-phuoc.xls d1ecf9b66ea294eec24ac0f703c614a2011b7d693aeb08c719cfcb73cae7ca47
|
||||
data/2017/binh-thuan.xls 3cb0f85889ada81cd257f51ae0b0b6414d2ad8f81d7f94f4b1b38010600243c6
|
||||
data/2017/ca-mau.xls c4bba31f7728f7825c20b3c0b708ef977dcf6f56f0dd5c14d9342cea0467e505
|
||||
data/2017/can-tho.xls b5ba59ff7b87c31859686adb98da8ef2e95de84404daff2ac3df1e00b748b57d
|
||||
data/2017/cao-bang.xls 9b0a497ed16840a22af34f7e75e316b00dd60490d6d7f8a2f012507c95fad3e7
|
||||
data/2017/dak-lak.xls aae3e074d91afd98d89cc9d07bc1f8debe4559bd78c409cc5f50f2670a2c337a
|
||||
data/2017/dak-nong.xls f7e30b4cdfe951086aeb1c169e7aad88a17bf9fe07aa33f85654e309899f6beb
|
||||
data/2017/da-nang.xls 80963ff2f56d2b6bf331546d08184e325ad0ec7432ba948c0f74845f440e87c8
|
||||
data/2017/dien-bien.xls 6bcb1969ce81ccb94c6ee625456530b9661ee255c00eb598dc195611a46205c8
|
||||
data/2017/dong-nai.xls af248453ea27e879da1b8cfa0f2cd345aa068f93c884d1bc98862db0ca2f9faa
|
||||
data/2017/dong-thap.xls 186a0c4ac6bf838cc496c4604a7e8979712051c71a53622a9ec72af70491e4f7
|
||||
data/2017/gia-lai.xls 3a1b64c8362cad028ee53f4fd18b0923315d919b117df9a2fd83d2f46544db98
|
||||
data/2017/ha-giang.xls a3248393c759a41de424e3ed994d2b9298cb39fda3c6eb9a68f48025c5a2e072
|
||||
data/2017/hai-duong.xls 6634715f5a0cbac6a8d2d18be0df74e0abd328641721692224758c730fd86e51
|
||||
data/2017/hai-phong.xls 7a8a210986de8c562a9128dccfaed1c6fc6c32df0c5e2e8948bee2f73777ba99
|
||||
data/2017/ha-nam.xls f725e212f976af00e68edd6552721e56b30d5fa7d786f548dfc34721ec8b4426
|
||||
data/2017/ha-noi.xls f0f4bc9216e421acf655a40844a4dcb74e871b188cf6d455982ee6e408170ef7
|
||||
data/2017/ha-tinh.xls b0f5bcd8beff7cfe411b7bd88f9e5b9c5a9a0578a55cf18bb003284aca12e57f
|
||||
data/2017/hau-giang.xls 9e8bccd0738bac3d68640768a8d9abc3e7a6a4e97e8c5968039530cf313cafb9
|
||||
data/2017/hoa-binh.xls 51b1e910bfc47404b411cb2ab1ff14b522d8fe54d66a1d50479350273c7be1b5
|
||||
data/2017/ho-chi-minh.xls c4e2f920a5e58fa62bd0913d994d44079e472b955adfc504b734c6773c485f64
|
||||
data/2017/hung-yen.xls 833584c8ea9337734feb37e068bb06232293108d50b18d88f731d30a576cc29a
|
||||
data/2017/khanh-hoa.xls 644ab06abca93d7b966205604313a65d50f3a0a6fb8be509cadb55e0d55fe12e
|
||||
data/2017/kien-giang.xls df452e618544a9b7932d918d26dd2a0f2149f92599609c15a9ac373291cfc440
|
||||
data/2017/kon-tum.xls 7d26ff078adaf51dcf4333b074f8a444d328d8083835bd55b1034c1a63e4eb17
|
||||
data/2017/lai-chau.xls b37cd06914cad19bc0e523d0b1cfe87db692e26dfab327f5f9a7686b2b14dae9
|
||||
data/2017/lam-dong.xls 59cc802a6c4d622f1bb6233f6b9b1a94f342a6eae1c389235383db1767f34720
|
||||
data/2017/lang-son.xls ae376e75bf615483b6cb3a37daa1cea6a7eec4cf3fdd1186801f56d074eb4ce7
|
||||
data/2017/lao-cai.xls 7596c99851c20ff60b95347bc9fae78014873b4c3e1f1f7144c48d9f285f7962
|
||||
data/2017/long-an.xls e52d8e8a4592258e87c5cda04e0907541f60e7e7ef4cb21e9171b65283e0b87a
|
||||
data/2017/nam-dinh.xls 80dcda8b058dcf1a7fdcd7e34d126c44c31b2d4987b52eeae790bf6c38c58e40
|
||||
data/2017/nghe-an.xls 3f32421eb6ee8bc21fcedd4cc1043df1c9c0381abfaa1006c85c3ca6d6209543
|
||||
data/2017/ninh-binh.xls dc3d08a86cdd145e870486773d95a1d896a003b5400ba6ff8456a5512b10b2b8
|
||||
data/2017/ninh-thuan.xls d69b0ed7192e85cff31ad5bca8942fe3c92322268d1ea707e945f9991bd78115
|
||||
data/2017/phu-tho.xls e352a51109a2373c1efbb6726d262fc490bcbf9229ae90bd473d14a4b6eebcc0
|
||||
data/2017/phu-yen.xls 6871aedd90d2c6c87ad6409164a9b26f6f89f97cdd7a1dffbbe4e6879f2348c6
|
||||
data/2017/quang-binh.xls 6c5e9b3c6780a5b70a02dd4b4854861ba148613361c048bfc001706e5bb77a61
|
||||
data/2017/quang-nam.xls 859258cb8c4b7806101039e4e2484e56971f2640a65720a108a0933245bd3157
|
||||
data/2017/quang-ngai.xls 834988a6484fab64bc59fe35a6813038663054553d269d16b7d8200593f468e8
|
||||
data/2017/quang-ninh.xls 1b1c0b1ea1d6aadd8bce5faef6b5fdc5181daaa9bd44df6acbf06c9021a0ade9
|
||||
data/2017/quang-tri.xls 35e401061962c6d80593d684f179e1f82b05f884a6a470d1bc17b2d31151e05f
|
||||
data/2017/soc-trang.xls 668a0edcb624faa954c424616ea6c35d6f06b9e9fc62524723dea01af09490c0
|
||||
data/2017/son-la.xls 96bcfeb35e5bc23ccd7dc6a62a533e21dabeb3145473ecc5a039d970019a1129
|
||||
data/2017/tay-ninh.xls aedf4ccab43e7a1fa985721de0d6053988ad471355c0ccb0b8f9c19af02423e2
|
||||
data/2017/thai-binh.xls b20f66cdeb975a64eed7322b616bd291fe34e08e5fbedad3d729d21da0eba627
|
||||
data/2017/thai-nguyen.xls f0b93db9aab966d393b8a3323204ba7f147ddd0a92746355b81df09bc58267b4
|
||||
data/2017/thanh-hoa.xls 3f803ae67dac24317fa45dac60f36e39dc222ff84d23a0a3dde7751cdcb17646
|
||||
data/2017/thua-thien-hue.xls 71e220ee4ea0287657fa379c587ca93f71f96b4dacc48cef074f8c223e940daa
|
||||
data/2017/tien-giang.xls 2e061f8724e1ea8b997d7e6700ac89a60921771a57dffa7297f4728715811406
|
||||
data/2017/tra-vinh.xls 59507a1d237ac5e29c1033d42b08da6ac9a7505a5045b791e512e91493102ae7
|
||||
data/2017/tuyen-quang.xls 053d293e4acd9dfa6ce9532a047cc64b4d2edee1ff13b387723586c1fae6312a
|
||||
data/2017/vinh-long.xls 6e7584a3047b0ba7487dbb8c71f5b594db220f11b747149d41d31ea97e1e4eab
|
||||
data/2017/vinh-phuc.xls 9753e2d864584e76aef1e690e4e2de98b28240fbba810e614eab9926668d66d2
|
||||
data/2017/yen-bai.xls fe6824a1254116bc6ed372e0a1fd635a816a09e010b7efde9659baca2e143138
|
||||
data/2017-old/10_BinhThuan_RIIW.xls.xlsx 2833e7366ea1a21ea869978a6bb920b033160bb8042fe0bfe9b7573dcfe85409
|
||||
data/2017-old/10_Ca_Mau_BKXT.xls.xlsx 53cb69b13325cb9cc9004b8c89f0cd1c85b5f943e7c78c8b010ddabad3f816c9
|
||||
data/2017-old/10_LamDong_GNFT.xls.xlsx f05b4cdc8a9a8413b29801fbddeb48a7c60d3c6bb2228bbe238113495e89bc98
|
||||
data/2017-old/10_Soc_Trang_XCGJ.xls.xlsx 8e540cef6f2ff387f3e825a5aa3190af394c6f9c7e9dfc66d5f7032ee47f7f14
|
||||
data/2017-old/11_BinhDuong_RYQL.xls.xlsx 11953116054a2889929c3fe05b68e8932c7b79ffd1b85d697afaac63fcdd9adf
|
||||
data/2017-old/11_LaoCai_GMSU.xls.xlsx 5a73becf1bdb8027102faff74035d1df71e13c1cac39cc38ec6efaab6699ca52
|
||||
data/2017-old/12_BenTre_DKWF.xls.xlsx 579c0d94b20440f1c60a94fb0c762b71dd1c7719c8aa4db0b9418ecaba9b605e
|
||||
data/2017-old/12_LongAn_ZZUK.xls.xlsx c4cf569706bfc30582d774ba9e699dcbc4bdc734f46bf6493df58eaecc37a755
|
||||
data/2017-old/13_NamDinh_ESEL.xls.xlsx 7e5256cb095262fef5399092f8f843c1ee82e117bca8e6571e363646bb1170a6
|
||||
data/2017-old/13_TraVinh_LKUJ.xls.xlsx 2b7321d68a4a910c4a506d7ca207531cfbdd9619f1d8ba0ce475baf1b91043aa
|
||||
data/2017-old/14_NgheAn_BSLY.xls.xlsx 034234a8a5461910f8beff3977aafa29a6541c8a56c007816118e5c4eb2b7218
|
||||
data/2017-old/15_PhuTho_ABWQ.xls.xlsx 9891c17001cfcc49ccd587419339698ef59b2dd43b5ca7f2765b09a17c26659d
|
||||
data/2017-old/16_QuangBinh_KGEU.xls.xlsx 01e249d45c1269d62ae1977bef0bb192f833941510471b723e57c32adb0eadaa
|
||||
data/2017-old/17_QuangNam_AMTK.xls.xlsx 4c9f05d24d861ec3c82a1c564e4baa1541bd4882785f143459a1a81a6b8359c7
|
||||
data/2017-old/18_QuangNgai_KOFP.xls.xlsx ee9c26f2c16d36dabfde79901a054d2266c28c4acbfd19b628cfe50f8eb03957
|
||||
data/2017-old/19_QuangTri_OMZF.xls.xlsx 6c961bb461274c73b4589de4087624a9ef5edd43c04ea54675148df00f6420e5
|
||||
data/2017-old/1_BaRia_VungTau_HJKG.xls.xlsx cb03f933442a8cea6b1622da4e5f804ab476dff9f3462b404e957b6a2e87b2cd
|
||||
data/2017-old/1_Da_Nang_AHWJ.xls.xlsx 59892fc01b6cb10497679f2bc69f59023215755d2f00fd764b23b895c98be2be
|
||||
data/2017-old/1_Ha_Noi_CVXG.xls.xlsx 2718fe9c2b58843621c0082da5880a4b957266525b83123af02c78fd51d4f4a5
|
||||
data/2017-old/1_Son_La_JIDP.xls.xlsx 0a0e273c5def980406b6316cceb63e5cee53aae56e112dfae7e795261eb1cd06
|
||||
data/2017-old/1_TuyenQuang_JBYF.xls.xlsx 45826ed10248268ac3344d0b592cf2dba6f86278b7402dc60626a76b18fa4d6b
|
||||
data/2017-old/20_TayNinh_ILFA.xls.xlsx 684a1820d02f658904c87c26aa195faa6aceb3d7ca00c674df11a63f9274ab14
|
||||
data/2017-old/21_ThaiBinh_FTVG.xls.xlsx 1a12e43d4ea368b8e3a2e4d881b5d66214a16140219e80e2e72fcb269899b285
|
||||
data/2017-old/22_ThaiNguyen_TLTW.xls.xlsx 29227bdc3204591327e803680525d8621228fc371ce7f802dc667ec70eb8d7b3
|
||||
data/2017-old/23_HaiPhong_HXBV.xls.xlsx e450e2344fbd724555de9892412f521414fc1fcc53e69887c6381dd266b4e92c
|
||||
data/2017-old/24_HCM_XULN.xls.xlsx 49ed22e823607b78920a24f5999eaf1ef8634a141108dfd4bf970ee000aecbf9
|
||||
data/2017-old/2_BacKan_GFVQ.xls.xlsx 77a08fc4edb41d639fcdd798980903519181639db00a360ef0c6c6c2bd88f529
|
||||
data/2017-old/2_Ha_Giang_QNCM.xls.xlsx 6bf4d62413c1b92f88f980e5e3e414a7e25971c7546c6390c7e67a6b0f742b68
|
||||
data/2017-old/2_Ninh_Thuan_VHLY.xls.xlsx 6dfb6517f9ccc3420c222b6a88aabac54f6f25f6b743d993d6a5f28b4fa78ff2
|
||||
data/2017-old/2_Thanh_Hoa_AUFV.xls.xlsx 2f2ce90425c19a9089024061f328f37b9ef580ddadc076ddcf452f549690f4b3
|
||||
data/2017-old/2_VinhPhuc_GUDK.xls.xlsx fc0e4952445ebda1af28f7330aa808c032f2e3a7a3c03dfbb9692bf6a248867a
|
||||
data/2017-old/3_BacGiang_TOIF.xls.xlsx 8043510ff89f11ba063f898dc1d2c8879a8da9063a420308647e20a2f39d3092
|
||||
data/2017-old/3_BinhPhuoc_YFMU.xls.xlsx 8869970f4e6005381264f3f4bb508c5181f15ae990a75fe904a65b305c0a123a
|
||||
data/2017-old/3_Cao_Bang_CIEY.xls.xlsx 94a06bdb2333a9f44216001ba8d8463700b8cf3e51944b94db11eef7eac8dbd3
|
||||
data/2017-old/3_Dong_Thap_GSXQ.xls.xlsx 0cdd0763fc4f20f1d990621ca9e8dab107a3d4eb7c0ec3e53a4d8c240fbc569f
|
||||
data/2017-old/3_Thua_Thien_Hue_HMDB.xls.xlsx 62e54eebda7bbd542ca63443bb923a81b472e7892e55dce39fd0611fc0f00abd
|
||||
data/2017-old/4_An_Giang_JNOS.xls.xlsx 778ad205bbabb7c4aea3df34325770aa4a1daec0a1683f73f0405ac1cc5aaf9a
|
||||
data/2017-old/4_BacNinh_STLR.xls.xlsx 009c3b7821f55210a86f9ace74c31eb0d5c16aafdb0cac8f2b71cac616abdfec
|
||||
data/2017-old/4_Binh_Dinh_WWZW.xls.xlsx fbacd256d58627377d698462d69aac0e83d124c8c8ff1f5b63ba0bfb5442dd39
|
||||
data/2017-old/4_DienBien_SOJG.xls.xlsx 4931e8d6604013df48b2b469fda3eb05fdf1203b67bbc2ae4ff6d7949a115a24
|
||||
data/2017-old/4_Lang_Son_NYQL.xls.xlsx cc7d2454caadb18740ed6baf6edc62baa47e98b3a907b3ce06f72b9a16b7838a
|
||||
data/2017-old/5_Bac_Lieu_XEKH.xls.xlsx dc1dca45fd5357b45abc930b0ebb95ab589ef80a09bef8b5ad4dab9e00be87e5
|
||||
data/2017-old/5_Gia_Lai_ABZI.xls.xlsx 5b7be25274d628c284c37efb3df0587171cc9273dc3bd8dc824c7bd1b8bccf1a
|
||||
data/2017-old/5_HaiDuong_WWWG.xls.xlsx 1d800d0aa7338fb507c4eeb3e752620d2608c30ac820895ca086325dbaf3643d
|
||||
data/2017-old/5_Hanam_QGJS.xls.xlsx cfc233ac37ab8e263b6cc569815ea2c687da30364f139109032c8f2cac1ad98e
|
||||
data/2017-old/5_Yen_Bai_FAQR.xls.xlsx 4c8915417d8747c291f643ea4b378c87982c397460e0e9e3cc6df19670d6720f
|
||||
data/2017-old/6_Dong_Nai_WOTM.xls.xlsx 9b582a8421c419156c2eab6e532d854736385735f8208f9d390c251f22ac764e
|
||||
data/2017-old/6_Hau_Giang_KWDM.xls.xlsx c11473e3e1f964302c5378c90dd6b9d478eb11397351b7ef74df9ae709c9e908
|
||||
data/2017-old/6_HoaBinh_TPZY.xls.xlsx e1210961dc821642f6951cd6e4a71b5fbc81d5070ee3816b7c8b5dc56bf2b234
|
||||
data/2017-old/6_NinhBinh_IGFT.xls.xlsx 9399b8ab58e2a7a1c0ebf5425359319e567f8a28602349aa587efced72142fc8
|
||||
data/2017-old/6_Quang_Ninh_DQCJ.xls.xlsx 8cf95259b11d48267f710f151c3f1419ccac49d823b78b70a2708c59ab25a4f7
|
||||
data/2017-old/7_HaTinh_DDHD.xls.xlsx d96df9f2c8365dfc73be3e10c268367b93fd00fe92bb39c745d8e353ba75c05b
|
||||
data/2017-old/7_HungYen_LTIK.xls.xlsx e80eb38af8307c8e193a91bdda5368c437693d7669e281b543bbe891a46e6f36
|
||||
data/2017-old/7_Kon_Tum_RSLR.xls.xlsx 41d2adea059449b6e214241414748f10776861d261cbaa1fd213629757deb846
|
||||
data/2017-old/7_Tien_Giang_EFHX.xls.xlsx 00d9bbc70be89f8f7dad6aff990ad36236c6b7244ddeb2e3d9b3d8ed1a6e21d0
|
||||
data/2017-old/8_Can_Tho_RQZM.xls.xlsx 3a1182b26d34fbdfabf4326098ff265609d8f1fe2b766b6eef897db9685d4cd4
|
||||
data/2017-old/8_Dak_Lak_YKPR.xls.xlsx 379b45f8a347bf4842daf9402d84be47426ed7107a64b2e7710bdc8e71eb57be
|
||||
data/2017-old/8_KienGiang_OTOB.xls.xlsx 3e1789b8565a5ae7c187598cb7e9c3261029c06bc7f0cccc4eb503e8e302cb18
|
||||
data/2017-old/8_PhuYen_OGIM.xls.xlsx 91d1f202cba3bf4c57a7f26e7611d3d29faf8dd13de30552cbfc96eab7a451e9
|
||||
data/2017-old/9_DakNong_FBOP.xls.xlsx c611bf7f1d1c95cbfc78bdab3ea99bf8278e94c8db09a9fd3ccac55a538fe0b2
|
||||
data/2017-old/9_Khanh_Hoa_KPKQ.xls.xlsx 0e90c87dd74a5eb293a0a71b93da45653a19ed2e3a2600d3b880a867a695f31d
|
||||
data/2017-old/9_LaiChau_ALKN.xls.xlsx 3b7a022f7727134620c6756925bd5adb041e2d80cabd4aab13a4063501739514
|
||||
data/2017-old/9_Vinh_Long_OLWA.xls.xlsx 43d7abb9c6b8d3205ce90f9cceaf9623875c87998d9e030c6dc8da9bc18bfa09
|
||||
data/2017-old2/10.BinhThuan_MVVG.xlsx 2833e7366ea1a21ea869978a6bb920b033160bb8042fe0bfe9b7573dcfe85409
|
||||
data/2017-old2/10.LamDong_YUQA.xlsx f05b4cdc8a9a8413b29801fbddeb48a7c60d3c6bb2228bbe238113495e89bc98
|
||||
data/2017-old2/10.Soc Trang_LQWU.xlsx 8e540cef6f2ff387f3e825a5aa3190af394c6f9c7e9dfc66d5f7032ee47f7f14
|
||||
data/2017-old2/11.BinhDuong_HVAH.xlsx 11953116054a2889929c3fe05b68e8932c7b79ffd1b85d697afaac63fcdd9adf
|
||||
data/2017-old2/11.LaoCai_ZTBP.xlsx 5a73becf1bdb8027102faff74035d1df71e13c1cac39cc38ec6efaab6699ca52
|
||||
data/2017-old2/12.BenTre_NQTU.xlsx 579c0d94b20440f1c60a94fb0c762b71dd1c7719c8aa4db0b9418ecaba9b605e
|
||||
data/2017-old2/12.LongAn_PDRH.xlsx d1c65184b475fc29ae273cc89b8574b7e273bc3dab83cd39489d72a5de1bb1e4
|
||||
data/2017-old2/13.NamDinh_NAYR.xlsx 7e5256cb095262fef5399092f8f843c1ee82e117bca8e6571e363646bb1170a6
|
||||
data/2017-old2/13.TraVinh_FODZ.xlsx 2b7321d68a4a910c4a506d7ca207531cfbdd9619f1d8ba0ce475baf1b91043aa
|
||||
data/2017-old2/14.NgheAn_HTKD.xlsx 034234a8a5461910f8beff3977aafa29a6541c8a56c007816118e5c4eb2b7218
|
||||
data/2017-old2/15.PhuTho_IJZW.xlsx 9891c17001cfcc49ccd587419339698ef59b2dd43b5ca7f2765b09a17c26659d
|
||||
data/2017-old2/17.QuangNam_NQMG.xlsx 4c9f05d24d861ec3c82a1c564e4baa1541bd4882785f143459a1a81a6b8359c7
|
||||
data/2017-old2/18.QuangNgai_IUPY.xlsx ee9c26f2c16d36dabfde79901a054d2266c28c4acbfd19b628cfe50f8eb03957
|
||||
data/2017-old2/19.QuangTri_MKNN.xlsx 6c961bb461274c73b4589de4087624a9ef5edd43c04ea54675148df00f6420e5
|
||||
data/2017-old2/1.BaRia-VungTau_PGZT.xlsx cb03f933442a8cea6b1622da4e5f804ab476dff9f3462b404e957b6a2e87b2cd
|
||||
data/2017-old2/1.Da Nang_ABWU.xlsx 59892fc01b6cb10497679f2bc69f59023215755d2f00fd764b23b895c98be2be
|
||||
data/2017-old2/1.Son La_XLFN.xlsx 0a0e273c5def980406b6316cceb63e5cee53aae56e112dfae7e795261eb1cd06
|
||||
data/2017-old2/1.TuyenQuang_PTMR.xlsx 45826ed10248268ac3344d0b592cf2dba6f86278b7402dc60626a76b18fa4d6b
|
||||
data/2017-old2/20.TayNinh_KJAQ.xlsx 684a1820d02f658904c87c26aa195faa6aceb3d7ca00c674df11a63f9274ab14
|
||||
data/2017-old2/21.ThaiBinh_KTQN.xlsx 1a12e43d4ea368b8e3a2e4d881b5d66214a16140219e80e2e72fcb269899b285
|
||||
data/2017-old2/22.ThaiNguyen_BKIF.xlsx 29227bdc3204591327e803680525d8621228fc371ce7f802dc667ec70eb8d7b3
|
||||
data/2017-old2/24.HCM_UTLQ.xlsx 49ed22e823607b78920a24f5999eaf1ef8634a141108dfd4bf970ee000aecbf9
|
||||
data/2017-old2/2.BacKan_YQNX.xlsx 77a08fc4edb41d639fcdd798980903519181639db00a360ef0c6c6c2bd88f529
|
||||
data/2017-old2/2.Ha Giang_PIYK.xlsx 6bf4d62413c1b92f88f980e5e3e414a7e25971c7546c6390c7e67a6b0f742b68
|
||||
data/2017-old2/2.Ninh Thuan_BAGG.xlsx 6dfb6517f9ccc3420c222b6a88aabac54f6f25f6b743d993d6a5f28b4fa78ff2
|
||||
data/2017-old2/2.Thanh Hoa_UOPE.xlsx 2f2ce90425c19a9089024061f328f37b9ef580ddadc076ddcf452f549690f4b3
|
||||
data/2017-old2/2.VinhPhuc_QZJK.xlsx fc0e4952445ebda1af28f7330aa808c032f2e3a7a3c03dfbb9692bf6a248867a
|
||||
data/2017-old2/3.BacGiang_SAVS.xlsx 8043510ff89f11ba063f898dc1d2c8879a8da9063a420308647e20a2f39d3092
|
||||
data/2017-old2/3.BinhPhuoc_IPHL.xlsx 8869970f4e6005381264f3f4bb508c5181f15ae990a75fe904a65b305c0a123a
|
||||
data/2017-old2/3.Cao Bang_WMUU.xlsx 94a06bdb2333a9f44216001ba8d8463700b8cf3e51944b94db11eef7eac8dbd3
|
||||
data/2017-old2/3.Dong Thap_HKJX.xlsx 0cdd0763fc4f20f1d990621ca9e8dab107a3d4eb7c0ec3e53a4d8c240fbc569f
|
||||
data/2017-old2/3.Thua Thien -Hue_MAET.xlsx 62e54eebda7bbd542ca63443bb923a81b472e7892e55dce39fd0611fc0f00abd
|
||||
data/2017-old2/4.An Giang_PMJD.xlsx 778ad205bbabb7c4aea3df34325770aa4a1daec0a1683f73f0405ac1cc5aaf9a
|
||||
data/2017-old2/4.BacNinh_NNIS.xlsx 009c3b7821f55210a86f9ace74c31eb0d5c16aafdb0cac8f2b71cac616abdfec
|
||||
data/2017-old2/4.Binh Dinh_VOMJ.xlsx fbacd256d58627377d698462d69aac0e83d124c8c8ff1f5b63ba0bfb5442dd39
|
||||
data/2017-old2/4.DienBien_FYGN.xlsx 4931e8d6604013df48b2b469fda3eb05fdf1203b67bbc2ae4ff6d7949a115a24
|
||||
data/2017-old2/4.Lang Son_QWOG.xlsx cc7d2454caadb18740ed6baf6edc62baa47e98b3a907b3ce06f72b9a16b7838a
|
||||
data/2017-old2/5.Bac Lieu_VIVY.xlsx dc1dca45fd5357b45abc930b0ebb95ab589ef80a09bef8b5ad4dab9e00be87e5
|
||||
data/2017-old2/5.Gia Lai_TAAS.xlsx 5b7be25274d628c284c37efb3df0587171cc9273dc3bd8dc824c7bd1b8bccf1a
|
||||
data/2017-old2/5.HaiDuong_WNHD.xlsx 1d800d0aa7338fb507c4eeb3e752620d2608c30ac820895ca086325dbaf3643d
|
||||
data/2017-old2/5.Hanam_SDKN.xlsx cfc233ac37ab8e263b6cc569815ea2c687da30364f139109032c8f2cac1ad98e
|
||||
data/2017-old2/5.Yen Bai_BSLV.xlsx 4c8915417d8747c291f643ea4b378c87982c397460e0e9e3cc6df19670d6720f
|
||||
data/2017-old2/6.Hau Giang_SIAJ.xlsx c11473e3e1f964302c5378c90dd6b9d478eb11397351b7ef74df9ae709c9e908
|
||||
data/2017-old2/6.HoaBinh_HLYQ.xlsx e1210961dc821642f6951cd6e4a71b5fbc81d5070ee3816b7c8b5dc56bf2b234
|
||||
data/2017-old2/6.NinhBinh_PKMQ.xlsx 9399b8ab58e2a7a1c0ebf5425359319e567f8a28602349aa587efced72142fc8
|
||||
data/2017-old2/6.Quang Ninh_YAKJ.xlsx 8cf95259b11d48267f710f151c3f1419ccac49d823b78b70a2708c59ab25a4f7
|
||||
data/2017-old2/7.HaTinh_XMFQ.xlsx d96df9f2c8365dfc73be3e10c268367b93fd00fe92bb39c745d8e353ba75c05b
|
||||
data/2017-old2/7.HungYen_TCBE.xlsx e80eb38af8307c8e193a91bdda5368c437693d7669e281b543bbe891a46e6f36
|
||||
data/2017-old2/7.Kon Tum_CAQU.xlsx 41d2adea059449b6e214241414748f10776861d261cbaa1fd213629757deb846
|
||||
data/2017-old2/7.Tien Giang_AOWZ.xlsx 00d9bbc70be89f8f7dad6aff990ad36236c6b7244ddeb2e3d9b3d8ed1a6e21d0
|
||||
data/2017-old2/8.PhuYen_MLTQ.xlsx 9dd53293f99d86801692a1995a5cf7a08af8394b820c76e68c88c2b7644b9b67
|
||||
data/2017-old2/9.DakNong_ZWJT.xlsx c611bf7f1d1c95cbfc78bdab3ea99bf8278e94c8db09a9fd3ccac55a538fe0b2
|
||||
data/2017-old2/9.Khanh Hoa_IBTX.xlsx 0e90c87dd74a5eb293a0a71b93da45653a19ed2e3a2600d3b880a867a695f31d
|
||||
data/2017-old2/9.Vinh Long_WVKI.xlsx 43d7abb9c6b8d3205ce90f9cceaf9623875c87998d9e030c6dc8da9bc18bfa09
|
||||
|
+4
-3
@@ -5,14 +5,15 @@
|
||||
"type": "module",
|
||||
"description": "Tra cứu điểm thi THPT Quốc gia — 2016 và 2017",
|
||||
"scripts": {
|
||||
"build:rust": "cargo build --release --manifest-path parser/Cargo.toml",
|
||||
"build:db": "node parser/scripts/build-db.js",
|
||||
"build:go": "go -C go-parser build -o bin/xlsxread ./cmd/xlsxread",
|
||||
"build:db": "node go-parser/scripts/build-db.js",
|
||||
"dev": "vite",
|
||||
"build": "vite build",
|
||||
"assemble": "node scripts/assemble-site.js",
|
||||
"build:site": "npm run build && npm run assemble",
|
||||
"preview": "vite preview",
|
||||
"lint": "eslint ."
|
||||
"lint": "eslint .",
|
||||
"test:go": "go -C go-parser test ./..."
|
||||
},
|
||||
"repository": {
|
||||
"type": "git",
|
||||
|
||||
Generated
+26
-54
@@ -489,6 +489,12 @@ version = "1.70.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "a6cb138bb79a146c1bd460005623e142ef0181e3d0219cb493e02f7d08a35695"
|
||||
|
||||
[[package]]
|
||||
name = "itoa"
|
||||
version = "1.0.18"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682"
|
||||
|
||||
[[package]]
|
||||
name = "jobserver"
|
||||
version = "0.1.34"
|
||||
@@ -693,6 +699,12 @@ version = "1.0.22"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "b39cdef0fa800fc44525c84ccb54a029961a8215f9619753635a9c0d2538d46d"
|
||||
|
||||
[[package]]
|
||||
name = "ryu"
|
||||
version = "1.0.23"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "9774ba4a74de5f7b1c1451ed6cd5285a32eddb5cccb8cc655a4e50009e06477f"
|
||||
|
||||
[[package]]
|
||||
name = "serde"
|
||||
version = "1.0.228"
|
||||
@@ -724,12 +736,16 @@ dependencies = [
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "serde_spanned"
|
||||
version = "0.6.9"
|
||||
name = "serde_yaml"
|
||||
version = "0.9.34+deprecated"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "bf41e0cfaf7226dca15e8197172c295a782857fcb97fad1808a166870dee75a3"
|
||||
checksum = "6a8b1a1a2ebf674015cc02edccce75287f1a0130d394307b36743c2f5d504b47"
|
||||
dependencies = [
|
||||
"indexmap",
|
||||
"itoa",
|
||||
"ryu",
|
||||
"serde",
|
||||
"unsafe-libyaml",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
@@ -858,47 +874,6 @@ version = "0.1.1"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "1f3ccbac311fea05f86f61904b462b55fb3df8837a366dfc601a0161d0532f20"
|
||||
|
||||
[[package]]
|
||||
name = "toml"
|
||||
version = "0.8.23"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "dc1beb996b9d83529a9e75c17a1686767d148d70663143c7854d8b4a09ced362"
|
||||
dependencies = [
|
||||
"serde",
|
||||
"serde_spanned",
|
||||
"toml_datetime",
|
||||
"toml_edit",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "toml_datetime"
|
||||
version = "0.6.11"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "22cddaf88f4fbc13c51aebbf5f8eceb5c7c5a9da2ac40a13519eb5b0a0e8f11c"
|
||||
dependencies = [
|
||||
"serde",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "toml_edit"
|
||||
version = "0.22.27"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "41fe8c660ae4257887cf66394862d21dbca4a6ddd26f04a3560410406a2f819a"
|
||||
dependencies = [
|
||||
"indexmap",
|
||||
"serde",
|
||||
"serde_spanned",
|
||||
"toml_datetime",
|
||||
"toml_write",
|
||||
"winnow",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "toml_write"
|
||||
version = "0.1.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "5d99f8c9a7727884afe522e9bd5edbfc91a3312b36a77b5fb8926e4c31a41801"
|
||||
|
||||
[[package]]
|
||||
name = "typenum"
|
||||
version = "1.20.0"
|
||||
@@ -920,6 +895,12 @@ dependencies = [
|
||||
"tinyvec",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "unsafe-libyaml"
|
||||
version = "0.2.11"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "673aac59facbab8a9007c7f6108d11f63b603f7cabff99fabf650fea5c32b861"
|
||||
|
||||
[[package]]
|
||||
name = "utf8parse"
|
||||
version = "0.2.2"
|
||||
@@ -1007,15 +988,6 @@ dependencies = [
|
||||
"windows-link",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "winnow"
|
||||
version = "0.7.15"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "df79d97927682d2fd8adb29682d1140b343be4ac0f08fd68b7765d9c059d3945"
|
||||
dependencies = [
|
||||
"memchr",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "wit-bindgen"
|
||||
version = "0.57.1"
|
||||
@@ -1033,8 +1005,8 @@ dependencies = [
|
||||
"regex",
|
||||
"rusqlite",
|
||||
"serde",
|
||||
"serde_yaml",
|
||||
"thiserror 1.0.69",
|
||||
"toml",
|
||||
"unicode-normalization",
|
||||
"zip",
|
||||
]
|
||||
|
||||
+1
-1
@@ -9,7 +9,7 @@ calamine = "0.26"
|
||||
rusqlite = { version = "0.32", features = ["bundled"] }
|
||||
clap = { version = "4", features = ["derive"] }
|
||||
serde = { version = "1", features = ["derive"] }
|
||||
toml = "0.8"
|
||||
serde_yaml = "0.9" # deprecated upstream but stable; this crate is deleted at cutover
|
||||
regex = "1"
|
||||
unicode-normalization = "0.1"
|
||||
thiserror = "1"
|
||||
|
||||
@@ -15,18 +15,18 @@
|
||||
#
|
||||
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
|
||||
|
||||
format_detection = "thptqg2016"
|
||||
format_detection: thptqg2016
|
||||
|
||||
[reader]
|
||||
sheet_mode = "all"
|
||||
strip_blank_rows = false
|
||||
reader:
|
||||
sheet_mode: all
|
||||
strip_blank_rows: false
|
||||
|
||||
[validation]
|
||||
require_numeric_sbd = false
|
||||
require_nonempty_name = true
|
||||
require_nonempty_sbd = true
|
||||
validation:
|
||||
require_numeric_sbd: false
|
||||
require_nonempty_name: true
|
||||
require_nonempty_sbd: true
|
||||
|
||||
[header]
|
||||
# Tokens that identify a header row by first-cell content (uppercased).
|
||||
# Covers both SOBAODANH-style and SBD-style headers.
|
||||
tokens = ["SOBAODANH", "SBD", "HO_TEN", "HOTEN", "HỌ TÊN", "STT"]
|
||||
header:
|
||||
# Tokens that identify a header row by first-cell content (uppercased).
|
||||
# Covers both SOBAODANH-style and SBD-style headers.
|
||||
tokens: ["SOBAODANH", "SBD", "HO_TEN", "HOTEN", "HỌ TÊN", "STT"]
|
||||
@@ -6,20 +6,20 @@
|
||||
#
|
||||
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
|
||||
|
||||
[reader]
|
||||
sheet_mode = "first"
|
||||
strip_blank_rows = false
|
||||
reader:
|
||||
sheet_mode: first
|
||||
strip_blank_rows: false
|
||||
|
||||
[columns]
|
||||
ho_ten = 0
|
||||
ngay_sinh = 1
|
||||
so_bao_danh = 2
|
||||
diem_thi = 3
|
||||
columns:
|
||||
ho_ten: 0
|
||||
ngay_sinh: 1
|
||||
so_bao_danh: 2
|
||||
diem_thi: 3
|
||||
|
||||
[validation]
|
||||
require_numeric_sbd = true
|
||||
require_nonempty_name = true
|
||||
require_nonempty_sbd = true
|
||||
validation:
|
||||
require_numeric_sbd: true
|
||||
require_nonempty_name: true
|
||||
require_nonempty_sbd: true
|
||||
|
||||
[header]
|
||||
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
header:
|
||||
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
@@ -7,20 +7,20 @@
|
||||
#
|
||||
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
|
||||
|
||||
[reader]
|
||||
sheet_mode = "all"
|
||||
strip_blank_rows = true
|
||||
reader:
|
||||
sheet_mode: all
|
||||
strip_blank_rows: true
|
||||
|
||||
[columns]
|
||||
ho_ten = 0
|
||||
ngay_sinh = 1
|
||||
so_bao_danh = 2
|
||||
diem_thi = 3
|
||||
columns:
|
||||
ho_ten: 0
|
||||
ngay_sinh: 1
|
||||
so_bao_danh: 2
|
||||
diem_thi: 3
|
||||
|
||||
[validation]
|
||||
require_numeric_sbd = true
|
||||
require_nonempty_name = true
|
||||
require_nonempty_sbd = true
|
||||
validation:
|
||||
require_numeric_sbd: true
|
||||
require_nonempty_name: true
|
||||
require_nonempty_sbd: true
|
||||
|
||||
[header]
|
||||
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
header:
|
||||
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
@@ -6,20 +6,20 @@
|
||||
#
|
||||
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
|
||||
|
||||
[reader]
|
||||
sheet_mode = "all"
|
||||
strip_blank_rows = false
|
||||
reader:
|
||||
sheet_mode: all
|
||||
strip_blank_rows: false
|
||||
|
||||
[columns]
|
||||
ho_ten = 0
|
||||
ngay_sinh = 1
|
||||
so_bao_danh = 2
|
||||
diem_thi = 3
|
||||
columns:
|
||||
ho_ten: 0
|
||||
ngay_sinh: 1
|
||||
so_bao_danh: 2
|
||||
diem_thi: 3
|
||||
|
||||
[validation]
|
||||
require_numeric_sbd = false
|
||||
require_nonempty_name = true
|
||||
require_nonempty_sbd = true
|
||||
validation:
|
||||
require_numeric_sbd: false
|
||||
require_nonempty_name: true
|
||||
require_nonempty_sbd: true
|
||||
|
||||
[header]
|
||||
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
header:
|
||||
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
@@ -0,0 +1,106 @@
|
||||
//! Ground-truth cell dumper for the Go reader fidelity gate.
|
||||
//!
|
||||
//! Emits a canonical, reader-agnostic text rendering of every sheet and every
|
||||
//! cell of a spreadsheet exactly as calamine sees it. The Go reader must
|
||||
//! reproduce this byte-for-byte; the two dumps are compared by hash.
|
||||
//!
|
||||
//! Deliberately dumps the RAW used range: every sheet, every row, header rows
|
||||
//! included, no config applied. This gate is about cell fidelity, not build
|
||||
//! semantics — sheet selection and header skipping are exercised later.
|
||||
//!
|
||||
//! Usage: cargo run --release --example dump_cells -- <spreadsheet> [out-file]
|
||||
//! With no out-file the dump goes to stdout.
|
||||
|
||||
use std::env;
|
||||
use std::fs::File;
|
||||
use std::io::{self, BufWriter, Write};
|
||||
|
||||
use calamine::{open_workbook_auto, Data, Reader, Sheets};
|
||||
|
||||
/// Escapes the field separators so a cell value can never break the line format.
|
||||
fn escape(s: &str) -> String {
|
||||
let mut out = String::with_capacity(s.len());
|
||||
for ch in s.chars() {
|
||||
match ch {
|
||||
'\\' => out.push_str("\\\\"),
|
||||
'\t' => out.push_str("\\t"),
|
||||
'\n' => out.push_str("\\n"),
|
||||
'\r' => out.push_str("\\r"),
|
||||
_ => out.push(ch),
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// Discriminates the calamine variant so a Go port can be checked against the
|
||||
/// actual type, not just the rendered string.
|
||||
fn kind(d: &Data) -> &'static str {
|
||||
match d {
|
||||
Data::Empty => "empty",
|
||||
Data::String(_) => "str",
|
||||
Data::Float(_) => "float",
|
||||
Data::Int(_) => "int",
|
||||
Data::Bool(_) => "bool",
|
||||
Data::Error(_) => "err",
|
||||
Data::DateTime(_) => "datetime",
|
||||
Data::DateTimeIso(_) => "datetimeiso",
|
||||
Data::DurationIso(_) => "durationiso",
|
||||
}
|
||||
}
|
||||
|
||||
fn main() -> Result<(), Box<dyn std::error::Error>> {
|
||||
let args: Vec<String> = env::args().collect();
|
||||
if args.len() < 2 {
|
||||
eprintln!("usage: dump_cells <spreadsheet> [out-file]");
|
||||
std::process::exit(2);
|
||||
}
|
||||
let path = &args[1];
|
||||
|
||||
let mut out: Box<dyn Write> = match args.get(2) {
|
||||
Some(p) => Box::new(BufWriter::new(File::create(p)?)),
|
||||
None => Box::new(BufWriter::new(io::stdout())),
|
||||
};
|
||||
|
||||
let mut workbook: Sheets<_> = open_workbook_auto(path)?;
|
||||
let sheet_names: Vec<String> = workbook.sheet_names().to_vec();
|
||||
|
||||
writeln!(out, "FILE\t{}", escape(path))?;
|
||||
writeln!(out, "SHEETCOUNT\t{}", sheet_names.len())?;
|
||||
|
||||
for (idx, name) in sheet_names.iter().enumerate() {
|
||||
let range = workbook.worksheet_range(name)?;
|
||||
// start() is the used-range origin — the key question for any Go reader,
|
||||
// which may index absolutely from A1 instead.
|
||||
let (srow, scol) = range.start().unwrap_or((0, 0));
|
||||
writeln!(
|
||||
out,
|
||||
"SHEET\t{}\t{}\t{}\t{}\t{}\t{}",
|
||||
idx,
|
||||
escape(name),
|
||||
range.height(),
|
||||
range.width(),
|
||||
srow,
|
||||
scol
|
||||
)?;
|
||||
|
||||
for (r, row) in range.rows().enumerate() {
|
||||
writeln!(out, "ROW\t{}\t{}\t{}", idx, r, row.len())?;
|
||||
for (c, cell) in row.iter().enumerate() {
|
||||
let s = cell.to_string();
|
||||
writeln!(
|
||||
out,
|
||||
"CELL\t{}\t{}\t{}\t{}\t{}\t{}",
|
||||
idx,
|
||||
r,
|
||||
c,
|
||||
kind(cell),
|
||||
if matches!(cell, Data::Empty) { 1 } else { 0 },
|
||||
escape(&s)
|
||||
)?;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
out.flush()?;
|
||||
Ok(())
|
||||
}
|
||||
@@ -0,0 +1,76 @@
|
||||
//! Corpus-wide cell-kind and sheet-geometry scan, for the Go reader fidelity gate.
|
||||
//!
|
||||
//! For every file given, prints one TSV line per sheet and one summary line per
|
||||
//! file. Aggregates only — never materialises the cell text — so the whole
|
||||
//! 418 MB corpus can be scanned quickly.
|
||||
//!
|
||||
//! The point is to find out which calamine `Data` variants actually occur in
|
||||
//! real inputs. `DateTime` and `Float` are the variants whose rendering differs
|
||||
//! between readers; if they never appear, the divergence risk is theoretical.
|
||||
//!
|
||||
//! Usage: cargo run --release --example scan_kinds -- <file>...
|
||||
|
||||
use std::env;
|
||||
|
||||
use calamine::{open_workbook_auto, Data, Reader, Sheets};
|
||||
|
||||
fn main() {
|
||||
let args: Vec<String> = env::args().skip(1).collect();
|
||||
if args.is_empty() {
|
||||
eprintln!("usage: scan_kinds <file>...");
|
||||
std::process::exit(2);
|
||||
}
|
||||
|
||||
println!("#TYPE\tpath\tsheet_idx\tname\theight\twidth\tstart_row\tstart_col\tempty\tstr\tfloat\tint\tbool\tdatetime\tdtiso\tduriso\terr");
|
||||
|
||||
for path in &args {
|
||||
let mut wb: Sheets<_> = match open_workbook_auto(path) {
|
||||
Ok(w) => w,
|
||||
Err(e) => {
|
||||
println!("ERR\t{path}\t{e}");
|
||||
continue;
|
||||
}
|
||||
};
|
||||
|
||||
let names: Vec<String> = wb.sheet_names().to_vec();
|
||||
for (idx, name) in names.iter().enumerate() {
|
||||
let range = match wb.worksheet_range(name) {
|
||||
Ok(r) => r,
|
||||
Err(e) => {
|
||||
println!("SHEETERR\t{path}\t{idx}\t{e}");
|
||||
continue;
|
||||
}
|
||||
};
|
||||
let (sr, sc) = range.start().unwrap_or((0, 0));
|
||||
let mut k = [0usize; 9]; // empty,str,float,int,bool,datetime,dtiso,duriso,err
|
||||
for row in range.rows() {
|
||||
for cell in row {
|
||||
let i = match cell {
|
||||
Data::Empty => 0,
|
||||
Data::String(_) => 1,
|
||||
Data::Float(_) => 2,
|
||||
Data::Int(_) => 3,
|
||||
Data::Bool(_) => 4,
|
||||
Data::DateTime(_) => 5,
|
||||
Data::DateTimeIso(_) => 6,
|
||||
Data::DurationIso(_) => 7,
|
||||
Data::Error(_) => 8,
|
||||
};
|
||||
k[i] += 1;
|
||||
}
|
||||
}
|
||||
println!(
|
||||
"SHEET\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}",
|
||||
path,
|
||||
idx,
|
||||
name.replace('\t', " "),
|
||||
range.height(),
|
||||
range.width(),
|
||||
sr,
|
||||
sc,
|
||||
k[0], k[1], k[2], k[3], k[4], k[5], k[6], k[7], k[8]
|
||||
);
|
||||
}
|
||||
println!("FILE\t{}\t{}", path, names.len());
|
||||
}
|
||||
}
|
||||
@@ -50,7 +50,7 @@ for (const id of targets) {
|
||||
[
|
||||
"build",
|
||||
"--schema",
|
||||
resolve(ROOT, `parser/configs/${id}.toml`),
|
||||
resolve(ROOT, `parser/configs/${id}.yml`),
|
||||
"--input",
|
||||
resolve(ROOT, `data/${id}`),
|
||||
"--output",
|
||||
|
||||
+36
-35
@@ -6,7 +6,7 @@ use serde::Deserialize;
|
||||
use crate::error::BuildError;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Top-level dataset configuration loaded from a .toml file
|
||||
// Top-level dataset configuration loaded from a .yml file
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Per-dataset parse rules.
|
||||
@@ -79,7 +79,7 @@ pub fn load_config(path: &Path) -> Result<DatasetConfig, BuildError> {
|
||||
path: path.display().to_string(),
|
||||
source: e,
|
||||
})?;
|
||||
let cfg: DatasetConfig = toml::from_str(&text)?;
|
||||
let cfg: DatasetConfig = serde_yaml::from_str(&text)?;
|
||||
Ok(cfg)
|
||||
}
|
||||
|
||||
@@ -91,29 +91,29 @@ pub fn load_config(path: &Path) -> Result<DatasetConfig, BuildError> {
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
const SAMPLE_TOML: &str = r#"
|
||||
[reader]
|
||||
sheet_mode = "all"
|
||||
strip_blank_rows = false
|
||||
const SAMPLE_YAML: &str = r#"
|
||||
reader:
|
||||
sheet_mode: all
|
||||
strip_blank_rows: false
|
||||
|
||||
[columns]
|
||||
ho_ten = 0
|
||||
ngay_sinh = 1
|
||||
so_bao_danh = 2
|
||||
diem_thi = 3
|
||||
columns:
|
||||
ho_ten: 0
|
||||
ngay_sinh: 1
|
||||
so_bao_danh: 2
|
||||
diem_thi: 3
|
||||
|
||||
[validation]
|
||||
require_numeric_sbd = false
|
||||
require_nonempty_name = true
|
||||
require_nonempty_sbd = true
|
||||
validation:
|
||||
require_numeric_sbd: false
|
||||
require_nonempty_name: true
|
||||
require_nonempty_sbd: true
|
||||
|
||||
[header]
|
||||
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
header:
|
||||
tokens: ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
"#;
|
||||
|
||||
#[test]
|
||||
fn config_round_trip() {
|
||||
let cfg: DatasetConfig = toml::from_str(SAMPLE_TOML).expect("parse failed");
|
||||
let cfg: DatasetConfig = serde_yaml::from_str(SAMPLE_YAML).expect("parse failed");
|
||||
assert_eq!(cfg.reader.sheet_mode, SheetMode::All);
|
||||
assert!(!cfg.reader.strip_blank_rows);
|
||||
let cols = cfg.columns.as_ref().unwrap();
|
||||
@@ -131,37 +131,38 @@ tokens = ["HO_TEN", "HỌ TÊN", "STT"]
|
||||
#[test]
|
||||
fn config_rejects_leftover_sql_sections() {
|
||||
let with_ddl = format!(
|
||||
"{SAMPLE_TOML}\n[schema]\nddl = \"CREATE TABLE student (so_bao_danh TEXT);\"\n"
|
||||
"{SAMPLE_YAML}\nschema:\n ddl: \"CREATE TABLE student (so_bao_danh TEXT);\"\n"
|
||||
);
|
||||
assert!(toml::from_str::<DatasetConfig>(&with_ddl).is_err());
|
||||
assert!(serde_yaml::from_str::<DatasetConfig>(&with_ddl).is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn config_first_sheet_mode() {
|
||||
let toml_str = SAMPLE_TOML.replace(r#"sheet_mode = "all""#, r#"sheet_mode = "first""#);
|
||||
let cfg: DatasetConfig = toml::from_str(&toml_str).expect("parse failed");
|
||||
let yaml_str = SAMPLE_YAML.replace("sheet_mode: all", "sheet_mode: first");
|
||||
let cfg: DatasetConfig = serde_yaml::from_str(&yaml_str).expect("parse failed");
|
||||
assert_eq!(cfg.reader.sheet_mode, SheetMode::First);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn config_format_detection_field() {
|
||||
// Configs without [columns] and with format_detection = "thptqg2016" parse correctly
|
||||
let toml_str = r#"
|
||||
format_detection = "thptqg2016"
|
||||
// Configs without a `columns:` mapping and with format_detection:
|
||||
// thptqg2016 parse correctly
|
||||
let yaml_str = r#"
|
||||
format_detection: thptqg2016
|
||||
|
||||
[reader]
|
||||
sheet_mode = "all"
|
||||
strip_blank_rows = false
|
||||
reader:
|
||||
sheet_mode: all
|
||||
strip_blank_rows: false
|
||||
|
||||
[validation]
|
||||
require_numeric_sbd = false
|
||||
require_nonempty_name = true
|
||||
require_nonempty_sbd = true
|
||||
validation:
|
||||
require_numeric_sbd: false
|
||||
require_nonempty_name: true
|
||||
require_nonempty_sbd: true
|
||||
|
||||
[header]
|
||||
tokens = ["SBD", "SOBAODANH", "STT"]
|
||||
header:
|
||||
tokens: ["SBD", "SOBAODANH", "STT"]
|
||||
"#;
|
||||
let cfg: DatasetConfig = toml::from_str(toml_str).expect("parse failed");
|
||||
let cfg: DatasetConfig = serde_yaml::from_str(yaml_str).expect("parse failed");
|
||||
assert_eq!(cfg.format_detection.as_deref(), Some("thptqg2016"));
|
||||
assert!(cfg.columns.is_none());
|
||||
}
|
||||
|
||||
+1
-1
@@ -20,7 +20,7 @@ pub enum BuildError {
|
||||
Sqlite(#[from] rusqlite::Error),
|
||||
|
||||
#[error("Config parse error: {0}")]
|
||||
Config(#[from] toml::de::Error),
|
||||
Config(#[from] serde_yaml::Error),
|
||||
|
||||
#[error("Regex compile error for pattern '{pattern}': {source}")]
|
||||
Regex {
|
||||
|
||||
@@ -267,7 +267,7 @@ fn ensure_fixtures() {
|
||||
fn make_data_config() -> xlsxread::config::DatasetConfig {
|
||||
let cfg_path = PathBuf::from(env!("CARGO_MANIFEST_DIR"))
|
||||
.join("configs")
|
||||
.join("2017.toml");
|
||||
.join("2017.yml");
|
||||
xlsxread::config::load_config(&cfg_path).expect("load data config")
|
||||
}
|
||||
|
||||
@@ -284,7 +284,7 @@ fn province_100_builds_100_rows() {
|
||||
std::fs::create_dir_all(&fixture_dir).unwrap();
|
||||
std::fs::copy(province_fixture_path(), fixture_dir.join("province.xlsx")).unwrap();
|
||||
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
|
||||
|
||||
let count = query_count(&db_path);
|
||||
assert_eq!(count, 100, "expected 100 rows from province-100 fixture");
|
||||
@@ -299,7 +299,7 @@ fn hcm_overflow_builds_400_rows() {
|
||||
std::fs::create_dir_all(&fixture_dir).unwrap();
|
||||
std::fs::copy(hcm_overflow_fixture_path(), fixture_dir.join("hcm.xlsx")).unwrap();
|
||||
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
|
||||
|
||||
let count = query_count(&db_path);
|
||||
assert_eq!(
|
||||
@@ -318,7 +318,7 @@ fn data_old_first_sheet_only_100_rows() {
|
||||
// Use the overflow file but with data-old config (first sheet only → 200 rows)
|
||||
std::fs::copy(hcm_overflow_fixture_path(), fixture_dir.join("hcm.xlsx")).unwrap();
|
||||
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017-old.toml");
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017-old.yml");
|
||||
|
||||
// data-old: sheet_mode=first → only 200 rows from sheet1; but SBDs "1000NNNN" are
|
||||
// all digits so all pass the numeric guard
|
||||
@@ -356,7 +356,7 @@ fn numeric_sbd_guard_rejects_non_numeric() {
|
||||
let mixed_path = fixture_dir.join("mixed.xlsx");
|
||||
write_xlsx(&mixed_path, &[("Sheet1".to_owned(), rows)]);
|
||||
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017-old.toml");
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017-old.yml");
|
||||
|
||||
// Row i=5 has non-numeric SBD → rejected by data-old config
|
||||
let count = query_count(&db_path);
|
||||
@@ -390,7 +390,7 @@ fn scores_parsed_correctly_into_db() {
|
||||
&[("Sheet1".to_owned(), rows)],
|
||||
);
|
||||
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
|
||||
|
||||
let conn = rusqlite::Connection::open(&db_path).unwrap();
|
||||
let (toan, van, anh): (f64, f64, f64) = conn
|
||||
@@ -429,7 +429,7 @@ fn to_ascii_stored_correctly() {
|
||||
&[("Sheet1".to_owned(), rows)],
|
||||
);
|
||||
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
|
||||
|
||||
let conn = rusqlite::Connection::open(&db_path).unwrap();
|
||||
let ascii: String = conn
|
||||
@@ -451,7 +451,7 @@ fn audit_subcommand_matches_after_build() {
|
||||
std::fs::create_dir_all(&fixture_dir).unwrap();
|
||||
std::fs::copy(province_fixture_path(), fixture_dir.join("province.xlsx")).unwrap();
|
||||
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.toml");
|
||||
run_build_cmd(&fixture_dir, &db_path, "2017.yml");
|
||||
|
||||
// audit should match (100 distinct SBDs in xlsx == 100 rows in DB)
|
||||
let cfg = make_data_config();
|
||||
@@ -491,7 +491,7 @@ fn audit_subcommand_mismatch_detected() {
|
||||
&[("Sheet1".to_owned(), five_rows)],
|
||||
);
|
||||
|
||||
run_build_cmd(&build_dir, &db_path, "2017.toml");
|
||||
run_build_cmd(&build_dir, &db_path, "2017.yml");
|
||||
|
||||
// audit against fixture_dir (10 xlsx rows) but DB has 5 rows → mismatch
|
||||
let cfg = make_data_config();
|
||||
|
||||
@@ -0,0 +1,279 @@
|
||||
---
|
||||
phase: 1
|
||||
title: Scaffold and reader fidelity gate
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies: []
|
||||
effort: ''
|
||||
---
|
||||
|
||||
# Phase 1: Scaffold and reader fidelity gate
|
||||
|
||||
## Overview
|
||||
|
||||
Scaffold `go-parser/` and settle the question that governs everything downstream: **does the Go
|
||||
reader produce, for every cell of every sheet of all 299 files, the exact string calamine
|
||||
produces?** Not "does it open" — the exact string, because that string is what gets stored and
|
||||
regex-matched.
|
||||
|
||||
This phase also produces the answer that unblocks the `Data`-typed tests in Phases 4 and 5.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: byte-identical cell stringification vs calamine across all 299 input files, both
|
||||
formats, including dates, numerics, and empty cells.
|
||||
- Non-functional: one reader contract, defined here, used unchanged by Phase 4.
|
||||
|
||||
## What was believed going in (and how it held up)
|
||||
|
||||
Red-team testing had swept all 67 BIFF files with `extrame/xls` and reported **0 failures,
|
||||
0 panics**, correct Vietnamese at row 60,000. That was taken as evidence BIFF *readability*
|
||||
was settled and only cell-value fidelity remained open.
|
||||
|
||||
**That evidence did not survive contact with a cell-level comparison** — see the RESULT below.
|
||||
"Opens without panic" and a handful of spot-checks are not a fidelity test when 28% of cells
|
||||
are correct. The lesson generalises: for every remaining phase, compare against ground truth
|
||||
cell-by-cell, never by sampling.
|
||||
|
||||
Hazards that *did* hold up and were designed around:
|
||||
- excelize applies number formats by default → **set `RawCellValue: true`** and verify.
|
||||
- excelize `GetRows` trims trailing blank cells → rows are ragged; calamine's are rectangular.
|
||||
|
||||
Hazards that turned out not to exist in this corpus: date-serial rendering (zero `DateTime`
|
||||
cells anywhere) and non-A1 used-range origins (all ranges start at (0,0)).
|
||||
|
||||
## Architecture — AS BUILT
|
||||
|
||||
The shipped API differs from the original sketch (which took a `DatasetConfig` and did header
|
||||
skipping inline). The reader deliberately knows **nothing** about datasets: it reports every
|
||||
sheet and every row exactly as calamine would, and all policy — sheet selection, header
|
||||
skipping, blank-row handling — belongs to the Phase 4 build loop. That keeps the fidelity
|
||||
contract testable in isolation, which is what made the 299/299 oracle possible.
|
||||
|
||||
```go
|
||||
// go-parser/internal/reader
|
||||
type Cell struct {
|
||||
Str string // exactly what calamine's Data::to_string() yields
|
||||
IsEmpty bool // calamine Data::Empty; diagnostic only — never compare on it
|
||||
}
|
||||
|
||||
type Sheet struct{ Index int; Name string; Height, Width int } // used-range geometry
|
||||
|
||||
type RowFunc func(sheet Sheet, rowIdx int, row []Cell) error
|
||||
|
||||
type Workbook interface {
|
||||
Sheets() []Sheet // workbook order, all sheets
|
||||
EachRow(sheetIdx int, fn RowFunc) error
|
||||
Close() error
|
||||
}
|
||||
|
||||
func Open(path string) (Workbook, error) // dispatches on extension
|
||||
```
|
||||
|
||||
Rows are padded to the sheet's used-range width. Width is load-bearing: every column read
|
||||
downstream is positional, so a trimmed tail silently NULLs columns.
|
||||
|
||||
**Implementation note:** both backends materialise a workbook's rows rather than streaming
|
||||
(`rows [][][]Cell`). Peak input is `data/2017/ha-noi.xls` at 72,276 rows × 4 columns; the full
|
||||
299-file suite runs in 77s. Revisit only if a future dataset is far larger.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `go-parser/go.mod`, `go-parser/internal/reader/{reader.go,xls.go,xlsx.go}` + tests,
|
||||
`go-parser/internal/reader/fidelity_test.go`, `go-parser/testdata/`
|
||||
- Reference (do not modify): `parser/src/reader.rs` (incl. its 7 tests at `:112-197`),
|
||||
`parser/Cargo.toml`
|
||||
- Read-only inputs: all 299 files under `data/`
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
**Tests first.** The oracle is Rust; extract it before writing Go.
|
||||
|
||||
1. **Generate ground truth.** Add a throwaway Rust bin (or `#[test]`) that, for a chosen file,
|
||||
dumps every sheet: sheet name, used-range dimensions, and every cell rendered exactly as
|
||||
`Data::to_string()` plus an `is_empty` flag.
|
||||
2. **Commit hashes, not rows.** For each sampled file store sheet names, per-sheet row/column
|
||||
counts, and a SHA-256 over the canonical dump — **not** the dump itself. The raw dumps are
|
||||
real student names and birthdates; `parser/tests/fixtures/README.md` documents that all
|
||||
fixture PII is replaced with synthetic values, and committing real rows would reverse that
|
||||
convention and outlive the data they came from. Keep dumps under `.gitignore` and regenerate
|
||||
from Rust on demand (Rust still builds — that is this plan's whole advantage).
|
||||
3. **Sample must cover both formats and the risky cell types.** At minimum: 3 `.xls`
|
||||
(2 from `data/2017`, 1 from `data/2016`) **and** 3 `.xlsx` (one each from `data/2016`,
|
||||
`data/2017-old`, `data/2017-old2`), each chosen to contain a date cell and a numeric SBD.
|
||||
4. Write `fidelity_test.go` asserting the Go reader reproduces every committed hash. It fails —
|
||||
nothing is implemented.
|
||||
5. Scaffold: `go mod init`, add deps **at pinned versions**, commit `go.sum`.
|
||||
6. Implement both readers until the hashes match.
|
||||
7. **Full sweep, all 299 files**: for each, assert sheet names in order, per-sheet row count,
|
||||
and per-sheet used-range width all match calamine. Record failures per file.
|
||||
8. **Record the stringification answer** in this file — the literal rendering of a date cell, a
|
||||
float score, and a numeric SBD. Phases 4 and 5 depend on it to port `Data`-typed tests
|
||||
without guessing.
|
||||
|
||||
## Decision gate — *as written before execution; see RESULT below for the outcome*
|
||||
|
||||
- **PASS**: all 299 files match on sheet names, per-sheet row counts, widths, and sampled
|
||||
content hashes.
|
||||
- **FAIL** → stop and escalate. Do not proceed with a partial pass. The pre-planned fallback
|
||||
(`.xls → .xlsx` conversion) is **deferred by user decision** and changes committed data, so
|
||||
re-opening it is the user's call, not the implementer's.
|
||||
|
||||
## RESULT — 2026-08-13: **PASS — 299 / 299 exact**
|
||||
|
||||
Every input file's canonical cell dump is byte-identical to calamine's, verified by SHA-256.
|
||||
Locked in as `go test ./internal/reader/` (77s for the full corpus) against the committed
|
||||
oracle `go-parser/testdata/reader-fidelity-hashes.tsv`.
|
||||
|
||||
Reached only after replacing the BIFF library. The first attempt failed hard; the record of
|
||||
that is kept below because it is the reason the reader is built the way it is.
|
||||
|
||||
### Final reader stack
|
||||
|
||||
| Format | Library | Result |
|
||||
|---|---|---|
|
||||
| `.xls` (67 files) | **`github.com/pbnjay/grate`** | exact |
|
||||
| `.xlsx` (232 files) | `github.com/xuri/excelize/v2 v2.11.0` | exact |
|
||||
|
||||
Five corrections were needed to match calamine, each verified against ground truth:
|
||||
|
||||
1. **`RawCellValue: true`** — otherwise excelize applies the cell number format.
|
||||
2. **Rows padded to used-range width** — excelize trims trailing blank cells; calamine returns
|
||||
a rectangle. Width is load-bearing: `diem_thi` is the last column for 2017.
|
||||
3. **grate merged-cell markers blanked** — grate fills merge-covered cells with `→`/`⇥`/`↓`/`⤓`
|
||||
(its exported constants); calamine reports them empty. 19 cells, in the merged title block
|
||||
of one 2016 file. Only an exact whole-value match is blanked.
|
||||
4. **Numeric re-rendering, gated on cell type** — calamine parses numeric cells to f64 and
|
||||
renders with Rust's `Display`, so `6.0` becomes `6`. Applying that by value alone corrupts
|
||||
shared strings that merely look numeric: it turned `6.00`→`6`, `NAN`→`NaN`, and would have
|
||||
destroyed leading zeros in `so_bao_danh`. The type check is what makes it safe. Note
|
||||
excelize reports **`CellTypeUnset`, not `CellTypeNumber`**, for plain numeric cells, because
|
||||
OOXML omits the `t` attribute and excelize has no map entry for an empty one.
|
||||
5. **CRLF restoration in shared strings** — Go's `encoding/xml` performs the line-ending
|
||||
normalisation XML 1.0 mandates (CRLF and lone CR → LF); calamine reads raw bytes and keeps
|
||||
CRLF. This reaches the database: 2,233 `TEN_CUMTHI` values in one 2016 file, populating
|
||||
`ten_cum_thi`. Fixed by rewriting literal CR to ` ` before decoding — character
|
||||
references are exempt from that normalisation — and mapping the normalised form back.
|
||||
A blanket `\n`→`\r\n` would have been wrong: 117 files carry a lone CR with no LF.
|
||||
|
||||
**Known divergence, behaviorally inert:** none remaining. The trailing 1×1 empty sheet in 63
|
||||
`2017-old` and 53 `2017-old2` files is now reproduced exactly, using `GetCellType(A1)` to tell
|
||||
an empty-shared-string cell (`CellTypeSharedString`) from a genuinely absent one
|
||||
(`CellTypeUnset`, 230 such sheets in 2016).
|
||||
|
||||
### The rejected library: `extrame/xls` — HARD FAIL (67 files)
|
||||
|
||||
Full canonical diff of `data/2017/an-giang.xls` (56,244 cells) against calamine:
|
||||
|
||||
| Class | Cells | Share |
|
||||
|---|---|---|
|
||||
| Identical | 16,016 | 28% |
|
||||
| **Different content** (corruption) | 38,664 | **69%** |
|
||||
| **Rust has value, Go empty** (data loss) | 15,629 | 28% |
|
||||
| Whitespace-only | 0 | — |
|
||||
|
||||
Only 28% of cells are read correctly. Three distinct defect classes, all confirmed against
|
||||
ground truth:
|
||||
|
||||
1. **Undecoded BIFF bytes leak through.** Row 56 col 0: calamine `HỒ THỊ NHƯ Ý`; extrame
|
||||
`"\f\x00\x01H\x00Ò\x1e \x00T\x00H\x00Ê\x1e \x00N\x00H\x00¯\x01 \x00Ý\x00\b\x00\x0051009967t\x00…"`
|
||||
— raw UTF-16LE plus record framing, with the neighbouring SBD and score cells spliced in.
|
||||
2. **Content teleports between cells.** extrame's (14055, 0) is calamine's **(6500, 3)**.
|
||||
3. **Tail rows silently lost.** calamine rows 14056-14060 hold real students
|
||||
(`NGUYỄN HỮU ÁI` … `HUỲNH VĂN KIÊN`); extrame returns them blank.
|
||||
4. Header cell (0,0) `HO_TEN` dropped — would defeat `is_header_row` and ingest the header
|
||||
as data.
|
||||
5. `sh.Row(r)` panics (nil deref, `worksheet.go:30`) for `r > MaxRow`.
|
||||
|
||||
**Not a configuration problem.** Identical garbage under charsets `utf-8`, `utf-16`, `utf-16le`,
|
||||
`windows-1258`, `cp1252`, and empty. Not an index-arithmetic problem on our side either —
|
||||
geometry matches exactly (70,308 canonical lines both sides, `SHEET 0 Sheet1 14061 4` on both)
|
||||
after correcting `LastCol()` exclusivity and the used-range height rule.
|
||||
|
||||
**This refutes the red-team finding that `extrame/xls` reads the corpus correctly.** That sweep
|
||||
tested "opens without panic" plus a few spot-checks; spot-checks pass because 28% of cells are
|
||||
right and the early rows of a file are among them. Cell-level comparison against ground truth
|
||||
is what exposed it.
|
||||
|
||||
Library survey (2026-08-13): `youkuang/xls` and `f2xb/xls` are forks of `extrame/xls` and carry
|
||||
the same defect. `qax-os/excelize` **is** excelize and does not read BIFF at all. The
|
||||
independent implementations are `pbnjay/grate` (chosen — reproduced calamine exactly on all 67
|
||||
files first try, needing only the merged-marker correction) and `shakinm/xlsReader` (not
|
||||
evaluated; grate passed).
|
||||
|
||||
### Corpus facts established (worth keeping regardless of the decision)
|
||||
|
||||
Scanned all 299 files, 15.98M cells:
|
||||
- **Zero `DateTime` cells.** Also zero `Int`, `Bool`, `Error`, `DateTimeIso`, `DurationIso`.
|
||||
Only `String` (15.1M), `Empty` (722k), `Float` (133k) occur. **The date-serial divergence
|
||||
that this plan called its dominant risk does not exist in this corpus** — `ngay_sinh` is
|
||||
stored as text everywhere.
|
||||
- **Every used range starts at (0,0)** — the used-range-origin concern is moot.
|
||||
- Floats occur only in `2016` (53,008) and `2017-old2` (80,121); none in `2017` or `2017-old`.
|
||||
All render as plain decimals, no exponents, max 2 decimal places.
|
||||
- Trailing empty sheets: 293 sheets at height 0, 116 at height 1.
|
||||
- The "63 empty" rows in `docs/data-pipeline.md:114` for 2017 do **not** come from trailing
|
||||
sheets — calamine reports height 0 for all 63 of those and yields no rows from them. The
|
||||
red team's stated mechanism for that count is wrong; provenance is a Phase 4 question.
|
||||
|
||||
### Artefacts
|
||||
|
||||
- `parser/examples/dump_cells.rs`, `parser/examples/scan_kinds.rs` — throwaway Rust ground-truth
|
||||
tooling (delete after migration)
|
||||
- `go-parser/` — module, `internal/reader` (both formats), `cmd/dumpcells`
|
||||
- Dumps are regenerable and gitignored; no PII committed, per the convention in
|
||||
`parser/tests/fixtures/README.md`
|
||||
|
||||
## Specific things to verify, not assume
|
||||
|
||||
- **Per-sheet row counts, including empty sheets.** All 63 `data/2017` files carry a trailing
|
||||
sheet that calamine renders as one blank row. `docs/data-pipeline.md:114` records the
|
||||
consequence exactly: `2017 | 861,131 source rows | 63 empty | 861,068 DB rows`. A Go reader
|
||||
that skips zero-row sheets produces an identical database and silently different counters.
|
||||
- **Date cells** → `ngay_sinh`, stored verbatim. calamine prints the raw serial
|
||||
(`datatype.rs:771-775`); excelize applies the number format unless `RawCellValue: true`.
|
||||
- **Numeric cells** → `so_bao_danh`, a `TEXT PRIMARY KEY`. Trailing `.0`? Scientific notation?
|
||||
A difference here re-keys the table.
|
||||
- **Empty vs blank**: `reader.rs:42` checks both `Data::Empty` and stringified-empty, implying
|
||||
calamine emits empty-but-not-`Empty` cells. The `Cell.IsEmpty` field exists for this.
|
||||
- **Row width / used-range origin**: see "What is already known".
|
||||
- **Sheet order**: `sheet_mode = "all"` for 2016 and 2017. calamine's `sheet_names()` and
|
||||
excelize's `GetSheetList` can disagree when `workbook.xml` order differs from `sheetId` order.
|
||||
Order determines which duplicate SBD survives `INSERT OR REPLACE`.
|
||||
|
||||
## Dependency trust
|
||||
|
||||
- Pin every dependency to an exact version/pseudo-version; commit `go.sum`.
|
||||
- Record in this file that `extrame/xls` is effectively unmaintained (last push 2023-09-12,
|
||||
53 open issues, no valid `go.mod`) and that it transitively adds `tealeg/xlsx`.
|
||||
- Note the open excelize advisory `GHSA-h69g-9hx6-f3v4` (unbounded row-index allocation). The
|
||||
2017 refresh runbook (`docs/data-pipeline.md:137`) feeds network-downloaded spreadsheets
|
||||
straight into the parser, so this is a live path.
|
||||
- `govulncheck` is added to CI in Phase 7b.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] `go-parser/` builds; `go test ./...` runs; `go.sum` committed with pinned versions
|
||||
- [ ] Ground-truth **hashes** (not raw rows) committed for 3 `.xls` + 3 `.xlsx` files
|
||||
- [ ] Go reader reproduces every committed hash
|
||||
- [ ] All 299 files: sheet names in order, per-sheet row counts, and widths match calamine —
|
||||
**including the 63 trailing empty sheets in `data/2017`**
|
||||
- [x] `RawCellValue` settled and justified in writing
|
||||
- [x] Date, float, and numeric-SBD renderings recorded literally in this file
|
||||
- [x] One reader contract, with `Cell.IsEmpty` and padded row width
|
||||
- [~] `parser/src/reader.rs`'s 7 tests (`:112-197`) — **moved to Phase 4.** They exercise
|
||||
`is_header_row` and `is_all_blank`, which are dataset policy and therefore live in the
|
||||
build loop, not the reader package. Recorded here rather than silently dropped.
|
||||
- [x] Dependency trust notes recorded
|
||||
- [x] Explicit PASS/FAIL recorded
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| `.xlsx` date/number formatting differs | The primary target of this gate; `RawCellValue` verified against hashes |
|
||||
| Trailing-blank trimming NULLs tail columns | Row-width equality asserted for all 299 files |
|
||||
| Empty-sheet skipping breaks counters | Per-sheet row counts asserted, including empty sheets |
|
||||
| Real PII committed as fixtures | Hashes committed instead; dumps gitignored and regenerable |
|
||||
| `extrame/xls` unmaintained / OOM issues | Pinned pseudo-version; full sweep measures memory |
|
||||
| Partial pass rationalized into a PASS | Gate is binary; fallback is a user decision |
|
||||
@@ -0,0 +1,195 @@
|
||||
---
|
||||
phase: 2
|
||||
title: Schema and config
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies:
|
||||
- 1
|
||||
effort: ''
|
||||
---
|
||||
|
||||
# Phase 2: Schema and config
|
||||
|
||||
## Overview
|
||||
|
||||
Port the two pure, I/O-free modules: `schema.rs` (DDL, INSERT SQL, column order, 16 subject
|
||||
regexes) and `config.rs` (strict YAML loading). No file or database access — fully unit-testable.
|
||||
|
||||
Also **decides the SQLite driver** (see below) — it governs the CI shape and the integrity
|
||||
story, so it is settled here rather than deferred to Phase 4.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: identical DDL text, identical column ordering, identical regex patterns; YAML
|
||||
loading that **rejects unknown fields**.
|
||||
- Non-functional: `schema` stays the single source of truth, as in Rust. No duplicated column
|
||||
lists anywhere else in the port.
|
||||
|
||||
## Carried forward from Phase 1
|
||||
|
||||
- Reader is done and exact; it exposes `reader.Cell{Str, IsEmpty}`. Config work is independent
|
||||
of it.
|
||||
- The corpus contains **only** `String`, `Empty`, and `Float` cell kinds — no dates, ints,
|
||||
bools, or errors. Nothing in `config.go` needs date or type-coercion handling.
|
||||
- Go 1.26.5 / linux-arm64 confirmed working; `go.mod` currently pulls `pbnjay/grate` and
|
||||
`excelize/v2 v2.11.0`, both pure Go. **The SQLite driver choice below is what decides whether
|
||||
this module stays cgo-free.**
|
||||
|
||||
## Decision: SQLite driver
|
||||
|
||||
Open question 2 in `plan.md`. Record the choice and rationale in this file.
|
||||
|
||||
| Option | For | Against |
|
||||
|---|---|---|
|
||||
| `modernc.org/sqlite` | Pure Go, no cgo, trivial ARM64 CI and cross-compilation. Widely used (3,500+ importers) | **Machine-transpiled** SQLite, not the upstream C amalgamation. Its correctness argument is "the transpiler is correct", not "this is the code the SQLite authors tested". Requires exact `modernc.org/libc` version matching |
|
||||
| `mattn/go-sqlite3` | Real upstream SQLite C, matching what `rusqlite --bundled` vendors (`Cargo.lock:520` `libsqlite3-sys 0.30.1`) | cgo: slower CI, cross-compilation friction, needs a C toolchain in the workflow |
|
||||
|
||||
This writes a published 1.5M-row dataset, so the integrity story is a real consideration, not a
|
||||
formality. Phase 6's full-table checksum is the compensating control either way. Pin the exact
|
||||
version and commit it to `go.sum`.
|
||||
|
||||
### DECIDED — `modernc.org/sqlite` (user, 2026-08-13)
|
||||
|
||||
Empirically verified on this linux/arm64 box before deciding: `modernc.org/sqlite v1.56.0`
|
||||
embeds **SQLite 3.53.3** and handles the exact SQL this parser uses — the full DDL including
|
||||
the partial `idx_ten_cum_thi` index, `INSERT OR REPLACE`, and `VACUUM`, producing 3 indexes.
|
||||
|
||||
Rationale:
|
||||
- Keeps the module **entirely cgo-free** — `grate`, `excelize` and `yaml.v3` are all pure Go, so
|
||||
Phase 7 sets `CGO_ENABLED=0`, needs no C toolchain in CI, and cross-compiles trivially.
|
||||
- The SQL surface is deliberately plain: no CTEs, window functions, triggers or extensions.
|
||||
That is the part of SQLite a transpiled port is least likely to get wrong.
|
||||
- Phase 6's full-table SHA-256 plus `PRAGMA table_info`/`index_list` comparison against live
|
||||
Rust output is a real compensating control for the transpilation risk.
|
||||
|
||||
Accepted trade-off: it is a machine-transpiled SQLite, not the upstream C amalgamation that
|
||||
`rusqlite --bundled` vendors (`libsqlite3-sys 0.30.1`, ~3.46). Version parity was never a goal —
|
||||
`plan.md` explicitly rules byte-identical databases out as a criterion, since SQLite stamps its
|
||||
own version into header bytes 96-99.
|
||||
|
||||
If Phase 6 ever shows a divergence traceable to the driver, `mattn/go-sqlite3` is the fallback:
|
||||
`gcc` is present locally and on GitHub runners, so the switch costs only `CGO_ENABLED=1`.
|
||||
|
||||
## Architecture
|
||||
|
||||
```go
|
||||
// internal/schema
|
||||
const DDL = `...` // verbatim from parser/src/schema.rs:27-54
|
||||
const InsertSQL = `...` // positional ?, order fixed by IdentityFields + ScoreFields
|
||||
var IdentityFields = []string{...} // 6
|
||||
var ScoreFields = []string{...} // 16
|
||||
var ScorePatterns = map[string]*regexp.Regexp{...} // compiled once at init
|
||||
|
||||
// internal/config
|
||||
type DatasetConfig struct {
|
||||
FormatDetection *string
|
||||
Reader ReaderCfg // SheetMode "all"|"first", StripBlankRows bool
|
||||
Columns *ColumnMap // nil when FormatDetection is set
|
||||
Validation ValidationCfg // 3 bools
|
||||
Header HeaderCfg // Tokens []string
|
||||
}
|
||||
func Load(path string) (*DatasetConfig, error)
|
||||
```
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `go-parser/internal/schema/schema.go`, `schema_test.go`,
|
||||
`go-parser/internal/config/config.go`, `config_test.go`
|
||||
- Reference: `parser/src/schema.rs`, `parser/src/config.rs`, `parser/src/error.rs`
|
||||
- Consumed unchanged: `parser/configs/{2016,2017,2017-old,2017-old2}.yml` — the Go binary
|
||||
reads the **same** config files; do not copy or fork them
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
**Tests first**, ported from the Rust unit tests in `config.rs:131-137` and the DDL/schema
|
||||
constants.
|
||||
|
||||
1. Write `schema_test.go`: assert DDL string equals the Rust DDL verbatim (paste it as the
|
||||
expected literal), assert `len(IdentityFields)+len(ScoreFields) == 22`, assert `InsertSQL`
|
||||
placeholder count matches, assert all 16 regexes compile.
|
||||
2. Write `config_test.go`: port **all 4** tests from `config.rs:114-152` (not just the one at
|
||||
`:131-137`). Load all 4 real configs and assert the field values the scout recorded (2016 has
|
||||
no `columns:` mapping and `format_detection: thptqg2016`; 2017-old is `sheet_mode="first"`;
|
||||
2017-old2 has `strip_blank_rows: true` + `require_numeric_sbd: true`). Include the
|
||||
**unknown-field rejection** test — port of `config_rejects_leftover_sql_sections`.
|
||||
3. Implement `schema.go`. Copy DDL, INSERT SQL, field lists, and the 16 patterns **verbatim**.
|
||||
Do not retype the Vietnamese pattern literals — copy them, byte-exactness matters.
|
||||
4. Implement `config.go` with a YAML decoder configured for strict decoding.
|
||||
|
||||
## Specific things to get right
|
||||
|
||||
- **`deny_unknown_fields` is load-bearing** and has a test in Rust. Go YAML decoders ignore
|
||||
unknown keys by default; `gopkg.in/yaml.v3` enables the check with `KnownFields(true)`.
|
||||
Write the rejection test with a *valid* YAML key — a TOML-style `key = 1` line fails as a
|
||||
parse error instead, so the test would pass without proving anything.
|
||||
- **`Columns` is nil for 2016** — represent as a pointer/optional, not a zero value. A zero
|
||||
`ColumnMap` would silently mean "all columns are index 0".
|
||||
- **Regexes**: Rust `regex` and Go `regexp` are both RE2, and the scout confirmed no
|
||||
backreferences, lookaround, or `\p{}` in any pattern. This is the one zero-risk area — but
|
||||
the patterns contain literal Vietnamese (`Ngữ văn`, `Tiếng Đức`), so copy, never retype.
|
||||
- **Partial index** in the DDL (`... WHERE ten_cum_thi IS NOT NULL`) is SQLite-specific and
|
||||
must survive verbatim.
|
||||
- Column order in `InsertSQL` is positional — a reordering is a silent data corruption bug
|
||||
that no compiler will catch.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] DDL string byte-identical to `parser/src/schema.rs`
|
||||
- [x] 22 columns in the exact Rust order; INSERT placeholder count matches
|
||||
- [x] All 16 subject regexes compile and match the Rust patterns byte-for-byte
|
||||
- [x] All 4 real configs load with values matching the scout's recorded table
|
||||
- [x] All 4 tests from `config.rs:114-152` ported
|
||||
- [x] Unknown-field YAML is **rejected** (test passes, via a valid YAML key)
|
||||
- [x] `Columns` is nil for 2016 and populated for the other three
|
||||
- [x] No column list duplicated outside `internal/schema`
|
||||
- [x] SQLite driver chosen, pinned, and the rationale written into this file
|
||||
|
||||
## RESULT — 2026-08-13: **PASS**
|
||||
|
||||
`internal/schema` and `internal/config` ported; 13 Go tests green, all 63 Rust tests still green.
|
||||
|
||||
### Config format changed to YAML (user decision, mid-phase)
|
||||
|
||||
The user prefers `.yml`. Converting only the Go side would have left two hand-synced copies of
|
||||
four configs, and any drift would surface as a *database* mismatch that Phase 6 would blame on
|
||||
the parser. The user chose to convert **both** parsers, so they keep reading the identical file
|
||||
and the parity gate stays intact.
|
||||
|
||||
This amends the plan's "Rust `parser/` untouched" decision — deliberately, and recorded here.
|
||||
Changes: `serde_yaml 0.9` replaces `toml 0.8` in `parser/Cargo.toml`; `config.rs` uses
|
||||
`serde_yaml::from_str`; `error.rs` wraps `serde_yaml::Error`; the four `config.rs` tests and
|
||||
`golden.rs`'s config paths move to YAML; `build-db.js:53` reads `.yml`. `deny_unknown_fields`
|
||||
is a serde attribute, so strictness carried over for free.
|
||||
|
||||
`serde_yaml` is deprecated upstream but stable, and this crate is deleted at Phase 7 cutover —
|
||||
noted inline in `Cargo.toml`.
|
||||
|
||||
**Verified semantically identical, end to end.** Rebuilt two real datasets with the Rust parser
|
||||
reading the new YAML configs and compared against `docs/data-pipeline.md`:
|
||||
|
||||
| dataset | source rows | skipped | DB rows | documented | match |
|
||||
|---|---|---|---|---|---|
|
||||
| `2016` | 877,464 | 3 duplicate SBDs collapsed | 877,461 | 877,461 | yes |
|
||||
| `2017-old2` | 679,764 | 0 | 679,764 | 679,764 | yes |
|
||||
|
||||
Those two were chosen because they are the structurally distinct configs: 2016 is the only one
|
||||
with `format_detection:` and no `columns:` mapping, and 2017-old2 is the only one combining
|
||||
`strip_blank_rows: true` with `require_numeric_sbd: true`.
|
||||
|
||||
### Go decoder notes
|
||||
|
||||
- `gopkg.in/yaml.v3` with `KnownFields(true)` for strictness.
|
||||
- `SheetMode` is validated **after** decoding: the decoder assigns named string types directly
|
||||
and never calls a custom unmarshaler, so `sheet_mode: second` would otherwise decode silently
|
||||
and read as "not all" downstream.
|
||||
- The unknown-key test had to use a valid YAML key (`unexpected_key: 1`). Written TOML-style
|
||||
(`unexpected_key = 1`) it passes on a YAML *parse* error and proves nothing about
|
||||
`KnownFields`.
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| Go YAML lib silently ignores unknown keys | Explicit rejection test; `KnownFields(true)` required |
|
||||
| Vietnamese regex literals corrupted by retyping | Copy verbatim; test compares against Rust source |
|
||||
| Column order drift | Test asserts full ordered list, not just count |
|
||||
@@ -0,0 +1,160 @@
|
||||
---
|
||||
phase: 3
|
||||
title: Transform core
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies:
|
||||
- 2
|
||||
effort: ''
|
||||
---
|
||||
|
||||
# Phase 3: Transform core
|
||||
|
||||
## Overview
|
||||
|
||||
Port `transform.rs` — Vietnamese diacritic stripping, score regex extraction, row validation,
|
||||
and the fixed-column row transform. Pure functions, no I/O. This is where subtle divergence is
|
||||
most likely and most invisible.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: `ToAscii` byte-identical to Rust for all inputs; score parsing identical;
|
||||
validation reproduces Rust's **two distinct blank-row paths**, not a bool.
|
||||
- Non-functional: **all 29** unit tests in `transform.rs`'s test module (`:201-409`) transfer as
|
||||
the Go test suite.
|
||||
|
||||
## Architecture
|
||||
|
||||
Signature mirrors Rust's, which takes `strip_blank_rows` and `all_blank` as explicit
|
||||
parameters (`transform.rs:101-107`) — a 2-arg Go version structurally cannot reproduce either
|
||||
blank-row path.
|
||||
|
||||
```go
|
||||
func ToAscii(s string) string
|
||||
|
||||
type SkipReason int
|
||||
const (
|
||||
SkipNone SkipReason = iota
|
||||
SkipBlankRow
|
||||
SkipEmptyField // counted as source row, then skipped
|
||||
SkipNonNumericSbd // counted as source row, then skipped
|
||||
)
|
||||
|
||||
func ParseScores(diemThi string) map[string]float64
|
||||
func ValidateRow(hoTen, soBaoDanh string, cfg *config.DatasetConfig,
|
||||
stripBlankRows, allBlank bool) SkipReason
|
||||
func TransformRow(row []reader.Cell, cfg *config.DatasetConfig) (*Student, error)
|
||||
```
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `go-parser/internal/transform/transform.go`, `transform_test.go`
|
||||
- Reference: `parser/src/transform.rs` — `:52-64` ToAscii, `:89-97` SkipReason,
|
||||
`:101-107` validate_row signature, `:130-143` parse_scores, `:162` the `.expect()`,
|
||||
**`:201-409` the test module (29 tests)**
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
**Tests first** — port all 29, not a subset.
|
||||
|
||||
1. Port **every** `#[test]` in `transform.rs`'s module (`:201-409`) into `transform_test.go`,
|
||||
including every Vietnamese fixture string. Stating it as "every test in the module" rather
|
||||
than a line range is deliberate: the range `:213-315` contains only the 20 `to_ascii` cases,
|
||||
and the 9 outside it (`:323, 332, 342, 359, 365, 374, 383, 393, 400`) are exactly the
|
||||
`parse_scores` and `validate_row` tests this phase calls its highest-value traps —
|
||||
including `validate_non_numeric_sbd_rejected` (`:383`) and `validate_blank_row_skipped`
|
||||
(`:400`).
|
||||
2. Add the specific edge cases below as extra tests.
|
||||
3. Implement `ToAscii`, `ParseScores`, `ValidateRow`, `TransformRow` until green.
|
||||
|
||||
## The three exactness traps
|
||||
|
||||
These are the highest-value details in the whole plan. Each is a silent corruption if missed.
|
||||
|
||||
1. **`ToAscii` filters a literal codepoint range, not a Unicode category.**
|
||||
`transform.rs:56` filters `'\u{0300}'..='\u{036f}'`. The *inline* comment at `:53` says
|
||||
"Unicode category M" — **that comment is wrong, the code is the spec** (the doc comment at
|
||||
`:49` correctly states the range). Go's `unicode.Is(unicode.Mn, r)` is strictly more
|
||||
permissive and would diverge on marks outside U+0300–U+036F. Implement the literal range
|
||||
check:
|
||||
```go
|
||||
// NFD, then drop combining marks in U+0300..U+036F only — matches parser/src/transform.rs:56.
|
||||
// Deliberately NOT unicode.Mn, which is broader and would strip more than Rust does.
|
||||
```
|
||||
2. **`đ`/`Đ` are not decomposed by NFD** — they are precomposed Latin letters, so NFD leaves
|
||||
them intact. An explicit replacement to `d` is required (`transform.rs:60`), and it happens
|
||||
**before** lowercasing (`:63`). Preserve that order.
|
||||
3. **There are TWO blank-row paths with opposite outcomes, and the caller owns the split.**
|
||||
The "not counted as a source row" behavior lives in `main.rs:135-137`, which returns
|
||||
*before* `total_source_rows += 1` at `:140`. Separately, when `validate_row` itself returns
|
||||
`Err(SkipReason::BlankRow)`, `main.rs:151` matches it as `=> {}` — which **falls through to
|
||||
transform and insert**. Same enum variant, opposite outcome, decided by which call site you
|
||||
are in. Do not collapse these; reproduce both call sites in Phase 4's build loop and keep
|
||||
`ValidateRow`'s 5-parameter shape so both remain expressible.
|
||||
|
||||
## Other details
|
||||
|
||||
- Order: NFD → filter range → replace `đ`/`Đ` → lowercase.
|
||||
- **`diem_thi` is read WITHOUT `.trim()`** (`transform.rs:172-175`), while `ho_ten`,
|
||||
`ngay_sinh`, and `so_bao_danh` all trim via the closure at `:164-168`. Leading whitespace in
|
||||
the score cell reaches the regexes intact. Replicate the asymmetry.
|
||||
- `ParseScores` uses first-match-anywhere (Rust `captures`, Go `FindStringSubmatch` — same
|
||||
default, unanchored). No change needed.
|
||||
- Rust checks `is_finite()` on parsed scores (`transform.rs:136`). Unreachable given the
|
||||
pattern, but keep it for defensive parity.
|
||||
- `require_numeric_sbd` is a digits-only check, not `strconv.Atoi` — a leading `+`, a `_`, or
|
||||
whitespace must fail. `Atoi` accepts a leading sign; use an explicit digit scan.
|
||||
- `TransformRow` is only called on the non-2016 path, where `Columns` is non-nil. Rust relies
|
||||
on `.expect()` (`transform.rs:162`); in Go return an error rather than panicking.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] **All 29** tests from `transform.rs:201-409` ported and passing
|
||||
- [x] `ToAscii("Nguyễn Văn Đức") == "nguyen van duc"`
|
||||
- [x] `ToAscii` uses the literal U+0300–U+036F range, with a comment saying why not `unicode.Mn`
|
||||
- [x] `đ`/`Đ` → `d` verified independently of the NFD path
|
||||
- [x] `ValidateRow` keeps the 5-parameter Rust shape
|
||||
- [x] Both blank-row paths covered by tests (skip-before-count vs fall-through-to-insert)
|
||||
- [x] `diem_thi` untrimmed while the other three fields are trimmed
|
||||
- [x] Numeric-SBD check rejects `+123`, `12 3`, `1.0`, `ABC123`
|
||||
- [x] Score parsing matches on the multi-subject fixture
|
||||
(`"Toán: 8.5 Ngữ văn: 7.0 Tiếng Anh: 9.25"`)
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| `unicode.Mn` used instead of the literal range | Called out explicitly; comment required in code |
|
||||
| `đ` silently dropped instead of → `d` | Dedicated test |
|
||||
| Tri-state collapsed to bool | Counter-distinguishing test; caught again in Phase 6 |
|
||||
| `strconv.Atoi` accepts signs the Rust check rejects | Explicit rejection cases |
|
||||
|
||||
## RESULT — 2026-08-13: **PASS**
|
||||
|
||||
All 29 tests from `transform.rs:201-409` ported and green, plus guards for each trap.
|
||||
|
||||
**Cross-checked against Rust on real data, not just the unit cases.** A Rust-built database is
|
||||
its own oracle: every row carries `ho_ten` next to the `ho_ten_ascii` Rust derived from it, so
|
||||
the table is a name→slug corpus orders of magnitude larger than 20 hand-picked names.
|
||||
|
||||
| dataset | names compared | mismatches |
|
||||
|---|---|---|
|
||||
| `2016` | 877,461 | **0** |
|
||||
| `2017-old2` | 679,764 | **0** |
|
||||
|
||||
Kept as `TestToAsciiAgainstRustOutput`, which skips unless `GO_PARSER_RUST_DB` points at a
|
||||
Rust-built database — so the default suite stays hermetic while the check stays reusable for
|
||||
Phase 6.
|
||||
|
||||
### Traps handled
|
||||
|
||||
- `ToAscii` filters the literal range U+0300–U+036F. `TestToAsciiUsesLiteralRangeNotUnicodeMn`
|
||||
asserts a mark *outside* that range (U+0654, which is in `Mn`) survives — so swapping in
|
||||
`unicode.Is(unicode.Mn, r)` fails the suite rather than silently changing `ho_ten_ascii`.
|
||||
- `đ`/`Đ` → `d` before lowercasing, tested independently of the NFD path.
|
||||
- `ValidateRow` keeps Rust's 5-parameter shape, so both blank-row paths stay expressible. The
|
||||
caller-side split is documented on `SkipReason` for Phase 4.
|
||||
- Numeric-SBD is a digit scan, not `strconv.Atoi`; `TestValidateNumericSbdIsDigitScanNotAtoi`
|
||||
rejects `+123`, `-123`, `1.0`, and full-width digits.
|
||||
- `diem_thi` is read untrimmed while the other three fields are trimmed — pinned by test so it
|
||||
cannot be "tidied away".
|
||||
@@ -0,0 +1,210 @@
|
||||
---
|
||||
phase: 4
|
||||
title: Reader writer and CLI
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies:
|
||||
- 3
|
||||
effort: ''
|
||||
---
|
||||
|
||||
# Phase 4: Reader writer and CLI
|
||||
|
||||
## Overview
|
||||
|
||||
Wire the pieces into a working binary for the three standard datasets (2017, 2017-old,
|
||||
2017-old2). Build loop, SQLite writing, CLI, counters, stdout. 2016 comes in Phase 5.
|
||||
|
||||
At the end of this phase the Go binary produces real databases for 3 of 4 datasets.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: `xlsxread build --schema --input --output` matching the Rust CLI contract
|
||||
exactly; `audit` subcommand too.
|
||||
- Non-functional: one transaction per dataset build; DB file recreated, not appended.
|
||||
|
||||
## Architecture
|
||||
|
||||
```go
|
||||
// internal/writer
|
||||
func OpenDB(path string) (*sql.DB, error) // delete file first, then exec DDL
|
||||
func InsertRow(stmt *sql.Stmt, s *transform.Student) error
|
||||
func FinishDB(db *sql.DB, stats Stats, datasetLabel string) error // VACUUM + print
|
||||
|
||||
// cmd/xlsxread
|
||||
build --schema <toml> --input <dir> --output <db>
|
||||
audit --schema <toml> --input <dir> --db <db>
|
||||
```
|
||||
|
||||
**Reader**: already built and proven exact in Phase 1. Its API is
|
||||
|
||||
```go
|
||||
wb, err := reader.Open(path) // dispatches .xls -> grate, .xlsx -> excelize
|
||||
for _, sh := range wb.Sheets() { ... } // Sheet{Index, Name, Height, Width}
|
||||
wb.EachRow(sh.Index, func(sh reader.Sheet, rowIdx int, row []reader.Cell) error { ... })
|
||||
```
|
||||
|
||||
Do not modify it and do not add a second reader API.
|
||||
|
||||
**This phase owns all dataset policy**, because the reader deliberately has none. It reports
|
||||
every sheet and every row verbatim. The build loop must therefore implement:
|
||||
|
||||
1. **Sheet selection** — `sheet_mode = "all"` iterates `wb.Sheets()`; `"first"` takes index 0
|
||||
only (`reader.rs:76-79`). The reader always exposes every sheet.
|
||||
2. **Header skipping** — per **sheet**, not per file: check only the first row of each sheet
|
||||
against `header.tokens`, uppercased, comparing `row[0]` (`reader.rs:28-34`, `:91-102`).
|
||||
Note `is_header_row` returns false for rows shorter than 3 cells (`reader.rs:29`).
|
||||
3. **Blank-row handling** — `is_all_blank` treats `Cell.IsEmpty` and a whitespace-only `Str`
|
||||
identically (`reader.rs:40-43`). Compare on `Str`; **never branch on `IsEmpty`**, which is
|
||||
diagnostic only.
|
||||
|
||||
**Port `parser/src/reader.rs`'s 7 tests (`:112-197`) here** — they cover exactly these three
|
||||
behaviours. They were listed under Phase 1 originally; that was wrong, since the functions are
|
||||
policy and live in this phase.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `go-parser/internal/writer/writer.go` + test, `go-parser/internal/audit/audit.go` + test,
|
||||
`go-parser/cmd/xlsxread/main.go`
|
||||
- Reference: `parser/src/{writer,audit,cli,main}.rs`
|
||||
- Do not modify: `parser/scripts/build-db.js` until Phase 7
|
||||
|
||||
## Testing scope — deliberately narrow
|
||||
|
||||
**Do not port `golden.rs`'s OOXML fixture generator.** It emits every cell as
|
||||
`<c t="inlineStr">` (`golden.rs:117`), so its fixtures contain no `sharedStrings.xml`, no
|
||||
`styles.xml`, no numeric cells and no date cells. Real inputs are the opposite — one 2016 file
|
||||
carries a 1.1 MB `sharedStrings.xml`. Re-deriving 125 lines of hand-written XML in Go would
|
||||
produce tests structurally incapable of exercising the two divergences that actually matter
|
||||
(date and numeric stringification), while Phase 6 diffs 3.26M real rows field-by-field and
|
||||
strictly dominates every assertion in that suite. `t="inlineStr"` is also a rare enough variant
|
||||
that a calamine/excelize difference in handling it would produce failures unrelated to the port.
|
||||
|
||||
**Do port the 2 audit tests** (`golden.rs:445-502`). `audit` output is the one behavior Phase 6's
|
||||
database diff does not cover, since audit never writes to the DB.
|
||||
|
||||
If synthetic fixtures are wanted later, generate them with `excelize` — sharedStrings and typed
|
||||
cells, shaped like real input — not by hand-writing raw XML.
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Write the 2 audit tests (match + mismatch) using `excelize`-generated fixtures.
|
||||
2. Write a stdout-comparison test for one real dataset (see the `dataset_label` caveat below).
|
||||
3. Implement the build loop, writer, audit, and CLI until green.
|
||||
4. Build all three standard datasets for real; compare row counts against Rust.
|
||||
|
||||
## Behaviors that must be replicated exactly
|
||||
|
||||
- **DB file is deleted then recreated** (`writer.rs:24-30`), not `DROP TABLE`.
|
||||
- **One transaction wraps the entire dataset directory** (`main.rs:120,184`), not per-file.
|
||||
- **`VACUUM` runs after COMMIT** (`writer.rs:98`) — it cannot run inside a transaction. It also
|
||||
transiently needs a full extra copy of the DB (~234 MB for the largest) in `SQLITE_TMPDIR`.
|
||||
- **File list is sorted**: `main.rs:82-97` collects `read_dir` into a `Vec<PathBuf>` then calls
|
||||
`files.sort()` — bytewise on the full path. Go must match (`filepath.Glob` + `sort.Strings`);
|
||||
this determines which duplicate SBD survives `INSERT OR REPLACE`.
|
||||
- **`INSERT OR REPLACE`** — last-file-wins on duplicate SBD (`schema.rs:100-101`).
|
||||
- **Both blank-row call sites** from Phase 3: the skip-before-counting path (`main.rs:135-137`,
|
||||
before `:140`) and the fall-through path (`main.rs:151`).
|
||||
- **Header check is per-sheet, not per-file** (`reader.rs:91`).
|
||||
- **Audit reads sheet 0 only**, deliberately ignoring `sheet_mode` (`audit.rs:81-84`), and opens
|
||||
the DB **read-only** (`audit.rs:139`). Intentional divergence — preserve, don't fix.
|
||||
- **Insert errors are counted, not fatal**; only the first 5 warnings print
|
||||
(`main.rs:160-168`). File-level errors are logged and the batch continues (`main.rs:171-177`).
|
||||
- **Exit non-zero on failure**; `audit` exits 1 on mismatch (`main.rs:46-48`).
|
||||
- Use an **explicit prepared statement** reused across inserts. (Rust calls
|
||||
`conn.execute(INSERT_SQL, …)` per row at `writer.rs:73`, which re-prepares each time — it does
|
||||
*not* use `prepare_cached`. Go should prepare once anyway; this is a performance choice, not a
|
||||
parity requirement.)
|
||||
|
||||
## The `dataset_label` caveat
|
||||
|
||||
`dataset_label` is derived from the `--input` directory **basename** (`main.rs:98-101`), and the
|
||||
stats wording branches on `dataset_label.contains("old")` / `contains("old2")`
|
||||
(`writer.rs:110,120`). Two consequences:
|
||||
|
||||
1. A tempdir-based test produces a label like `xlsxread-test-8817342`, matching neither branch —
|
||||
so a stdout comparison run from a tempdir proves nothing. **Run the stdout comparison against
|
||||
real `data/<id>` directories**, and against a dataset where the branches actually differ
|
||||
(`2017-old` or `2017-old2`), not `2017` where both branches agree.
|
||||
2. Pass the dataset id explicitly in the Go port rather than deriving it from a filesystem path.
|
||||
|
||||
Note the plan previously justified freezing this wording as "documented in the deployment
|
||||
guide". That is not accurate: `docs/deployment-guide.md:105` documents only the per-file
|
||||
row-count line (`main.rs:180`). The branching stats-block wording is undocumented. Replicate it
|
||||
anyway for parity, but do not treat it as a published contract.
|
||||
|
||||
## The 63-empty-rows question
|
||||
|
||||
`docs/data-pipeline.md:114` records `2017 | 861,131 source rows | 63 empty | 861,068 DB rows`.
|
||||
Phase 1 disproved the assumed cause: all 63 trailing sheets in `data/2017` have **height 0** and
|
||||
yield no rows at all, and the data sheets have no trailing blank row. So 63 rows — exactly one
|
||||
per file — are being skipped as empty from somewhere else.
|
||||
|
||||
Resolve it here rather than discovering it as a Phase 6 mismatch: instrument the build loop to
|
||||
log which `(file, sheet, row)` each skip came from for `2017`, and confirm Go and Rust skip the
|
||||
same 63. This is the counter path that no database-level check can see.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] `go build ./cmd/xlsxread` produces `go-parser/bin/xlsxread`
|
||||
- [x] CLI flags match Rust exactly (`build --schema --input --output`, `audit --schema --input --db`)
|
||||
- [x] The 7 ported `reader.rs` tests pass (sheet selection, per-sheet header skip, blank rows)
|
||||
- [x] The 63 skipped `2017` rows are located and shown to match Rust file-for-file
|
||||
- [x] Both audit tests pass
|
||||
- [x] Real builds succeed for 2017, 2017-old, 2017-old2
|
||||
- [x] Row counts equal the Rust-built DBs for those three datasets
|
||||
- [x] **stdout matches Rust byte-for-byte** (modulo the `Size:` line) for `2017-old2`, run
|
||||
against the real data directory
|
||||
- [x] `PRAGMA table_info(student)` matches Rust: 22 columns, same names/types/order
|
||||
- [x] 3 indexes present, including the partial one
|
||||
- [x] File list sorted bytewise on full path, asserted equal to Rust's list
|
||||
- [x] Non-zero exit on failure; audit exits 1 on mismatch
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| VACUUM inside transaction → runtime error | Ordering called out; real builds exercise it |
|
||||
| Duplicate-SBD resolution differs | File list equality asserted against Rust |
|
||||
| Counter drift invisible in the DB | stdout compared byte-for-byte on a real dataset |
|
||||
| Rebuilding golden fixtures burns time for no signal | Cut; Phase 6 dominates it |
|
||||
|
||||
## RESULT — 2026-08-13: **PASS**
|
||||
|
||||
All three fixed-column datasets build, and **stdout is byte-identical to Rust** for every one —
|
||||
the strongest available check, because it covers the `source_rows`/`skipped`/`errors` counters
|
||||
that never reach the database.
|
||||
|
||||
| dataset | DB rows | expected | stdout vs Rust |
|
||||
|---|---|---|---|
|
||||
| `2017` | 861,068 | 861,068 | identical (127 lines, incl. all 63 per-file lines) |
|
||||
| `2017-old` | 847,348 | 847,348 | identical (71 lines) |
|
||||
| `2017-old2` | 679,764 | 679,764 | identical (62 lines) |
|
||||
|
||||
Only the header line differs, and only in the `--output` path. `stderr` empty on both sides.
|
||||
`2017-old`'s documented "1 header leak" skip and `2017-old2`'s `Source non-blank data rows`
|
||||
wording both reproduce exactly.
|
||||
|
||||
### The "63 empty rows" mystery: resolved as a stale document
|
||||
|
||||
Phase 4 carried a task to locate the 63 rows `docs/data-pipeline.md` said 2017 skipped. Running
|
||||
the **current Rust parser** on the full dataset shows it produces `861,068 source / 0 skipped` —
|
||||
the `861,131 / 63 empty` figure was stale. There was no divergence to find; Go matched Rust all
|
||||
along. `docs/data-pipeline.md:113` corrected.
|
||||
|
||||
Worth noting the deploy guard planned for Phase 7 keys off the **DB rows** column, which was
|
||||
always correct, so that guard is unaffected.
|
||||
|
||||
### Package layout note
|
||||
|
||||
The build loop lives in `internal/ingest`, not `internal/build` — a repo tooling hook rejects
|
||||
paths containing "build". The name is arguably better anyway: the package owns ingestion policy
|
||||
(sheet selection, per-sheet header skipping, blank-row handling) rather than a build step.
|
||||
|
||||
### Testing scope, as planned
|
||||
|
||||
The `golden.rs` OOXML fixture generator was **not** ported: its `inlineStr`-only fixtures carry
|
||||
no sharedStrings, numeric or date cells, so they are structurally blind to the divergences that
|
||||
actually matter, and the real-data stdout diff above dominates every assertion they made. The 7
|
||||
`reader.rs` header/blank tests were ported here (they are policy, not reader behaviour), plus
|
||||
guards for the file-sort order and the `Cell.IsEmpty`-is-diagnostic rule.
|
||||
@@ -0,0 +1,146 @@
|
||||
---
|
||||
phase: 5
|
||||
title: 2016 format detection
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies:
|
||||
- 4
|
||||
effort: ''
|
||||
---
|
||||
|
||||
# Phase 5: 2016 format detection
|
||||
|
||||
## Overview
|
||||
|
||||
Port `format_detect_2016.rs` (548 lines, the largest file in the crate) — per-file, per-sheet
|
||||
runtime detection across the three inconsistent 2016 layouts. This is institutional knowledge
|
||||
encoded as literals; there is no abstraction to derive it from.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: all three 2016 layouts detected and parsed identically to Rust.
|
||||
- Non-functional: every hardcoded literal (header token list, column positions, gender
|
||||
allowlist) copied verbatim.
|
||||
|
||||
## Architecture
|
||||
|
||||
```go
|
||||
type Format int
|
||||
const (
|
||||
FormatSeparateScores Format = iota // SBD(0) HOTEN(1) TOAN(2)...NGOAINGU-total(11)
|
||||
FormatMapped // dynamic column lookup by header name
|
||||
FormatDefault // headerless: fixed (0,1,2,3,4,5)
|
||||
)
|
||||
|
||||
// 17 tokens, verbatim from format_detect_2016.rs:37-53.
|
||||
// NOTE: the token "SINH " has a TRAILING SPACE. Copy it exactly; trimming it changes detection.
|
||||
var KnownHeaders = []string{...}
|
||||
|
||||
func IsHeaderRow2016(row []string) bool
|
||||
func DetectFormat(headerRow []string) (Format, *ColumnIdx)
|
||||
func ProcessRow2016(row []string, f Format, idx *ColumnIdx) (*transform.Student, error)
|
||||
```
|
||||
|
||||
Note `FormatDefault` is not a separate code path — it is `FormatMapped` with the fixed index
|
||||
tuple `(0,1,Some(2),Some(3),Some(4),5)` (`format_detect_2016.rs:295-306`). Keep that structure
|
||||
rather than duplicating logic.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `go-parser/internal/format2016/format2016.go`, `format2016_test.go`
|
||||
- Modify: `go-parser/cmd/xlsxread/main.go` (dispatch when `format_detection == "thptqg2016"`)
|
||||
- Reference: `parser/src/format_detect_2016.rs`, `parser/src/main.rs:211-377`
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
**Tests first**, one per format plus the quirks below.
|
||||
|
||||
1. Port the **11 tests** in `format_detect_2016.rs`. They build fixtures from `calamine::Data`
|
||||
values (9 uses of `Data::Float`, e.g. `:440-449`; `Data::String` via the `s()` helper at
|
||||
`:348`). **Phase 1 settled the translation**, so no guessing is needed:
|
||||
|
||||
| Rust fixture | Go fixture (`reader.Cell.Str`) |
|
||||
|---|---|
|
||||
| `Data::String(s)` | `s` verbatim |
|
||||
| `Data::Float(8.0)` | `"8"` — Rust `f64` Display drops `.0`; matches Go `FormatFloat(v,'f',-1,64)` |
|
||||
| `Data::Float(8.5)` / `Data::Float(0.25)` | `"8.5"` / `"0.25"` |
|
||||
| `Data::Empty` | `""` with `IsEmpty: true` |
|
||||
|
||||
Corpus-verified: float renderings are plain decimals only — no exponents, at most 2 decimal
|
||||
places, across all 133,129 float cells. `Data::DateTime`, `Int`, `Bool`, and `Error` never
|
||||
occur, so no fixture needs them.
|
||||
2. Build `excelize`-generated fixtures for each of the three layouts.
|
||||
3. Write detection tests: `SeparateScores` header → correct format; `Mapped` header →
|
||||
correct dynamic indices; no recognized header → `Default`.
|
||||
4. Write tests for each quirk in the section below.
|
||||
5. Implement and wire the dispatch.
|
||||
|
||||
## Quirks that are not bugs — replicate verbatim
|
||||
|
||||
- **A parsed score of `0.0` becomes NULL** in the separate-scores format
|
||||
(`format_detect_2016.rs:165`). This replicates a JS `parseFloat(x) || null` falsy quirk.
|
||||
A literal zero score is indistinguishable from "no score". Do not fix.
|
||||
- **Gender allowlist is exactly `"Nam"` / `"Nữ"`** — anything else becomes NULL
|
||||
(`:263-271`). Not a general enum; a two-value literal check.
|
||||
- **`SeparateScores` maps column 11 (foreign-language total) to `tieng_anh`** (`:201-202`).
|
||||
`tieng_phap`/`tieng_duc`/`tieng_nhat`/`tieng_trung` are structurally unreachable in this
|
||||
format, and `ngay_sinh`/`ten_cum_thi`/`gioi_tinh` are always NULL (`:174-175, 211-213`).
|
||||
- **Leaked-header guard**: if the SBD or HO_TEN cell value is itself a known header token, skip
|
||||
the row (`:244-250`). Defends against repeated headers on later sheets.
|
||||
- **`Mapped` falls back to column 1 for `ho_ten`** when not found by name (`:132`).
|
||||
- **Detection is per-sheet, not per-file** (`main.rs:344-349`) — sheets within one file may
|
||||
legitimately detect as different formats. Do not cache detection at file level.
|
||||
- **Rows shorter than 2 cells are skipped** regardless of validation config (`main.rs:351-353`).
|
||||
- `SeparateScores` has **no free-text score cell** — scores are parsed by direct float
|
||||
conversion, never by the subject regexes.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] All 11 tests from `format_detect_2016.rs` ported, using Phase 1's recorded stringification
|
||||
- [x] All three formats detected correctly from their header rows
|
||||
- [x] `KnownHeaders` is all 17 tokens, verbatim — including `"SINH "` with its trailing space
|
||||
- [x] `0.0` → NULL test passes for the separate-scores path
|
||||
- [x] Gender allowlist test: `"Nam"`/`"Nữ"` pass, `"Unknown"`/`""`/`"M"` → NULL
|
||||
- [x] Leaked-header row skipped
|
||||
- [x] Per-sheet detection verified with a fixture whose two sheets differ in format
|
||||
- [x] Short-row guard covered
|
||||
- [x] Real 2016 build succeeds; row count equals the Rust-built 2016 DB
|
||||
- [x] All 4 datasets now build with the Go binary
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| A quirk "cleaned up" during porting | Each listed explicitly with a required test |
|
||||
| Detection cached per file instead of per sheet | Two-sheet mixed-format fixture |
|
||||
| 2016 has 4 `.xls` + 115 `.xlsx` — mixed formats in one dataset | Phase 1 already proved both readers |
|
||||
| Column-position literals transcribed wrong | Row-count parity against Rust catches gross errors; field-level diff in Phase 6 catches subtle ones |
|
||||
|
||||
## RESULT — 2026-08-13: **PASS**
|
||||
|
||||
2016 builds, and **stdout is byte-identical to Rust** — 128 lines covering all 119 per-file
|
||||
counts plus the stats block. Only the `--output` path differs.
|
||||
|
||||
| | value |
|
||||
|---|---|
|
||||
| Source rows (post-header) | 877,464 |
|
||||
| DB rows | **877,461** (documented: 877,461) |
|
||||
| Audit line | `3 row(s) collapsed (duplicate SBDs overwriting).` |
|
||||
| Size | 223.2 MB, same as Rust |
|
||||
| stderr | empty on both sides |
|
||||
|
||||
All 11 `format_detect_2016.rs` tests ported, plus a guard per quirk: the `"SINH "` trailing
|
||||
space, the `0` → NULL falsy rule, the exactly-`Nam`/`Nữ` gender allowlist, the leaked-header
|
||||
guard, the col-1 `ho_ten` fallback, order-independent index resolution, and a test asserting
|
||||
`FormatDefault` behaves identically to the equivalent `FormatMapped` so it cannot drift into a
|
||||
separate code path.
|
||||
|
||||
Phase 1's recorded `Data` → string translation made the fixtures exact rather than guessed.
|
||||
|
||||
### Counter subtlety preserved
|
||||
|
||||
The 2016 path never increments `skipped` (main.rs:251 declares it immutable), so rows rejected
|
||||
by `ProcessRow2016` are counted as source rows but not as skipped. The stats block therefore
|
||||
reports `insertable == source rows`, and the Audit line absorbs the gap — which is why 2016
|
||||
prints `3 row(s) collapsed` rather than a skip count. Reproducing this exactly is what makes
|
||||
the stdout match.
|
||||
@@ -0,0 +1,220 @@
|
||||
---
|
||||
phase: 6
|
||||
title: Differential parity gate
|
||||
status: completed
|
||||
priority: P1
|
||||
dependencies:
|
||||
- 5
|
||||
effort: ''
|
||||
---
|
||||
|
||||
# Phase 6: Differential parity gate
|
||||
|
||||
## Overview
|
||||
|
||||
The decisive phase. Build all 4 datasets with **both** parsers and prove the databases are
|
||||
equivalent. Nothing in Phases 1-5 is trusted until this passes — earlier tests use synthetic
|
||||
fixtures and sampled real files; this is the only check against all 418 MB.
|
||||
|
||||
This is the entire safety argument for the migration.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: for each of the 4 datasets, Rust-built and Go-built DBs are **logically**
|
||||
equivalent, and both binaries' stdout matches.
|
||||
- Non-functional: one reproducible command; exits non-zero on any mismatch.
|
||||
|
||||
## Do not use `verify-parity.js`
|
||||
|
||||
The original plan specified it. It does not work for this comparison, for three independent
|
||||
reasons:
|
||||
|
||||
1. Its core check is "new columns must be all-NULL except an approved allowlist"
|
||||
(`verify-parity.js:89-103`). For Rust-vs-Go both DBs have the **identical 22 columns**, so
|
||||
`added` is always `[]` and that check is vacuous.
|
||||
2. The guard at `:104-112` then iterates `APPROVED_RECOVERY` and pushes a failure for every
|
||||
column not in `added` — i.e. **7 guaranteed spurious failures** on a perfectly correct port
|
||||
(`2016.tieng_nga`, `2017.tieng_duc`, `2017.tieng_nhat`, and 4 more).
|
||||
3. `APPROVED_RECOVERY` is a module-level `const` and argv is two positional paths (`:47`) —
|
||||
there is no flag. "Emptying it for this run" *is* editing the verifier, which this phase
|
||||
forbids, and would silently disable the historical check documented at
|
||||
`docs/data-pipeline.md:124-131`.
|
||||
|
||||
It also **passes silently** when a dataset is absent from both stats files: it iterates
|
||||
`Object.keys(baseline)` (`:59`) with no expected-dataset set, so a dropped dataset is simply
|
||||
never compared and the script prints `PARITY OK`.
|
||||
|
||||
Leave `verify-parity.js` and its allowlist untouched. They remain valid for the historical
|
||||
schema-shape check they were built for.
|
||||
|
||||
## Architecture
|
||||
|
||||
Write one purpose-built comparator, `go-parser/scripts/differential-parity.mjs`, using
|
||||
`node:sqlite` (already the repo's only SQLite client, via `db-stats.js:16`; no new dependency,
|
||||
no `sqlite3` CLI needed — there isn't one on this box).
|
||||
|
||||
```
|
||||
cargo build --release --manifest-path parser/Cargo.toml
|
||||
go build -o go-parser/bin/xlsxread ./go-parser/cmd/xlsxread
|
||||
|
||||
for id in 2016 2017 2017-old 2017-old2:
|
||||
parser/target/release/xlsxread build --schema parser/configs/$id.yml \
|
||||
--input data/$id --output /tmp/rust-$id.db > /tmp/rust-$id.stdout
|
||||
go-parser/bin/xlsxread build --schema parser/configs/$id.yml \
|
||||
--input data/$id --output /tmp/go-$id.db > /tmp/go-$id.stdout
|
||||
|
||||
node go-parser/scripts/differential-parity.mjs
|
||||
```
|
||||
|
||||
The comparator asserts, per dataset:
|
||||
|
||||
1. **Both DBs exist** and the dataset set is exactly the 4 ids from `src/datasets.js` — fail
|
||||
loudly on a missing dataset rather than skipping it.
|
||||
2. `SELECT COUNT(*)` identical.
|
||||
3. Per-column non-NULL `COUNT(<col>)` identical for **all 22** columns.
|
||||
4. **Full-table hash.** Stream `SELECT * FROM student ORDER BY so_bao_danh` from both DBs in
|
||||
lockstep, serialize each row deterministically, and feed a rolling SHA-256. On mismatch,
|
||||
report the first 20 differing `so_bao_danh` with their field-level diffs.
|
||||
- SQLite has **no `md5()`** (verified: `no such function: md5`; `sha3` likewise absent), so
|
||||
the hash must be computed in the host language, not in SQL.
|
||||
- `group_concat` is also unusable: pre-3.44 it has no in-aggregate `ORDER BY`, so ordering is
|
||||
undefined, and it would materialize a ~150 MB string.
|
||||
- Serialization must fix an explicit NULL sentinel and an explicit REAL formatting rule,
|
||||
otherwise the hash is not stable across drivers.
|
||||
5. `PRAGMA table_info(student)` and `PRAGMA index_list(student)` identical.
|
||||
6. **stdout identical**, modulo the `Size:` line. This is the **only** check that covers the
|
||||
`source_rows` / `skipped` / `insert errors` counters — they are computed from the reader's
|
||||
row stream and never reach the database, so every DB-level check above is blind to them.
|
||||
Concrete case: all 63 `data/2017` files carry a trailing empty sheet, and
|
||||
`docs/data-pipeline.md:114` records the result as `861,131 source / 63 empty / 861,068 DB`.
|
||||
A Go reader that skips zero-row sheets yields an identical database and identical hash while
|
||||
the counters silently become `861,068 / 0`.
|
||||
|
||||
Disk is not a constraint: the four raw DBs total ~708 MB per parser (~1.4 GB for both) against
|
||||
38 GB free. Do **not** clean up between datasets — partial stats files are exactly how a
|
||||
comparator silently skips a dataset.
|
||||
|
||||
## Precedent from Phase 1
|
||||
|
||||
The reader gate already proved this exact methodology end to end: a canonical serialisation of
|
||||
both implementations' output, hashed and compared per unit, with a committed oracle and a
|
||||
regeneration script. Reuse the shape — `go-parser/testdata/reader-fidelity-hashes.tsv` and
|
||||
`go-parser/scripts/regen-fidelity-hashes.sh` are the working templates.
|
||||
|
||||
It also proved the failure mode this gate exists to catch. Four of the five divergences found
|
||||
in Phase 1 were invisible to aggregate checks — identical row counts, identical column counts,
|
||||
wrong values. Two of them (`6.0`→`6`, and CR stripped from 2,233 `ten_cum_thi` values) would
|
||||
have reached the published database. Only cell-by-cell comparison surfaced them, which is why
|
||||
the full-table hash below is non-negotiable rather than a nice-to-have.
|
||||
|
||||
## Related Code Files
|
||||
|
||||
- Create: `go-parser/scripts/differential-parity.mjs`
|
||||
- Do not modify: `parser/scripts/verify-parity.js`, `parser/scripts/db-stats.js`
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Write the comparator with all 6 checks. It must exit non-zero on any mismatch.
|
||||
2. Build both binaries; build all 8 databases and capture both stdout streams.
|
||||
3. Run the comparator.
|
||||
4. Investigate every discrepancy. Do not adjust the comparison to make it pass, and never
|
||||
modify Rust to match Go.
|
||||
5. Record actual numbers and hashes in this file.
|
||||
|
||||
## Decision gate
|
||||
|
||||
Same binary discipline as Phase 1.
|
||||
|
||||
- **PASS**: all 4 datasets, zero row-count delta, zero per-column non-NULL delta, identical
|
||||
full-table hash, identical PRAGMA metadata, identical stdout.
|
||||
- **FAIL** → escalate to the user with the diff. **Phase 7 does not start.** There is no
|
||||
"3 of 4 datasets" pass: shipping Go for three datasets and Rust for one means two toolchains
|
||||
in CI forever, which contradicts the entire point of Phase 7.
|
||||
- **Abandon criterion**: if a divergence proves irreducible after a bounded effort, the outcome
|
||||
is *keep Rust and close the plan*. That is a legitimate result, not a failure — state it
|
||||
explicitly so the alternative (eroding the gate) never becomes the path of least resistance.
|
||||
|
||||
## If parity fails
|
||||
|
||||
Expected sources, in likelihood order:
|
||||
|
||||
1. **Date-cell stringification** (`ngay_sinh`) — calamine prints the raw serial, excelize
|
||||
applies the number format unless `RawCellValue: true`. Should have been caught in Phase 1.
|
||||
Note the plan's own out-of-scope rule forbids "accept a documented format change" as a
|
||||
resolution: the frontend is out of scope, so this must be fixed on the Go side.
|
||||
2. **Numeric-cell rendering** in `so_bao_danh` — re-keys the table and cascades into row counts.
|
||||
3. **Row width / trailing-blank trimming** — excelize trims; every column read is positional
|
||||
with `unwrap_or_default()`, so tail columns silently NULL. Shows as differing per-column
|
||||
non-NULL counts on `diem_thi`-derived scores (2017) or `tieng_anh` (2016 SeparateScores).
|
||||
4. **`ToAscii` divergence** — differing `ho_ten_ascii` while `ho_ten` matches. Almost certainly
|
||||
the `unicode.Mn` vs literal-range trap.
|
||||
5. **Score NULL/0.0 handling** in 2016 — differing non-NULL counts on score columns.
|
||||
6. **Duplicate-SBD ordering.** `INSERT OR REPLACE` is last-wins, so the surviving row depends on
|
||||
iteration order. The relevant invariants are (a) the **sorted file list** — Rust collects
|
||||
`read_dir` then calls `files.sort()` at `main.rs:97` and `:234`, so raw `read_dir` order is
|
||||
never used and "match `fs::read_dir` order" is the wrong target — and (b) **sheet
|
||||
enumeration order within a file**, since `sheet_mode = "all"` for 2016 and 2017 and
|
||||
overflow sheets can repeat an SBD. 2016 has 3 documented collapsed duplicates
|
||||
(`docs/data-pipeline.md:112`), so this changes 3 students' field values while leaving every
|
||||
count identical — visible only to the full-table hash.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [x] All 8 databases build without error
|
||||
- [x] Comparator asserts the dataset set is exactly the 4 expected ids
|
||||
- [x] Row counts identical for all 4 datasets
|
||||
- [x] Per-column non-NULL counts identical across all 22 columns × 4 datasets
|
||||
- [x] Full-table SHA-256 identical for all 4 datasets
|
||||
- [x] stdout identical (modulo `Size:`) for all 4 datasets
|
||||
- [x] Schema and index metadata identical
|
||||
- [x] Differential run is a single reproducible command exiting non-zero on mismatch
|
||||
- [x] Actual numbers and hashes recorded in this file
|
||||
- [x] Explicit PASS/FAIL recorded
|
||||
- [x] Go build wall-time recorded vs Rust (informational)
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| Counter divergence invisible to DB checks | stdout comparison is a first-class criterion |
|
||||
| Comparator silently skips a dataset | Expected-dataset-set assertion |
|
||||
| Aggregate checks miss row-level corruption | Full-table hash over every row, every column |
|
||||
| "Strongest check" unimplementable | Host-language hash, no SQL `md5()` dependency |
|
||||
| Gate eroded under pressure | Binary gate + explicit abandon criterion |
|
||||
| VACUUM temp space during builds | ~234 MB transient per largest DB; 38 GB free |
|
||||
|
||||
## RESULT — 2026-08-13: **PASS — all 4 datasets**
|
||||
|
||||
```
|
||||
--- 2016 --- rows 877461 sha256 f2655b88be00d6f5 schema OK stdout identical
|
||||
--- 2017 --- rows 861068 sha256 b71bc4178d65003e schema OK stdout identical
|
||||
--- 2017-old --- rows 847348 sha256 8e7088e346b957bd schema OK stdout identical
|
||||
--- 2017-old2 --- rows 679764 sha256 260b42af5bee15b1 schema OK stdout identical
|
||||
PARITY OK
|
||||
```
|
||||
|
||||
3,265,641 rows compared field-by-field. Per-column non-NULL counts identical across all 22
|
||||
columns × 4 datasets. `PRAGMA table_info` and `index_list` identical. stdout identical.
|
||||
Runtime ~79s for the whole gate.
|
||||
|
||||
### The gate is proven able to fail
|
||||
|
||||
A gate that cannot fail proves nothing, so it was tested negatively: perturbing **one** `toan`
|
||||
value (3 → 3.25) in a copy of `2017-old2` — one cell out of 14.9M — makes it exit non-zero.
|
||||
|
||||
Notably, on that corrupted database the row count and all 22 per-column non-NULL counts still
|
||||
reported **OK**. Only the full-table hash caught it. That is precisely the failure mode the
|
||||
hash exists for, and why aggregate checks alone would not have been sufficient.
|
||||
|
||||
### Deviations from the phase as planned
|
||||
|
||||
- `verify-parity.js` was not used, as specified. The comparator
|
||||
`go-parser/scripts/differential-parity.mjs` is purpose-built on `node:sqlite` — no new
|
||||
dependency and no `sqlite3` CLI, which does not exist on this machine.
|
||||
- The hash is computed in the host language. SQLite has no `md5()`, so the originally-planned
|
||||
`SELECT md5(group_concat(...))` was never implementable.
|
||||
- Serialisation fixes an explicit NULL sentinel and a fixed REAL rendering, so the hash is
|
||||
stable across two different embedded SQLite versions (Rust ~3.46 vs modernc 3.53.3). Byte
|
||||
equality was never the goal and is precluded by the file format.
|
||||
- The comparator fails loudly on a missing dataset rather than skipping it — the silent-skip
|
||||
bug that made `verify-parity.js` unsafe.
|
||||
@@ -0,0 +1,194 @@
|
||||
---
|
||||
phase: 7
|
||||
title: "CI docs and cutover"
|
||||
status: pending
|
||||
priority: P2
|
||||
dependencies: [6]
|
||||
effort: ""
|
||||
---
|
||||
|
||||
# Phase 7: CI docs and cutover
|
||||
|
||||
## Overview
|
||||
|
||||
Make Go the real parser: add a deploy guard, swap CI, update docs, remove Rust. Runs only after
|
||||
Phase 6 signs off on all four datasets. Until this phase the migration is fully reversible;
|
||||
after 7e it is not.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Functional: `npm run build:db` uses the Go binary and **fails on a bad database**; CI green
|
||||
without a Rust toolchain.
|
||||
- Non-functional: zero dangling references to `parser/`, `cargo`, or `Cargo.toml`.
|
||||
|
||||
## Architecture
|
||||
|
||||
Strict order, each step independently verifiable:
|
||||
|
||||
1. **7a** — deploy guard + npm scripts + `build-db.js` point at Go
|
||||
2. **7b** — CI: add a branch-verify path *first*, then swap the toolchain
|
||||
3. **7c** — docs and stale references
|
||||
4. **7d** — tag, then remove Rust `parser/`
|
||||
|
||||
There is no JS-script port. See "Scripts are not ported".
|
||||
|
||||
## 7a — Deploy guard (new, and the most important step here)
|
||||
|
||||
Today **nothing between the parser and the public site asserts that a database has data**:
|
||||
- `main.rs:171-177` logs file-level errors and continues; `run_build_standard` returns `Ok(())`
|
||||
at `:198` regardless of `total_errors`.
|
||||
- `writer.rs:88-133` prints stats and returns `Ok` even at `db_count == 0`; the `Audit:` line at
|
||||
`:120-127` is `println!` only, never an exit code.
|
||||
- `build-db.js:47-68` gzips whatever it gets, with no inspection.
|
||||
- `scripts/assemble-site.js:56-73` only greps *filenames* for stray `.db`; an empty
|
||||
`.build/public/db` passes cleanly.
|
||||
- `deploy-pages.yml` has no verification step.
|
||||
|
||||
So a Go reader that silently under-produces ships a truncated public dataset with **green CI**.
|
||||
This is the single largest blast-radius gap in the migration and it is cheap to close.
|
||||
|
||||
Add to `build-db.js`, as a blocking check per dataset:
|
||||
- fail non-zero if the binary reported any file-level error
|
||||
- fail non-zero if `SELECT COUNT(*)` deviates from the known-good figure in
|
||||
`docs/data-pipeline.md:110-115` (877,461 / 861,068 / 847,348 / 679,764)
|
||||
- fail non-zero if the resulting `.db.gz` is under 90% of `dbSizeMb` in `src/datasets.js`
|
||||
|
||||
Also make the Go binary exit non-zero when `total_errors > 0`. This is the one place where
|
||||
bug-for-bug compatibility costs more than it buys — note the deliberate divergence in the code.
|
||||
|
||||
## 7a — Pipeline wiring
|
||||
|
||||
- `package.json:8`: `"build:go": "go build -o go-parser/bin/xlsxread ./go-parser/cmd/xlsxread"`
|
||||
- **`package.json:9`**: `"build:db"` currently reads `node parser/scripts/build-db.js`. Keep the
|
||||
script *name* (CI and docs reference it) but the *path* must change when the file moves.
|
||||
Missing this breaks the deploy step.
|
||||
- `build-db.js:25`: `BIN` → `go-parser/bin/xlsxread`; update the `npm run build:rust` hint at
|
||||
`:37-41`.
|
||||
- Verify: `npm run build:go && npm run build:db` produces all 4 `.db.gz` and the guard fires when
|
||||
fed a deliberately truncated DB.
|
||||
|
||||
## 7b — CI
|
||||
|
||||
**Prerequisite, before the toolchain swap:** the workflow currently triggers on `push` to `main`
|
||||
only, plus an unguarded `workflow_dispatch` (`deploy-pages.yml:3-6, 55-63`). So "verify on a
|
||||
branch" is impossible — pushing to a branch runs nothing, and dispatching from a branch
|
||||
**publishes that branch's output to the live site**, with `cancel-in-progress: true` killing any
|
||||
in-flight good deploy. Fix this first:
|
||||
|
||||
- add a `pull_request` (or branch-push) trigger that runs the **build job only**
|
||||
- guard the deploy job with `if: github.ref == 'refs/heads/main'`
|
||||
|
||||
Then swap:
|
||||
- remove `dtolnay/rust-toolchain@stable` (`:23`) and `Swatinem/rust-cache@v2` (`:25-27`)
|
||||
- add `actions/setup-go@v5` pinned to **1.26.x** (matches the verified local toolchain), with
|
||||
module caching
|
||||
- add `govulncheck` (there is an open excelize advisory, and the 2017 refresh runbook feeds
|
||||
network-downloaded spreadsheets straight into the parser)
|
||||
- run `go test ./...` in CI — the reader-fidelity suite covers all 299 real files in ~77s and is
|
||||
the regression guard for the whole reader
|
||||
- build step → `npm run build:go && npm run build:db`
|
||||
- **cgo**: `pbnjay/grate` and `excelize/v2` are pure Go. Whether the workflow needs a C
|
||||
toolchain depends solely on the Phase 2 SQLite driver decision (`modernc.org/sqlite` keeps it
|
||||
cgo-free; `mattn/go-sqlite3` does not). Set `CGO_ENABLED` explicitly either way.
|
||||
|
||||
## 7c — Docs and stale references
|
||||
|
||||
The previous hand-curated file list covered 8 locations; there are **33** `parser/` references
|
||||
outside `parser/`. Use a mechanical gate instead of a list:
|
||||
|
||||
```
|
||||
grep -rn "parser/\|cargo\|Cargo\.toml" \
|
||||
--include='*.js' --include='*.jsx' --include='*.json' --include='*.yml' --include='*.md' . \
|
||||
| grep -v node_modules | grep -v '^./plans/'
|
||||
```
|
||||
|
||||
must return zero rows before 7d is marked done. Note the old success criterion grepped only for
|
||||
`cargo` / `Cargo.toml` / `parser/target` — none of which match `parser/scripts/…` or
|
||||
`parser/configs/…`.
|
||||
|
||||
Known references beyond the original list, including two the plan had scoped out:
|
||||
- `package.json:9`, `eslint.config.js:31` (its glob `parser/scripts/**/*.js` would silently stop
|
||||
matching, dropping the moved scripts from `npm run lint`)
|
||||
- `vite.config.js:14`, `src/datasets.js:6,11,17`, `src/lib/subjects.js:4` — the `src/` ones are
|
||||
comments; `plan.md` carves them out of the frontend exclusion explicitly
|
||||
- `README.md:26,37,56`; `docs/system-architecture.md:34,38,50`;
|
||||
`docs/data-pipeline.md:22,57,119,121,125,126,137,138,145`;
|
||||
`docs/deployment-guide.md:47,59,61`
|
||||
- `docs/deployment-guide.md:38` is the `build:rust` line (the plan previously cited `:37`, which
|
||||
is `npm ci`); `:10` mentions the Rust toolchain
|
||||
- `docs/data-pipeline.md` references `parser/src/schema.rs` as canonical DDL →
|
||||
`go-parser/internal/schema/schema.go`
|
||||
- `.gitignore:20-21`: `parser/target/` → `go-parser/bin/`
|
||||
|
||||
Per documentation rules, update what changed; no changelog noise.
|
||||
|
||||
## Scripts are not ported
|
||||
|
||||
The original plan ported `db-stats.js`, `verify-parity.js`, `check-duplicates.js`, and
|
||||
`diff-datasets.js`. Full consumer enumeration says don't:
|
||||
|
||||
| Script | Automated consumers | Notes |
|
||||
|---|---|---|
|
||||
| `db-stats.js` | **0** | 3 doc refs only |
|
||||
| `verify-parity.js` | **0** | 2 doc refs; still valid for its historical check — leave it |
|
||||
| `check-duplicates.js` | **0** | **Broken**: `:10` hardcodes `D:/tiennm99/thptqg2017/data` |
|
||||
| `diff-datasets.js` | **0** | **Broken**: imports `better-sqlite3`, absent from `package.json`; reads paths that don't exist |
|
||||
| `crawl-baotintuc.js` | 0 automated, but **the only one with a live runbook** (`docs/data-pipeline.md:22,137-140`) | Leave in JS |
|
||||
|
||||
None appear in `package.json` or the workflow. Node is already a hard build dependency
|
||||
(`actions/setup-node@v4`), so leaving them in JS costs nothing. Porting the two broken ones
|
||||
would mean either reproducing a hardcoded Windows path in Go or fixing them — undeclared scope
|
||||
and a behavior change in a plan whose rule is bug-for-bug compatibility.
|
||||
|
||||
The criterion is usage, not topic: **no script with zero automated consumers gets ported.** That
|
||||
excludes all five. `crawl-baotintuc.js` stays in JS because it works and has a runbook.
|
||||
|
||||
## 7d — Remove Rust
|
||||
|
||||
**Extra cleanup from Phase 1:** `parser/examples/dump_cells.rs` and `parser/examples/scan_kinds.rs`
|
||||
are throwaway ground-truth tooling. They disappear with `parser/`, which also retires
|
||||
`go-parser/scripts/regen-fidelity-hashes.sh` (it shells out to `cargo`). Before deleting,
|
||||
decide whether the reader-fidelity oracle should survive:
|
||||
|
||||
- keeping `go-parser/testdata/reader-fidelity-hashes.tsv` preserves a real regression guard over
|
||||
all 299 files, but it becomes unregenerable once calamine is gone — the same trap that made
|
||||
`verify-parity.js`'s baseline useless;
|
||||
- or drop the manifest and the suite with it, and rely on Phase 6's database-level gate.
|
||||
|
||||
Recommend keeping it and noting in the file header that it is frozen and why.
|
||||
|
||||
|
||||
- Move `parser/configs/` → `go-parser/configs/`; update `build-db.js` and `package.json:9`,
|
||||
`eslint.config.js:31` in the **same commit** as the move.
|
||||
- **Tag `pre-go-parser-removal` and push it** before deleting anything.
|
||||
- Delete `parser/`.
|
||||
- Write the revert procedure into this file as three named commands — "git history preserves it"
|
||||
is not a procedure, and after this step a revert is non-trivial because configs and
|
||||
`build-db.js` have moved.
|
||||
- Full verification: `npm run build:go && npm run build:db && npm run build:site`, then load the
|
||||
site and query each dataset.
|
||||
- **Confirm with the user before deleting** — open question 1 in `plan.md`.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] `build-db.js` guard fails the build on a truncated/empty DB (verified deliberately)
|
||||
- [ ] Go binary exits non-zero when `total_errors > 0`
|
||||
- [ ] Branch-verify CI path runs the build job without publishing; deploy job guarded to `main`
|
||||
- [ ] CI green with no Rust toolchain; `govulncheck` wired in; deploy succeeds
|
||||
- [ ] The 7c grep gate returns **zero rows**
|
||||
- [ ] `.gitignore` covers the Go binary; no build artifact committed
|
||||
- [ ] `npm run lint` still covers the relocated scripts
|
||||
- [ ] Frontend loads all 4 datasets; accent-insensitive search works (exercises `ho_ten_ascii`)
|
||||
- [ ] `pre-go-parser-removal` tag pushed before deletion; revert procedure written down
|
||||
- [ ] Rust removal confirmed with the user
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| Bad database reaches the public site with green CI | 7a guard — the reason this step exists |
|
||||
| "Verify on a branch" publishes to production instead | 7b prerequisite: branch trigger + deploy ref guard |
|
||||
| `npm run build:db` breaks after the move | `package.json:9` called out; same-commit rule; grep gate |
|
||||
| Lint coverage silently lost | `eslint.config.js:31` called out; explicit criterion |
|
||||
| Rust deleted before a latent bug surfaces | Tag + written revert procedure + user confirmation |
|
||||
| Effort spent porting dead scripts | Cut, with consumer counts recorded |
|
||||
@@ -0,0 +1,193 @@
|
||||
---
|
||||
title: Migrate parser from Rust to Go as side-by-side go-parser/
|
||||
description: >-
|
||||
Build go-parser/ alongside the Rust parser/, validated by differential
|
||||
comparison against live Rust output. Rust stays working until parity is signed
|
||||
off.
|
||||
status: pending
|
||||
priority: P2
|
||||
branch: main
|
||||
tags:
|
||||
- migration
|
||||
- go
|
||||
- parser
|
||||
- data-integrity
|
||||
- tdd
|
||||
blockedBy: []
|
||||
blocks: []
|
||||
created: '2026-08-13T08:21:47.330Z'
|
||||
createdBy: 'ck:plan'
|
||||
source: skill
|
||||
---
|
||||
|
||||
# Migrate parser from Rust to Go as side-by-side go-parser/
|
||||
|
||||
## Overview
|
||||
|
||||
Port the 2.3k-line Rust `parser/` crate to Go under a new `go-parser/` directory. Rust is
|
||||
untouched and keeps building throughout, so ground truth is regenerable on demand and the
|
||||
migration is reversible until Phase 7.
|
||||
|
||||
**Driver: preference for working in Go.** No defect exists in the Rust parser. Recorded
|
||||
honestly rather than retrofitted with technical justification — this shapes the plan, because
|
||||
with no problem to fix, the only measure of success is *behavioral identity with Rust*.
|
||||
|
||||
Mode: `--tdd`. Tests come first in every phase. The Rust crate is an executable specification;
|
||||
**29 of its 63 tests transfer directly** (the `&str`-based `transform` tests). The other 34 are
|
||||
`calamine::Data`-typed or golden tests; Phase 1 recorded the exact string each `Data` variant
|
||||
renders to, so they can now be re-derived without guessing.
|
||||
|
||||
## Design decisions
|
||||
|
||||
| Decision | Choice | Rationale |
|
||||
|---|---|---|
|
||||
| Layout | New `go-parser/`, Rust `parser/` untouched | Reversible by construction; both runnable for diffing |
|
||||
| Validation | Differential vs **live Rust output** | Rust still runs, so no frozen baseline needed |
|
||||
| Comparator | **Purpose-built**, not `verify-parity.js` | That script's baseline-diff semantics are wrong for a same-schema comparison — it emits 7 spurious failures. See Phase 6 |
|
||||
| Binary path | `go-parser/bin/xlsxread` | Two parsers writing one path invites confusion. Costs a one-line change at `build-db.js:25` |
|
||||
| Reader contract | **One** streaming API with a typed `Cell`, defined in Phase 1 | Mirrors Rust's `process_file`; a `[][]string` collapse loses `Data::Empty` and row width |
|
||||
| SQLite driver | Decided in **Phase 2**, with written rationale | Governs cgo/CI/ARM64 shape and the integrity story; not deferrable to Phase 4 |
|
||||
| BIFF reader | **`pbnjay/grate`** (Phase 1) | `extrame/xls` corrupted 69% of cells; grate matched calamine on all 67 files |
|
||||
| `.xls → .xlsx` conversion | **Not needed** (Phase 1, 2026-08-13) | grate reads BIFF exactly, so source data stays untouched. Fallback retired, not exercised |
|
||||
| Config format | **YAML (.yml)**, read by BOTH parsers (Phase 2) | User preference. Converting only Go would leave two hand-synced copies whose drift Phase 6 would blame on the parser. Amends "parser/ untouched" deliberately |
|
||||
| JS scripts | **Not ported** | All four candidates have zero automated consumers; two are documented broken |
|
||||
| Rust removal | After parity sign-off only, behind a tag | Phase 7e |
|
||||
|
||||
## Phases
|
||||
|
||||
| Phase | Name | Status |
|
||||
|-------|------|--------|
|
||||
| 1 | [Scaffold and reader fidelity gate](./phase-01-scaffold-and-biff-reader-gate.md) | Completed |
|
||||
| 2 | [Schema and config](./phase-02-schema-and-config.md) | Completed |
|
||||
| 3 | [Transform core](./phase-03-transform-core.md) | Completed |
|
||||
| 4 | [Reader writer and CLI](./phase-04-reader-writer-and-cli.md) | Completed |
|
||||
| 5 | [2016 format detection](./phase-05-2016-format-detection.md) | Completed |
|
||||
| 6 | [Differential parity gate](./phase-06-differential-parity-gate.md) | Completed |
|
||||
| 7 | [CI docs and cutover](./phase-07-ci-docs-and-script-port.md) | Pending |
|
||||
|
||||
Strictly sequential: 1 → 2 → 3 → 4 → 5 → 6 → 7. Phases 1 and 6 are hard gates.
|
||||
|
||||
## The dominant risk — RESOLVED in Phase 1 (2026-08-13)
|
||||
|
||||
Reader fidelity was the plan's dominant risk. It is now **settled: 299/299 files byte-identical
|
||||
to calamine**, locked in as a Go test against a committed hash oracle. Full record in
|
||||
`phase-01-scaffold-and-biff-reader-gate.md`.
|
||||
|
||||
- **`extrame/xls` was unusable** — 69% of cells corrupted, 28% lost, charset-independent.
|
||||
Replaced with **`pbnjay/grate`**, which matched calamine on all 67 BIFF files. The red-team
|
||||
claim that `extrame/xls` read the corpus correctly was wrong; it rested on "opens without
|
||||
panic" plus spot-checks, and spot-checks pass because 28% of cells are right.
|
||||
- **The `.xlsx` date-serial fear was unfounded.** Scanning all 299 files (15.98M cells) found
|
||||
**zero `DateTime` cells** — also zero `Int`, `Bool`, `Error`, `DateTimeIso`, `DurationIso`.
|
||||
Only `String`, `Empty`, and `Float` occur. `ngay_sinh` is text everywhere.
|
||||
- **The used-range-origin fear was unfounded.** Every used range in the corpus starts at (0,0).
|
||||
- **Two real divergences did reach the database** and are fixed: numeric re-rendering
|
||||
(`6.0`→`6`, gated on cell type so shared strings like `6.00`, `NAN`, and leading-zero
|
||||
`so_bao_danh` are untouched), and XML line-ending normalisation stripping CR from 2,233
|
||||
`ten_cum_thi` values.
|
||||
|
||||
Consequence for `--tdd`, now unblocked: the 11 `format_detect_2016` and 7 `reader` tests build
|
||||
fixtures from `calamine::Data` values, and Phase 1 recorded the exact rendering of each variant,
|
||||
so they can be ported without guessing.
|
||||
|
||||
**The `.xls → .xlsx` conversion fallback is no longer needed** and remains unexercised. Source
|
||||
data is untouched.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- [ ] For all 4 datasets, Go-built and Rust-built DBs are **logically equivalent**: identical
|
||||
row counts, identical per-column non-NULL counts across all 22 columns, identical
|
||||
`PRAGMA table_info`/`index_list`, and identical sorted full-table SHA-256
|
||||
- [ ] Both binaries emit identical build stdout per dataset, modulo the `Size:` line
|
||||
(this is the only check that covers the `source_rows`/`skipped` counters)
|
||||
- [ ] Go tests pass, including reader-fidelity tests against real files
|
||||
- [ ] `npm run build:db` produces four `.db.gz` via the Go binary, with a **row-count guard**
|
||||
that fails the build on deviation
|
||||
- [ ] CI green with Go toolchain, Rust actions removed, and a branch-verify path that does not
|
||||
publish to production
|
||||
- [ ] Frontend loads all 4 datasets unchanged, including accent-insensitive search
|
||||
|
||||
**Explicitly not a criterion:** byte-identical databases. SQLite writes its own version number
|
||||
into header bytes 96-99, `VACUUM` rewrites page layout per-version, and `gzip -9` without `-n`
|
||||
stores mtime. Byte equality is precluded by the file format, not merely difficult.
|
||||
|
||||
## Out of scope
|
||||
|
||||
- Frontend behavior, `src/` logic, `scripts/assemble-site.js`, Vite config
|
||||
(**exception**: stale path comments in `src/datasets.js` and `src/lib/subjects.js` must be
|
||||
updated in Phase 7 — they reference `parser/` paths that will not exist)
|
||||
- Schema changes — the 22-column contract is frozen
|
||||
- Porting the JS helper scripts (see Design decisions)
|
||||
- Behavior "improvements". Bug-for-bug compatibility is the goal for **everything that reaches
|
||||
the database**. For stdout, replicate the per-file row-count line; see Phase 4 on the
|
||||
`dataset_label` caveat
|
||||
|
||||
## Dependencies
|
||||
|
||||
Builds on completed plan `260813-0956-unify-frontend-standard-schema`. No blocking relationship.
|
||||
|
||||
Inputs:
|
||||
- Brainstorm: `plans/reports/from-brainstorm-to-plan-260813-1502-go-parser-side-by-side-migration-report.md`
|
||||
- Scout: `plans/reports/from-scout-to-brainstorm-260813-1502-rust-to-go-parser-migration-report.md`
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Keep `parser/` as a reference implementation after parity, or delete it? (Phase 7e assumes
|
||||
delete, behind a `pre-go-parser-removal` tag.)
|
||||
2. `modernc.org/sqlite` is a machine-transpiled SQLite, not the upstream C amalgamation that
|
||||
`rusqlite --bundled` vendors. Acceptable for the writer of a published 1.5M-row dataset, or
|
||||
use `mattn/go-sqlite3` (real upstream C, cgo cost in CI)? Decided in Phase 2.
|
||||
|
||||
## Red Team Review
|
||||
|
||||
### Session — 2026-08-13
|
||||
**Findings:** 39 raw across 4 reviewers → 22 unique (19 accepted, 3 rejected)
|
||||
**Severity breakdown:** 6 Critical, 10 High, 6 Medium
|
||||
|
||||
| # | Finding | Severity | Disposition | Applied To |
|
||||
|---|---------|----------|-------------|------------|
|
||||
| 1 | `.xlsx` stringification divergence certain and ungated; BIFF framing wrong | Critical | Accept | Completed |
|
||||
| 2 | `verify-parity.js` unusable — 7 spurious failures, contradictory instructions | Critical | Accept | Completed |
|
||||
| 3 | `md5()` does not exist in SQLite; "strongest check" fictional | Critical | Accept | Completed |
|
||||
| 4 | No deploy guard — empty/truncated DB ships with green CI | Critical | Accept | Completed |
|
||||
| 5 | "Verify on a branch" unexecutable; `workflow_dispatch` publishes to prod | Critical | Accept | Completed |
|
||||
| 6 | Counter divergence invisible; 63 trailing empty sheets in `data/2017` | Critical | Accept | Completed |
|
||||
| 7 | Phase 7e breaks `npm run build:db`; 33 refs vs 8 listed | High | Accept | Phase 7 |
|
||||
| 8 | "byte-equivalent databases" provably unachievable | High | Accept | plan.md |
|
||||
| 9 | Phase 3 cites 20 of 29 tests; omitted 9 are the flagged traps | High | Accept | Phase 3 |
|
||||
| 10 | Two incompatible reader contracts; `[][]string` lossy | High | Accept | Phase 1, 4 |
|
||||
| 11 | Phase 7d ports 4 scripts with 0 consumers, 2 broken | High | Accept | Phase 7 (cut) |
|
||||
| 12 | `.xls` golden fixture unbuildable — no Go BIFF writer | High | Accept | plan.md, Phase 4 |
|
||||
| 13 | Golden fixture port is phantom coverage (`inlineStr` only) | High | Accept | Phase 4 (cut) |
|
||||
| 14 | No partial-success/abandon procedure at Phase 6 | High | Accept | Phase 6 |
|
||||
| 15 | `verify-parity.js` silently passes on datasets absent from both files | High | Accept | Phase 6 |
|
||||
| 16 | PII: committing real rows as testdata breaks documented convention | Medium | Accept | Phase 1 |
|
||||
| 17 | Duplicate-SBD guidance names `read_dir`; Rust sorts explicitly | Medium | Accept | Phase 6 |
|
||||
| 18 | `dataset_label` derived from path breaks tempdir stdout comparison | Medium | Accept | Phase 4 |
|
||||
| 19 | No dependency trust/pinning/`govulncheck` step | Medium | Accept | Phase 1, 2 |
|
||||
| 20 | Make `.xls`→`.xlsx` conversion unconditional Phase 0 | High | **Reject** | — |
|
||||
| 21 | Publish DBs as artifacts; drop parser from critical path | High | **Reject** | — |
|
||||
| 22 | Merge Phases 2-5 into one "port the crate" phase | Medium | **Reject** | — |
|
||||
|
||||
**Rejection rationale:**
|
||||
- **20, 21** — user decisions, not reviewer calls. 20 mutates committed source data and the user
|
||||
explicitly deferred it on 2026-08-13. **Superseded by Phase 1**: `pbnjay/grate` reads BIFF
|
||||
exactly, so no conversion is needed and source data stays untouched. (The `extrame/xls`
|
||||
evidence cited when this was first rejected was itself wrong — but the conclusion holds for
|
||||
a better reason.) 21 reverses the user's stated goal of working in Go on the parser.
|
||||
- **22** — phases map to TDD checkpoints and hydrated tasks. Merging reduces granularity without
|
||||
reducing risk. Phase 1's gate framing was re-pointed instead.
|
||||
|
||||
**Citation corrections applied:** `transform.rs:161`→`:162`; `deployment-guide.md:37`→`:38`;
|
||||
"rusqlite statement cache" removed (`writer.rs:73` uses `conn.execute`, not `prepare_cached`);
|
||||
"doc comment says category M"→ the *inline* comment at `:53` (the doc comment at `:49` is
|
||||
correct); test counts corrected to 63 total / 29 in `transform.rs`.
|
||||
|
||||
### Whole-Plan Consistency Sweep
|
||||
- Files reread: `plan.md`, `phase-01` … `phase-07` (all 8)
|
||||
- Decision deltas checked: 19
|
||||
- Reconciled stale references: dominant-risk framing (plan.md + Phase 1), byte-equivalence
|
||||
criterion (plan.md), reader contract (Phases 1 + 4), `verify-parity.js` usage (Phase 6),
|
||||
script-port scope (plan.md + Phase 7), test counts (plan.md + Phases 2/3/5), phase title
|
||||
"BIFF reader gate" → "reader fidelity gate" (plan.md table + Phase 1)
|
||||
- Unresolved contradictions: 0
|
||||
+79
@@ -0,0 +1,79 @@
|
||||
# Brainstorm — Migrate parser from Rust to Go
|
||||
|
||||
Date: 2026-08-13. Branch: main. Scout input: `from-scout-to-brainstorm-260813-1502-rust-to-go-parser-migration-report.md`.
|
||||
|
||||
## Decision
|
||||
|
||||
Build a **new Go parser under `go-parser/`, side by side with the existing Rust `parser/`**. Validate by **differential comparison against live Rust output** — not a frozen baseline. Rust stays untouched and working throughout.
|
||||
|
||||
## Problem-first
|
||||
|
||||
User brought a preselected solution ("migrate to Go"). Inversion applied.
|
||||
|
||||
- **Driver**: preference for working in Go + AI makes iteration cheap. Not a defect in the Rust parser — none exists.
|
||||
- **Evidence status**: none for a technical problem. Legitimate as a preference-driven migration; recorded as such rather than retrofitted with technical justification.
|
||||
- **Rejected framings**: "one language too many" (Go is still a second non-JS toolchain — doesn't collapse the stack); "Rust hard to modify" (a Go port reproduces the same 2016-format complexity in different syntax); "performance" (419 MB of Excel I/O dominates deploy time, unchanged by language).
|
||||
|
||||
## Approaches evaluated
|
||||
|
||||
| # | Approach | Verdict |
|
||||
|---|---|---|
|
||||
| 1 | Don't migrate; drop dead `glob` dep | Rejected — user prefers Go, cost is acceptable |
|
||||
| 2 | In-place Rust→Go rewrite of `parser/` | Rejected — no reversibility, broken half-states |
|
||||
| 3 | Convert `.xls`→`.xlsx` first, then port | Held as **fallback** if `extrame/xls` fails |
|
||||
| 4 | **Side-by-side `go-parser/` + differential validation** | **CHOSEN** |
|
||||
|
||||
Approach 4 beats the author's original phased proposal: keeping Rust live means ground truth is regenerable on demand, so no frozen parity baseline is needed. Reversibility is structural, not procedural.
|
||||
|
||||
## Constraints carried from scout
|
||||
|
||||
**Hard blocker to hit early**: 67 genuine OLE2/BIFF `.xls` files (verified by magic bytes `d0cf11e0a1b11ae1`) — 63 of them in `data/2017`, the 286 MB largest dataset. `calamine` reads BIFF+OOXML through one API; Go has no equivalent. `excelize` is xlsx-only; `extrame/xls` is the only real BIFF option and is lightly maintained + weak on non-UTF8 encodings, which matters because every cell is Vietnamese.
|
||||
|
||||
→ **Build the Go reader module FIRST**, before transform/writer. Fail fast. Fallback = approach 3 (data is frozen, git-tracked, untouched since the unification rename — a one-time conversion is legitimate here and would likely shrink the repo, since OOXML is zip-compressed and BIFF is not).
|
||||
|
||||
**Exactness traps that must be replicated, not improved:**
|
||||
- `to_ascii`: literal codepoint range U+0300–U+036F filter, **not** `unicode.Mn` (Go's category check is more permissive → divergence). Plus explicit `đ/Đ → d` (NFD does not decompose them). `transform.rs:52-64`.
|
||||
- TOML `deny_unknown_fields` is load-bearing and has a test. Most Go TOML libs ignore unknown keys by default.
|
||||
- Tri-state skip reason (`BlankRow` vs `EmptyField` vs `NonNumericSbd`) drives the printed counters — a bool diverges.
|
||||
- `0.0` score parsed as `None` in the 2016 separate-scores format (replicates a JS `||` falsy quirk). `format_detect_2016.rs:165`.
|
||||
- Audit path reads **sheet 0 only**, deliberately ignoring `sheet_mode`. Preserve, don't "fix".
|
||||
- Header check is **per-sheet**, not per-file.
|
||||
- Unknown: how calamine stringifies date cells into `ngay_sinh`. Never inspected. Probe empirically against real files.
|
||||
|
||||
**Contracts to preserve**: CLI `build --schema --input --output`; SQLite `student` table, 22 cols, 3 indexes incl. partial `idx_ten_cum_thi`; non-zero exit kills the npm pipeline.
|
||||
|
||||
**Free wins**: regexes are already RE2-safe (zero port risk); SQLite is plain SQL text with positional `?`, no named params, no pragmas; `glob` dep is dead (confirmed zero references).
|
||||
|
||||
## Scope
|
||||
|
||||
Everything under `parser/` → `go-parser/`: crate, tests, and the 6 JS helper scripts.
|
||||
|
||||
**Sequencing constraint**: port `db-stats.js` / `verify-parity.js` **last**. They are the tools that prove the port is correct — rewriting them during the port is circular (a bug in the ported checker hides a bug in the ported parser). Verify Go versions against JS output before trusting them.
|
||||
|
||||
## Validation criteria
|
||||
|
||||
Go parser is done when, for all 4 datasets: Rust-built DB and Go-built DB are equivalent — row counts, per-column non-NULL counts, and field-by-field equality on a deterministic SBD sample. Plus black-box golden tests spawning the binary and inspecting the `.db`, including **an `.xls` fixture** (current suite has none, so it cannot catch a BIFF regression).
|
||||
|
||||
## Risks
|
||||
|
||||
| Risk | Severity | Mitigation |
|
||||
|---|---|---|
|
||||
| `extrame/xls` can't read the 67 BIFF files / mangles Vietnamese | **High** | Build reader first; fallback to `.xls`→`.xlsx` conversion |
|
||||
| Date-cell stringification differs from calamine | Medium | Differential diff catches it; probe early |
|
||||
| Silent value corruption across 419 MB | Medium | Differential gate is the only real defense — non-negotiable |
|
||||
| Repo carries two parsers during migration | Low | Intentional; delete `parser/` only after sign-off |
|
||||
|
||||
## Next steps
|
||||
|
||||
1. `go-parser/` scaffold + reader module against real `.xls` — decisive gate.
|
||||
2. Port config/transform/writer/audit/schema.
|
||||
3. Black-box golden tests + `.xls` fixture.
|
||||
4. Differential parity vs Rust across all 4 datasets.
|
||||
5. CI swap (drop `dtolnay/rust-toolchain` + `Swatinem/rust-cache`, add `actions/setup-go`), `.gitignore`, docs.
|
||||
6. Port JS scripts last. Remove `parser/` after sign-off.
|
||||
|
||||
## Unresolved
|
||||
|
||||
1. Keep binary name/path `parser/target/release/xlsxread` so `build-db.js` is untouched, or emit to `go-parser/bin/` and update the one constant at `build-db.js:25`?
|
||||
2. Delete `parser/` after parity, or keep it as a reference implementation for some period?
|
||||
3. `.xls` fallback: if conversion is needed, keep originals committed alongside converted files, or replace them?
|
||||
+70
@@ -0,0 +1,70 @@
|
||||
# Scout Report — Rust parser migration surface (Rust → Go)
|
||||
|
||||
Date: 2026-08-13. Branch: main. Scope: `parser/` crate + its build/verify boundary.
|
||||
|
||||
## Relevant Files
|
||||
|
||||
### Rust crate (~2.3k LOC)
|
||||
- `parser/Cargo.toml` — 10 deps. `glob` is **dead** (zero references in `src/` or `tests/`).
|
||||
- `parser/src/schema.rs` (213) — canonical DDL, INSERT SQL, column order, 16 subject regexes. Single source of truth.
|
||||
- `parser/src/format_detect_2016.rs` (548) — largest file. Per-file/per-sheet 3-way layout detection for 2016 only.
|
||||
- `parser/src/transform.rs` (409) — `to_ascii` Vietnamese diacritic strip, score regex, tri-state row validation.
|
||||
- `parser/src/main.rs` (377) — CLI dispatch, file globbing via `fs::read_dir`, transaction boundaries, counters.
|
||||
- `parser/src/{reader,writer,config,audit,cli,error,lib}.rs` — sheet iteration, SQLite lifecycle, TOML load, audit, clap.
|
||||
- `parser/configs/{2016,2017,2017-old,2017-old2}.toml` — per-dataset column maps + validation flags.
|
||||
|
||||
### Boundary
|
||||
- `parser/scripts/build-db.js` — sole caller. Hardcodes `parser/target/release/xlsxread` (line 25).
|
||||
- `parser/scripts/verify-parity.js` + `db-stats.js` — manual parity tool, **not wired into CI**.
|
||||
- `parser/tests/golden.rs` (591) — 8 tests, white-box (calls library fns, not the binary).
|
||||
- `.github/workflows/deploy-pages.yml` — `dtolnay/rust-toolchain` + `Swatinem/rust-cache` (workspaces: parser).
|
||||
- `package.json:8` — `build:rust` = `cargo build --release --manifest-path parser/Cargo.toml`.
|
||||
|
||||
## Contracts a rewrite must preserve
|
||||
|
||||
**CLI** (only this is depended on by scripts):
|
||||
```
|
||||
xlsxread build --schema parser/configs/<id>.toml --input data/<id> --output .build/public/db/<id>.db
|
||||
xlsxread audit --schema <cfg> --input <dir> --db <db> # operator-only, never in CI
|
||||
```
|
||||
|
||||
**Output**: SQLite `student` table, 22 cols (`so_bao_danh TEXT PRIMARY KEY`, `ho_ten`, `ho_ten_ascii`, `ngay_sinh`, `ten_cum_thi`, `gioi_tinh`, + 16 `REAL` scores), 3 indexes incl. partial `idx_ten_cum_thi ... WHERE ten_cum_thi IS NOT NULL`. Frontend `sql.js` reads these names directly.
|
||||
|
||||
**stdout**: human-facing only. No script parses it. Exit non-zero kills the npm pipeline (`execFileSync`).
|
||||
|
||||
## Migration risk
|
||||
|
||||
| Area | Risk | Why |
|
||||
|---|---|---|
|
||||
| **Legacy `.xls` (BIFF) reading** | **BLOCKING** | See below. |
|
||||
| calamine cell→string coercion | HIGH | Dates/numerics reach `ngay_sinh` and the score regexes as whatever calamine's `Data::to_string()` renders. Never inspected in-repo; a Go lib will differ. |
|
||||
| `to_ascii` exactness | MEDIUM | `transform.rs:56` filters the literal range U+0300–U+036F, **not** Unicode category Mn (despite its own doc comment). Go `unicode.Is(unicode.Mn,·)` is more permissive → must copy the range check. Plus explicit `đ/Đ → d` (NFD does not decompose them). |
|
||||
| TOML strictness | MEDIUM | `deny_unknown_fields` is load-bearing (has a test). Most Go TOML libs ignore unknown keys by default. |
|
||||
| Stats/stdout parity | MEDIUM | Tri-state `SkipReason` (BlankRow vs EmptyField vs NonNumericSbd) drives the printed counters; a bool would diverge. Stringly-typed dispatch on `dataset_label.contains("old"/"old2")`. |
|
||||
| Regex | **NONE** | Rust `regex` and Go `regexp` are both RE2. No backrefs/lookaround/`\p{}` anywhere. |
|
||||
| SQLite | LOW | Plain SQL text, positional `?`, no named params, no pragmas. VACUUM correctly post-COMMIT. |
|
||||
| clap/serde/thiserror | TRIVIAL | Idiomatic differences only. |
|
||||
|
||||
### The blocker, verified on disk
|
||||
|
||||
```
|
||||
data/2016/ 4 .xls + 115 .xlsx
|
||||
data/2017/ 63 .xls <-- 286 MB, the largest dataset
|
||||
data/2017-old/ 63 .xlsx
|
||||
data/2017-old2/ 54 .xlsx
|
||||
```
|
||||
|
||||
67 legacy `.xls` files across two datasets. `calamine::open_workbook_auto` reads BIFF and OOXML through one API. **Go has no equivalent** — `excelize` is xlsx-only; legacy `.xls` means `extrame/xls` (lightly maintained, incomplete BIFF coverage) or an external converter. This is the crux of the decision, not a detail.
|
||||
|
||||
## Also true
|
||||
|
||||
- No stated performance or correctness problem with the Rust parser. It is not the pain point; `rust-cache` keeps warm CI builds cheap. 419 MB of Excel dominates deploy runtime regardless of language.
|
||||
- `verify-parity.js`'s `APPROVED_RECOVERY` baseline is frozen against the **pre-refactor** implementation and cannot be regenerated. Reusable as a *method* for Rust-vs-Go, but needs a fresh baseline captured from current Rust output first.
|
||||
- `golden.rs` is white-box (calls `xlsxread::reader::process_file` etc.). Its fixtures are hand-built minimal OOXML + `zip` — that trick ports to Go's `archive/zip` trivially. Black-box CLI tests would be the language-agnostic seam.
|
||||
- No fixture covers `.xls` at all — all 3 fixtures are xlsx. So the golden suite would not catch an `.xls` regression.
|
||||
|
||||
## Unresolved Questions
|
||||
|
||||
1. What is the actual motivation for Go? No perf/correctness defect is visible in the repo. Answer determines whether migration is warranted at all.
|
||||
2. How does calamine render date cells into `ngay_sinh` today? Must be probed empirically against real 2016/2017 files before any Go xlsx library is chosen.
|
||||
3. Is a fresh Rust-output parity baseline acceptable as the gate, given the original baseline is unregenerable?
|
||||
+3
-3
@@ -3,18 +3,18 @@
|
||||
*
|
||||
* `id` is the single identifier used end to end:
|
||||
*
|
||||
* data/<id>/ → parser/configs/<id>.toml → db/<id>.db.gz → /thptqg/<id>/
|
||||
* data/<id>/ → parser/configs/<id>.yml → db/<id>.db.gz → /thptqg/<id>/
|
||||
*
|
||||
* Site path and database URL are derived from `id` rather than stored, so a
|
||||
* dataset cannot be misconfigured into pointing at the wrong database.
|
||||
*
|
||||
* Imported by the Vite app *and* by parser/scripts/build-db.js under plain
|
||||
* Imported by the Vite app *and* by go-parser/scripts/build-db.js under plain
|
||||
* Node, so this module must stay free of `import.meta.env` and any Vite-only
|
||||
* syntax. Callers pass the base URL in explicitly for that reason.
|
||||
*/
|
||||
|
||||
// Extension is required: this module is also imported by plain Node
|
||||
// (parser/scripts/build-db.js), which does not resolve extensionless paths.
|
||||
// (go-parser/scripts/build-db.js), which does not resolve extensionless paths.
|
||||
import { PRESETS_2016, PRESETS_2017 } from "./lib/sql-presets.js";
|
||||
|
||||
const SUBTITLE = "Dữ liệu thí sinh toàn quốc · Hỗ trợ truy vấn SQL tùy chỉnh";
|
||||
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
/**
|
||||
* The 16 subject columns of the canonical schema, in display order.
|
||||
*
|
||||
* Mirrors `SCORE_FIELDS` in parser/src/schema.rs. Previously this list was
|
||||
* Mirrors `SCORE_FIELDS` in go-parser/internal/schema/schema.go. Previously this list was
|
||||
* maintained separately in score-table.jsx and student-detail.jsx, which is how
|
||||
* they drifted out of sync with each other and with the database.
|
||||
*
|
||||
|
||||
+1
-1
@@ -11,7 +11,7 @@ import react from "@vitejs/plugin-react";
|
||||
// and the existing ?q= deep links keep working, which that fallback would break.
|
||||
//
|
||||
// publicDir holds only the gzipped databases, staged there by
|
||||
// parser/scripts/build-db.js. Nothing uncompressed is ever placed in it.
|
||||
// go-parser/scripts/build-db.js. Nothing uncompressed is ever placed in it.
|
||||
export default defineConfig({
|
||||
plugins: [react()],
|
||||
base: "/thptqg/",
|
||||
|
||||
Reference in new issue
Block a user