refactor(parser): define one canonical student schema for all datasets

The 2016 and 2017 parsers were the same crate with divergent SQL. DDL, INSERT
and the subject regex table lived in each dataset's TOML config, so four copies
had to be kept in step by hand — which is how the two schemas drifted apart.

Move all of it into src/schema.rs as a single 22-column definition: 6 identity
columns (adding ten_cum_thi and gioi_tinh, previously 2016-only) and 16 subject
columns (the union of both exam years). Columns a dataset carries no data for
bind NULL, costing ~1 byte per row.

This collapses writer.rs to one insert path: insert_row_2016, SCORE_FIELDS_2016,
SCORE_FIELDS_2017 and the SCORE_FIELDS alias all go away. Configs shrink to the
parse rules that genuinely vary per dataset — sheet mode, column indices, SBD
validation, header tokens, blank-row stripping — and carry no SQL at all.

Config parsing now denies unknown fields, so a leftover [schema] block fails
loudly instead of looking effective while schema.rs drives the build.

idx_ten_cum_thi is partial, so it holds zero entries on the three 2017 datasets
where the column is always NULL.

Adds db-stats.js and verify-parity.js to prove no data moved: row counts,
per-column non-NULL counts and a deterministic student sample are compared
against databases built from the previous code.

Verified across all four datasets — row counts, every pre-existing column count,
and all sampled students are identical. Database size grows 1.4-2.2%.
This commit is contained in:
tiennm99 committed 2026-08-13 11:16:24 +07:00
1 parent 944ee14b36
commit cd4b07f9cb
22 files changed
+10218 -341

No files matched your search

@@ -1,4 +1,4 @@
# Config for thptqg2016 data/ — 4 .xls + 115 .xlsx mixed files.
# thptqg2016 data/ — 4 .xls + 115 .xlsx mixed files.
#
# Three column layouts exist across the 119 files; the binary selects the
# right one per-file at runtime via format_detection = "thptqg2016":
@@ -7,15 +7,13 @@
# mapped header SOBAODANH|SBD + DIEM_THI — most provinces
# default no header; positional 6-col layout — remaining files
#
# Schema differences from thptqg2017:
# + ten_cum_thi TEXT (exam-cluster name, TEN_CUMTHI column)
# + gioi_tinh TEXT (gender: "Nam"/"Nữ", GIOI_TINH column)
# + tieng_duc REAL (German)
# + tieng_nhat REAL (Japanese)
# - khtn, khxh, tieng_nga (not in 2016 dataset)
# This is the only dataset that populates ten_cum_thi and gioi_tinh, and the
# only one whose files carry Tiếng Đức / Tiếng Nhật scores.
#
# sheet_mode = "all": several provinces overflow into Sheet2 (65k Excel row cap).
# strip_blank_rows = false: no blank-row anomaly observed in this dataset.
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
format_detection = "thptqg2016"
@@ -32,55 +30,3 @@ require_nonempty_sbd = true
# Tokens that identify a header row by first-cell content (uppercased).
# Covers both SOBAODANH-style and SBD-style headers.
tokens = ["SOBAODANH", "SBD", "HO_TEN", "HOTEN", "HỌ TÊN", "STT"]
[schema]
ddl = """
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
ngay_sinh TEXT,
ten_cum_thi TEXT,
gioi_tinh TEXT,
toan REAL,
ngu_van REAL,
vat_ly REAL,
hoa_hoc REAL,
sinh_hoc REAL,
lich_su REAL,
dia_ly REAL,
tieng_anh REAL,
tieng_phap REAL,
tieng_duc REAL,
tieng_nhat REAL,
tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi);
"""
[scores]
toan = 'Toán:\s*(\d+(?:\.\d+)?)'
ngu_van = 'Ngữ văn:\s*(\d+(?:\.\d+)?)'
vat_ly = 'Vật lí:\s*(\d+(?:\.\d+)?)'
hoa_hoc = 'Hóa học:\s*(\d+(?:\.\d+)?)'
sinh_hoc = 'Sinh học:\s*(\d+(?:\.\d+)?)'
lich_su = 'Lịch sử:\s*(\d+(?:\.\d+)?)'
dia_ly = 'Địa lí:\s*(\d+(?:\.\d+)?)'
tieng_anh = 'Tiếng Anh:\s*(\d+(?:\.\d+)?)'
tieng_phap = 'Tiếng Pháp:\s*(\d+(?:\.\d+)?)'
tieng_duc = 'Tiếng Đức:\s*(\d+(?:\.\d+)?)'
tieng_nhat = 'Tiếng Nhật:\s*(\d+(?:\.\d+)?)'
tieng_trung = 'Tiếng Trung:\s*(\d+(?:\.\d+)?)'
[insert]
sql = """
INSERT OR REPLACE INTO student
(so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh, ten_cum_thi, gioi_tinh,
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc,
lich_su, dia_ly,
tieng_anh, tieng_phap, tieng_duc, tieng_nhat, tieng_trung)
VALUES
(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
"""
@@ -1,8 +1,10 @@
# Test-only config used by golden integration tests (data-old variant).
# Same column layout as thptqg2017-data.toml but with:
# - sheet_mode = "first" (reads only first sheet)
# - require_numeric_sbd = true (rejects non-digit SBDs)
# Schema and INSERT match SCORE_FIELDS_2017 (14 score cols).
# thptqg2017 data-old/ — 63 .xlsx files (pre-baotintuc refresh).
#
# sheet_mode = "first": single-sheet workbooks, never hit the 65k row cap.
# SBD validation: require ^\d+$ (build-database-old.js:55 guard).
# strip_blank_rows = false: no explicit blank-skip in build-database-old.js.
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
[reader]
sheet_mode = "first"
@@ -21,56 +23,3 @@ require_nonempty_sbd = true
[header]
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
[schema]
ddl = """
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
ngay_sinh TEXT,
toan REAL,
ngu_van REAL,
vat_ly REAL,
hoa_hoc REAL,
sinh_hoc REAL,
khtn REAL,
lich_su REAL,
dia_ly REAL,
gdcd REAL,
khxh REAL,
tieng_anh REAL,
tieng_phap REAL,
tieng_nga REAL,
tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
"""
[scores]
toan = 'Toán:\s*(\d+(?:\.\d+)?)'
ngu_van = 'Ngữ văn:\s*(\d+(?:\.\d+)?)'
vat_ly = 'Vật lí:\s*(\d+(?:\.\d+)?)'
hoa_hoc = 'Hóa học:\s*(\d+(?:\.\d+)?)'
sinh_hoc = 'Sinh học:\s*(\d+(?:\.\d+)?)'
khtn = 'KHTN:\s*(\d+(?:\.\d+)?)'
lich_su = 'Lịch sử:\s*(\d+(?:\.\d+)?)'
dia_ly = 'Địa lí:\s*(\d+(?:\.\d+)?)'
gdcd = 'GDCD:\s*(\d+(?:\.\d+)?)'
khxh = 'KHXH:\s*(\d+(?:\.\d+)?)'
tieng_anh = 'Tiếng Anh:\s*(\d+(?:\.\d+)?)'
tieng_phap = 'Tiếng Pháp:\s*(\d+(?:\.\d+)?)'
tieng_nga = 'Tiếng Nga:\s*(\d+(?:\.\d+)?)'
tieng_trung = 'Tiếng Trung:\s*(\d+(?:\.\d+)?)'
[insert]
sql = """
INSERT OR REPLACE INTO student
(so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh,
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn,
lich_su, dia_ly, gdcd, khxh,
tieng_anh, tieng_phap, tieng_nga, tieng_trung)
VALUES
(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
"""
@@ -0,0 +1,26 @@
# thptqg2017 data-old2/ — 54 .xlsx files (corrected-export set).
#
# sheet_mode = "all": 24.HCM_UTLQ.xlsx overflows into Sheet2 (+6,446 rows).
# SBD validation: require ^\d+$ (build-database-old2.js:57 guard).
# strip_blank_rows = true: skip fully blank rows BEFORE counting sourceRows
# (build-database-old2.js:50-51 checks blank before sourceRows++).
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
[reader]
sheet_mode = "all"
strip_blank_rows = true
[columns]
ho_ten = 0
ngay_sinh = 1
so_bao_danh = 2
diem_thi = 3
[validation]
require_numeric_sbd = true
require_nonempty_name = true
require_nonempty_sbd = true
[header]
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
@@ -1,8 +1,10 @@
# Test-only config used by golden integration tests.
# Matches the fixture row layout produced by tests/golden.rs: write_xlsx()
# header: HO_TEN(0) NGAY_SINH(1) SO_BAO_DANH(2) DIEM_THI(3)
# Schema and INSERT match SCORE_FIELDS_2017 (14 score cols) so run_build_cmd
# in golden.rs can call insert_row(..., SCORE_FIELDS) without modification.
# thptqg2017 data/ — 63 .xls files from baotintuc.vn (current generation).
#
# sheet_mode = "all": Hà Nội and HCM overflow into Sheet2 (65k Excel row cap).
# SBD validation: no numeric guard — build-database.js did not apply ^\d+$.
# strip_blank_rows = false: no blank-row anomaly in this dataset.
#
# Table shape, INSERT and subject regexes are canonical — see src/schema.rs.
[reader]
sheet_mode = "all"
@@ -21,56 +23,3 @@ require_nonempty_sbd = true
[header]
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
[schema]
ddl = """
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
ngay_sinh TEXT,
toan REAL,
ngu_van REAL,
vat_ly REAL,
hoa_hoc REAL,
sinh_hoc REAL,
khtn REAL,
lich_su REAL,
dia_ly REAL,
gdcd REAL,
khxh REAL,
tieng_anh REAL,
tieng_phap REAL,
tieng_nga REAL,
tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
"""
[scores]
toan = 'Toán:\s*(\d+(?:\.\d+)?)'
ngu_van = 'Ngữ văn:\s*(\d+(?:\.\d+)?)'
vat_ly = 'Vật lí:\s*(\d+(?:\.\d+)?)'
hoa_hoc = 'Hóa học:\s*(\d+(?:\.\d+)?)'
sinh_hoc = 'Sinh học:\s*(\d+(?:\.\d+)?)'
khtn = 'KHTN:\s*(\d+(?:\.\d+)?)'
lich_su = 'Lịch sử:\s*(\d+(?:\.\d+)?)'
dia_ly = 'Địa lí:\s*(\d+(?:\.\d+)?)'
gdcd = 'GDCD:\s*(\d+(?:\.\d+)?)'
khxh = 'KHXH:\s*(\d+(?:\.\d+)?)'
tieng_anh = 'Tiếng Anh:\s*(\d+(?:\.\d+)?)'
tieng_phap = 'Tiếng Pháp:\s*(\d+(?:\.\d+)?)'
tieng_nga = 'Tiếng Nga:\s*(\d+(?:\.\d+)?)'
tieng_trung = 'Tiếng Trung:\s*(\d+(?:\.\d+)?)'
[insert]
sql = """
INSERT OR REPLACE INTO student
(so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh,
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn,
lich_su, dia_ly, gdcd, khxh,
tieng_anh, tieng_phap, tieng_nga, tieng_trung)
VALUES
(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
"""
+18 -38
View File
@@ -1,4 +1,3 @@
use std::collections::HashMap;
use std::fs;
use std::path::Path;
@@ -10,7 +9,14 @@ use crate::error::BuildError;
// Top-level dataset configuration loaded from a .toml file
// ---------------------------------------------------------------------------
/// Per-dataset parse rules.
///
/// Deliberately carries no SQL. The table shape, the INSERT and the subject
/// regexes are identical for every dataset and live in `crate::schema` — keeping
/// them here meant four copies of the same DDL, which is how the 2016 and 2017
/// schemas drifted apart.
#[derive(Debug, Deserialize, Clone)]
#[serde(deny_unknown_fields)]
pub struct DatasetConfig {
pub reader: ReaderCfg,
/// Fixed column indices. Optional when format_detection handles per-file mapping.
@@ -18,10 +24,6 @@ pub struct DatasetConfig {
pub columns: Option<ColumnMap>,
pub validation: ValidationCfg,
pub header: HeaderCfg,
pub schema: SchemaCfg,
/// field name → regex source string (one entry per scoreable subject)
pub scores: HashMap<String, String>,
pub insert: InsertCfg,
/// When set to "thptqg2016", enables per-file format auto-detection.
/// Each file's header row is inspected at runtime to choose the right
/// column layout (separate-scores / mapped / default-positional).
@@ -68,18 +70,6 @@ pub struct HeaderCfg {
pub tokens: Vec<String>,
}
#[derive(Debug, Deserialize, Clone)]
pub struct SchemaCfg {
/// DDL executed verbatim before inserts (CREATE TABLE + CREATE INDEX)
pub ddl: String,
}
#[derive(Debug, Deserialize, Clone)]
pub struct InsertCfg {
/// Parameterised INSERT OR REPLACE SQL using :named_param style
pub sql: String,
}
// ---------------------------------------------------------------------------
// Loader
// ---------------------------------------------------------------------------
@@ -119,16 +109,6 @@ require_nonempty_sbd = true
[header]
tokens = ["HO_TEN", "HỌ TÊN", "STT"]
[schema]
ddl = "CREATE TABLE student (so_bao_danh TEXT PRIMARY KEY);"
[scores]
toan = 'Toán:\s*(\d+(?:\.\d+)?)'
ngu_van = 'Ngữ văn:\s*(\d+(?:\.\d+)?)'
[insert]
sql = "INSERT OR REPLACE INTO student (so_bao_danh) VALUES (:so_bao_danh)"
"#;
#[test]
@@ -142,11 +122,20 @@ sql = "INSERT OR REPLACE INTO student (so_bao_danh) VALUES (:so_bao_danh)"
assert!(!cfg.validation.require_numeric_sbd);
assert!(cfg.validation.require_nonempty_name);
assert_eq!(cfg.header.tokens.len(), 3);
assert!(cfg.scores.contains_key("toan"));
assert!(cfg.scores.contains_key("ngu_van"));
assert!(cfg.format_detection.is_none());
}
/// A config carrying leftover SQL sections must be rejected rather than
/// silently ignored — otherwise a stale [schema] block would look effective
/// while `crate::schema` was actually driving the build.
#[test]
fn config_rejects_leftover_sql_sections() {
let with_ddl = format!(
"{SAMPLE_TOML}\n[schema]\nddl = \"CREATE TABLE student (so_bao_danh TEXT);\"\n"
);
assert!(toml::from_str::<DatasetConfig>(&with_ddl).is_err());
}
#[test]
fn config_first_sheet_mode() {
let toml_str = SAMPLE_TOML.replace(r#"sheet_mode = "all""#, r#"sheet_mode = "first""#);
@@ -171,15 +160,6 @@ require_nonempty_sbd = true
[header]
tokens = ["SBD", "SOBAODANH", "STT"]
[schema]
ddl = "CREATE TABLE student (so_bao_danh TEXT PRIMARY KEY);"
[scores]
toan = 'Toán:\s*(\d+(?:\.\d+)?)'
[insert]
sql = "INSERT OR REPLACE INTO student (so_bao_danh) VALUES (?)"
"#;
let cfg: DatasetConfig = toml::from_str(toml_str).expect("parse failed");
assert_eq!(cfg.format_detection.as_deref(), Some("thptqg2016"));
@@ -337,20 +337,13 @@ pub fn process_row_2016(
#[cfg(test)]
mod tests {
use super::*;
use std::collections::HashMap;
fn s(v: &str) -> Data {
Data::String(v.to_string())
}
fn make_patterns() -> CompiledPatterns {
let mut map = HashMap::new();
map.insert("toan".into(), r"Toán:\s*(\d+(?:\.\d+)?)".into());
map.insert("ngu_van".into(), r"Ngữ văn:\s*(\d+(?:\.\d+)?)".into());
map.insert("tieng_anh".into(), r"Tiếng Anh:\s*(\d+(?:\.\d+)?)".into());
map.insert("tieng_duc".into(), r"Tiếng Đức:\s*(\d+(?:\.\d+)?)".into());
map.insert("tieng_nhat".into(), r"Tiếng Nhật:\s*(\d+(?:\.\d+)?)".into());
CompiledPatterns::new(&map).unwrap()
CompiledPatterns::new().unwrap()
}
// --- detect_format: branch 1 — separate-scores ---
+1
View File
@@ -6,5 +6,6 @@ pub mod config;
pub mod error;
pub mod format_detect_2016;
pub mod reader;
pub mod schema;
pub mod transform;
pub mod writer;
+7 -9
View File
@@ -25,9 +25,7 @@ use xlsxread::format_detect_2016::{
};
use xlsxread::reader::{is_all_blank, process_file};
use xlsxread::transform::{validate_row, CompiledPatterns, SkipReason};
use xlsxread::writer::{
finish_db, insert_row, insert_row_2016, open_db, SCORE_FIELDS, SCORE_FIELDS_2016,
};
use xlsxread::writer::{finish_db, insert_row, open_db};
fn main() -> Result<()> {
let cli = Cli::parse();
@@ -79,7 +77,7 @@ fn run_build_standard(
output_path: &Path,
) -> Result<()> {
let patterns =
CompiledPatterns::new(&cfg.scores).with_context(|| "Failed to compile score regexes")?;
CompiledPatterns::new().with_context(|| "Failed to compile score regexes")?;
let mut files: Vec<std::path::PathBuf> = std::fs::read_dir(input_dir)
.with_context(|| format!("Cannot read input dir: {}", input_dir.display()))?
@@ -109,7 +107,7 @@ fn run_build_standard(
files.len()
);
let conn = open_db(output_path, cfg)
let conn = open_db(output_path)
.with_context(|| format!("Failed to open DB: {}", output_path.display()))?;
let mut total_source_rows: u64 = 0;
@@ -159,7 +157,7 @@ fn run_build_standard(
let parsed = xlsxread::transform::transform_row(raw, cfg, &patterns);
match insert_row(&conn, &cfg.insert.sql, &parsed, SCORE_FIELDS) {
match insert_row(&conn, &parsed) {
Ok(()) => file_rows += 1,
Err(e) => {
file_errors += 1;
@@ -216,7 +214,7 @@ fn run_build_2016(
output_path: &Path,
) -> Result<()> {
let patterns =
CompiledPatterns::new(&cfg.scores).with_context(|| "Failed to compile score regexes")?;
CompiledPatterns::new().with_context(|| "Failed to compile score regexes")?;
let mut files: Vec<std::path::PathBuf> = std::fs::read_dir(input_dir)
.with_context(|| format!("Cannot read input dir: {}", input_dir.display()))?
@@ -246,7 +244,7 @@ fn run_build_2016(
files.len()
);
let conn = open_db(output_path, cfg)
let conn = open_db(output_path)
.with_context(|| format!("Failed to open DB: {}", output_path.display()))?;
let mut total_source_rows: u64 = 0;
@@ -361,7 +359,7 @@ fn process_file_2016(
// Row was empty/invalid — skipped (mirrors JS `if (!record) continue`)
}
Some(parsed) => {
match insert_row_2016(conn, &cfg.insert.sql, &parsed, SCORE_FIELDS_2016) {
match insert_row(conn, &parsed) {
Ok(()) => file_rows += 1,
Err(e) => {
*total_errors += 1;
+213
View File
@@ -0,0 +1,213 @@
//! Canonical `student` table definition — the single source of truth for the
//! SQL shape of every dataset.
//!
//! Every dataset (2016, 2017, 2017-old, 2017-old2) is written into this same
//! 22-column table. Columns a dataset has no data for bind NULL, which costs
//! ~1 byte per row in SQLite's record header.
//!
//! Column provenance:
//! ten_cum_thi, gioi_tinh, tieng_duc, tieng_nhat → 2016 only
//! khtn, khxh, gdcd, tieng_nga → 2017 datasets only
//! everything else → both
//!
//! Before this module existed the DDL, the INSERT statement and the subject
//! regex table were duplicated across four TOML configs, which is how the 2016
//! and 2017 schemas drifted apart in the first place. The configs now carry only
//! per-dataset parse rules.
// ---------------------------------------------------------------------------
// DDL
// ---------------------------------------------------------------------------
/// Executed verbatim after the output database is (re)created.
///
/// `idx_ten_cum_thi` is partial so it holds zero entries on the three 2017
/// datasets — where the column is always NULL — while staying fully useful for
/// the 2016 cluster-grouping queries.
pub const DDL: &str = "
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
ngay_sinh TEXT,
ten_cum_thi TEXT,
gioi_tinh TEXT,
toan REAL,
ngu_van REAL,
vat_ly REAL,
hoa_hoc REAL,
sinh_hoc REAL,
khtn REAL,
lich_su REAL,
dia_ly REAL,
gdcd REAL,
khxh REAL,
tieng_anh REAL,
tieng_phap REAL,
tieng_nga REAL,
tieng_duc REAL,
tieng_nhat REAL,
tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
";
// ---------------------------------------------------------------------------
// Column order
// ---------------------------------------------------------------------------
/// Identity columns, in INSERT parameter order.
pub const IDENTITY_FIELDS: &[&str] = &[
"so_bao_danh",
"ho_ten",
"ho_ten_ascii",
"ngay_sinh",
"ten_cum_thi",
"gioi_tinh",
];
/// Subject columns, in INSERT parameter order. Bound as NULL when a row has no
/// score for that subject.
pub const SCORE_FIELDS: &[&str] = &[
"toan",
"ngu_van",
"vat_ly",
"hoa_hoc",
"sinh_hoc",
"khtn",
"lich_su",
"dia_ly",
"gdcd",
"khxh",
"tieng_anh",
"tieng_phap",
"tieng_nga",
"tieng_duc",
"tieng_nhat",
"tieng_trung",
];
/// Total bound parameters per row.
pub const PARAM_COUNT: usize = IDENTITY_FIELDS.len() + SCORE_FIELDS.len();
// ---------------------------------------------------------------------------
// INSERT
// ---------------------------------------------------------------------------
/// Positional INSERT matching IDENTITY_FIELDS followed by SCORE_FIELDS.
///
/// `OR REPLACE` preserves the pre-existing behaviour where a repeated SBD
/// overwrites the earlier row rather than aborting the transaction.
pub const INSERT_SQL: &str = "
INSERT OR REPLACE INTO student
(so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh, ten_cum_thi, gioi_tinh,
toan, ngu_van, vat_ly, hoa_hoc, sinh_hoc, khtn,
lich_su, dia_ly, gdcd, khxh,
tieng_anh, tieng_phap, tieng_nga, tieng_duc, tieng_nhat, tieng_trung)
VALUES
(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
";
// ---------------------------------------------------------------------------
// Subject score patterns
// ---------------------------------------------------------------------------
/// Regex per subject, applied to the DIEM_THI cell text.
///
/// Every pattern runs against every dataset. A subject absent from a given exam
/// year simply never matches and stays NULL — 2016 source files contain no
/// "KHTN:" or "Tiếng Nga:" tokens, and 2017 files contain no "Tiếng Đức:" or
/// "Tiếng Nhật:". The parity check asserts those counts are exactly zero rather
/// than assuming it.
///
/// Order here is irrelevant (matching is by name); SCORE_FIELDS fixes the
/// INSERT order.
pub const SCORE_PATTERNS: &[(&str, &str)] = &[
("toan", r"Toán:\s*(\d+(?:\.\d+)?)"),
("ngu_van", r"Ngữ văn:\s*(\d+(?:\.\d+)?)"),
("vat_ly", r"Vật lí:\s*(\d+(?:\.\d+)?)"),
("hoa_hoc", r"Hóa học:\s*(\d+(?:\.\d+)?)"),
("sinh_hoc", r"Sinh học:\s*(\d+(?:\.\d+)?)"),
("khtn", r"KHTN:\s*(\d+(?:\.\d+)?)"),
("lich_su", r"Lịch sử:\s*(\d+(?:\.\d+)?)"),
("dia_ly", r"Địa lí:\s*(\d+(?:\.\d+)?)"),
("gdcd", r"GDCD:\s*(\d+(?:\.\d+)?)"),
("khxh", r"KHXH:\s*(\d+(?:\.\d+)?)"),
("tieng_anh", r"Tiếng Anh:\s*(\d+(?:\.\d+)?)"),
("tieng_phap", r"Tiếng Pháp:\s*(\d+(?:\.\d+)?)"),
("tieng_nga", r"Tiếng Nga:\s*(\d+(?:\.\d+)?)"),
("tieng_duc", r"Tiếng Đức:\s*(\d+(?:\.\d+)?)"),
("tieng_nhat", r"Tiếng Nhật:\s*(\d+(?:\.\d+)?)"),
("tieng_trung", r"Tiếng Trung:\s*(\d+(?:\.\d+)?)"),
];
// ---------------------------------------------------------------------------
// Tests
// ---------------------------------------------------------------------------
#[cfg(test)]
mod tests {
use super::*;
/// The INSERT placeholder count, the named column list and the field
/// constants must agree, or rows silently land in the wrong columns.
#[test]
fn insert_matches_field_order() {
assert_eq!(PARAM_COUNT, 22);
assert_eq!(INSERT_SQL.matches('?').count(), PARAM_COUNT);
let named = INSERT_SQL
.split_once('(')
.and_then(|(_, rest)| rest.split_once(')'))
.map(|(cols, _)| cols)
.expect("INSERT must contain a column list");
let listed: Vec<&str> = named
.split(',')
.map(str::trim)
.filter(|s| !s.is_empty())
.collect();
let expected: Vec<&str> = IDENTITY_FIELDS
.iter()
.chain(SCORE_FIELDS.iter())
.copied()
.collect();
assert_eq!(listed, expected);
}
/// Every subject column must have a pattern and vice versa.
#[test]
fn score_patterns_cover_score_fields() {
assert_eq!(SCORE_PATTERNS.len(), SCORE_FIELDS.len());
for (field, _) in SCORE_PATTERNS {
assert!(
SCORE_FIELDS.contains(field),
"pattern {field} has no column"
);
}
for field in SCORE_FIELDS {
assert!(
SCORE_PATTERNS.iter().any(|(f, _)| f == field),
"column {field} has no pattern"
);
}
}
/// Every column named in the DDL must be bound by the INSERT.
#[test]
fn ddl_columns_match_insert() {
for field in IDENTITY_FIELDS.iter().chain(SCORE_FIELDS.iter()) {
assert!(DDL.contains(field), "DDL missing column {field}");
}
}
#[test]
fn score_patterns_compile() {
for (field, src) in SCORE_PATTERNS {
regex::Regex::new(src).unwrap_or_else(|e| panic!("{field}: {e}"));
}
}
}
+12 -13
View File
@@ -15,22 +15,25 @@ use crate::error::BuildError;
// ---------------------------------------------------------------------------
pub struct CompiledPatterns {
/// Ordered list so INSERT column order is deterministic
/// One compiled regex per subject column in `schema::SCORE_PATTERNS`.
pub patterns: Vec<(String, Regex)>,
}
impl CompiledPatterns {
pub fn new(scores: &HashMap<String, String>) -> Result<Self, BuildError> {
let mut patterns = Vec::with_capacity(scores.len());
for (field, src) in scores {
/// Compile the canonical subject patterns once at startup.
///
/// All 16 patterns run against every dataset. A subject that did not exist
/// in a given exam year simply never matches and stays NULL — the parity
/// check asserts those counts are exactly zero rather than assuming it.
pub fn new() -> Result<Self, BuildError> {
let mut patterns = Vec::with_capacity(crate::schema::SCORE_PATTERNS.len());
for (field, src) in crate::schema::SCORE_PATTERNS {
let re = Regex::new(src).map_err(|e| BuildError::Regex {
pattern: src.clone(),
pattern: (*src).to_string(),
source: e,
})?;
patterns.push((field.clone(), re));
patterns.push(((*field).to_string(), re));
}
// Sort for deterministic order across HashMap iteration
patterns.sort_by(|a, b| a.0.cmp(&b.0));
Ok(Self { patterns })
}
}
@@ -314,11 +317,7 @@ mod tests {
// --- Score parsing tests ---
fn make_patterns() -> CompiledPatterns {
let mut map = HashMap::new();
map.insert("toan".into(), r"Toán:\s*(\d+(?:\.\d+)?)".into());
map.insert("ngu_van".into(), r"Ngữ văn:\s*(\d+(?:\.\d+)?)".into());
map.insert("vat_ly".into(), r"Vật lí:\s*(\d+(?:\.\d+)?)".into());
CompiledPatterns::new(&map).unwrap()
CompiledPatterns::new().unwrap()
}
#[test]
+21 -89
View File
@@ -3,27 +3,24 @@
/// Mirrors build-lib.js createDb + the transaction loop in each build-database*.js.
/// Stats output lines match the JS stdout exactly so existing CI log-greps still work.
///
/// thptqg2016 differences vs thptqg2017:
/// - INSERT includes ten_cum_thi and gioi_tinh after ngay_sinh
/// - Score columns are toan/ngu_van/vat_ly/hoa_hoc/sinh_hoc/lich_su/dia_ly/
/// tieng_anh/tieng_phap/tieng_duc/tieng_nhat/tieng_trung
/// (no khtn, khxh, tieng_nga)
/// Every dataset writes the same canonical table (see `crate::schema`), so there
/// is exactly one insert path. Columns a dataset carries no data for bind NULL.
use std::fs;
use std::path::Path;
use rusqlite::{params_from_iter, Connection, ToSql};
use crate::config::DatasetConfig;
use crate::error::BuildError;
use crate::schema;
use crate::transform::ParsedRow;
// ---------------------------------------------------------------------------
// DB initialisation — mirrors build-lib.js createDb (delete + recreate)
// ---------------------------------------------------------------------------
/// Open (or recreate) the output SQLite database, execute the DDL from config,
/// Open (or recreate) the output SQLite database, execute the canonical DDL,
/// and return the open connection ready for inserts.
pub fn open_db(db_path: &Path, cfg: &DatasetConfig) -> Result<Connection, BuildError> {
pub fn open_db(db_path: &Path) -> Result<Connection, BuildError> {
// Mirror Node behaviour: delete existing file before creating (build-lib.js:54)
if db_path.exists() {
fs::remove_file(db_path).map_err(|e| BuildError::Io {
@@ -43,92 +40,22 @@ pub fn open_db(db_path: &Path, cfg: &DatasetConfig) -> Result<Connection, BuildE
}
let conn = Connection::open(db_path)?;
conn.execute_batch(&cfg.schema.ddl)?;
conn.execute_batch(schema::DDL)?;
Ok(conn)
}
// ---------------------------------------------------------------------------
// Score field lists — one per dataset schema
// Insert a single parsed row
// ---------------------------------------------------------------------------
/// thptqg2017 score columns (14 fields including khtn/khxh/tieng_nga).
pub const SCORE_FIELDS_2017: &[&str] = &[
"toan",
"ngu_van",
"vat_ly",
"hoa_hoc",
"sinh_hoc",
"khtn",
"lich_su",
"dia_ly",
"gdcd",
"khxh",
"tieng_anh",
"tieng_phap",
"tieng_nga",
"tieng_trung",
];
/// Bind every field from `row` into the canonical INSERT and execute it.
///
/// Parameter order is `schema::IDENTITY_FIELDS` followed by
/// `schema::SCORE_FIELDS`. Subjects absent from `row.scores` — and the two
/// identity columns only the thptqg2016 layouts populate — bind NULL.
pub fn insert_row(conn: &Connection, row: &ParsedRow) -> Result<(), BuildError> {
let mut params: Vec<Box<dyn ToSql>> = Vec::with_capacity(schema::PARAM_COUNT);
/// thptqg2016 score columns (12 fields: tieng_duc/tieng_nhat present; no khtn/khxh/tieng_nga).
pub const SCORE_FIELDS_2016: &[&str] = &[
"toan",
"ngu_van",
"vat_ly",
"hoa_hoc",
"sinh_hoc",
"lich_su",
"dia_ly",
"tieng_anh",
"tieng_phap",
"tieng_duc",
"tieng_nhat",
"tieng_trung",
];
/// Alias kept so thptqg2017 callers that import SCORE_FIELDS continue to compile.
pub const SCORE_FIELDS: &[&str] = SCORE_FIELDS_2017;
// ---------------------------------------------------------------------------
// Insert a single parsed row (thptqg2017 schema — no ten_cum_thi / gioi_tinh)
// ---------------------------------------------------------------------------
/// Bind all fields from `row` into the prepared statement and execute it.
/// Positional params: so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh, <scores...>
pub fn insert_row(
conn: &Connection,
sql: &str,
row: &ParsedRow,
score_fields: &[&str],
) -> Result<(), BuildError> {
let mut params: Vec<Box<dyn ToSql>> = Vec::with_capacity(4 + score_fields.len());
params.push(Box::new(row.so_bao_danh.clone()));
params.push(Box::new(row.ho_ten.clone()));
params.push(Box::new(row.ho_ten_ascii.clone()));
params.push(Box::new(row.ngay_sinh.clone()));
for field in score_fields {
let val: Option<f64> = row.scores.get(*field).copied();
params.push(Box::new(val));
}
conn.execute(sql, params_from_iter(params.iter().map(|p| p.as_ref())))?;
Ok(())
}
// ---------------------------------------------------------------------------
// Insert a single parsed row (thptqg2016 schema — includes ten_cum_thi + gioi_tinh)
// ---------------------------------------------------------------------------
/// Bind all fields from `row` into the prepared statement for the thptqg2016 schema.
/// Positional params: so_bao_danh, ho_ten, ho_ten_ascii, ngay_sinh, ten_cum_thi,
/// gioi_tinh, <scores...>
pub fn insert_row_2016(
conn: &Connection,
sql: &str,
row: &ParsedRow,
score_fields: &[&str],
) -> Result<(), BuildError> {
let mut params: Vec<Box<dyn ToSql>> = Vec::with_capacity(6 + score_fields.len());
params.push(Box::new(row.so_bao_danh.clone()));
params.push(Box::new(row.ho_ten.clone()));
params.push(Box::new(row.ho_ten_ascii.clone()));
@@ -136,12 +63,17 @@ pub fn insert_row_2016(
params.push(Box::new(row.ten_cum_thi.clone()));
params.push(Box::new(row.gioi_tinh.clone()));
for field in score_fields {
for field in schema::SCORE_FIELDS {
let val: Option<f64> = row.scores.get(*field).copied();
params.push(Box::new(val));
}
conn.execute(sql, params_from_iter(params.iter().map(|p| p.as_ref())))?;
debug_assert_eq!(params.len(), schema::PARAM_COUNT);
conn.execute(
schema::INSERT_SQL,
params_from_iter(params.iter().map(|p| p.as_ref())),
)?;
Ok(())
}
+3 -9
View File
@@ -581,8 +581,7 @@ fn run_build_cmd(input_dir: &Path, db_path: &Path, config_name: &str) {
let cfg = xlsxread::config::load_config(&cfg_path)
.unwrap_or_else(|e| panic!("load config {config_name}: {e}"));
let patterns =
xlsxread::transform::CompiledPatterns::new(&cfg.scores).expect("compile patterns");
let patterns = xlsxread::transform::CompiledPatterns::new().expect("compile patterns");
// Collect files
let mut files: Vec<PathBuf> = std::fs::read_dir(input_dir)
@@ -602,7 +601,7 @@ fn run_build_cmd(input_dir: &Path, db_path: &Path, config_name: &str) {
.collect();
files.sort();
let conn = xlsxread::writer::open_db(db_path, &cfg).expect("open db");
let conn = xlsxread::writer::open_db(db_path).expect("open db");
conn.execute_batch("BEGIN").unwrap();
for file in &files {
@@ -632,12 +631,7 @@ fn run_build_cmd(input_dir: &Path, db_path: &Path, config_name: &str) {
return;
}
let row = xlsxread::transform::transform_row(raw, &cfg, &patterns);
let _ = xlsxread::writer::insert_row(
&conn,
&cfg.insert.sql,
&row,
xlsxread::writer::SCORE_FIELDS,
);
let _ = xlsxread::writer::insert_row(&conn, &row);
})
.expect("process file");
}
+79
View File
@@ -0,0 +1,79 @@
#!/usr/bin/env node
/**
* Dump per-dataset statistics from built SQLite databases as JSON.
*
* Used twice: once against the pre-refactor databases to capture a baseline,
* and again after the schema unification. Comparing the two outputs is what
* proves no score data was lost or invented.
*
* Schema-agnostic on purpose — columns come from PRAGMA table_info, so the same
* script runs against the old 18/20-column tables and the new 22-column one.
*
* Usage:
* node db-stats.js <label>=<db-path> [<label>=<db-path> ...] > stats.json
*/
import { DatabaseSync } from "node:sqlite";
import { statSync } from "node:fs";
// Deterministic value-level sample: every SBD ending in these digits.
// Re-running against a rebuilt DB compares the same students.
const SAMPLE_SUFFIX = "0000";
function collect(dbPath) {
const db = new DatabaseSync(dbPath, { readOnly: true });
const columns = db
.prepare("PRAGMA table_info(student)")
.all()
.map((c) => c.name);
const rowCount = db.prepare("SELECT COUNT(*) AS c FROM student").get().c;
// One pass over the table counting non-NULLs for every column at once.
const sums = columns
.map((c) => `SUM(CASE WHEN "${c}" IS NOT NULL THEN 1 ELSE 0 END) AS "${c}"`)
.join(", ");
const nonNull = db.prepare(`SELECT ${sums} FROM student`).get();
const sample = db
.prepare(
`SELECT * FROM student WHERE so_bao_danh LIKE '%${SAMPLE_SUFFIX}'
ORDER BY so_bao_danh`,
)
.all();
db.close();
return {
rowCount,
columns,
nonNull: Object.fromEntries(columns.map((c) => [c, Number(nonNull[c])])),
sizeBytes: statSync(dbPath).size,
sampleCount: sample.length,
// Store the sample keyed by SBD so a later diff can report which student
// and which field changed, not just that something did.
sample: Object.fromEntries(sample.map((r) => [r.so_bao_danh, r])),
};
}
const args = process.argv.slice(2);
if (args.length === 0) {
console.error("usage: db-stats.js <label>=<db-path> [...]");
process.exit(2);
}
const out = {};
for (const arg of args) {
const idx = arg.indexOf("=");
if (idx === -1) {
console.error(`bad argument (expected label=path): ${arg}`);
process.exit(2);
}
const label = arg.slice(0, idx);
const path = arg.slice(idx + 1);
process.stderr.write(`collecting ${label} from ${path}\n`);
out[label] = collect(path);
}
process.stdout.write(JSON.stringify(out, null, 2) + "\n");
+157
View File
@@ -0,0 +1,157 @@
#!/usr/bin/env node
/**
* Compare rebuilt databases against the pre-refactor parity baseline.
*
* The schema deliberately changed shape, so a whole-file hash is meaningless.
* What must hold instead:
*
* 1. Row count per dataset is unchanged.
* 2. Every column that existed before has the same non-NULL count.
* 3. Every column newly added to a dataset has a non-NULL count of exactly 0.
* This is the check that catches the union-regex risk — if the 16-pattern
* map starts matching text the narrower per-year map ignored, it shows up
* here rather than silently corrupting the dataset.
* 4. A deterministic sample of students is identical field by field.
*
* Exits non-zero on any mismatch.
*
* Usage:
* node verify-parity.js <baseline.json> <current.json>
*/
import { readFileSync } from "node:fs";
/**
* Foreign-language scores the pre-refactor configs silently discarded.
*
* The 2016 config listed 12 subject regexes and the 2017 configs listed 14;
* neither list was complete. Candidates could sit German, Japanese and Russian
* in both exam years, so every one of these students previously ended up with
* no foreign-language score at all.
*
* Unifying to the canonical 16 patterns recovers them. Verified real, not
* spurious matches: across all four datasets every student holds either zero
* or exactly one foreign language — never two — and each affected student had
* all language columns NULL beforehand.
*
* These exact counts are approved. Any other newly-populated column, or any
* drift in these numbers, still fails the gate.
*/
const APPROVED_RECOVERY = {
"2016": { tieng_nga: 182 },
"2017": { tieng_duc: 93, tieng_nhat: 512 },
"2017-old": { tieng_duc: 85, tieng_nhat: 484 },
"2017-old2": { tieng_duc: 22, tieng_nhat: 313 },
};
const [baselinePath, currentPath] = process.argv.slice(2);
if (!baselinePath || !currentPath) {
console.error("usage: verify-parity.js <baseline.json> <current.json>");
process.exit(2);
}
const baseline = JSON.parse(readFileSync(baselinePath, "utf8"));
const current = JSON.parse(readFileSync(currentPath, "utf8"));
const failures = [];
const notes = [];
for (const dataset of Object.keys(baseline)) {
const b = baseline[dataset];
const c = current[dataset];
if (!c) {
failures.push(`${dataset}: missing from current stats`);
continue;
}
// 1. Row count
if (b.rowCount !== c.rowCount) {
failures.push(
`${dataset}: row count ${b.rowCount} → ${c.rowCount} (${c.rowCount - b.rowCount >= 0 ? "+" : ""}${c.rowCount - b.rowCount})`,
);
}
// 2. Pre-existing columns keep their non-NULL counts
for (const col of b.columns) {
if (!(col in c.nonNull)) {
failures.push(`${dataset}.${col}: column disappeared from schema`);
continue;
}
if (b.nonNull[col] !== c.nonNull[col]) {
failures.push(
`${dataset}.${col}: non-NULL ${b.nonNull[col]} → ${c.nonNull[col]} (${c.nonNull[col] - b.nonNull[col] >= 0 ? "+" : ""}${c.nonNull[col] - b.nonNull[col]})`,
);
}
}
// 3. Newly added columns must be NULL — except the approved recoveries,
// which must match their approved count exactly.
const approved = APPROVED_RECOVERY[dataset] ?? {};
const added = c.columns.filter((col) => !b.columns.includes(col));
const recovered = [];
for (const col of added) {
const expected = approved[col] ?? 0;
if (c.nonNull[col] !== expected) {
failures.push(
expected === 0
? `${dataset}.${col}: new column has ${c.nonNull[col]} non-NULL values, expected 0`
: `${dataset}.${col}: recovered ${c.nonNull[col]} values, approved count is ${expected}`,
);
} else if (expected > 0) {
recovered.push(`${col}=${expected}`);
}
}
// An approved recovery that vanished means the union patterns regressed.
for (const [col, expected] of Object.entries(approved)) {
if (!added.includes(col)) {
failures.push(
`${dataset}.${col}: expected ${expected} recovered values but column is not new`,
);
}
}
if (added.length) {
const nulls = added.length - recovered.length;
notes.push(
`${dataset}: +${added.length} new columns (${nulls} all-NULL as expected` +
(recovered.length ? `, recovered ${recovered.join(", ")}` : "") +
")",
);
}
// 4. Deterministic sample compared field by field
if (b.sampleCount !== c.sampleCount) {
failures.push(
`${dataset}: sample size ${b.sampleCount} → ${c.sampleCount}`,
);
}
for (const [sbd, bRow] of Object.entries(b.sample)) {
const cRow = c.sample[sbd];
if (!cRow) {
failures.push(`${dataset}: sampled student ${sbd} missing after rebuild`);
continue;
}
for (const [field, bVal] of Object.entries(bRow)) {
if (cRow[field] !== bVal) {
failures.push(
`${dataset}: student ${sbd} field ${field}: ${JSON.stringify(bVal)} → ${JSON.stringify(cRow[field])}`,
);
}
}
}
const sizeDelta = ((c.sizeBytes - b.sizeBytes) / b.sizeBytes) * 100;
notes.push(
`${dataset}: ${c.rowCount} rows, ${b.columns.length} → ${c.columns.length} cols, size ${sizeDelta >= 0 ? "+" : ""}${sizeDelta.toFixed(1)}%`,
);
}
for (const n of notes) console.log(` ${n}`);
if (failures.length) {
console.error(`\nPARITY FAILED — ${failures.length} mismatch(es):\n`);
for (const f of failures) console.error(` ✗ ${f}`);
process.exit(1);
}
console.log("\nPARITY OK — row counts, per-column non-NULL counts and samples all match.");
@@ -0,0 +1,170 @@
---
phase: 1
title: "Standard schema and unified parser"
status: pending
priority: P1
dependencies: []
effort: ""
---
# Phase 1: Standard schema and unified parser
## Overview
Move the SQL schema out of the four TOML configs into one Rust module, widen it
to the 22-column union, and prove the four rebuilt databases still match their
current contents. Done in place under the existing `2016/`/`2017/` layout so
parity is measurable before anything moves.
## Context
`2016/tools/xlsxread` and `2017/tools/xlsxread` are the same crate. Verified by
`diff -rq`: `cli.rs`, `error.rs`, `reader.rs` are byte-identical; `audit.rs`,
`config.rs`, `lib.rs`, `main.rs`, `transform.rs`, `writer.rs` differ only where
2016 adds a branch; `format_detect_2016.rs` (549 lines) exists only in 2016.
The 2016 crate is therefore the base to keep — it already builds 2017 datasets
(its `configs/` holds 2017 fixtures used by `tests/golden.rs`).
Current per-dataset schema divergence:
| Column group | 2016 | 2017 / old / old2 |
|---|---|---|
| `ten_cum_thi`, `gioi_tinh` | present | absent |
| `khtn`, `khxh`, `gdcd`, `tieng_nga` | absent | present |
| `tieng_duc`, `tieng_nhat` | present | absent |
## Requirements
**Functional**
- One DDL, one INSERT, one ordered subject list, one regex map — used by all four datasets
- Configs keep only genuinely per-dataset settings
- Rebuilt DBs preserve every value currently produced
**Non-functional**
- No measurable build-time regression (currently a few minutes per dataset)
- DB size growth from the added NULL columns stays under ~2%
## Architecture
**New `parser/src/schema.rs`** (written into `2016/tools/xlsxread/src/` this
phase; relocated in Phase 2) exports:
- `pub const DDL: &str` — the 22-column CREATE TABLE + 3 indexes
- `pub const INSERT_SQL: &str` — 22 positional placeholders
- `pub const IDENTITY_FIELDS: &[&str]` — 6 identity columns in INSERT order
- `pub const SCORE_FIELDS: &[&str]` — 16 subject columns in INSERT order
- `pub fn score_patterns() -> Vec<(&'static str, &'static str)>` — the union of
all 16 subject regexes (2016's 12 ∪ 2017's 14; both sets share 10)
**`writer.rs` collapses.** `SCORE_FIELDS_2016`, `SCORE_FIELDS_2017`, the
`SCORE_FIELDS` alias, and `insert_row_2016` all disappear. A single `insert_row`
binds identity fields then subject fields, always 22 params.
**`config.rs` shrinks.** `[schema]` and `[insert]` sections are removed from
`DatasetConfig`; `[scores]` becomes optional and unused (delete the field once
no config sets it). `columns` stays `Option<Columns>` — the 2016 detection path
does not use it.
**`main.rs`** keeps the `run_build_2016` / `run_build_standard` split; both now
call the same `insert_row` with `schema::SCORE_FIELDS`.
### Union regex risk — RESOLVED, and it found a real bug
Applying all 16 patterns to every dataset means a source cell containing a
subject the old config never looked for now populates a column it previously
could not.
**Outcome: the gate fired, and the matches were real data, not false positives.**
The pre-refactor configs listed 12 subject regexes (2016) and 14 (2017). Neither
list was complete — Vietnamese candidates could sit German, Japanese and Russian
in *both* exam years. Every affected student previously ended up with **no**
foreign-language score at all. Unifying to 16 patterns recovered 1,691 scores:
| Dataset | Recovered |
|---|---|
| 2016 | 182 × `tieng_nga` |
| 2017 | 93 × `tieng_duc`, 512 × `tieng_nhat` |
| 2017-old | 85 × `tieng_duc`, 484 × `tieng_nhat` |
| 2017-old2 | 22 × `tieng_duc`, 313 × `tieng_nhat` |
Verified genuine, not spurious:
- Across all four datasets every student holds **zero or exactly one** foreign
language — never two. So the new columns duplicate nothing.
- Each affected student had *all* language columns NULL beforehand
(e.g. SBD `01003198`: all NULL → `tieng_duc = 8`).
- Counts track dataset size consistently across the three 2017 generations
(93/85/22 German, 512/484/313 Japanese).
- Score ranges and distributions are ordinary exam values (0–10).
Row counts, all 18 pre-existing per-column non-NULL counts, and the
deterministic student samples were **identical** — nothing was lost.
Accepted as a data-quality fix. The earlier guidance to add a per-config subject
allowlist assumed spurious matches and does **not** apply; suppressing these
would knowingly discard real scores. The exact counts are now encoded as
`APPROVED_RECOVERY` in `verify-parity.js`, so any *other* newly-populated column
— or any drift in these numbers — still fails the gate.
## Related Code Files
- Create: `2016/tools/xlsxread/src/schema.rs`
- Modify: `2016/tools/xlsxread/src/writer.rs` (drop dual insert paths + field lists)
- Modify: `2016/tools/xlsxread/src/config.rs` (drop `[schema]`/`[insert]`/`[scores]`)
- Modify: `2016/tools/xlsxread/src/main.rs` (single insert call site per path)
- Modify: `2016/tools/xlsxread/src/lib.rs` (declare `schema` module)
- Modify: `2016/tools/xlsxread/src/transform.rs` (source patterns from `schema`)
- Modify: `2016/tools/xlsxread/tests/golden.rs` (assert against canonical schema)
- Create: `2016/tools/xlsxread/configs/thptqg2017-data-old2.toml` (copy from 2017 crate)
- Modify: all 4 configs — strip DDL/INSERT/scores down to parse rules
- Delete (Phase 2): `2017/tools/`
## Implementation Steps
1. **Capture the baseline first.** Build all four DBs with the *current* code and
record, per dataset: `SELECT COUNT(*) FROM student`, and for each column
`SUM(col IS NOT NULL)`. Store as `plans/reports/parser-parity-baseline.json`.
Nothing else in this plan is verifiable without this artifact.
2. Copy `thptqg2017-data-old2.toml` into the 2016 crate's `configs/`.
3. Write `schema.rs` with DDL, INSERT, field orders, and the union regex map.
4. Rewrite `writer.rs` to a single `insert_row`; delete `insert_row_2016` and
the three score-field constants.
5. Strip `[schema]`, `[insert]`, `[scores]` from all four configs; update
`config.rs` structs to match. Each config should end at ~15–20 lines.
6. Update `main.rs` and `transform.rs` call sites.
7. Update `tests/golden.rs` to assert the canonical column set.
8. `cargo test` — golden tests must pass.
9. Rebuild all four DBs; regenerate the same stats and diff against the baseline.
## Tests / Validation
- `cargo test` — 63 tests (55 unit + 8 golden)
- `cargo clippy --all-targets -- -D warnings`
- Row count per dataset equals baseline exactly
- Per-column non-NULL count equals baseline for every previously-existing column
- Newly-added columns read 0 non-NULL, except the approved recoveries above
- DB file size within 2% of baseline
**Note on the clippy gate:** it was already red before this phase — measured at
**8 errors on the branch base**. This phase's changes introduced none. The
pre-existing lints were cleared in a separate commit so the gate is genuinely
green from here on.
## Success Criteria
- [x] `schema.rs` is the only place DDL/INSERT/subject-order/regexes are written
- [x] All 4 configs under 35 lines, containing no SQL
- [x] `writer.rs` has exactly one insert function
- [x] `cargo test` (63 passing) and `cargo clippy --all-targets -D warnings` green
- [x] All 4 DBs match baseline row counts and per-column non-NULL counts
- [x] Baseline JSON committed under `plans/reports/`
- [x] Config parsing rejects leftover SQL sections (`deny_unknown_fields`)
## Risk Assessment
| Risk | Mitigation |
|---|---|
| Union regexes populate unexpected columns | Baseline diff catches it; fall back to per-config subject allowlist |
| INSERT param order drifts from DDL order | Single `IDENTITY_FIELDS`/`SCORE_FIELDS` source drives both; golden test asserts round-trip |
| Rebuilding 419 MB of source is slow | Run once per verification pass, not per edit; iterate against golden fixtures |
| Baseline skipped under time pressure | Phase 1 is unverifiable without it — treat step 1 as blocking |
@@ -0,0 +1,149 @@
---
phase: 2
title: "Repo restructure"
status: pending
priority: P1
dependencies: [1]
effort: ""
---
# Phase 2: Repo restructure
## Overview
Move files into the target layout and migrate pnpm → npm. Relocation plus
package-manager swap — no application logic changes. Kept as its own commit so
the 419 MB data move is trivially revertable and reviewable separately from
behavior changes.
## Requirements
- All moves via `git mv` so blobs are reused and history follows
- Working tree builds after the move (paths updated, nothing dangling)
- Repository size does not grow
- npm is the only package manager; `package-lock.json` committed, no pnpm files remain
## Architecture
Git stores blobs content-addressed, so renaming 299 tracked data files
(~419 MB working-tree) adds no new objects. The cost is local I/O and one large
tree rewrite, not repository growth.
Data directory sizes being moved: `2016/data` 62 MB, `2017/data` 286 MB,
`2017/data-old` 39 MB, `2017/data-old2` 32 MB.
### pnpm → npm
A single package with 3 runtime and 9 dev dependencies, no workspace. pnpm's
advantages (strict resolution, shared store) buy nothing at this size, and two
pieces of scaffolding disappear with it:
- `pnpm-workspace.yaml` exists **only** to whitelist `better-sqlite3`'s native
build (`allowBuilds`). npm runs postinstall by default, so the file has no npm
equivalent — it is deleted, not translated. Verified: this is the file's
entire content in both projects.
- The `pnpm/action-setup@v4` CI step and `cache: 'pnpm'` both drop out.
Lockfiles cannot be converted; `package-lock.json` is generated fresh from
`package.json`. Migration direction is safe: pnpm's strict `node_modules` layout
forbids phantom dependencies, so anything that resolved under pnpm also resolves
under npm's flat tree. The reverse would not hold.
Latent issue surfaced while checking: `2017/scripts/diff-datasets.js` imports
`better-sqlite3`, which is **not declared in any `package.json`** — that script
cannot run today without an ad-hoc install. Do not add the dependency; Phase 5
uses the built-in `node:sqlite` instead (verified working on Node 24, no flag).
Either port `diff-datasets.js` to `node:sqlite` in this phase or leave it broken
exactly as it is today and note it — do not silently half-fix it.
## Related Code Files
**Moves**
| From | To |
|---|---|
| `2016/tools/xlsxread/` | `parser/` |
| `2016/tools/xlsxread/configs/thptqg2016-data.toml` | `parser/configs/2016.toml` |
| `2017/tools/xlsxread/configs/thptqg2017-data.toml` | `parser/configs/2017.toml` |
| `2017/tools/xlsxread/configs/thptqg2017-data-old.toml` | `parser/configs/2017-old.toml` |
| `2017/tools/xlsxread/configs/thptqg2017-data-old2.toml` | `parser/configs/2017-old2.toml` |
| `2017/scripts/` | `parser/scripts/` |
| `2016/data/` | `data/2016/` |
| `2017/data/` | `data/2017/` |
| `2017/data-old/` | `data/2017-old/` |
| `2017/data-old2/` | `data/2017-old2/` |
| `2017/src/` | `src/` |
| `2017/index.html` | `index.html` (overwrites the old static landing page) |
| `2017/package.json`, `eslint.config.js`, `vite.config.js` | repo root |
| `2016/docs/*`, `2017/docs/*` | `docs/` |
| `2017/LICENSE` | `LICENSE` |
The old root `index.html` (the static hub) is **not moved** — it becomes
`src/components/hub.jsx` in Phase 3. Read it before deleting; its four links,
candidate counts, and Vietnamese copy are the source material for that
component.
**Deletes**
- `2016/` and `2017/` directories entirely (after moves)
- `2016/src/` — superseded by `src/` (its unique columns and SQL presets are ported in Phase 3)
- `2016/tools/xlsxread/configs/thptqg2017-*.toml` — test fixtures, superseded by real configs
- `2016/pnpm-lock.yaml`, `2017/pnpm-lock.yaml`, both `pnpm-workspace.yaml`
- Duplicate `2016/package.json`, `2016/eslint.config.js`, `2016/vite.config.js`, `2016/LICENSE`
- Root `index.html` (static hub) — only after its content is captured for `hub.jsx`
**Merges**
- `.gitignore` — one root file. 2016's is the verbose GitHub Node template,
2017's is terse and accurate. Take 2017's as the base, add `parser/target/`,
`.build/`, and `dist/`.
- `package.json` — root file keeps 2017's dependency set (identical to 2016's).
Drop the `"packageManager": "pnpm@11.1.1"` field.
## Implementation Steps
1. Branch: `refactor/unify-frontend-and-schema`.
2. Copy the old root `index.html` content somewhere durable for Phase 3
(`plans/reports/` or the phase-03 file itself) — it is the hub's source copy.
3. `git mv 2017/src src`, then the 2017 root config files, then
`git mv 2017/index.html index.html`.
4. `git mv 2016/tools/xlsxread parser`, then rename the four configs.
5. `git mv 2017/scripts parser/scripts`.
6. Move the four data directories.
7. Merge docs into `docs/`, prefixing 2016-specific filenames where they collide
(`deployment-guide.md` and `system-architecture.md` exist in both — read both
before merging; they describe different pipelines).
8. Delete the emptied `2016/` and `2017/` trees plus both `pnpm-workspace.yaml`
and both `pnpm-lock.yaml`.
9. Write the merged root `.gitignore`.
10. Drop `"packageManager"` from `package.json`; run `npm install` to generate
`package-lock.json`; commit the lockfile.
11. Fix paths inside moved files: `parser/Cargo.toml` package name/paths,
config `--input`/`--output` defaults, `parser/scripts/*.js` relative paths.
12. Replace `pnpm` with `npm run` in every `package.json` script body.
13. `cargo test` from `parser/` — must still be green.
## Tests / Validation
- `cargo test --manifest-path parser/Cargo.toml`
- `npm ci && npm run lint` at root — `npm ci` proves the lockfile is in sync
- `git status` shows renames (R), not delete+add pairs
- No file outside `plans/` still references `2016/tools`, `2017/tools`,
`2017/src`, `2017/data`, or `pnpm` — grep to confirm
## Success Criteria
- [ ] Target layout matches `plan.md` exactly
- [ ] `2016/` and `2017/` no longer exist
- [ ] Old hub page content captured before deletion
- [ ] Git reports renames, repository size unchanged
- [ ] `package-lock.json` committed; `npm ci` succeeds; zero pnpm files remain
- [ ] `cargo test` green from the new location
## Risk Assessment
| Risk | Mitigation |
|---|---|
| Old hub page deleted before its copy is captured | Step 2 runs before any move; content also recoverable from git history |
| Docs collide silently on merge | Read both copies before merging; two same-named files describe different pipelines |
| Git records delete+add instead of rename | Use `git mv`; verify with `git status` before commit |
| Stale path references in scripts/CI | Grep sweep in validation; Phase 4 rewrites the workflow |
| Fresh npm lockfile resolves different transitive versions than pnpm did | All deps are caret-ranged and already floating; `npm run lint` + the Phase 4 local build run are the check |
@@ -0,0 +1,249 @@
---
phase: 3
title: "Unified frontend and dataset registry"
status: pending
priority: P1
dependencies: [2]
effort: ""
---
# Phase 3: Unified frontend and dataset registry
## Overview
Make the single 2017-derived frontend serve every page on the site: the four
dataset views plus the `/thptqg/` hub. Three kinds of work — port 2016's only
unique feature (cluster + gender columns) into the shared components, extract
everything genuinely per-dataset into a runtime registry, and add a small
pathname router with a hub route.
## Context — what actually differs
The 2017 frontend is a superset in every respect except two columns. Verified by
reading both trees:
**2017 has, 2016 lacks:** debounced live search with input-mode hints
(`search-form.jsx`, 127 vs 46 lines), `?q=` URL deep links, `student-detail.jsx`
(183 lines, single-result view), `lib/admission-blocks.js` (49 blocks +
`scoreTier` ladder), all-NULL column hiding, `/` and Ctrl+Enter shortcuts,
footer total count, load progress bar, richer SQL presets (335 vs 219 lines).
**2016 has, 2017 lacks:** `ten_cum_thi` and `gioi_tinh` table columns with a
`cumthi-cell` title tooltip, and a simpler 3-tier `scoreClass` (superseded by
2017's 6-tier `scoreTier`).
## Requirements
**Functional**
- 2016 site keeps cluster + gender columns; 2017 sites must not show them
- 2016 site gains every 2017 feature listed above
- Per-dataset chrome (title, subtitle, source, DB size, search examples, SQL presets) is data, not code
- Admission-block computation works for both exam years without branching
- `/thptqg/` renders a hub listing the four datasets; no DB is fetched there
- Deep links (`?q=`) keep working on every dataset route
**Non-functional**
- Dataset identity resolves from `location.pathname` — no fetch, no env plumbing
- No routing library; five static routes do not justify a dependency
- Hub route must not pull the 47 MB DB into its critical path
## Architecture
### `src/datasets.js` — the registry
```js
export const DATASETS = [
{
id: "2016", // === URL segment === data dir === config === db file
dbSizeMb: 48,
title: "Tra cứu điểm thi THPT Quốc gia 2016",
subtitle: "Dữ liệu thí sinh toàn quốc · Hỗ trợ truy vấn SQL tùy chỉnh",
source: "Bộ GD&ĐT",
examples: ["<real 2016 SBD>", "Nguyễn Minh Tiến"],
presets: PRESETS_2016, // from 2016/src/components/custom-query.jsx
blocks: BLOCKS_2016, // see admission blocks below
},
{ id: "2017", /* presets: PRESETS_2017 */ },
{ id: "2017-old", /* same presets as 2017, own label */ },
{ id: "2017-old2", /* same presets as 2017, own label */ },
];
export const pathOf = (d) => `${import.meta.env.BASE_URL}${d.id}/`;
export const dbOf = (d) => `${import.meta.env.BASE_URL}db/${d.id}.db.gz`;
```
Plain runtime data — no `import.meta.env.VITE_DATASET` inlining, no build
variants. All four entries ship in the one bundle; the preset SQL totals a few
KB gzipped. Because URLs are flat, path and DB URL are *derived* from `id`
rather than stored, so a dataset cannot be misconfigured into pointing at the
wrong database.
### `src/router.js`
Flat URLs make this an exact match on a single segment — the segment **is** the
dataset ID:
```js
export function resolveRoute(pathname = location.pathname) {
const seg = pathname
.slice(import.meta.env.BASE_URL.length) // strip "/thptqg/"
.replace(/\/$/, "");
return DATASETS.find((d) => d.id === seg) ?? null; // null → hub
}
```
No prefix sorting, no ambiguity between `2017` and `2017-old` — that entire
class of bug is designed out by the flat scheme. No `react-router`.
### Legacy path redirects
The two old nested URLs are kept alive by a small map, so existing links and
bookmarks resolve instead of 404ing:
```js
const LEGACY = { "2017/old": "2017-old", "2017/old2": "2017-old2" };
```
On a legacy match, `history.replaceState` to the canonical flat URL **preserving
`location.search`** (the `?q=` deep link must survive), then resolve normally.
Phase 4 emits `index.html` at both legacy paths so Pages serves them at all.
Droppable if you'd rather let the old URLs 404 — it costs ~5 lines and two file
copies, and nothing else in the plan depends on it.
### `src/components/hub.jsx`
Renders the four dataset links from `DATASETS` via `pathOf()`. Content ported
from the old static root `index.html` (captured in Phase 2 step 2): heading, the
sql.js one-liner explanation, candidate counts, and the "phiên bản cũ" grouping
of the two 2017 variants. Styled with the app's existing CSS instead of the old
inline `<style>` block.
Note the links now point at `/thptqg/2017-old/`, not `/thptqg/2017/old/`.
`App.jsx` becomes: resolve route → hub, or dataset view. `useSqlite` is only
mounted on a dataset route, so the hub never touches the DB.
### Accepted trade-off
The hub was static HTML that painted instantly and worked without JS; it now
waits on the bundle (~150 KB gzipped). Accepted for the pipeline simplification.
If it ever matters, the four links can be inlined as static markup in
`index.html` so they paint pre-hydration.
### Column visibility — no new mechanism needed
`score-table.jsx` already computes `visibleColumns` by dropping columns where
every row is NULL. Extending that same filter to `ten_cum_thi` and `gioi_tinh`
makes them appear on 2016 and vanish on 2017 automatically, with no dataset
conditional in the component. This is the cheapest correct approach and it is
already the file's established pattern.
Identity columns need their own render path (they are text, not scored cells),
so add an `IDENTITY_COLUMNS` list alongside `SUBJECT_COLUMNS` and apply the same
all-NULL filter to both.
### Admission blocks — one union list
`computeBlocks()` already skips any block where a subject score is missing.
So a single list covering both years needs no branching: GDCD/KHTN blocks
self-exclude on 2016 rows, and the 2016-only foreign-language blocks
self-exclude on 2017 rows.
Add to `ADMISSION_BLOCKS`: `D05` (Toán+Văn+Đức), `D06` (Toán+Văn+Nhật), and the
other Đức/Nhật combinations from Circular 03/2017 that the current list drops
with the comment "neither language appears in any source file" — that comment
becomes false once 2016 data uses the same schema, so update it.
Verify the 2016 block list against the 2016 admission regulation rather than
assuming the 2017 circular's codes applied unchanged that year. If they differ
materially, key the block list by exam year via the registry's `blocks` field;
if they do not, drop that field and keep the single union list.
### Subject labels
`student-detail.jsx`'s `SUBJECT_LABELS`/`SUBJECT_ORDER` and `score-table.jsx`'s
`SUBJECT_COLUMNS` are two hand-maintained copies of the same subject list. Merge
into one `src/lib/subjects.js` exporting the 16-subject ordered list with labels;
both components consume it. This mirrors what Phase 1 does on the Rust side.
## Related Code Files
- Create: `src/datasets.js`
- Create: `src/router.js`
- Create: `src/components/hub.jsx`
- Create: `src/lib/subjects.js`
- Modify: `src/App.jsx` — resolve route; hub or dataset view; chrome from the registry entry
- Modify: `src/components/score-table.jsx` — add identity columns, use `subjects.js`
- Modify: `src/components/student-detail.jsx` — use `subjects.js`, show cluster/gender when present
- Modify: `src/components/search-form.jsx` — `EXAMPLES` from the active dataset
- Modify: `src/components/custom-query.jsx` — `PRESET_GROUPS` from the active dataset
- Modify: `src/lib/admission-blocks.js` — add Đức/Nhật blocks, update stale comment
- Modify: `src/App.css` — port `.cumthi-cell` from `2016/src/App.css`; add hub styles
- Reference (deleted in Phase 2): old root `index.html` for hub content,
`2016/src/components/custom-query.jsx` for the 2016 preset SQL
## Implementation Steps
1. Extract `src/lib/subjects.js`; repoint both components at it.
2. Add `IDENTITY_COLUMNS` + all-NULL filtering to `score-table.jsx`; port the
`.cumthi-cell` style.
3. Write `src/datasets.js` with all four entries. Lift the 2016 preset SQL
verbatim from the old 2016 `custom-query.jsx` (cluster averages, gender
breakdown, cluster+gender listing, language-count summary).
4. Note: the 2017 presets contain a hardcoded `so_bao_danh LIKE '49%'` Long An
query. Keep it only in the 2017 entries; write a 2016 equivalent or drop it.
5. Write `src/router.js`: exact segment match, plus the legacy redirect map.
6. Write `src/components/hub.jsx` from the captured static hub content.
7. Restructure `App.jsx`: resolve route → hub or dataset view; mount `useSqlite`
only on dataset routes; read chrome from the resolved entry.
8. Repoint `search-form.jsx` and `custom-query.jsx` at the active dataset.
9. Extend `ADMISSION_BLOCKS` with the Đức/Nhật blocks; verify against the 2016
regulation before committing the list.
10. Show cluster/gender in `student-detail.jsx` when non-null.
11. `npm run lint`.
## Tests / Validation
Manual, against a local preview of the single build (no test harness exists in
this repo today):
- `/thptqg/` renders the hub; network tab shows **no** `.db.gz` request
- Each of the four dataset routes loads its own DB — confirm `2017-old` fetches
`db/2017-old.db.gz`, not `db/2017.db.gz`
- `/thptqg/2017/old/?q=Nguyen` redirects to `/thptqg/2017-old/?q=Nguyen` with the
query intact, and lands on the right dataset
- Hub links point at flat URLs
- 2016 route: cluster + gender columns render; KHTN/KHXH/GDCD/Nga columns absent
- 2017 routes: no cluster/gender columns; KHTN/KHXH/GDCD present
- Single-result search opens `student-detail` on all four
- `?q=` deep link hydrates search on all four
- SQL tab presets execute without error on their own dataset
- Admission blocks: a 2016 student shows D05/D06 where applicable and no GDCD
blocks; a 2017 student shows GDCD blocks and no Đức/Nhật blocks
- `npm run lint` clean
## Success Criteria
- [ ] One `src/` serving four datasets **and** the hub, zero `if (dataset === ...)` in components
- [ ] Subject list defined once in `src/lib/subjects.js`
- [ ] Hub route fetches no database
- [ ] Dataset path and DB URL are derived from `id`, not stored per entry
- [ ] Legacy `/2017/old/` and `/2017/old2/` redirect to flat URLs, `?q=` preserved
- [ ] 2016 route renders cluster + gender; 2017 routes do not
- [ ] 2016 route has live search, deep links, student detail, score tiers
- [ ] Admission blocks correct for both exam years
- [ ] No routing library added
- [ ] `npm run lint` green
## Risk Assessment
| Risk | Mitigation |
|---|---|
| 2016 SQL presets lost when `2016/src` is deleted | Step 3 lifts them verbatim; recoverable from git history if ordering slips |
| 2017 block list assumed valid for 2016 | Step 9 requires checking the 2016 regulation; registry can key blocks per year if they differ |
| All-NULL filter hides a column on a legitimately sparse result set | Filter is per result set, matching today's 2017 behavior — accepted existing trade-off, not a new one |
| Hardcoded Long An preset leaks into 2016 | Explicit step 4 |
| Router sends `/2017-old/` to the `2017` dataset | Designed out: exact segment match on a flat scheme, no prefix logic exists |
| Existing links to nested URLs break | Legacy redirect map + Phase 4 stub pages; `?q=` preserved through the rewrite |
| Hub paints slower than the old static HTML | Accepted and documented; inline-links fallback available if it bites |
@@ -0,0 +1,185 @@
---
phase: 4
title: "Build and deploy pipeline"
status: pending
priority: P1
dependencies: [3]
effort: ""
---
# Phase 4: Build and deploy pipeline
## Overview
Rewire Vite, npm scripts, and GitHub Actions to produce the whole site from
**one** frontend build and **one** parser binary. Because the app now owns
routing (Phase 3), the four Vite build variants collapse into a single build
plus a copy step.
## Requirements
**Functional**
- One Vite build produces every page: hub + four dataset routes
- Flat published paths: `/thptqg/`, `/thptqg/2016/`, `/thptqg/2017/`, `/thptqg/2017-old/`, `/thptqg/2017-old2/`
- Legacy `/thptqg/2017/old/` and `/thptqg/2017/old2/` still resolve (redirect stubs)
- Deep links work on every route without a 404 fallback
- Uncompressed `.db` files never ship — only `.db.gz`
**Non-functional**
- One `cargo build`, one `npm ci`, one `vite build` per CI run
- npm only; no pnpm steps or caches remain
## Architecture
### Vite
One config. No variants, no `DATASET` env, no `emptyOutDir` ordering problem —
all three of those existed only to work around the missing router.
```js
export default defineConfig({
plugins: [react()],
base: "/thptqg/",
publicDir: ".build/public", // gitignored; holds db/*.db.gz only
});
```
### Why the entry-point copies work
With an absolute `base`, the emitted `index.html` references
`/thptqg/assets/index-HASH.js` no matter which directory it is served from. So
the same file is a valid entry point at every depth, and GitHub Pages serves
each as a directory index:
```bash
for ds in "${DATASETS[@]}"; do
mkdir -p "_site/$ds" && cp dist/index.html "_site/$ds/index.html"
done
```
With flat URLs the loop iterates the same `DATASETS` array used to build the
databases — no path translation between dataset ID and URL path, because they
are the same string.
This is what removes the need for the usual `404.html` SPA-fallback hack — which
matters concretely here, because that hack rewrites the URL and would interfere
with the existing `?q=` deep-link handling.
Also emit `dist/index.html` as `_site/404.html` so unknown paths render the hub
instead of Pages' default 404.
### Database staging
The parser writes into `.build/public/db/`, gzips in place, and the raw `.db` is
deleted before Vite copies `publicDir`. Today's pipeline instead ships the
uncompressed DB into `dist` and deletes it afterwards (`rm -f dist/*.db`) —
staging makes shipping a 47 MB uncompressed file structurally impossible rather
than dependent on a cleanup step running.
### Workflow
Current workflow compiles the same Rust crate twice and runs two `pnpm install`s.
Collapse to one of each, then loop the four datasets:
```yaml
- uses: actions/setup-node@v4
with:
node-version: '24'
cache: 'npm'
cache-dependency-path: package-lock.json
- name: Build databases
run: |
set -euo pipefail
cargo build --release --manifest-path parser/Cargo.toml
mkdir -p .build/public/db
for ds in 2016 2017 2017-old 2017-old2; do
./parser/target/release/xlsxread build \
--schema "parser/configs/$ds.toml" \
--input "data/$ds" \
--output ".build/public/db/$ds.db"
gzip -9 ".build/public/db/$ds.db" # no -k: raw file must not survive
done
- name: Build site
run: |
npm ci
npm run build
- name: Assemble
run: |
set -euo pipefail
mkdir -p _site && cp -r dist/* _site/
cp dist/index.html _site/404.html
# one entry point per dataset — ID and URL segment are the same string
for ds in 2016 2017 2017-old 2017-old2; do
mkdir -p "_site/$ds" && cp dist/index.html "_site/$ds/index.html"
done
# legacy nested URLs — router rewrites these to the flat form
for legacy in 2017/old 2017/old2; do
mkdir -p "_site/$legacy" && cp dist/index.html "_site/$legacy/index.html"
done
```
The dataset list appears in both steps. Define it once as a job-level env var
(`DATASETS: "2016 2017 2017-old 2017-old2"`) rather than repeating the literal.
Cache changes: `Swatinem/rust-cache` workspaces → `parser`; `setup-node` cache →
`npm` keyed on `package-lock.json`; the `pnpm/action-setup` step is deleted.
## Related Code Files
- Modify: `vite.config.js` — single build, `base: /thptqg/`, `.build/public` publicDir
- Modify: `package.json` — replace the six `build:db*` and three `build:*` variant
scripts with one `build:db` (takes a dataset argument) and one `build`
- Modify: `.github/workflows/deploy-pages.yml` — single toolchain setup, dataset
loop, npm caches, new assemble step
- Modify: `.gitignore` — add `.build/`, `parser/target/`, keep `dist/`
## Implementation Steps
1. Write the single-build `vite.config.js`.
2. Collapse `package.json` scripts. The current `build:old`/`build:old2` shell
out through `node -e` + `spawnSync` purely to set an env var — both delete
outright rather than getting converted.
3. Rewrite the workflow: one cargo build, one `npm ci`, one `npm run build`,
dataset loop, new assemble step.
4. Update the Rust and npm caches; delete the pnpm setup step.
5. Run the whole pipeline locally — `cargo` and `npm` are both available — and
inspect `_site/` before pushing.
6. Serve `_site/` locally and click through all five routes.
## Tests / Validation
- Local run produces `_site/index.html`, `_site/404.html`, and
`index.html` under `2016/`, `2017/`, `2017-old/`, `2017-old2/`,
plus legacy `2017/old/` and `2017/old2/`
- `find _site -name '*.db'` returns nothing (only `.db.gz` present)
- All emitted `index.html` files are byte-identical
- Asset URLs in them are absolute `/thptqg/assets/...`
- Serving `_site/` locally: `/thptqg/2017-old/?q=...` loads
`db/2017-old.db.gz` and hydrates the query with no redirect
- `/thptqg/2017/old/?q=...` rewrites to the flat URL with the query intact
- Workflow run on the branch deploys all five flat URLs plus the two legacy paths
## Success Criteria
- [ ] One `vite.config.js`, no build variants, no `DATASET` env
- [ ] Workflow compiles Rust once, installs Node deps once, builds the site once
- [ ] All five flat URLs live; deep links intact
- [ ] Both legacy nested URLs resolve rather than 404
- [ ] Dataset list written once in the workflow, not repeated per step
- [ ] No SPA 404-redirect hack in the repo
- [ ] No uncompressed DB anywhere in the artifact
- [ ] Generated DBs live in gitignored `.build/`, not in source directories
- [ ] Zero pnpm references in the workflow
## Risk Assessment
| Risk | Mitigation |
|---|---|
| 47 MB uncompressed DB ships | `gzip -9` without `-k` leaves no raw file; validation greps `_site` for `*.db` |
| Relative asset path sneaks in and breaks nested entry points | Absolute `base`; validation asserts `/thptqg/assets/` in all five copies |
| Nested route 404s on Pages | Real `index.html` at each path, verified against a local static server before push |
| Total artifact size | Unchanged — all four DBs already ship today; no new Pages size exposure |
| Stale cache keys silently rebuild everything | Cosmetic; verify first workflow run's timing |
@@ -0,0 +1,128 @@
---
phase: 5
title: "Parity verification and docs"
status: pending
priority: P1
dependencies: [4]
effort: ""
---
# Phase 5: Parity verification and docs
## Overview
The release gate. Prove no data was lost or invented by the schema unification,
then update the documentation that describes a two-project repo which no longer
exists.
## Requirements
**Functional**
- Every dataset's row count and per-column non-NULL count matches the Phase 1 baseline
- Newly-added columns read exactly zero on datasets that never had them
- All four sites work end-to-end against their rebuilt DBs
**Non-functional**
- The comparison is a committed, re-runnable script, not a one-off shell session
## Architecture
### Why counts, not checksums
The schema changed shape, so the DB files cannot be byte-identical and a whole-file
hash is meaningless. The meaningful invariants are:
1. **Row count** per dataset — unchanged
2. **Per-column non-NULL count** for every column that existed before — unchanged
3. **Per-column non-NULL count** for every column newly added to a dataset — zero
4. **Value-level spot check** — a stable sample of SBDs compared field by field
Invariant 3 is what catches the union-regex risk flagged in Phase 1: if the
16-subject regex map starts matching text in 2016 files that the 12-subject map
ignored, `khtn`/`khxh`/`gdcd`/`tieng_nga` will be non-zero on 2016 and the gate
fails loudly rather than silently corrupting the dataset.
### Script
`parser/scripts/verify-parity.js` — reads
`plans/reports/parser-parity-baseline.json`, opens each rebuilt DB, recomputes
the same statistics, and exits non-zero on any mismatch with a per-column diff
table.
Use **`node:sqlite`** (`DatabaseSync`), not `better-sqlite3` or `sql.js`.
Verified working on this repo's Node 24 with no flag and no dependency:
```js
const { DatabaseSync } = require("node:sqlite");
```
That choice matters beyond convenience — `better-sqlite3` is a native module
whose postinstall is exactly what the deleted `pnpm-workspace.yaml` `allowBuilds`
entry existed to permit. Using the built-in keeps the dependency count at zero
and leaves nothing for a future package-manager change to trip over.
For invariant 4, sample deterministically — e.g. every SBD ending in `0000` —
so reruns compare the same students.
## Related Code Files
- Create: `parser/scripts/verify-parity.js`
- Reference: `plans/reports/parser-parity-baseline.json` (from Phase 1)
- Create: `plans/reports/parser-parity-result.md` (the gate's output)
- Modify: `README.md` — new layout, new commands, four datasets
- Modify: `docs/system-architecture.md` — merged, single-project architecture
- Modify: `docs/deployment-guide.md` — merged, new workflow
- Modify: `docs/data-pipeline.md` — canonical schema, one parser, four configs
- Delete: `docs/codebase-summary.md`, `docs/project-overview-pdr.md` if they
describe only the old 2016 project — read before deciding
## Implementation Steps
1. Write `verify-parity.js`.
2. Run against all four rebuilt DBs; capture output to
`plans/reports/parser-parity-result.md`.
3. Investigate any mismatch before touching docs. A failure here means Phase 1's
schema or regex unification is wrong — fix it there, do not adjust the gate.
4. Manual pass on all four preview builds against the checklist in Phase 3.
5. Update `README.md`: layout tree, npm commands (`npm ci`, `npm run build`),
parser invocation, the four dataset descriptions and their flat URLs. Call
out the URL change from `/2017/old/` to `/2017-old/` and that the old paths
redirect.
6. Merge the duplicated docs. `deployment-guide.md` and `system-architecture.md`
exist in both old projects and describe different pipelines — read both
copies fully before writing the merged version.
7. Rewrite `docs/data-pipeline.md` around the canonical schema: the 22 columns,
which datasets populate which, and how `format_detection` selects the 2016
column-layout path.
8. Document the canonical schema itself in one place, referenced from the others.
## Tests / Validation
- `node parser/scripts/verify-parity.js` exits 0
- Row counts: 2016 ≈ 877,461 (per the current landing page), 2017 ≈ 861,000 —
confirm against the baseline, not against these approximations
- 2016 DB: `khtn`, `khxh`, `gdcd`, `tieng_nga` all read 0 non-NULL
- 2017 DBs: `ten_cum_thi`, `gioi_tinh`, `tieng_duc`, `tieng_nhat` all read 0 non-NULL
- All four dataset routes: search by SBD, search by name, deep link, SQL preset,
student detail; plus the hub route linking to all four
- No doc references `2016/tools`, `2017/src`, `public-old/`, `build:all`, `pnpm`,
`landing/`, or `VITE_DATASET`
## Success Criteria
- [ ] `verify-parity.js` committed and exiting 0 on all four datasets
- [ ] Parity result report committed under `plans/reports/`
- [ ] Zero unexpected non-NULL columns
- [ ] Manual checklist passes on the hub and all four dataset routes
- [ ] `verify-parity.js` has zero npm dependencies (`node:sqlite` only)
- [ ] `README.md` and `docs/` describe the actual repo, with npm commands
- [ ] No stale path, pnpm, or build-variant references anywhere outside `plans/`
## Risk Assessment
| Risk | Mitigation |
|---|---|
| Parity failure discovered late, after the data move | Phase 1 runs the same comparison in place before Phase 2 moves anything |
| Gate weakened to make it pass | Explicit step 3: a mismatch is a Phase 1 bug, not a gate-tuning problem |
| Docs merged by picking one copy and discarding the other | Step 6 requires reading both; they document different pipelines |
| Baseline missing because Phase 1 step 1 was skipped | Phase 1 treats it as blocking; without it this phase cannot run |
@@ -0,0 +1,184 @@
---
title: "Unify frontend, standardize SQL schema, restructure repo"
description: "One 2017-based frontend, one canonical 22-column student schema, one parser crate, four datasets under data/"
status: pending
priority: P2
branch: "main"
tags: [refactor, schema, frontend, parser]
blockedBy: []
blocks: []
created: "2026-08-13T03:01:04.073Z"
createdBy: "ck:plan"
source: skill
---
# Unify frontend, standardize SQL schema, restructure repo
## Overview
Today the repo holds two near-duplicate projects (`2016/`, `2017/`), each with its
own React frontend, its own copy of the same Rust parser crate, and its own SQL
schema. The 2017 frontend is a strict feature superset; the 2016 parser is a
strict code superset. Both duplications are drift hazards, not real divergence.
Collapse to one of each:
- **One canonical schema** — 22-column `student` table (6 identity + 16 subject),
defined once in `parser/src/schema.rs`. Absent columns bind NULL.
- **One parser crate** — `parser/`, four config files carrying only per-dataset
parse rules.
- **One frontend** — repo root, built from `2017/src/`, plus the two identity
columns 2016 renders today. The app owns every page GitHub Pages serves,
including the `/thptqg/` hub, from a **single Vite build**.
- **Four datasets** — `data/2016`, `data/2017`, `data/2017-old`, `data/2017-old2`.
Published URLs are **flat**, one segment per dataset:
```text
/thptqg/ hub
/thptqg/2016/
/thptqg/2017/
/thptqg/2017-old/ was /thptqg/2017/old/
/thptqg/2017-old2/ was /thptqg/2017/old2/
```
The two old-generation URLs change. That is deliberate: the flat form makes the
URL segment **identical to the dataset ID**, which is already the name of the
data directory, the config file, and the DB file. One identifier end to end:
```text
data/2017-old/ → parser/configs/2017-old.toml → db/2017-old.db.gz → /thptqg/2017-old/
```
Phase 3 keeps the two legacy paths working via redirect so existing links do not
break (see that phase; droppable if you don't care).
Package manager: **npm**. pnpm is dropped (see Phase 2).
## Target Layout
```text
/
├── index.html Vite entry — serves every route
├── vite.config.js ONE build, base: /thptqg/
├── package.json npm; package-lock.json committed
├── src/
│ ├── datasets.js runtime registry: title, source, examples, SQL presets
│ ├── router.js pathname → dataset (or hub)
│ ├── components/
│ │ └── hub.jsx the /thptqg/ landing route
│ ├── hooks/
│ └── lib/
├── data/
│ ├── 2016/ 2017/ 2017-old/ 2017-old2/
├── parser/ single Rust crate (from 2016/tools/xlsxread)
│ ├── src/schema.rs canonical DDL + INSERT + 16 subject regexes
│ ├── configs/*.toml per-dataset parse rules only
│ ├── scripts/ crawl-baotintuc.js, diff-datasets.js, check-duplicates.js,
│ │ verify-parity.js
│ └── tests/golden.rs
└── docs/ merged from 2016/docs + 2017/docs
```
## Published Artifact
One bundle, five entry points. Because `base` is absolute (`/thptqg/`), the
emitted `index.html` references `/thptqg/assets/index-HASH.js` regardless of the
directory it sits in — so copying it to each dataset path yields a real static
file at every URL. GitHub Pages serves them as directory indexes. **No SPA
404-fallback hack is required**, which matters because the existing `?q=`
deep-link behaviour would not survive one.
```text
_site/
├── index.html hub route
├── 404.html copy of index.html
├── assets/index-HASH.js ONE bundle, cached across all five pages
├── db/2016.db.gz 2017.db.gz 2017-old.db.gz 2017-old2.db.gz
├── 2016/index.html ┐
├── 2017/index.html │ byte-identical copies of the root index.html
├── 2017-old/index.html │
├── 2017-old2/index.html ┘
└── 2017/old/index.html ┐ legacy redirect stubs (same file again)
2017/old2/index.html ┘
```
## Canonical Schema
```sql
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
ngay_sinh TEXT,
ten_cum_thi TEXT, -- 2016 only; NULL elsewhere
gioi_tinh TEXT, -- 2016 only; NULL elsewhere
toan REAL, ngu_van REAL, vat_ly REAL, hoa_hoc REAL, sinh_hoc REAL,
khtn REAL, -- 2017 only
lich_su REAL, dia_ly REAL,
gdcd REAL, khxh REAL, -- 2017 only
tieng_anh REAL, tieng_phap REAL,
tieng_nga REAL, -- 2017 only
tieng_duc REAL, tieng_nhat REAL, -- 2016 only
tieng_trung REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
CREATE INDEX idx_ho_ten_ascii ON student(ho_ten_ascii);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
```
`idx_ten_cum_thi` is partial so it costs nothing on the three 2017 datasets
(zero entries) while staying fully useful for 2016's cluster grouping queries.
## Phases
| Phase | Name | Status |
|-------|------|--------|
| 1 | [Standard schema and unified parser](./phase-01-standard-schema-and-unified-parser.md) | Pending |
| 2 | [Repo restructure](./phase-02-repo-restructure.md) | Pending |
| 3 | [Unified frontend and dataset registry](./phase-03-unified-frontend-and-dataset-registry.md) | Pending |
| 4 | [Build and deploy pipeline](./phase-04-build-and-deploy-pipeline.md) | Pending |
| 5 | [Parity verification and docs](./phase-05-parity-verification-and-docs.md) | Pending |
## Dependencies
Strictly sequential. Phase 1 must produce parity-verified DBs under the old
layout before Phase 2 moves 419 MB of tracked data. Phase 5's parity gate is
the release gate — nothing merges until all four DBs match their pre-refactor
row counts and per-column non-NULL counts.
No cross-plan dependencies (`plans/` was empty before this plan).
## Acceptance Criteria
- [ ] One Rust crate; `2016/tools/` and `2017/tools/` gone
- [ ] One `src/`; `2016/src/` and `2017/src/` gone
- [ ] One Vite build producing all five pages
- [ ] All four DBs built from the same DDL, same INSERT, same 16 regexes
- [ ] Per-dataset row count and per-column non-NULL count identical to pre-refactor baseline
- [ ] All five published URLs functional under the flat scheme, deep links included
- [ ] Legacy `/2017/old/` and `/2017/old2/` redirect to their flat equivalents, preserving `?q=`
- [ ] 2016 site still shows `ten_cum_thi` + `gioi_tinh`; 2017 sites do not
- [ ] 2016 site gains 2017's features (deep links, student detail, tiers, live search)
- [ ] npm only: `package-lock.json` committed, no pnpm files or CI steps remain
- [ ] `cargo test` green; `npm run lint` green
## Rollback
Every phase is a separate commit on a feature branch. Phase 2's `git mv` is the
only hard-to-undo step; it is content-preserving, so `git revert` restores the
old layout exactly. Do not squash before the Phase 5 gate passes.
## Open Questions
None outstanding. Four decisions were taken before planning:
1. Frontend lives at repo root.
2. Schema is defined in parser code, not duplicated across the TOML configs.
3. The hub page is a route inside the app, not a separate static file — one
Vite build covers every page on GitHub Pages. Trade-off accepted: the hub
now needs the JS bundle (~150 KB gzipped) to paint, where today it is 25
lines of static HTML.
4. npm replaces pnpm.
5. URLs are flat (`/thptqg/2017-old/`), not nested (`/thptqg/2017/old/`), so the
URL segment equals the dataset ID everywhere. Legacy paths redirect.
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff