Commit Graph
60 Commits
Author SHA1 Message Date
tiennm99 00a08d5fab refactor(parser): remove the Rust crate now that Go is at parity
The Go parser has matched the Rust one field-by-field across all four datasets,
so the Rust crate is retired and CI builds the Go binary instead.

crawl-baotintuc.js moves to go-parser/scripts/ — it is the only mechanism for
refreshing data/2017 and has a documented runbook. check-duplicates.js and
diff-datasets.js are dropped: both had been broken since before the repo was
unified, and neither had a caller. db-stats.js and verify-parity.js are dropped
as superseded by differential-parity.mjs, which compares more and cannot
silently skip a dataset.

The reader-fidelity oracle is kept and marked frozen. It was produced by the
Rust reader so it can no longer be regenerated, but it still fails if any single
cell of any of the 299 input files reads differently.

Source comments cite the original Rust by file and line; those paths resolve at
tag pre-go-parser-removal, recorded in go-parser/README.md.
2026-08-13 20:38:25 +07:00
tiennm99 0eb174721d refactor(parser): reimplement the parser in Go alongside the Rust crate
Adds go-parser/, a Go reimplementation of the xlsxread parser, verified
byte-for-byte against the Rust original before any cutover.

Reader fidelity is exact across all 299 input files: the canonical cell dump
of every sheet matches calamine's, locked in as a test against a committed
hash oracle. Reaching that required replacing extrame/xls, which corrupted
69% of cells and dropped a further 28% on the BIFF corpus, with pbnjay/grate;
correcting excelize's number-format application and trailing-cell trimming;
restoring carriage returns that XML line-ending normalisation strips from
2,233 ten_cum_thi values; and gating numeric re-rendering on cell type so
shared strings that merely look numeric keep their leading zeros.

The differential gate compares both parsers over all four datasets:
3,265,641 rows with identical full-table SHA-256, identical per-column
non-NULL counts, identical schema metadata and identical stdout.

Config moves from TOML to YAML for both parsers, so they keep reading the
same files and the gate stays meaningful. Verified by rebuilding 2016 and
2017-old2 with Rust under the new configs and matching the recorded counts.

build-db.js now refuses to publish a database whose row count does not match
the known figure, closing a path where an under-producing parser could ship a
truncated public dataset with green CI. The deploy workflow gains a
pull_request trigger and guards deploy to main, so branch verification can no
longer publish to production.
2026-08-13 20:27:44 +07:00
tiennm99 8902232747 Merge pull request #5 from tiennm99/refactor/unify-frontend-and-schema
refactor: one frontend, one parser, one schema
2026-08-13 14:38:55 +07:00
tiennm99 3fcfe4862d docs: rewrite for the unified repo and record the parity gate
The docs still described two standalone projects with separate frontends,
separate parsers and three Vite variants. Rewritten around what the repo now
is, merging both projects' copies rather than keeping one and discarding the
other — deployment-guide.md and system-architecture.md existed in both and
documented different pipelines.

  project-overview.md    goal, scope, constraints, the four datasets, history
  system-architecture.md data flow, canonical schema, routing, how one
                         frontend serves both exam years without branching
  data-pipeline.md       per-dataset Excel formats, the three 2016 layouts,
                         overflow-sheet gotcha, expected row counts
  deployment-guide.md    the single-build workflow, adding a dataset,
                         why no uncompressed database can ship

Records the release gate in plans/reports/parser-parity-result.md: row counts,
all 18 pre-existing per-column non-NULL counts and the deterministic student
samples are identical across all four datasets, against databases decompressed
from the exact bytes the pipeline publishes. The 1,691 recovered
foreign-language scores are documented with the evidence they are real.

Also drops a machine-specific absolute path from a comment in
format_detect_2016.rs. The build-database.js citations there are kept: that
file no longer exists in this repo, but the references explain why several
parsing rules look arbitrary.
2026-08-13 13:25:46 +07:00
tiennm99 a5904a5d78 build: produce the whole site from one Vite build
The four build variants existed only because the app could not resolve its own
dataset. Now that it can, vite.config.js is a single build with an absolute
base and no VARIANT switching, and the workflow compiles Rust once, installs
Node dependencies once, and builds the site once — it previously built the same
Rust crate twice and ran two separate pnpm installs.

Because base is absolute, the emitted index.html references /thptqg/assets/...
regardless of where it is served from, so the same file works as an entry point
at any depth. scripts/assemble-site.js copies it to each dataset path and to
the two legacy nested URLs, giving a real static file at every published route.
That is what removes the need for an SPA 404-fallback redirect, which would
otherwise have rewritten URLs and interfered with the ?q= deep links.

Generated databases move to a gitignored .build/public, which Vite consumes as
its publicDir. build-db.js gzips without -k, and the assemble step then refuses
to finish if any uncompressed database artefact reached the output — .db,
.db-journal, .db-wal or .db-shm. Previously a raw 100+ MB database was written
into the source tree and deleted afterwards by an rm in the workflow, so
shipping one was a missing cleanup step away.

Both build scripts import DATASET_IDS from src/datasets.js rather than
repeating the dataset list in workflow shell, so the four ids are declared in
exactly one place across the frontend, the database build and the assembly.

Verified by running the full pipeline locally and serving the artifact over
HTTP: all nine routes return 200, every entry point is byte-identical, asset
references are absolute, and the databases are fetchable. The guard was tested
by injecting a raw .db and .db-journal into the build, which failed the
assembly as intended.
2026-08-13 11:57:26 +07:00
tiennm99 003e7c8afd style(frontend): document the four one-shot setState-in-effect sites
eslint-plugin-react-hooks v7 flags setState inside an effect body. All four
occurrences predate this refactor and are one-shot initialisation or external
sync, not the cascading-render pattern the rule targets:

  - hydrating a ?q= deep link once the database is ready
  - reading the candidate count for the footer
  - auto-running the schema preset when the SQL tab first opens
  - mirroring the parent-owned query into SearchForm for deep links and clear

Each gets a scoped disable with a comment explaining why it is correct here,
rather than a blanket rule change. Restructuring these properly would mean
reworking deep-link hydration and auto-search with no test harness to catch a
regression; that is worth doing separately, not inside this refactor.

npm run lint is now clean.
2026-08-13 11:48:31 +07:00
tiennm99 83bbc597ce feat(frontend): serve all four datasets and the hub from one app
The 2017 frontend becomes the only frontend. It now resolves which dataset to
show from the URL, so the four dataset views and the /thptqg/ landing page all
come from a single build instead of four builds plus a hand-written static
index.html.

Routing is an exact match on one path segment, because the segment is the
dataset id: /thptqg/2017-old/ -> "2017-old". The nested form this replaces
(/thptqg/2017/old/) would have needed longest-prefix matching, since it also
starts with /thptqg/2017/. Previously published nested URLs are rewritten to
their flat equivalent with the query string intact, so ?q= deep links keep
working.

Restores what the 2016 site had and the 2017 site lacked:

  - Cụm thi and Giới tính columns, in the table and on the detail card
  - exam-ID lookup for letter-prefixed số báo danh

That second one would have been a severe regression. 616,593 of 877,461
candidates in the 2016 dataset — 70.3% — have an SBD like BAL000001, and the
2017 app matched /^\d+$/ only, so exam-ID search would have silently failed for
most of that year. Detection now lives in lib/query-mode.js, shared by App and
SearchForm, which had already drifted apart on exactly this rule. Letter
prefixes are upper-cased before lookup so "bal000001" resolves.

Neither dataset needs a conditional in a component. Columns where every row is
NULL are hidden, so 2016 rows surface Cụm thi / Đức / Nhật while 2017 rows
surface KHTN / KHXH / GDCD / Nga. computeBlocks() likewise skips blocks with a
missing subject, so one admission-block list covers both years — D05 and D06
are added now that German and Japanese scores exist.

Per-dataset content moves into the registry: title, source, database size,
search examples and SQL presets. The 2016 presets are lifted from the deleted
2016 component; the Long An (49xxx) queries stay on 2017 only, where that SBD
prefix means something.

The subject list, previously maintained separately in score-table.jsx and
student-detail.jsx, is now declared once in lib/subjects.js mirroring
SCORE_FIELDS in the parser.
2026-08-13 11:37:59 +07:00
tiennm99 6ff2ed99ec refactor: collapse the two projects into one tree and move to npm
The repo held two near-duplicate projects. 2016/ and 2017/ each carried their
own React frontend, their own copy of the same Rust crate, and their own
package manager setup. 2016/tools/sync-from-thptqg2017.sh existed purely to
copy the parser source between them.

New layout:

  index.html + src/  the 2017 frontend, now the only one
  data/<id>/         2016, 2017, 2017-old, 2017-old2
  parser/            the single Rust crate, configs renamed to <id>.toml
  docs/              both projects' docs, 2016 copies suffixed -2016-legacy
                     pending the merge pass

<id> is now one identifier end to end: data/<id>/ feeds parser/configs/<id>.toml
and produces db/<id>.db.gz.

pnpm gives way to npm. pnpm-workspace.yaml existed only to whitelist
better-sqlite3's native build, which npm permits by default, so it has no
equivalent and is simply gone. Lockfiles cannot be converted; package-lock.json
is generated fresh. The migration direction is safe — pnpm's strict layout
forbids phantom dependencies, so anything that resolved under pnpm resolves
under npm's flat tree.

Adds parser/scripts/build-db.js and src/datasets.js: the four dataset IDs are
declared once and read by both the build tooling and (from the next phase) the
frontend.

Follow-on fixes the move made necessary:
  - eslint's Node-globals override pointed at scripts/, now parser/scripts/
  - crawl-baotintuc.js wrote to <root>/data, now data/2017
  - golden tests loaded configs by their old thptqg*-data.toml names

Drops the #[ignore]d Rust-vs-Node golden test. It shelled out to pnpm to run
scripts/build-database.js, a file removed when the parser was ported to Rust,
so it could never pass. check-duplicates.js and diff-datasets.js were already
broken before this change and are annotated as such rather than half-fixed.

63 Rust tests pass and clippy is clean from the new location.
2026-08-13 11:27:01 +07:00
tiennm99 82db8295bb style(parser): clear pre-existing clippy warnings
cargo clippy --all-targets -- -D warnings failed with 8 errors before the schema
refactor and was never part of CI. Clearing them so the gate is meaningful from
here on.

All behaviour-preserving: slice contains() over iter().any(), is_multiple_of()
over a modulo, a needless borrow in a test helper, and a doc-list indent.
process_mapped_row keeps its 8 parameters behind an allow attribute — the
signature mirrors the JS column map one-for-one, and bundling the indices into
a struct would hide that correspondence.

Kept separate from the schema change so that diff stays readable.
2026-08-13 11:16:51 +07:00
tiennm99 cd4b07f9cb refactor(parser): define one canonical student schema for all datasets
The 2016 and 2017 parsers were the same crate with divergent SQL. DDL, INSERT
and the subject regex table lived in each dataset's TOML config, so four copies
had to be kept in step by hand — which is how the two schemas drifted apart.

Move all of it into src/schema.rs as a single 22-column definition: 6 identity
columns (adding ten_cum_thi and gioi_tinh, previously 2016-only) and 16 subject
columns (the union of both exam years). Columns a dataset carries no data for
bind NULL, costing ~1 byte per row.

This collapses writer.rs to one insert path: insert_row_2016, SCORE_FIELDS_2016,
SCORE_FIELDS_2017 and the SCORE_FIELDS alias all go away. Configs shrink to the
parse rules that genuinely vary per dataset — sheet mode, column indices, SBD
validation, header tokens, blank-row stripping — and carry no SQL at all.

Config parsing now denies unknown fields, so a leftover [schema] block fails
loudly instead of looking effective while schema.rs drives the build.

idx_ten_cum_thi is partial, so it holds zero entries on the three 2017 datasets
where the column is always NULL.

Adds db-stats.js and verify-parity.js to prove no data moved: row counts,
per-column non-NULL counts and a deterministic student sample are compared
against databases built from the previous code.

Verified across all four datasets — row counts, every pre-existing column count,
and all sampled students are identical. Database size grows 1.4-2.2%.
2026-08-13 11:16:24 +07:00
tiennm99 944ee14b36 chore: merge the 2016 and 2017 lookups into one repository and one Pages site
Each year keeps its full pipeline under its directory; vite bases move to
/thptqg/<year>/ (2017 keeps its old and old2 generations), one deploy
workflow builds both years' databases and bundles behind a root index.
2026-08-13 09:36:56 +07:00
tiennm99 8df6f2d46b chore: remove dependabot version-update config 2026-07-25 14:13:59 +07:00
tiennm99 447796aacf chore: remove dependabot version-update config 2026-07-25 14:13:37 +07:00
tiennm99 f5d56b4273 chore: add dependabot config (#2) 2026-05-23 11:43:29 +07:00
tiennm99 79a8a73cde chore: add dependabot config (#2) 2026-05-23 11:43:25 +07:00
tiennm99 6b21c625ac feat: replace xlsx (SheetJS) build pipeline with Rust xlsxread CLI (#1)
* feat(xlsxread): vendor Rust binary cloned from thptqg2017

Copies the xlsxread Rust crate from thptqg2017@8b4a755 (chore/xlsxread-rust).
Adds format_detect_2016 module with per-file column-layout auto-detection
mirroring detectFormat() in scripts/build-database.js (lines 63-87):
  - separate-scores: SBD/HOTEN/TOAN... fixed columns (dhhanghai files)
  - mapped: header-derived SOBAODANH|SBD + DIEM_THI dynamic indices
  - default: positional 6-col layout (no header)

Extends ParsedRow with ten_cum_thi and gioi_tinh fields.
Adds SCORE_FIELDS_2016 (12 cols: tieng_duc/tieng_nhat; no khtn/khxh/tieng_nga).
Adds thptqg2016-data.toml config with 18-column schema and format_detection flag.
58 tests pass (50 unit + 8 integration), 0 failures.

* feat(build): wire build:db to xlsxread CLI

Replaces the Node.js build:db script with the xlsxread Rust binary.
Adds build:rust script for the cargo compile step in isolation.

* chore: remove deprecated build-database.js

Superseded by the xlsxread Rust binary. All 119 source files (4 .xls +
115 .xlsx) are now processed by xlsxread with per-file format detection.

* ci: build xlsxread before running database build job

Adds dtolnay/rust-toolchain@stable and Swatinem/rust-cache@v2 steps
before the xlsxread build and database generation steps. Node/pnpm
steps now follow the Rust build rather than preceding it.

* chore: add Rust build artifacts to .gitignore

* docs: update README build instructions for xlsxread pipeline

* chore(deps): drop xlsx and better-sqlite3 from package.json and lockfile
2026-05-19 16:33:20 +07:00
tiennm99 99465fda59 feat: replace xlsx (SheetJS) build pipeline with Rust xlsxread CLI (#1)
* feat(xlsxread): stage 0 scaffold with pinned deps and clap CLI skeleton

* feat(xlsxread): stage 1 reader — calamine sheet enumeration and header skip

* feat(xlsxread): stage 2 transform — to_ascii, score regex, validation with 38 unit tests

* feat(xlsxread): stage 3 writer — SQLite DDL, INSERT OR REPLACE, VACUUM, stats output

* feat(xlsxread): stage 4 audit — distinct SBD scan vs DB count, mirrors audit-row-counts.js output

* feat(xlsxread): stage 5 golden tests — in-process xlsx fixtures, 8 integration tests pass

* chore(xlsxread): commit Cargo.lock for reproducible Rust builds

* feat(build): wire root build:db scripts to xlsxread CLI

Replace node scripts/build-database*.js invocations with the Rust
xlsxread binary. Each build:db* script now calls `pnpm build:rust`
(cargo build --release) before invoking the xlsxread build subcommand
with the matching per-dataset config.

Drop xlsx and better-sqlite3 from devDependencies — no Node script
consumes them anymore. sql.js (runtime DB reader in the SPA) is
unaffected and remains in dependencies.

* ci: build xlsxread before running database build jobs

Add dtolnay/rust-toolchain@stable and Swatinem/rust-cache@v2
(workspaces: tools/xlsxread) for warm incremental Rust builds.

Replace the single `pnpm build:db:all` step with explicit xlsxread
invocations so CI doesn't call pnpm build:rust redundantly three times.
The binary is built once, then each of the three datasets is processed
in sequence.

* chore: remove deprecated xlsx-based build scripts

Delete scripts/build-database.js, build-database-old.js,
build-database-old2.js, build-lib.js, and audit-row-counts.js.
Functionality replaced by the xlsxread Rust CLI configured via
tools/xlsxread/configs/*.toml. History preserved in git; one-click
revert available via the chore/migration-backup-260519 branch.

* docs: update README build instructions for xlsxread pipeline

Replace Node.js + xlsx references with Rust + xlsxread workflow.
Update requirements (Node 24+, pnpm, Rust stable), quickstart, scripts
table, and project layout tree to reflect the current state after the
xlsx-based build scripts were removed.

* chore(deps): drop xlsx and better-sqlite3 from package.json and lockfile

Remove xlsx (SheetJS, vulnerable: GHSA-4r6h-8v6p-xvw6, GHSA-5pgg-2g8v-p4x9)
and better-sqlite3 from devDependencies. Both were only used by the now-deleted
Node build scripts. The Rust xlsxread CLI vendors SQLite via rusqlite-bundled;
no Node-side SQLite dependency is needed. `pnpm audit` returns clean.
2026-05-19 16:33:16 +07:00
tiennm99 5b0cec3b27 chore(ci): bump node to 24 2026-05-13 10:52:42 +07:00
tiennm99 b5bd88bfbb chore(ci): bump node to 24 2026-05-13 10:52:39 +07:00
tiennm99 12c41c747e fix(ci): approve better-sqlite3 build via pnpm-workspace.yaml allowBuilds 2026-05-13 10:39:20 +07:00
tiennm99 a132847cda fix(ci): approve better-sqlite3 build via pnpm-workspace.yaml allowBuilds 2026-05-13 10:39:17 +07:00
tiennm99 b3b71f577c fix(ci): bump Node.js to 22 required by pnpm@11.1.1 2026-05-13 10:34:03 +07:00
tiennm99 75201ab06a fix(ci): bump Node.js to 22 required by pnpm@11.1.1 2026-05-13 10:33:59 +07:00
tiennm99 37cee58171 fix(ci): remove pnpm version override conflicting with packageManager 2026-05-13 10:29:03 +07:00
tiennm99 8732ee138d fix(ci): remove pnpm version override conflicting with packageManager 2026-05-13 10:28:59 +07:00
tiennm99 eb7409a503 chore: migrate from npm to pnpm
Replace package-lock.json with pnpm-lock.yaml, add packageManager field,
and update GitHub Actions workflow to use pnpm/action-setup@v4 with
frozen-lockfile installs.
2026-05-13 10:20:34 +07:00
tiennm99 08e55d98b1 chore: migrate from npm to pnpm
Replace package-lock.json with pnpm-lock.yaml, add packageManager field,
update package.json scripts from npm run to pnpm, and update GitHub Actions
workflow to use pnpm/action-setup@v4 with frozen-lockfile installs.
2026-05-13 10:20:29 +07:00
tiennm99 5bdcc01f60 docs: note original data source URL (now inaccessible) 2026-04-14 23:33:15 +07:00
tiennm99 5ffc26f177 docs: note data/ source URL; mark data-old/ and data-old2/ origin unrecorded 2026-04-14 23:30:12 +07:00
tiennm99 2c7ded1f39 docs: expand README and add project documentation 2026-04-14 23:28:56 +07:00
tiennm99 76ea795910 docs: translate narrative prose to English in README and docs/ 2026-04-14 23:28:55 +07:00
tiennm99 8ca6bfa7a8 docs: expand README, add architecture / deployment / data-pipeline docs
- README: live-site URLs, features, scripts, project layout, quickstart
- docs/system-architecture.md: data flow, 3-variant deploy mechanism,
  schema, parse-quirk matrix, score-tier model, admission-block model,
  frontend key behaviors (deep-link, share, keyboard shortcuts)
- docs/data-pipeline.md: source Excel shape, score regex, why three
  parsers, overflow-sheet gotcha, audit tooling, refresh flow
- docs/deployment-guide.md: CI workflow, adding a new variant, rollback,
  GH Pages file-size note
- docs/README.md: index for the docs dir
2026-04-14 23:25:51 +07:00
tiennm99 1940e8936d tweak: Long An best-block preset → top 10 instead of top 100 2026-04-14 23:19:36 +07:00
tiennm99 c4f043ab7a feat: full 49-block admission catalog + winning-block label in top-100 query
- admission-blocks.js: expand from 8 core blocks to all 49 official 2017
  blocks (A00-A11, B00-B08, C00-C20, D01-D15) the DB schema can support.
  Detail card now shows every qualifying block, sorted best-first
- custom-query: rewrite "Top 100 Long An" preset using UNION ALL +
  ROW_NUMBER() OVER PARTITION BY so we can label each row with the
  winning block code, not just the score
2026-04-14 23:17:58 +07:00
tiennm99 a5213c8667 feat: add 'Top 100 điểm khối cao nhất - Long An' SQL preset
Computes each student's max admission-block sum across 49 official 2017
blocks (A00-A11, B00-B08, C00-C20, D01-D15, restricted to subjects
present in our DB — no Đức/Nhật languages). One row per student, sorted
by their personal best, top 100.

SQL note: SQLite's MAX(x,y,...) scalar returns NULL if ANY arg is NULL.
Each block expression is wrapped in COALESCE(..., -1); NULLIF(..., -1)
restores NULL only for students with zero computable blocks.
2026-04-14 23:11:01 +07:00
tiennm99 d1410a4729 style: TFT rarity-ladder score tiers (6 levels, unique hues)
Old palette reused gold for both "Trung bình" and "Xuất sắc" — confusing.
Replaces with League of Legends / TFT rarity ladder where each rank gets
a distinct hue that also communicates relative rarity.

Tiers:
  ≤ 1    Điểm liệt   (common)    white / gray
  < 5    Chưa đạt    (uncommon)  green
  5-6.5  Trung bình  (rare)      blue
  6.5-8  Khá         (epic)      purple
  8-9    Giỏi        (legendary) gold
  9-10   Xuất sắc    (prismatic) multi-color gradient

- admission-blocks.js: scoreTier now returns 6 keys (common..prismatic)
  plus the "điểm liệt" tier (≤1) that didn't exist before
- student-detail.jsx: TIER_LEGEND updated to 6 entries with ranges
- App.css: --tier-{common,uncommon,rare,epic,legendary,prismatic}-*
  tokens, light + dark variants each AA-contrast verified; applied to
  .score-cell, .score-tile, .tier-legend-item
2026-04-14 23:03:08 +07:00
tiennm99 1d84148bc5 tweak: reduce Long An khối A preset from top 100 to top 10 2026-04-14 22:51:48 +07:00
tiennm99 f818003c3e feat: deep-link, share, score legend, web font, grouped SQL presets
- App.jsx owns query state, syncs to URL via ?q= (works for SBD or name);
  hydrates initial search from URL on DB ready
- Loading panel now shows DB size (~47 MB) and one-time download note
- Footer reports total student count from the live DB
- Keyboard shortcut '/' focuses the search input when not already typing
- StudentDetail: "Chia sẻ" button uses Web Share API or clipboard with a
  formatted multi-line summary including a deep-link URL
- StudentDetail: inline tier legend (▽ ○ ◆ ★ ✦) with score ranges
- ScoreTable cells now use background tint matching detail tiles
- CustomQuery: presets grouped into 4 categories; auto-runs PRAGMA
  table_info(student) on first DB ready so the tab opens with the schema
  visible instead of a blank textarea
- SearchForm accepts controlled value prop for URL hydration
- Add Be Vietnam Pro web font for consistent Vietnamese diacritic rendering
2026-04-14 22:44:00 +07:00
tiennm99 8cfc9116fb feat: Long An SQL presets, clickable search examples
- custom-query: two new presets filtered by Long An (SBD prefix 49):
  Top 10 điểm Toán, Top 100 khối A (Toán+Lý+Hóa)
- search-form: show clickable example pills ("49008235" /
  "Nguyễn Minh Tiến") next to the empty-state hint; clicking fills
  the input and triggers the debounced live search
- App.css: .example-btn pill styling matching preset buttons
2026-04-14 22:21:54 +07:00
tiennm99 8a43f7214f chore: remove completed static-score-lookup plan (fully implemented) 2026-04-14 22:13:20 +07:00
tiennm99 f82542f764 feat: build 3 parallel site variants, one per dataset
Sites:
  /thptqg2017/      → data/      861,068 students (baotintuc.vn .xls)
  /thptqg2017/old/  → data-old/  847,348 students (original xlsx)
  /thptqg2017/old2/ → data-old2/ 679,764 students (update/ overrides)

Build pipeline:
- scripts/build-lib.js: shared schema, score regex, toAscii, insert
  builder — keeps the 3 DBs drop-in compatible for the same frontend
- scripts/build-database.js: multi-sheet walk for baotintuc .xls (HN/HCM
  overflow past the 65k-per-sheet .xls cap)
- scripts/build-database-old.js: single-sheet xlsx + strict numeric SBD
  guard (rejects the header-row leak present in data-old)
- scripts/build-database-old2.js: multi-sheet + blank-row pre-filter for
  the 54 update files (HCM overflow present, HN not in this subset)
- Each build prints a source-vs-DB audit; verified 0 parse errors and
  exact row-count match across all 3 datasets

Frontend plumbing:
- vite.config.js: VARIANT env selects base path, publicDir, outDir
- package.json: build:db:*, build:db:all, build:old, build:old2,
  build:all
- .github/workflows/deploy.yml: CI builds all 3 DBs + sites
- .gitignore: cover public-old/ and public-old2/ outputs
- eslint.config.js: add Node globals for config/script files
2026-04-14 22:12:25 +07:00
tiennm99 4dd376b921 chore: preserve prior update/ overrides in data-old2/
Restore the 54 corrected-export files that previously lived in
data/raw/update/ (removed in 37cf9df). Kept alongside data-old/ for
historical reference; not consumed by build pipeline.
2026-04-14 21:46:42 +07:00
tiennm99 3ad0104565 feat: refresh data from baotintuc.vn source, fix overflow sheet loss
Dataset update:
- Crawl all 63 .xls province files from baotintuc.vn CDN (original source)
- Old xlsx dataset moved to data-old/ for reference
- Net: +13,719 students (Hà Nội +7,275, HCM +6,445) — the old .xls → xlsx
  conversion silently dropped rows beyond the 65,536 per-sheet cap
- Also removes 1 bogus header row that had leaked into the old DB
- 100% identical scores on the 847,348 SBDs present in both datasets

Build pipeline:
- build-database.js: iterate ALL sheets per workbook (fixes the overflow
  loss) and accept .xls in addition to .xlsx

Audit tooling:
- scripts/crawl-baotintuc.js: idempotent 63-province downloader
- scripts/diff-datasets.js: compares two DBs by SBD set and per-column
  score deltas
2026-04-14 21:42:29 +07:00
tiennm99 c832810b78 feat: live search, student detail card, score tier color coding
- SearchForm: 300ms debounced live search, visible label, aria-live mode
  hint, inline clear (×) button, minimum-length guards (3 digits / 2 chars)
- StudentDetail: rich single-result card with copy-SBD action, subject
  tiles with 5-tier color+icon coding, admission-block totals (A/A1/B/C/
  D1-D4) computed only when the student has all 3 required subject scores
- ScoreTable: per-cell tier color coding with tooltip
- App.jsx: switches to detail card when exactly 1 result; onClear wired
- App.css: semantic token system with light + dark (prefers-color-scheme)
  variants, :focus-visible ring, reduced-motion fallback, desktop-first
  breakpoints; removes hardcoded hex across components
- admission-blocks.js: pure data + calculator, no UI coupling
2026-04-14 21:21:46 +07:00
tiennm99 d968459b9d refactor: rename assets/ to data/ for raw Excel inputs 2026-04-14 21:20:10 +07:00
tiennm99 e165e3c5b6 refactor: flatten data layout to data/, drop update/ overrides
- Move 63 Excel files from data/raw/ to data/ (single flat dir)
- Remove all 53 files in data/raw/update/: verified identical SBD
  coverage to raw/ (847349 rows either way), so they added no new
  students — only potential score corrections that can be reintroduced
  later if source is recovered
- Update build-database.js to read data/ directly
- Add scripts/audit-row-counts.js: compares source row count to DB row
  count to verify zero-loss parsing
- Point check-duplicates.js at new data/ location
2026-04-14 21:02:47 +07:00
tiennm99 c1cd39e535 feat(ui): refresh styling with design tokens, dark mode, and a11y polish
- Add CSS variables and prefers-color-scheme dark theme
- Color-code score cells (high/low/empty) and add zebra rows
- Add ARIA tablist, search clear button, and focus halos
- Card-style loading, results table, and message blocks
2026-04-14 20:50:57 +07:00
tiennm99 641d0ed05a chore: remove duplicate Excel files, add md5 audit script
- Drop 10_LamDong_GNFT (1) and 2.BacKan_YQNX(1): identical row content to
  siblings (Excel metadata differs but file size & sheet rows match)
- Add scripts/check-duplicates.js to detect byte-identical and row-identical
  files across data/raw and data/raw/update
2026-04-14 20:49:41 +07:00
tiennm99 97519b4b78 feat: diacritics-insensitive search and SQL query tab
- Add ho_ten_ascii column + index for accent-folded name search
- Parse Tiếng Pháp/Nga/Trung scores (recovers ~2k students' foreign-language data)
- Loosen score regex to accept integer values (e.g. KHTN: 4)
- App.jsx: 3-mode search (SBD / ASCII / Vietnamese) and Tra cứu/SQL tabs
- New custom-query component: 8 presets adapted to 2017 schema, read-only safety, auto-LIMIT
- ScoreTable: render new foreign-lang columns, hide columns null across all results
2026-04-14 20:44:43 +07:00
tiennm99 f885aeb6cc feat: add diacritics-insensitive name search
Add ho_ten_ascii column with normalized names (no diacritics, lowercase)
so users can search "nguyen van a" to find "NGUYỄN VĂN A".

- ASCII input searches against ho_ten_ascii column
- Vietnamese input searches both ho_ten and ho_ten_ascii
- Indexed for fast lookups
2026-04-14 20:15:20 +07:00
tiennm99 4aa1f5bd2f update README with features, demo link, and dev instructions 2026-04-14 20:07:23 +07:00
tiennm99 372f1e04fb feat: add GitHub Pages app with custom SQL query support
- React + Vite + sql.js frontend with two-tab UI:
  - Quick search by exam ID (alphanumeric) or student name
  - Custom SQL query editor with 7 preset queries
- Build script handles all 5 Excel data formats (varied column
  orders, separate score columns, no-header, .xls/.xlsx)
- Database: 877,461 students with exam center, gender, 12 subjects
  (including 4 foreign languages: French, German, Japanese, Chinese)
- GitHub Actions CI/CD: build DB, gzip compress, deploy to Pages
- Safety: read-only queries only, auto LIMIT 1000, Ctrl+Enter shortcut
2026-04-14 20:05:38 +07:00
tiennm99 d81ed19a2c feat: add exam result data files for THPT QG 2016
Add 115 Excel files containing exam results from various provinces
and universities across Vietnam for the 2016 national high school exam.
2026-04-14 19:06:11 +07:00
tiennm99 2ff1f8b45d Initial commit 2026-04-14 19:02:17 +07:00
tiennm99 b9e16f63f8 refactor: remove Java code, move web app to project root
- Remove Gradle build, Java sources, Hibernate config, old database.sqlite
- Move Excel data files from src/main/resources/raw/ to data/raw/
- Move Vite+React app from web/ to project root
- Merge package.json into single root-level config
- Update build script paths and CI workflow accordingly
2026-04-13 00:06:22 +07:00
tiennm99 00148c2748 feat: add implementation plans and Vite scaffold files 2026-04-13 00:00:43 +07:00
tiennm99 2bed92547b feat: add static score lookup site with Node.js DB builder
- Node script parses 119 Excel files into SQLite (847K students)
- Vite + React frontend with sql.js for client-side querying
- Search by exam ID (số báo danh) or student name
- Gzipped DB (36MB) with download progress bar
- GitHub Actions workflow for GitHub Pages deployment
2026-04-12 23:54:06 +07:00
tiennm99 cee9df9f1a [Add] logics to insert from excel files into sqlite database 2023-11-13 23:22:59 +07:00
tiennm99 2d2850f20f [Add] lib to create database 2023-10-30 07:35:55 +07:00
tiennm99 17e238a76b [Add] init with idea & add Excel sources 2023-10-29 06:55:42 +07:00