chore: remove completed static-score-lookup plan (fully implemented)

This commit is contained in:
tiennm99 committed 2026-04-14 22:13:20 +07:00
1 parent f82542f764
commit 8a43f7214f
5 files changed
-846

No files matched your search

@@ -1,117 +0,0 @@
# Phase 01 — Project Scaffolding
## Context Links
- [plan.md](./plan.md)
- [Converter.java](../../src/main/java/dev/miti99/thptqg2017/Converter.java) — reference parsing logic
- [Student.java](../../src/main/java/dev/miti99/thptqg2017/entity/Student.java) — reference data model
## Overview
- **Priority**: P1 (blocker for all other phases)
- **Status**: Pending
- **Description**: Initialize Node.js project, Vite+React app, and folder structure
## Key Decisions
- **Monorepo with single package.json** at root — project is small, no need for workspaces
- **better-sqlite3** for build script (fast native writes), **sql.js** for browser runtime
- **Vite + React + TypeScript** for the static site
- Keep Java source intact — no deletions in this phase
## Architecture
```
thptqg2017/
├── scripts/ # Node.js build scripts
│ └── build-database.js # Excel -> SQLite converter
├── web/ # Vite + React app
│ ├── src/
│ │ ├── main.tsx
│ │ ├── app.tsx
│ │ └── components/
│ ├── public/ # .db file goes here after build
│ ├── index.html
│ └── vite.config.ts
├── package.json # Root: scripts + web deps
├── src/main/resources/raw/ # Existing Excel files (unchanged)
└── ... (existing Java files unchanged)
```
## Related Code Files
**Create:**
- `package.json` — root package with workspaces or flat deps
- `web/index.html` — Vite entry
- `web/vite.config.ts`
- `web/src/main.tsx`
- `web/src/app.tsx`
- `web/tsconfig.json`
- `scripts/` directory (empty, placeholder)
**No modifications to existing files.**
## Implementation Steps
1. Initialize `package.json` at project root
```bash
npm init -y
```
2. Install build-script dependencies
```bash
npm install --save-dev better-sqlite3 xlsx
```
3. Install web dependencies
```bash
npm install react react-dom sql.js
npm install --save-dev @types/react @types/react-dom typescript vite @vitejs/plugin-react
```
4. Create `web/` directory with Vite scaffold
- `web/index.html` — minimal HTML shell
- `web/vite.config.ts` — configure root, public dir, build output
- `web/src/main.tsx` — React entry
- `web/src/app.tsx` — placeholder App component
- `web/tsconfig.json` — strict TS config
5. Create `scripts/` directory
6. Add npm scripts to `package.json`:
```json
{
"scripts": {
"build:db": "node scripts/build-database.js",
"dev": "vite --config web/vite.config.ts",
"build": "vite build --config web/vite.config.ts",
"preview": "vite preview --config web/vite.config.ts"
}
}
```
7. Update `.gitignore` to add:
```
node_modules/
web/dist/
web/public/thptqg2017.db
```
8. Verify `npm run dev` starts without errors
## Todo List
- [ ] Create package.json with all dependencies
- [ ] Create web/ directory with Vite + React scaffold
- [ ] Create scripts/ directory
- [ ] Update .gitignore
- [ ] Verify `npm run dev` starts clean
## Success Criteria
- `npm install` completes without errors
- `npm run dev` opens a blank React page in browser
- No existing files modified (Java code untouched)
- All new files use kebab-case naming
## Risk Assessment
| Risk | Mitigation |
|------|------------|
| better-sqlite3 native build fails on Windows | Use prebuild binaries (default); fallback to sql.js for build script too |
| Vite config path issues with nested web/ dir | Set `root: 'web'` in vite.config.ts explicitly |
## Next Steps
- Phase 02 depends on `scripts/` dir and `better-sqlite3` being available
- Phase 03 depends on `web/` scaffold and React being configured
@@ -1,201 +0,0 @@
# Phase 02 — Excel Parser + DB Builder
## Context Links
- [plan.md](./plan.md)
- [phase-01](./phase-01-project-scaffolding.md) — prerequisite
- [Converter.java](../../src/main/java/dev/miti99/thptqg2017/Converter.java) — source parsing logic to port
## Overview
- **Priority**: P1 (produces the .db file needed by Phase 03)
- **Status**: Pending
- **Description**: Node.js script that reads ~119 Excel files, regex-parses scores, writes SQLite .db
## Key Insights
From Converter.java analysis:
- 4 columns per row: hoTen (0), ngaySinh (1), soBaoDanh (2), diemThi (3)
- Column 3 is a text blob with scores like `"Toán: 8.5 Ngữ văn: 7.0 Vật lí: 6.25 ..."`
- 11 subject patterns extracted via regex; not all subjects present for every student
- Some files have no header row (comment in Java: "Một số file lỗi nên không chắc có header")
- `session.merge()` = upsert behavior; update folder files should overwrite raw folder data
- ngaySinh format: `dd/MM/yyyy`
**Format edge cases to handle:**
- Row 0 might be header OR data — detect by checking if cell 2 (soBaoDanh) looks like an ID pattern
- Some cells may be numeric instead of string (Excel auto-detection)
- Duplicate .xlsx files exist: `10_LamDong_GNFT (1).xls.xlsx` and `10_LamDong_GNFT.xls.xlsx`
## Requirements
### Functional
- Parse all .xlsx from `src/main/resources/raw/` and `src/main/resources/raw/(update)/`
- Process `raw/` first, then `(update)/` so updates overwrite via INSERT OR REPLACE
- Extract all 11 subject scores via regex (matching Java patterns exactly)
- Store ngaySinh as text string (no date conversion)
- Output `web/public/thptqg2017.db`
### Non-Functional
- Complete in < 60 seconds on modern hardware
- Log progress: file count, row count, error count
- Skip malformed rows gracefully (log, don't crash)
## Architecture
```
scripts/build-database.js
├── reads: src/main/resources/raw/**/*.xlsx
├── reads: src/main/resources/raw/(update)/**/*.xlsx
├── uses: xlsx (SheetJS) for Excel parsing
├── uses: better-sqlite3 for fast SQLite writes
└── outputs: web/public/thptqg2017.db
```
### Data Flow Per File
```
.xlsx file
→ xlsx.readFile()
→ sheet_to_json({ header: 1, raw: false }) // array of arrays, all strings
→ for each row:
→ validate: row[2] matches soBaoDanh pattern (skip headers/junk)
→ extract: hoTen, ngaySinh, soBaoDanh from columns 0-2
→ regex match: diemThi (column 3) against 11 subject patterns
→ INSERT OR REPLACE into student table
```
## Related Code Files
**Create:**
- `scripts/build-database.js` — main parser script (~120 lines)
**Read (reference only):**
- `src/main/java/dev/miti99/thptqg2017/Converter.java`
**No modifications to existing files.**
## Implementation Steps
### 1. Create score-parsing utility
Port the 11 regex patterns from Converter.java:
```javascript
const SCORE_PATTERNS = {
toan: /Toán:\s*(\d*\.\d*)/,
ngu_van: /Ngữ văn:\s*(\d*\.\d*)/,
vat_ly: /Vật lí:\s*(\d*\.\d*)/,
hoa_hoc: /Hóa học:\s*(\d*\.\d*)/,
sinh_hoc: /Sinh học:\s*(\d*\.\d*)/,
khtn: /KHTN:\s*(\d*\.\d*)/,
lich_su: /Lịch sử:\s*(\d*\.\d*)/,
dia_ly: /Địa lí:\s*(\d*\.\d*)/,
gdcd: /GDCD:\s*(\d*\.\d*)/,
khxh: /KHXH:\s*(\d*\.\d*)/,
tieng_anh: /Tiếng Anh:\s*(\d*\.\d*)/,
};
```
### 2. Create database schema
```javascript
db.exec(`
CREATE TABLE IF NOT EXISTS student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ngay_sinh TEXT,
toan REAL, ngu_van REAL, vat_ly REAL,
hoa_hoc REAL, sinh_hoc REAL, khtn REAL,
lich_su REAL, dia_ly REAL, gdcd REAL,
khxh REAL, tieng_anh REAL
);
`);
```
### 3. Build the main parsing loop
```
for each folder in [raw/, raw/(update)/]:
for each .xlsx file:
workbook = xlsx.readFile(filePath)
sheet = workbook.Sheets[workbook.SheetNames[0]]
rows = xlsx.utils.sheet_to_json(sheet, { header: 1, raw: false })
for each row:
if row[2] doesn't look like a valid soBaoDanh → skip (header detection)
parse scores from row[3]
INSERT OR REPLACE
```
### 4. Header detection heuristic
```javascript
// soBaoDanh is typically a numeric string like "02000001"
function isDataRow(row) {
return row[2] && /^\d{6,}$/.test(String(row[2]).trim());
}
```
### 5. Wrap inserts in a transaction
```javascript
const insert = db.prepare(`INSERT OR REPLACE INTO student (...) VALUES (...)`);
const insertMany = db.transaction((rows) => {
for (const row of rows) insert.run(row);
});
```
### 6. Add indexes after all inserts
```javascript
db.exec('CREATE INDEX IF NOT EXISTS idx_ho_ten ON student(ho_ten)');
```
### 7. Add summary logging
```
console.log(`Processed ${fileCount} files, ${rowCount} rows, ${errorCount} errors`);
```
### 8. Ensure output directory exists
```javascript
fs.mkdirSync('web/public', { recursive: true });
```
## Todo List
- [ ] Create `scripts/build-database.js`
- [ ] Port all 11 regex patterns from Converter.java
- [ ] Implement header-detection heuristic
- [ ] Process raw/ folder first, then (update)/ folder
- [ ] Wrap all inserts in single transaction for speed
- [ ] Add index on ho_ten after insert
- [ ] Log file count, row count, error count
- [ ] Run script and verify row count matches Java output
- [ ] Verify .db file size is reasonable (< 60MB expected)
## Success Criteria
1. `node scripts/build-database.js` completes without unhandled errors
2. Output .db contains same number of unique students as existing database.sqlite
3. Spot-check: query 5 random soBaoDanh values → scores match between old and new DB
4. Script completes in < 60 seconds
5. Skipped/errored rows are logged with file name and row number
## Risk Assessment
| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|------------|
| xlsx library misreads Vietnamese characters | Low | High | Use `{ raw: false }` to get string values; test with known file |
| Regex patterns don't match all score formats | Medium | Medium | Test against sample rows; add `\d+\.?\d*` fallback if needed |
| better-sqlite3 won't install on Windows | Low | Medium | Prebuild binaries exist; fallback: use sql.js for build too |
| Memory pressure with 119 files open | Low | Low | Process files sequentially, one at a time |
| Duplicate students from overlapping files | Medium | Low | INSERT OR REPLACE handles this; (update) processed last wins |
## Security Considerations
- Public exam data, no PII concerns beyond names (already public)
- No network access needed during build
- Output .db is read-only in production
## Next Steps
- Phase 03 consumes the `web/public/thptqg2017.db` file produced here
- Verify row count against existing database before proceeding
@@ -1,201 +0,0 @@
# Phase 03 — React Static Site
## Context Links
- [plan.md](./plan.md)
- [phase-01](./phase-01-project-scaffolding.md) — provides Vite+React scaffold
- [phase-02](./phase-02-excel-parser-db-builder.md) — provides .db file
## Overview
- **Priority**: P1
- **Status**: Pending
- **Description**: Vite+React+TypeScript site that loads SQLite .db client-side via sql.js for score lookup
## Key Insights
- .db file is ~50-60MB — must show loading progress to user
- sql.js requires WASM file — load from CDN (cdnjs) or bundle in public/
- All queries are readonly SELECT — no write operations
- Vietnamese UI — all labels, placeholders, messages in Vietnamese
- Two search modes: by name (ho_ten LIKE) and by exam ID (so_bao_danh exact match)
## Requirements
### Functional
- Load .db file on page load with progress indicator
- Search by số báo danh (exact match) or họ tên (partial match, case-insensitive)
- Display results in a clean table with all score columns
- Show "không tìm thấy" when no results
- Limit results to 50 rows (prevent rendering 800K rows)
### Non-Functional
- First meaningful paint < 2s (before DB loads)
- Search response < 100ms after DB is loaded
- Mobile-responsive layout
- Works offline after initial load (static site + cached .db)
## Architecture
```
Browser
├── index.html (Vite entry)
├── main.tsx → App
│ ├── useSqlite hook (loads .db, exposes query fn)
│ ├── SearchBar component (input + mode toggle)
│ └── ResultTable component (score display)
└── sql.js WASM (from CDN)
```
### Data Flow
```
Page Load
→ fetch('/thptqg2017.db') with progress tracking
→ initSqlJs({ locateFile: cdnjs url })
→ new SQL.Database(arrayBuffer)
→ DB ready, enable search
User Types Query
→ debounce 300ms
→ if mode=id: SELECT * FROM student WHERE so_bao_danh = ?
→ if mode=name: SELECT * FROM student WHERE ho_ten LIKE ? LIMIT 50
→ render ResultTable with rows
```
## Related Code Files
**Create:**
- `web/src/app.tsx` — main app layout (~60 lines)
- `web/src/hooks/use-sqlite.ts` — sql.js loading + query hook (~70 lines)
- `web/src/components/search-bar.tsx` — search input + mode toggle (~40 lines)
- `web/src/components/result-table.tsx` — score results table (~60 lines)
- `web/src/types/student.ts` — TypeScript interface (~20 lines)
- `web/src/index.css` — minimal styling (~50 lines)
**Modify:**
- `web/src/main.tsx` — import App + CSS
## Implementation Steps
### 1. Define Student type
```typescript
// web/src/types/student.ts
export interface Student {
so_bao_danh: string;
ho_ten: string;
ngay_sinh: string | null;
toan: number | null;
ngu_van: number | null;
vat_ly: number | null;
hoa_hoc: number | null;
sinh_hoc: number | null;
khtn: number | null;
lich_su: number | null;
dia_ly: number | null;
gdcd: number | null;
khxh: number | null;
tieng_anh: number | null;
}
```
### 2. Implement use-sqlite hook
```typescript
// web/src/hooks/use-sqlite.ts
// States: loading (with progress %), ready, error
// On mount: fetch .db file → init sql.js → create Database instance
// Expose: { db, loading, progress, error, query(sql, params) }
```
Key details:
- Use `fetch()` with `response.body.getReader()` for progress tracking
- sql.js WASM from: `https://cdnjs.cloudflare.com/ajax/libs/sql.js/1.11.0/sql-wasm.wasm`
- Memoize the Database instance with useRef
### 3. Implement search-bar component
- Single text input with placeholder "Nhập số báo danh hoặc họ tên..."
- Auto-detect mode: if input is all digits → search by soBaoDanh; else → search by hoTen
- Debounce input by 300ms before triggering query
- Minimum 2 characters to trigger name search
### 4. Implement result-table component
- Responsive HTML table
- Columns: STT, Số báo danh, Họ tên, Ngày sinh, Toán, Ngữ văn, Vật lí, Hóa học, Sinh học, KHTN, Lịch sử, Địa lí, GDCD, KHXH, Tiếng Anh
- Show null scores as "-" (not 0)
- Highlight search match in name column (bold)
- "Không tìm thấy kết quả" message when empty
### 5. Implement app.tsx layout
```
<header>
<h1>Tra cứu điểm thi THPT QG 2017</h1>
</header>
<main>
{loading ? <ProgressBar /> : <SearchBar />}
<ResultTable />
</main>
<footer>
Dữ liệu sưu tầm, chỉ mang tính tham khảo
</footer>
```
### 6. Styling (index.css)
- Clean, minimal CSS — no framework needed for this scope
- CSS variables for colors
- Mobile-first responsive: table scrolls horizontally on small screens
- Loading bar animation
### 7. SQL query construction
```sql
-- By exam ID (exact)
SELECT * FROM student WHERE so_bao_danh = ?
-- By name (partial, accent-insensitive not feasible in SQLite, use LIKE)
SELECT * FROM student WHERE ho_ten LIKE ? LIMIT 50
-- param: `%${query}%`
```
Note: SQLite LIKE is case-insensitive for ASCII only. Vietnamese diacritics mean users must type exact accents. This is acceptable — Vietnamese users expect this.
## Todo List
- [ ] Create Student TypeScript interface
- [ ] Implement use-sqlite hook with progress tracking
- [ ] Implement search-bar with auto-detect mode + debounce
- [ ] Implement result-table with all score columns
- [ ] Implement app.tsx layout with loading state
- [ ] Add minimal CSS styling (mobile-responsive)
- [ ] Test with actual .db file from Phase 02
- [ ] Verify search by soBaoDanh returns exact match
- [ ] Verify search by hoTen returns partial matches (limit 50)
## Success Criteria
1. `npm run dev` shows app with loading indicator while .db fetches
2. Search by exact soBaoDanh returns single student with correct scores
3. Search by partial name returns up to 50 matching students
4. No results shows Vietnamese "không tìm thấy" message
5. Table is readable on mobile (horizontal scroll)
6. No console errors in Chrome/Firefox
## Risk Assessment
| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|------------|
| .db fetch takes too long (50MB+) | High | Medium | Show progress bar; consider gzip in Phase 04 |
| sql.js WASM CDN unreachable | Low | High | Bundle WASM in public/ as fallback |
| Vietnamese text search limitations | Medium | Low | Document that exact diacritics required; acceptable UX |
| 800K row render if query too broad | Medium | Medium | LIMIT 50 on all queries; show "refine search" message |
## Security Considerations
- All data is public exam scores — no auth needed
- sql.js runs client-side only — no server attack surface
- No user data collected or stored
## Next Steps
- Phase 04 adds build optimization, gzip, and deploy configuration
@@ -1,222 +0,0 @@
# Phase 04 — Integration + Deploy Config
## Context Links
- [plan.md](./plan.md)
- [phase-02](./phase-02-excel-parser-db-builder.md) — produces .db file
- [phase-03](./phase-03-react-static-site.md) — produces React app
## Overview
- **Priority**: P2
- **Status**: Pending
- **Description**: Wire everything together, optimize .db delivery, configure GitHub Pages deploy
## Key Insights
- .db file is ~50-60MB — gzip reduces SQLite files by ~70% typically → ~15-20MB
- GitHub Pages has 100MB file size limit; raw .db is borderline, gzipped is safe
- Vite build output goes to `web/dist/` — GitHub Actions deploys this folder
- No server needed — fully static
## Requirements
### Functional
- `npm run build:db` → `npm run build` produces deployable `web/dist/`
- .db file served with gzip (either pre-compressed or server-side)
- GitHub Actions workflow for automated deploy on push to main
### Non-Functional
- Total deploy artifact < 25MB (gzipped .db + app bundle)
- Build completes in CI in < 5 minutes
## Architecture
```
GitHub Push (main)
→ Actions workflow
→ npm ci
→ node scripts/build-database.js (produces web/public/thptqg2017.db)
→ npm run build (Vite builds web/dist/)
→ deploy web/dist/ to GitHub Pages
```
### .db Delivery Optimization
**Option A — Pre-gzip (recommended):**
- Build script outputs `thptqg2017.db`
- Post-build step: `gzip -k web/public/thptqg2017.db` → produces `.db.gz`
- Vite copies `.db.gz` to `dist/`
- Client fetches `.db.gz`, decompresses with `DecompressionStream` or pako
- Pros: works on any static host, no server config needed
**Option B — Rely on server gzip:**
- Serve raw .db, let CDN/server gzip on-the-fly
- Pros: simpler client code
- Cons: GitHub Pages may not gzip .db extension; unreliable
**Decision: Option A** — pre-gzip with client-side decompress. Use browser-native `DecompressionStream` (supported in all modern browsers).
## Related Code Files
**Create:**
- `.github/workflows/deploy.yml` — GitHub Actions workflow (~40 lines)
**Modify:**
- `scripts/build-database.js` — add gzip step after DB creation
- `web/src/hooks/use-sqlite.ts` — fetch .db.gz, decompress, then init sql.js
- `package.json` — add `build:all` script combining db + vite build
- `.gitignore` — ensure web/dist/ and .db files are excluded
## Implementation Steps
### 1. Add gzip post-processing to build script
```javascript
// At end of scripts/build-database.js
const zlib = require('zlib');
const dbBuffer = fs.readFileSync('web/public/thptqg2017.db');
const gzipped = zlib.gzipSync(dbBuffer);
fs.writeFileSync('web/public/thptqg2017.db.gz', gzipped);
console.log(`Compressed: ${(gzipped.length / 1024 / 1024).toFixed(1)}MB`);
```
### 2. Update use-sqlite hook for gzip fetch
```typescript
// Fetch .db.gz instead of .db
const response = await fetch('/thptqg2017.db.gz');
const reader = response.body.getReader();
// ... read chunks with progress ...
const compressed = new Uint8Array(allChunks);
// Decompress using DecompressionStream
const ds = new DecompressionStream('gzip');
const writer = ds.writable.getWriter();
writer.write(compressed);
writer.close();
const decompressed = await new Response(ds.readable).arrayBuffer();
// Init sql.js with decompressed buffer
const db = new SQL.Database(new Uint8Array(decompressed));
```
### 3. Add combined build script
```json
{
"scripts": {
"build:db": "node scripts/build-database.js",
"build:web": "vite build --config web/vite.config.ts",
"build:all": "npm run build:db && npm run build:web",
"dev": "vite --config web/vite.config.ts",
"preview": "vite preview --config web/vite.config.ts"
}
}
```
### 4. Create GitHub Actions deploy workflow
```yaml
# .github/workflows/deploy.yml
name: Deploy to GitHub Pages
on:
push:
branches: [main]
workflow_dispatch:
permissions:
contents: read
pages: write
id-token: write
jobs:
build-and-deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
lfs: true # in case Excel files are in LFS
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
- run: npm ci
- run: npm run build:all
- uses: actions/upload-pages-artifact@v3
with:
path: web/dist
- uses: actions/deploy-pages@v4
```
### 5. Configure Vite base path for GitHub Pages
```typescript
// web/vite.config.ts
export default defineConfig({
base: '/thptqg2017/', // matches GitHub repo name
// ...
});
```
### 6. Update .gitignore
```
node_modules/
web/dist/
web/public/thptqg2017.db
web/public/thptqg2017.db.gz
```
### 7. End-to-end verification
1. Run `npm run build:all`
2. Run `npm run preview`
3. Open in browser, search for known student
4. Verify scores display correctly
5. Check network tab: .db.gz transfer size < 25MB
## Todo List
- [ ] Add gzip step to build-database.js
- [ ] Update use-sqlite hook for .db.gz fetch + decompress
- [ ] Add build:all npm script
- [ ] Configure Vite base path for GitHub Pages
- [ ] Create .github/workflows/deploy.yml
- [ ] Update .gitignore
- [ ] End-to-end test: build:all → preview → search
- [ ] Verify GitHub Actions workflow passes
## Success Criteria
1. `npm run build:all` produces `web/dist/` with all assets + .db.gz
2. .db.gz file < 25MB
3. `npm run preview` serves working site from dist/
4. GitHub Actions workflow deploys successfully
5. Live site at `https://{username}.github.io/thptqg2017/` is functional
## Risk Assessment
| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|------------|
| Excel files not in git (too large for checkout) | Medium | High | Check if LFS needed; or commit .db.gz directly and skip build:db in CI |
| DecompressionStream not supported in old browsers | Low | Low | 95%+ browser support; show "update browser" message for others |
| GitHub Pages deploy fails on large artifact | Low | Medium | Gzipped artifact should be well under limits |
| CI build timeout (Excel parsing slow) | Low | Low | 119 files in < 60s locally; CI has 6h limit |
## Backwards Compatibility
- Existing `database.sqlite` (97MB) preserved — not deleted or modified
- Java source code preserved — can be removed in separate cleanup PR
- New site deployed to GitHub Pages — no impact on existing setup
## Rollback Plan
- Revert the deploy.yml workflow → Pages stops updating
- Previous Pages deployment (if any) is preserved in GitHub Pages history
- Local: `git revert` the integration commit; Phases 1-3 artifacts still work standalone
## Next Steps (Future, out of scope)
- Remove Java source code and Gradle files (cleanup PR)
- Add statistics page (average scores per province)
- PWA support for offline access
@@ -1,105 +0,0 @@
---
title: "Static Score Lookup - Excel to SQLite + Vite React Site"
description: "Replace Java/Hibernate pipeline with Node.js parser + client-side SQLite lookup site"
status: pending
priority: P1
effort: 6h
branch: main
tags: [node, vite, react, sql.js, sqlite, migration]
created: 2026-04-12
---
# Static Score Lookup
## Goal
Replace the Java-based Excel-to-SQLite pipeline with a Node.js script, and build a static Vite+React site that loads the .db file client-side via sql.js for readonly score queries.
## Architecture Overview
```
[~119 .xlsx files] --> [Node.js parser script] --> [thptqg2017.db]
|
[Vite build copies to public/]
|
[React app loads via sql.js]
|
[User searches by name/ID]
```
### Data Flow
1. **Parse**: Node script reads all .xlsx from `src/main/resources/raw/` and `raw/(update)/`
2. **Transform**: Extract 4 columns per row; regex-parse scores from column 4 text
3. **Load**: Insert rows into SQLite via better-sqlite3 (Node-native, fast)
4. **Serve**: Vite copies .db to `public/`; React app fetches + initializes sql.js
5. **Query**: User input -> SQL WHERE -> render results table
### Database Schema
```sql
CREATE TABLE student (
so_bao_danh TEXT PRIMARY KEY,
ho_ten TEXT NOT NULL,
ngay_sinh TEXT, -- stored as dd/MM/yyyy string
toan REAL,
ngu_van REAL,
vat_ly REAL,
hoa_hoc REAL,
sinh_hoc REAL,
khtn REAL,
lich_su REAL,
dia_ly REAL,
gdcd REAL,
khxh REAL,
tieng_anh REAL
);
CREATE INDEX idx_ho_ten ON student(ho_ten);
```
**Change from Java version**: `ngay_sinh` stored as TEXT (not DATE) — simpler, no timezone issues, display-only field.
## Phases
| # | Phase | Status | Effort | Details |
|---|-------|--------|--------|---------|
| 1 | Project scaffolding | Pending | 30m | [phase-01](./phase-01-project-scaffolding.md) |
| 2 | Excel parser + DB builder | Pending | 2h | [phase-02](./phase-02-excel-parser-db-builder.md) |
| 3 | React static site | Pending | 2.5h | [phase-03](./phase-03-react-static-site.md) |
| 4 | Integration + deploy config | Pending | 1h | [phase-04](./phase-04-integration-deploy.md) |
## Dependencies
```
Phase 1 --> Phase 2 --> Phase 3 --> Phase 4
| ^
+--(produces .db)----+
```
## Risk Assessment
| Risk | Likelihood | Impact | Mitigation |
|------|-----------|--------|------------|
| .db file too large for browser fetch (~97MB) | High | High | Compress with gzip; sql.js supports ArrayBuffer; lazy-load on first query |
| Excel format variations (missing headers, bad cells) | Medium | Medium | try/catch per row (same as Java); log skipped rows; validate after build |
| sql.js WASM loading fails on some browsers | Low | Medium | Fallback error message; test Chrome/Firefox/Safari |
| Duplicate soBaoDanh across files (raw + update) | Medium | Low | Use INSERT OR REPLACE (update folder overwrites raw) |
## Backwards Compatibility
- Existing `database.sqlite` (97MB) preserved until new pipeline verified
- Java source code left intact — can be removed in future cleanup
- New .db file placed at `public/thptqg2017.db` (different name/location)
## Rollback Plan
- Phase 1-2: Delete `scripts/` and `package.json`; Java pipeline still works
- Phase 3-4: Delete `web/` folder; .db file from phase 2 still standalone-usable
## Success Criteria
1. `node scripts/build-database.js` produces valid .db with ~800K+ rows (matching Java output count)
2. `npm run dev` serves site; search by soBaoDanh returns correct student
3. `npm run build` produces static dist/ deployable to any static host
4. .db file size < 60MB (better-sqlite3 is more compact than Hibernate output)