feat: download the database instead of reading it over HTTP

Reading the file where it lay never worked well enough. Two costs were
structural rather than bugs: reads are serial, because the worker uses
synchronous XHR, so a name search touching 390 pages waited 17 seconds
to move 608 KB — roughly one request per result row, which no page size
removes — and the first visitor after each deploy waited ~26 seconds for
the CDN to fill its cache with a 288 MB object.

The browser now downloads the whole database once and queries it in
memory with sql.js. A dataset page is gated behind that: the gate states
what it will cost, in transfer and in memory, and offers only the
download, because there is nothing to show without it.

Dropping the structures that existed to make range-request queries
index-driven halved the file. name_word carried one row per word of
every name, about 3.5 million of them, and with the partial score
indexes it was more than half of what every visitor would now download.
Measured on rebuilt databases: 2016 went 288.6 -> 142.5 MB (31 MB
gzipped on the wire), 2017 237.7 -> 119.3 MB, both with row counts and
audits unchanged. Queries on the result: an exam number is immediate, a
name scans all 877,460 rows in about 240 ms.

Alternatives were measured before choosing this. sqlite-wasm-http sizes
files from a HEAD Content-Length with no override, so on a host that
gzips it silently uses the compressed size. DuckDB-WASM ships 32-37 MB
of WebAssembly before its Parquet extension, more than this whole
download. Static pre-generated shards are the most robust option but
cannot answer arbitrary SQL, and cannot stop early the way LIMIT does.

The published name loses its chunk index, the byte budgets and the SQL
consent modal go with the range reads that made them necessary, and the
docs no longer describe a design the site does not use.
This commit is contained in:
tiennm99 committed 2026-08-14 17:43:10 +07:00
1 parent cbd262e5db
commit 4cb0a4f340
28 files changed
+449 -986

No files matched your search

+7 -10
View File
@@ -124,26 +124,23 @@ jobs:
- uses: actions/checkout@v7
# The site reads the databases a page at a time over range requests, so
# what matters is not that the host leaves the file alone in general — it
# gzips the un-ranged response, and browsers work around that by sending
# Accept-Encoding: identity whenever a request carries a Range header —
# but that a ranged read returns raw database bytes. Checking headers is
# what missed this before: a bare `curl -sI` advertises no encoding and
# so passes whatever the host does. Check the bytes instead.
- name: Verify ranged reads return database bytes
# Checks the published file is a database, not an error page or a
# truncated upload, by reading its first bytes. A range request is used
# only because it is the cheapest way to see them without pulling 142 MB;
# the site itself downloads the file whole.
- name: Verify the published databases are readable
env:
PAGE_URL: ${{ steps.deployment.outputs.page_url }}
run: |
set -euo pipefail
for id in $(jq -r '.datasets[].id' datasets.json); do
url="${PAGE_URL%/}/db/${id}.sqlite30"
url="${PAGE_URL%/}/db/${id}.sqlite3"
if ! magic=$(curl -sf -r 0-14 -H 'Accept-Encoding: identity;q=1, *;q=0' "$url"); then
echo "::error::$url is not fetchable"
exit 1
fi
if [ "$magic" != "SQLite format 3" ]; then
echo "::error::$url did not return database bytes over a range request"
echo "::error::$url does not start with the SQLite header"
exit 1
fi
echo "$url: SQLite format 3"
+19 -25
View File
@@ -57,31 +57,25 @@ hashes every real input file. That is the point of it; do not skip it.
prerendered to its own file and asset URLs stay absolute, so the copy of
`index.html` serving as `404.html` works at any depth. No SPA 404-fallback is
used — a fallback would break the `?q=` deep links.
- **`dbSizeMb` in `datasets.json` is a build guard, not just a label.** The
assembler refuses to publish an artifact that falls below a ratio of it.
- **The databases ship uncompressed, as `<id>.sqlite30`.** The trailing 0 is a
chunk index, not a typo — see the next point. The browser reads byte ranges
of the file, and a range of a gzip stream is not a range of the database.
- **The site uses chunked mode over a single chunk, and that is deliberate.**
GitHub Pages gzips `application/octet-stream`, so the HEAD request
`sql.js-httpvfs` sizes a file with reports the compressed length, and the
library refuses to open the file. Chunked mode is the only mode whose config
takes a length (`databaseLengthBytes`); in full mode the worker hardcodes it
to `undefined`, so a length passed there is silently dropped. One chunk holds
the database, so the index is always 0 and every request goes to
`<id>.sqlite3` + `0`. `web/src/lib/db-probe.js` supplies the length by
reading the file header over a range request.
- **Ranged reads were never affected by the compression**, because browsers
must send `Accept-Encoding: identity` whenever a request carries a `Range`
header. Verify the way a browser asks —
`curl -s -r 0-14 -H 'Accept-Encoding: identity' …` must print
`SQLite format 3` — never a bare `curl -sI`, which advertises no encoding
and so passes whatever the host does.
- **Every query the site runs must be index-driven.** Over range requests an
unindexed query fetches the whole table. Hence no index on `ho_ten` (nothing
can use one), `name_word` for name search, partial indexes for the score
presets, and the footer count read from `datasets.json` instead of
`COUNT(*)`.
- **`dbSizeMb` in `datasets.json` is a build guard and a user-facing figure.**
The assembler refuses to publish an artifact that falls below a ratio of it,
and the download gate shows it as the memory the tab will need.
- **The browser downloads the whole database and queries it in memory.** The
dataset page is gated behind that download: `download-gate.svelte` states the
transfer and memory cost and offers nothing but the download, because there
is no answer to give without it. `sql.js` holds the file in WebAssembly
memory for as long as the tab is open, so it is RAM, not disk, and it is gone
on reload.
- **The schema carries no secondary indexes, deliberately.** An index saves a
scan that already takes a few hundred milliseconds in memory, and costs every
visitor megabytes of download. An earlier design read the file over HTTP
range requests and needed the opposite — a `name_word` table with one row per
word of every name, plus partial score indexes — which was more than half the
published file: 288 MB became 142 MB when they went.
- **The databases ship uncompressed, as `<id>.sqlite3`.** GitHub Pages gzips
them on the wire anyway, which is where the transfer figure comes from
(142 MB stored, 31 MB delivered), so publishing a `.gz` would only mean
decompressing twice.
## Conventions
+3 -3
View File
@@ -1,8 +1,8 @@
# thptqg
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
school graduation exam. Client-side SQL over a SQLite database read in place by
HTTP range request, built
school graduation exam. Client-side SQL over a SQLite database the browser
downloads once and queries in memory, built
from the published `.xls`/`.xlsx` score files by the Go `parser` module. Where
those files come from: [data pipeline](./docs/data-pipeline.md#sources).
@@ -42,7 +42,7 @@ app both read it and neither needs a dependency to do so; presentation stays in
The dataset id is one identifier end to end:
```
data/2017/ → parser/configs/2017.yml → db/2017.sqlite30 → /thptqg/2017/
data/2017/ → parser/configs/2017.yml → db/2017.sqlite3 → /thptqg/2017/
```
## Build
+7 -15
View File
@@ -11,9 +11,9 @@
// The guards below close that: a build whose row count does not match the
// registry, or whose artifact is implausibly small, fails the pipeline.
//
// The databases ship uncompressed. The browser reads them a page at a time over
// HTTP range requests, and a range of a gzip stream is not a range of the
// database.
// The databases ship uncompressed. The browser downloads one whole and opens
// it in memory, and the host compresses it on the wire anyway — publishing a
// .gz would only mean decompressing twice.
package databases
import (
@@ -35,18 +35,10 @@ const driverName = "sqlite"
// even if the row count somehow passed.
const minSizeRatio = 0.9
// Extension is the published suffix. The trailing 0 is a chunk index, not a
// typo: the browser reads the file through sql.js-httpvfs in chunked mode,
// which is the only mode whose config accepts the file's length. That mode
// builds each request's URL as urlPrefix + chunkIndex, and one chunk holds the
// whole database, so the index is always 0 and the prefix is "<id>.sqlite3".
//
// The length has to come from the config because the library otherwise takes
// it from a HEAD request, which GitHub Pages answers with the gzipped size.
//
// Not ".db" either: keeping that name free lets the site assembly treat any
// stray .db or SQLite journal in the output as the leftover it is.
const Extension = ".sqlite30"
// Extension is the published suffix. Not ".db": keeping that name free lets
// the site assembly treat any stray .db or SQLite journal in the output as the
// leftover it is.
const Extension = ".sqlite3"
// Paths locates the pieces this package needs.
type Paths struct {
@@ -12,9 +12,9 @@ import (
func TestCleanRemovesOnlyDroppedDatasets(t *testing.T) {
dir := t.TempDir()
for _, name := range []string{
"2016.sqlite30", "2017.sqlite30",
"2017-old.sqlite30", // dropped from the registry
"2017-old2.sqlite30", // dropped from the registry
"2016.sqlite3", "2017.sqlite3",
"2017-old.sqlite3", // dropped from the registry
"2017-old2.sqlite3", // dropped from the registry
"2016.db-journal", // interrupted run
} {
if err := os.WriteFile(filepath.Join(dir, name), []byte("x"), 0o644); err != nil {
@@ -38,7 +38,7 @@ func TestCleanRemovesOnlyDroppedDatasets(t *testing.T) {
}
slices.Sort(left)
want := []string{"2016.sqlite30", "2017.sqlite30"}
want := []string{"2016.sqlite3", "2017.sqlite3"}
if !slices.Equal(left, want) {
t.Errorf("left %v, want %v", left, want)
}
+6 -7
View File
@@ -147,17 +147,16 @@ func checkDatabasesPresent(siteDir string, datasets []registry.Dataset) error {
return nil
}
// strayArtifact matches what must never reach the output: a SQLite journal from
// an interrupted run, a database under either older name — .db, or .sqlite3
// without the chunk index the client asks for — or a gzipped database from
// before the switch to range requests.
var strayArtifact = regexp.MustCompile(`(\.db|\.sqlite30?)(-journal|-wal|-shm)$|\.db$|\.sqlite3$|\.gz$`)
// strayArtifact matches what must never reach the output: a SQLite journal
// from an interrupted run, a database under the old .db name, or one under the
// .sqlite30 name a former range-request client asked for. Each is 100+ MB.
var strayArtifact = regexp.MustCompile(`(\.db|\.sqlite30?)(-journal|-wal|-shm)$|\.db$|\.sqlite30$|\.gz$`)
// checkNoStrayArtifacts rejects leftovers that would be published.
//
// The staging directory is copied wholesale, so anything an interrupted run left
// behind goes straight through — and each of these is 100+ MB. A gzipped
// database would also be unreadable to the site, which reads byte ranges.
// behind goes straight through — and each of these is 100+ MB, downloaded in
// full by anyone the site hands the wrong name to.
func checkNoStrayArtifacts(siteDir string) error {
var stray []string
err := filepath.WalkDir(siteDir, func(path string, d os.DirEntry, err error) error {
+15 -17
View File
@@ -39,7 +39,7 @@ func write(t *testing.T, path, body string) {
}
func TestAssembleProducesAPageForEveryDataset(t *testing.T) {
p := fakeBuild(t, "2016.sqlite30", "2017.sqlite30")
p := fakeBuild(t, "2016.sqlite3", "2017.sqlite3")
if err := Assemble(p, datasets); err != nil {
t.Fatal(err)
}
@@ -49,7 +49,7 @@ func TestAssembleProducesAPageForEveryDataset(t *testing.T) {
filepath.Join("2016", "index.html"),
filepath.Join("2017", "index.html"),
filepath.Join("_app", "immutable", "entry.js"),
filepath.Join("db", "2016.sqlite30"),
filepath.Join("db", "2016.sqlite3"),
} {
if _, err := os.Stat(filepath.Join(p.Site, want)); err != nil {
t.Errorf("missing from the artifact: %s", want)
@@ -61,7 +61,7 @@ func TestAssembleProducesAPageForEveryDataset(t *testing.T) {
// entry generator, which reads the same registry this does. If the two fall out
// of step, that dataset's URL 404s — so the build stops instead.
func TestMissingDatasetPageFailsTheBuild(t *testing.T) {
p := fakeBuild(t, "2016.sqlite30", "2017.sqlite30")
p := fakeBuild(t, "2016.sqlite3", "2017.sqlite3")
if err := os.RemoveAll(filepath.Join(p.Dist, "2017")); err != nil {
t.Fatal(err)
}
@@ -80,12 +80,12 @@ func TestMissingDatasetPageFailsTheBuild(t *testing.T) {
// renders, every query 404s, and CI stays green. The row-count and size guards
// cannot catch this — they only run when a database was built at all.
func TestMissingDatabaseFailsTheBuild(t *testing.T) {
p := fakeBuild(t, "2016.sqlite30") // 2017 never built
p := fakeBuild(t, "2016.sqlite3") // 2017 never built
err := Assemble(p, datasets)
if err == nil {
t.Fatal("expected an error when a database is missing")
}
if !strings.Contains(err.Error(), "2017.sqlite30") {
if !strings.Contains(err.Error(), "2017.sqlite3") {
t.Errorf("the error should name the missing database, got: %v", err)
}
}
@@ -93,25 +93,23 @@ func TestMissingDatabaseFailsTheBuild(t *testing.T) {
// TestEmptyDatabaseFailsTheBuild: a zero-byte file satisfies "exists" but is
// not a database.
func TestEmptyDatabaseFailsTheBuild(t *testing.T) {
p := fakeBuild(t, "2016.sqlite30", "2017.sqlite30")
write(t, filepath.Join(p.Dist, "db", "2017.sqlite30"), "")
p := fakeBuild(t, "2016.sqlite3", "2017.sqlite3")
write(t, filepath.Join(p.Dist, "db", "2017.sqlite3"), "")
if err := Assemble(p, datasets); err == nil {
t.Fatal("expected an error for a zero-byte database")
}
}
// TestStrayArtifactFailsTheBuild: a journal means an interrupted run, a .db or
// a chunk-index-less .sqlite3 means an older naming the client no longer asks
// for, a .gz means a database the site could not read a range of — and each is
// 100+ MB.
// a .sqlite30 means a name from an earlier design that no client asks for now,
// a .gz means a database left compressed — and each is 100+ MB.
func TestStrayArtifactFailsTheBuild(t *testing.T) {
for _, name := range []string{
"2016.db", "2016.sqlite3",
"2016.db", "2016.sqlite30",
"2016.sqlite3-journal", "2016.sqlite3-wal", "2016.sqlite3-shm", "2016.sqlite3.gz",
"2016.sqlite30-journal", "2016.sqlite30-wal", "2016.sqlite30-shm",
} {
t.Run(name, func(t *testing.T) {
p := fakeBuild(t, "2016.sqlite30", "2017.sqlite30")
p := fakeBuild(t, "2016.sqlite3", "2017.sqlite3")
write(t, filepath.Join(p.Dist, "db", name), "raw sqlite")
err := Assemble(p, datasets)
if err == nil {
@@ -127,12 +125,12 @@ func TestStrayArtifactFailsTheBuild(t *testing.T) {
// TestPublishedDatabaseIsNotMistakenForStray: the pattern must pass the one
// file the site is built to serve. Getting this wrong would fail every build.
func TestPublishedDatabaseIsNotMistakenForStray(t *testing.T) {
if strayArtifact.MatchString("2016.sqlite30") {
if strayArtifact.MatchString("2016.sqlite3") {
t.Error("the published database must not be treated as a stray artifact")
}
for _, name := range []string{
"2016.db", "x.sqlite3", "x.sqlite3-journal", "x.sqlite3-wal", "x.sqlite3-shm",
"x.sqlite30-journal", "x.sqlite30-wal", "x.sqlite30-shm", "x.sqlite3.gz", "x.db.gz",
"2016.db", "x.sqlite30", "x.sqlite3-journal", "x.sqlite3-wal", "x.sqlite3-shm",
"x.sqlite30-journal", "x.sqlite3.gz", "x.db.gz",
} {
if !strayArtifact.MatchString(name) {
t.Errorf("%s should be rejected", name)
@@ -151,7 +149,7 @@ func TestAssembleRejectsAMissingBuild(t *testing.T) {
// TestAssembleIsIdempotent: the site directory is rebuilt from scratch, so a
// previous run's leftovers cannot survive into the artifact.
func TestAssembleIsIdempotent(t *testing.T) {
p := fakeBuild(t, "2016.sqlite30", "2017.sqlite30")
p := fakeBuild(t, "2016.sqlite3", "2017.sqlite3")
if err := Assemble(p, datasets); err != nil {
t.Fatal(err)
}
+3 -3
View File
@@ -135,7 +135,7 @@ func compareOne(id, dirA, dirB string) (Result, error) {
return res, nil
}
// open finds <id>.sqlite30 in dir and returns a read-only handle.
// open finds <id>.sqlite3 in dir and returns a read-only handle.
//
// The older names are still accepted, and a gzipped database is expanded to a
// temporary file rather than rejected: the two sides of a comparison are often
@@ -143,7 +143,7 @@ func compareOne(id, dirA, dirB string) (Result, error) {
func open(dir, id string) (*sql.DB, func(), error) {
noop := func() {}
for _, name := range []string{id + ".sqlite30", id + ".sqlite3", id + ".db"} {
for _, name := range []string{id + ".sqlite3", id + ".sqlite30", id + ".db"} {
plain := filepath.Join(dir, name)
if _, err := os.Stat(plain); err == nil {
db, err := sql.Open(driverName, "file:"+plain+"?mode=ro")
@@ -154,7 +154,7 @@ func open(dir, id string) (*sql.DB, func(), error) {
gzPath := filepath.Join(dir, id+".db.gz")
f, err := os.Open(gzPath)
if err != nil {
return nil, noop, fmt.Errorf("no %s.sqlite30, %s.sqlite3, %s.db or %s.db.gz in %s", id, id, id, id, dir)
return nil, noop, fmt.Errorf("no %s.sqlite3, %s.sqlite30, %s.db or %s.db.gz in %s", id, id, id, id, dir)
}
defer f.Close()
+5 -5
View File
@@ -9,9 +9,9 @@
" means something changed unintentionally and the assembler",
" refuses to publish.",
" dbSizeMb usual size of the published database. Required and non-zero:",
" the assembler rejects a build that comes out far smaller. The",
" file is served uncompressed and read a page at a time over",
" HTTP range requests, so nothing downloads it whole.",
" the assembler rejects a build that comes out far smaller. It",
" is also the memory figure the download gate shows, since the",
" browser holds the whole file while the tab is open.",
"",
"Presentation (titles, labels, SQL presets) lives in web/src/datasets.js keyed",
"by id; that file throws at load if the two lists disagree."
@@ -20,12 +20,12 @@
{
"id": "2016",
"expectedRows": 877460,
"dbSizeMb": 288
"dbSizeMb": 142
},
{
"id": "2017",
"expectedRows": 861068,
"dbSizeMb": 237
"dbSizeMb": 119
}
]
}
+3 -3
View File
@@ -194,7 +194,7 @@ HTTP request, and those requests are serial. `CHUNK_BYTES` in
## Verifying a rebuild
The assembler verifies itself: each database's row count must match the
figure in the table above, and each `.sqlite30` must be at least 90% of its usual
figure in the table above, and each `.sqlite3` must be at least 90% of its usual
size, or the build fails rather than publishing. That guard is the reason a
truncated dataset cannot reach the site with a green pipeline.
@@ -208,8 +208,8 @@ go -C assembler run ./cmd/assemble db # rebuild
go -C assembler run ./cmd/assemble verify /tmp/before .build/public/db
```
Each side is a directory of `<id>.sqlite30`; a gzipped database from before the
switch to range requests is still expanded to a temporary file automatically. It exits non-zero on any mismatch,
Each side is a directory of `<id>.sqlite3`; a database under an older name, or
a gzipped one, is still opened automatically. It exits non-zero on any mismatch,
names the first differing rows and columns, and fails rather than skipping when a
dataset is absent from either side — silently comparing one of two datasets is
how a gate passes without proving anything.
+13 -22
View File
@@ -83,26 +83,18 @@ artifact — one missing line away from publishing it.
files committed to a repository; the databases are built in CI and uploaded
as a Pages artifact, and the documented Pages limits are a 1 GB published
site and 100 GB/month of bandwidth, with no per-file figure. The two
databases are 288 MB and 238 MB.
- **Total artifact is about 526 MB**, inside the 1 GB site limit but with less
headroom than before: a third dataset of this size would not fit. The fallback
is `sql.js-httpvfs`'s chunked mode, which splits a database into parts.
- **Pages does compress the databases, and that is survivable.** The extension
is unknown to Pages, so the file is served as `application/octet-stream`,
which is marked compressible in `mime-db` and gzipped: a plain request
returns `Content-Encoding: gzip` and the compressed length. Ranged reads are
not affected, because the Fetch standard makes browsers send
`Accept-Encoding: identity` on any request carrying a `Range` header. Only
the length probe breaks, so the site supplies the length itself instead of
trusting HEAD — see `web/src/lib/db-probe.js` and the chunked-mode note in
[system-architecture](./system-architecture.md).
- **Verify the way a browser asks.** A bare `curl -sI` advertises no encoding
and so reports success whatever the host does; it is what let this reach
production. Check ranged reads instead, and check the bytes, not the headers:
databases are 142 MB and 119 MB.
- **Total artifact is about 262 MB**, comfortably inside the 1 GB site limit.
Dropping the search index the range-request design needed halved both files.
- **Pages compresses the databases on the wire, which is a benefit here.** The
extension is unknown to Pages, so the file is served as
`application/octet-stream`, which `mime-db` marks compressible: 142 MB stored
becomes about 31 MB delivered, and the browser decompresses it transparently.
- **Verify the bytes, not the headers.** A bare `curl -sI` reports success
whatever the host does. Read the file's first bytes instead:
```bash
curl -s -r 0-15 -H 'Accept-Encoding: identity;q=1, *;q=0' \
https://<user>.github.io/thptqg/db/2016.sqlite30 | head -c 16
curl -s -r 0-15 https://<user>.github.io/thptqg/db/2016.sqlite3 | head -c 16
# must print: SQLite format 3
```
@@ -118,8 +110,7 @@ run rebuilds the older state. There is no data to migrate.
| Blank page, 404 on assets | `paths.base` in `svelte.config.js` does not match the repo name |
| `Failed to fetch database: 404` | Dataset id in `datasets.json` does not match the file in `db/` |
| A route 404s | The site step did not run, or the id is missing from `datasets.json` |
| `Length of the file not known` | The host gzipped the un-ranged response, so HEAD reports the compressed size. The site supplies `databaseLengthBytes` from a range probe; if this returns, the config is no longer reaching the worker in chunked mode |
| Database fails to open | A ranged read did not return raw database bytes. The range check above must print `SQLite format 3` |
| Every query is slow or huge | It is not using an index. `EXPLAIN QUERY PLAN` it: a `SCAN` means the browser is fetching the whole table |
| Deploy fails on assembly | An uncompressed database artefact reached the output; the error names the files |
| Download gate never finishes | The file is not being served, or the tab ran out of memory holding it. The check above must print `SQLite format 3` |
| Tab crashes on a phone | The database needs 142 MB of memory for 2016, 119 MB for 2017; a low-memory device may have the tab killed |
| Deploy fails on assembly | A stray database artefact reached the output — a journal, a `.gz`, or a name from an earlier design; the error names the files |
| Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints |
+1 -1
View File
@@ -22,7 +22,7 @@ running entirely in the browser and hosted for free on GitHub Pages. Covers the
## Constraints
- **Zero backend.** The database (238–289 MB per dataset) stays on the server
and the browser reads the pages a query touches over HTTP range requests.
and the browser downloads it once and queries it in memory.
- **Read-only.** `INSERT`/`UPDATE`/`DELETE` are rejected, so nobody is misled
into thinking edits persist. The file is fetched, never written.
- **Row caps.** 100 rows for lookups, 1000 for custom SQL, to prevent browser
+49 -50
View File
@@ -1,9 +1,13 @@
# System Architecture
Static site, no backend. The SQLite file stays on the server and the browser
reads the pages a query touches over HTTP range requests, via `sql.js-httpvfs`
(SQLite compiled to WebAssembly behind a virtual file system). A lookup costs a
few hundred KB; nothing downloads the database.
Static site, no backend. The browser downloads the whole SQLite file once and
queries it in memory with `sql.js` (SQLite compiled to WebAssembly). A dataset
costs about 31 MB to fetch and 142 MB of memory to hold; every query after that
is local — an exam-number lookup is immediate, and a name search scans all
877,460 rows in a few hundred milliseconds.
That is a reversal of the previous design, which read the file where it lay
over HTTP range requests. See [Considered and not taken](#considered-and-not-taken).
One frontend, one parser, one schema, two datasets.
@@ -17,11 +21,11 @@ through. `assembler/` sequences everything from the parser onwards.
data/<id>/*.xls(x)
│
▼ parser/ (Go, one binary, one config per dataset)
.build/public/db/<id>.sqlite30
.build/public/db/<id>.sqlite3
│
▼ assembler/ — row count and size must match datasets.json
.build/public/db/<id>.sqlite30 (uncompressed: ranges of a gzip stream
│ are not ranges of the database)
.build/public/db/<id>.sqlite3 (uncompressed: the host gzips it on the
│ wire, so a .gz would decompress twice)
▼ assembler/ → npm run build (SvelteKit static, assets = .build/public)
web/dist/
│
@@ -29,7 +33,7 @@ data/<id>/*.xls(x)
_site/ → GitHub Pages
│
▼ browser
sql.js-httpvfs asks for pages → HTTP range requests → results client-side
download gate → whole file → sql.js in memory → queries never leave the tab
```
## The dataset id
@@ -37,7 +41,7 @@ data/<id>/*.xls(x)
One identifier ties the whole pipeline together:
```
data/2017/ → parser/configs/2017.yml → db/2017.sqlite30 → /thptqg/2017/
data/2017/ → parser/configs/2017.yml → db/2017.sqlite3 → /thptqg/2017/
```
`datasets.json` at the repository root declares the ids once, with the row count
@@ -170,56 +174,51 @@ total descending.
| Concern | Choice | Rationale |
| --- | --- | --- |
| Storage | Static SQLite file, read by range request | No backend; the datasets are frozen, and a lookup needs a few pages of them |
| Reading mode | `serverMode: "chunked"` over a single chunk | The only mode whose config accepts the file length. In full mode the worker hardcodes it to `undefined` and falls back to a HEAD request, which Pages answers with the gzipped size. One chunk means the index is always 0, hence the published name `<id>.sqlite30` |
| Compression | None | A byte range of a gzip stream is not a byte range of the database |
| WASM hosting | Bundled with the app | `sql.js-httpvfs` ships its own build; one less third-party runtime dependency |
| Diacritics search | Pre-computed `ho_ten_ascii`, indexed word by word | `LOWER(REPLACE(...))` at query time defeats the index, and `LIKE '%x%'` reads the whole table |
| Row count in the footer | Read from `datasets.json` | `COUNT(*)` scans an index — 20 MB over range requests |
| Page size | 4 KiB, matched by `requestChunkSize` | One HTTP request is one page, and the worker issues them serially over synchronous XHR at ~40 ms each. Request count, not size, is what a search waits on: 390 requests moved 608 KB in 17 s. Four times the page is a quarter of the pages per scan. The 1 KiB both httpvfs libraries suggest optimises bytes per seek instead |
| SQL safety | Leading-keyword allowlist | `sql.js` is in-memory so writes cannot persist; the allowlist prevents confusion |
| Storage | Static SQLite file, downloaded whole | No backend; the datasets are frozen, and one large transfer is something browsers and CDNs are both good at |
| Download gate | Blocking, no dismiss | The page has no answers before the file arrives, and 31 MB of someone's mobile data should be asked for rather than spent silently |
| Compression | None published | The host gzips on the wire, so a `.gz` artifact would only be decompressed twice |
| WASM hosting | Bundled with the app | `sql.js` ships its own build; one less third-party runtime dependency |
| Diacritics search | Pre-computed `ho_ten_ascii` | `LOWER(REPLACE(...))` at query time is far slower over 877,460 rows than a column computed once at build |
| Secondary indexes | None | In memory a full scan costs a few hundred milliseconds; an index costs every visitor megabytes of download. The name index alone was 146 MB |
| Row count in the footer | Read from `datasets.json` | Costs nothing and is the same number the assembler enforces |
| Page size | 4 KiB | SQLite's default, and nothing on the client depends on it any more |
| SQL safety | Leading-keyword allowlist | The copy is the visitor's own, so this guards their session against a typo rather than protecting data |
| Row caps | 100 (lookup), 1000 (SQL) | Keeps DOM render sizes reasonable |
| Routing | SvelteKit file routes, prerendered | Each dataset gets a real HTML file with its own title |
| Styling | Tailwind, with tier colours as CSS variables | Tier classes are chosen at runtime, which no utility generator can see |
### Considered and not taken
- **Splitting the database into several chunks.** The site uses chunked mode,
but over one chunk (see above). Real splitting would let a CDN cache each part
whole; GitHub Pages serves everything with `Cache-Control: max-age=600`, and
every rebuild relays SQLite's pages so the file changes even when the data
does not, so that caching is cancelled by the host. Worth revisiting behind a
CDN with long TTLs, and it is the fallback if a single 288 MB file ever
becomes a problem.
- **`sqlite-wasm-http`.** Maintained, and built on the official SQLite WASM
rather than a 2022 fork, which is the better long-term footing. It does not
help here: its worker sizes the file from a HEAD request's `Content-Length`
exactly as `sql.js-httpvfs` does, and its `Options` has no field for the
length, so on Pages it would silently take the gzipped size instead of
failing. Its shared-cache backend needs COOP/COEP, which Pages cannot send,
but it ships a fallback backend that does not — so isolation is not the
blocker, the missing length option is.
- **Substring name search.** `LIKE '%x%'` cannot use an index, so it read the
whole student table. `name_word` keeps search by any word of a name without
it.
- **Reading the file over HTTP range requests** (`sql.js-httpvfs`), which this
site did until it was measured. Two costs killed it. Reads are serial —
the worker uses synchronous XHR — so a name search that touched 390 pages
waited 17 seconds to move 608 KB, and roughly one request per result row is a
floor no page size removes. And the first visitor after each deploy waited
~26 s for the CDN to fill its cache with a 288 MB object. The download pays
once, up front, visibly.
- **`sqlite-wasm-http`.** Built on the official SQLite WASM and maintained,
but it sizes the file from a HEAD request's `Content-Length` and exposes no
option to override it, so on a host that gzips it would silently use the
compressed size. Same class of problem, less recourse.
- **DuckDB-WASM over Parquet.** Genuinely maintained, async, parallel range
requests. Its binaries are 32–37 MB before the Parquet extension, which is
more than the entire database download for a phone looking up one score.
- **Static pre-generated shards**, one file per exam-number bucket. The most
robust option and the fastest single lookup, but it cannot answer arbitrary
SQL, and a static file cannot stop early the way `LIMIT` does — a common
Vietnamese surname would mean fetching a very large posting list.
## Risks and limitations
- **Unindexed queries are expensive.** The SQL tab can express a query that
walks the table, which over range requests means fetching 100+ MB. A byte
budget stops one before it gets that far, and the tab warns before it opens.
- **`Content-Encoding` on a ranged response would break everything.** A range
of a compressed body addresses the wrong bytes. In practice browsers prevent
it: the Fetch standard requires `Accept-Encoding: identity` on any request
carrying a `Range` header. GitHub Pages *does* gzip the un-ranged response —
`application/octet-stream` is compressible in `mime-db` — which is why the
file length is probed with a range request and passed as
`databaseLengthBytes` rather than left to the library's HEAD.
`db-probe.js` checks the returned bytes
start with the SQLite magic, so a host that ever compresses a ranged response
fails loudly instead of returning nonsense.
- **`sql.js-httpvfs` is unmaintained** (0.8.12, September 2022) and ships its
own SQLite WASM. `sqlite-wasm-http`, on the official build, is the fallback.
- **Memory is the binding constraint.** The database lives in the tab's
WebAssembly memory for as long as the page is open: 142 MB for 2016, 119 MB
for 2017. A low-memory phone may have the tab killed, which is why the gate
states the figure before the download starts.
- **The download is repeated every visit.** Nothing is persisted — a reload
fetches the file again. Caching it in IndexedDB or the Cache API would fix
that and has not been done.
- **A visitor who will not download cannot use the site.** That is the
deliberate shape of the gate, and it makes the first impression a 31 MB ask.
- **Hosted size.** 526 MB for both datasets against the 1 GB GitHub Pages
limit; a third dataset of this size would not fit.
- **Excel format drift.** A new source file with an unseen header layout needs a
+10 -52
View File
@@ -18,17 +18,17 @@ import "regexp"
// DDL is executed verbatim after the output database is (re)created.
//
// Every index here is chosen for a database read over HTTP range requests,
// where an unindexed query downloads the table. The rules that follow from
// that:
// One table and no secondary indexes. The browser downloads this file and
// queries it in memory, so an index buys a scan that already takes a few
// hundred milliseconds while costing tens of megabytes that every visitor
// pays for on the network.
//
// - No index on ho_ten or ho_ten_ascii. Neither substring nor prefix LIKE can
// use one (SQLite's LIKE optimisation needs a NOCASE index or
// case_sensitive_like), so both scanned the whole table. name_word replaces
// them.
// - idx_ten_cum_thi is partial, so it holds zero entries on the 2017 dataset
// — where the column is always NULL — while serving the 2016 cluster
// grouping. Partial indexes are SQLite-specific.
// That is a reversal. An earlier design read the file over HTTP range
// requests, where any unindexed query pulls the whole table down, and it
// carried a name_word table — one row per word of every name, ~3.5 million of
// them — plus partial indexes on the score columns. Those made range-request
// queries seek instead of scan, and together they were more than half the
// published file.
//
// This text is frozen: it decides the shape of every database the parser
// produces. TestDDLIsFrozen holds an independent copy so any edit has to be
@@ -58,48 +58,6 @@ CREATE TABLE student (
tieng_nhat REAL,
tieng_trung REAL
);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
CREATE TABLE name_word (
word TEXT NOT NULL,
so_bao_danh TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
PRIMARY KEY (word, so_bao_danh)
) WITHOUT ROWID;
CREATE TABLE name_word_freq (
word TEXT PRIMARY KEY,
n INTEGER NOT NULL
) WITHOUT ROWID;
`
// PostLoadSQL runs once the student rows are in, before VACUUM.
//
// The frequency table is what lets the site pick which word of a query to seek
// on: the vocabulary is about 4,400 words and the rarest word of a real query
// matches a few hundred rows, so seeking on it and filtering the rest inside
// name_word keeps a search to a few hundred kilobytes.
//
// The three score indexes are partial for the same reason idx_ten_cum_thi is:
// each covers only the exam year that has the column, and each costs about
// 4 MB. They exist so the SQL presets that rank by these columns seek instead
// of scanning 127 MB.
const PostLoadSQL = `
INSERT INTO name_word_freq (word, n)
SELECT word, COUNT(*) FROM name_word GROUP BY word;
CREATE INDEX idx_toan ON student(toan) WHERE toan IS NOT NULL;
CREATE INDEX idx_khtn ON student(khtn) WHERE khtn IS NOT NULL;
CREATE INDEX idx_khxh ON student(khxh) WHERE khxh IS NOT NULL;
`
// NameWordInsertSQL adds one row per distinct word of a candidate's ASCII name.
//
// ho_ten_ascii is carried along deliberately: a query with several words seeks
// on the rarest one and filters the others against this copy, so the whole
// match happens inside one b-tree and only the surviving rows are read from
// student.
const NameWordInsertSQL = `
INSERT OR IGNORE INTO name_word (word, so_bao_danh, ho_ten_ascii) VALUES (?, ?, ?)
`
// IdentityFields are the identity columns, in INSERT parameter order.
+6 -38
View File
@@ -109,50 +109,18 @@ CREATE TABLE student (
tieng_nhat REAL,
tieng_trung REAL
);
CREATE INDEX idx_ten_cum_thi ON student(ten_cum_thi) WHERE ten_cum_thi IS NOT NULL;
CREATE TABLE name_word (
word TEXT NOT NULL,
so_bao_danh TEXT NOT NULL,
ho_ten_ascii TEXT NOT NULL,
PRIMARY KEY (word, so_bao_danh)
) WITHOUT ROWID;
CREATE TABLE name_word_freq (
word TEXT PRIMARY KEY,
n INTEGER NOT NULL
) WITHOUT ROWID;
`
if DDL != want {
t.Errorf("DDL changed\n--- got ---\n%s\n--- want ---\n%s", DDL, want)
}
}
// TestNoIndexOnNameColumns: the databases are read over HTTP range requests, so
// an index that no query can use is dead weight in a file the browser pages
// through. Neither substring nor prefix LIKE can use one on these columns —
// name_word is what serves name search.
func TestNoIndexOnNameColumns(t *testing.T) {
for _, dead := range []string{"idx_ho_ten ", "idx_ho_ten_ascii"} {
if strings.Contains(DDL, dead) {
t.Errorf("DDL creates %q, which no query plan can use", dead)
}
}
}
// TestPostLoadBuildsTheSearchTables: the frequency table is what lets a search
// pick which word to seek on, and the score indexes are what keep the SQL
// presets off a full scan.
func TestPostLoadBuildsTheSearchTables(t *testing.T) {
for _, want := range []string{
"INSERT INTO name_word_freq",
"CREATE INDEX idx_toan",
"CREATE INDEX idx_khtn",
"CREATE INDEX idx_khxh",
} {
if !strings.Contains(PostLoadSQL, want) {
t.Errorf("PostLoadSQL is missing %q", want)
}
// TestNoSecondaryIndexes: the browser downloads this file whole and queries it
// in memory, where a scan of 877,460 rows takes a few hundred milliseconds. An
// index would save that and cost every visitor tens of megabytes of download.
func TestNoSecondaryIndexes(t *testing.T) {
if strings.Contains(DDL, "CREATE INDEX") {
t.Error("DDL creates an index; every one of them is paid for on the network")
}
}
+6 -83
View File
@@ -14,7 +14,6 @@ import (
"fmt"
"os"
"path/filepath"
"strings"
"github.com/tiennm99/thptqg/parser/internal/schema"
"github.com/tiennm99/thptqg/parser/internal/sqlitedb"
@@ -46,19 +45,10 @@ func OpenDB(dbPath string) (*sql.DB, error) {
// Before the DDL, because a page size cannot change once a table exists —
// only the VACUUM in Finish could rewrite it, and only to this same value.
//
// SQLite's default, kept because the browser reads this file a page at a
// time over HTTP and what costs time is the number of requests, not their
// size. The library reads with synchronous XHR, so requests are strictly
// serial at about 40 ms each; a name search that made 390 of them took 17
// seconds to move 608 KB. Four times the page size is a quarter of the
// pages for the same scan, and a shallower b-tree for each seek.
//
// The 1 KiB both sql.js-httpvfs and sqlite-wasm-http suggest optimises the
// other way — fewest bytes per seeked row — which is the wrong trade when
// one request costs more than the kilobyte it carries.
//
// web/src/lib/sqlite.svelte.js must request the same size, and
// web/src/lib/db-probe.js warns when a published database disagrees.
// SQLite's own default, set explicitly so the published file does not
// change shape if that default ever does. The browser downloads this file
// whole and queries it in memory, so the page size no longer has to match
// anything on the client.
if _, err := db.Exec("PRAGMA page_size = 4096"); err != nil {
db.Close()
return nil, fmt.Errorf("set page size: %w", err)
@@ -131,77 +121,10 @@ type Stats struct {
Errors uint64
}
// BuildNameIndex fills name_word from the student rows, one entry per distinct
// word of each ASCII name.
// Finish compacts the database and prints the stats block.
//
// A second pass rather than a write alongside each insert: a repeated exam
// number replaces its earlier row, and the words of the row it replaced would
// otherwise stay behind pointing at a name that is no longer there.
func BuildNameIndex(db *sql.DB) error {
rows, err := db.Query("SELECT so_bao_danh, ho_ten_ascii FROM student")
if err != nil {
return fmt.Errorf("read names: %w", err)
}
defer rows.Close()
tx, err := db.Begin()
if err != nil {
return fmt.Errorf("begin name index: %w", err)
}
stmt, err := tx.Prepare(schema.NameWordInsertSQL)
if err != nil {
tx.Rollback()
return fmt.Errorf("prepare name index: %w", err)
}
var words uint64
seen := make(map[string]struct{}, 8)
for rows.Next() {
var sbd, ascii string
if err := rows.Scan(&sbd, &ascii); err != nil {
tx.Rollback()
return fmt.Errorf("scan name: %w", err)
}
clear(seen)
for _, w := range strings.Fields(ascii) {
if _, dup := seen[w]; dup {
continue
}
seen[w] = struct{}{}
if _, err := stmt.Exec(w, sbd, ascii); err != nil {
tx.Rollback()
return fmt.Errorf("insert name word: %w", err)
}
words++
}
}
if err := rows.Err(); err != nil {
tx.Rollback()
return fmt.Errorf("read names: %w", err)
}
if err := stmt.Close(); err != nil {
tx.Rollback()
return err
}
if err := tx.Commit(); err != nil {
return fmt.Errorf("commit name index: %w", err)
}
if _, err := db.Exec(schema.PostLoadSQL); err != nil {
return fmt.Errorf("post-load statements: %w", err)
}
fmt.Printf("Name index: %d words\n", words)
return nil
}
// Finish builds the derived tables, runs VACUUM and prints the stats block.
//
// VACUUM must run AFTER the transaction commits — SQLite refuses it inside one
// — and after the name index, so the file is laid out in one pass.
// VACUUM must run AFTER the transaction commits — SQLite refuses it inside one.
func Finish(db *sql.DB, dbPath string, st Stats) error {
if err := BuildNameIndex(db); err != nil {
return err
}
if _, err := db.Exec("VACUUM"); err != nil {
return fmt.Errorf("vacuum: %w", err)
}
+6 -15
View File
@@ -8,7 +8,7 @@
"name": "thptqg-web",
"version": "1.0.0",
"dependencies": {
"sql.js-httpvfs": "^0.8.12"
"sql.js": "^1.14.1"
},
"devDependencies": {
"@eslint/js": "^9.39.4",
@@ -2090,12 +2090,6 @@
"dev": true,
"license": "MIT"
},
"node_modules/comlink": {
"version": "4.4.2",
"resolved": "https://registry.npmjs.org/comlink/-/comlink-4.4.2.tgz",
"integrity": "sha512-OxGdvBmJuNKSCMO4NTl1L47VRp6xn2wG4F/2hYzB6tiCb709otOxtEYCSvK80PtjODfXXZu8ds+Nw5kVCjqd2g==",
"license": "Apache-2.0"
},
"node_modules/concat-map": {
"version": "0.0.1",
"dev": true,
@@ -3242,14 +3236,11 @@
"node": ">=0.10.0"
}
},
"node_modules/sql.js-httpvfs": {
"version": "0.8.12",
"resolved": "https://registry.npmjs.org/sql.js-httpvfs/-/sql.js-httpvfs-0.8.12.tgz",
"integrity": "sha512-lcEBc2q0psFRfdCx8Di22oUIkkv5MUIaVO/fGCj/Jjx6YQDKVylQEcjd7NSSbmINHTRwVkm/vWP8uuevT7Rkkw==",
"license": "Apache-2.0",
"dependencies": {
"comlink": "^4.3.0"
}
"node_modules/sql.js": {
"version": "1.14.1",
"resolved": "https://registry.npmjs.org/sql.js/-/sql.js-1.14.1.tgz",
"integrity": "sha512-gcj8zBWU5cFsi9WUP+4bFNXAyF1iRpA3LLyS/DP5xlrNzGmPIizUeBggKa8DbDwdqaKwUcTEnChtd2grWo/x/A==",
"license": "MIT"
},
"node_modules/stackback": {
"version": "0.0.2",
+2 -2
View File
@@ -3,7 +3,7 @@
"private": true,
"version": "1.0.0",
"type": "module",
"description": "Frontend — one SvelteKit app serving every dataset and the hub",
"description": "Frontend \u2014 one SvelteKit app serving every dataset and the hub",
"scripts": {
"dev": "vite dev",
"build": "vite build",
@@ -12,7 +12,7 @@
"lint": "eslint ."
},
"dependencies": {
"sql.js-httpvfs": "^0.8.12"
"sql.js": "^1.14.1"
},
"devDependencies": {
"@eslint/js": "^9.39.4",
+5 -19
View File
@@ -1,7 +1,5 @@
<script>
import { formatBytes, isBudgetError } from "$lib/sqlite.svelte";
const MAX_ROWS = 1000;
import { MAX_ROWS } from "$lib/database.svelte";
const PLACEHOLDER = "Nhập truy vấn SQL...\nVí dụ: SELECT * FROM student WHERE toan >= 9 LIMIT 10";
let { db, disabled = false, presets = [] } = $props();
@@ -29,8 +27,8 @@
const trimmed = queryStr.trim();
if (!trimmed) return;
// Read-only statements only. The database is remote and read-only anyway,
// so this guards the user's own session against a typo.
// Read-only statements only. The copy in memory is the visitor's own, so
// this guards their session against a typo rather than protecting data.
const upper = trimmed.toUpperCase();
if (!["SELECT", "PRAGMA", "EXPLAIN", "WITH"].some((kw) => upper.startsWith(kw))) {
queryError = "Chỉ hỗ trợ truy vấn đọc (SELECT, PRAGMA, EXPLAIN, WITH).";
@@ -47,17 +45,12 @@
running = true;
const start = performance.now();
try {
const result = await source.query(finalSql, [], "SQL tab");
const result = source.query(finalSql);
execTime = (performance.now() - start).toFixed(1);
rows = result.slice(0, MAX_ROWS);
columns = rows.length > 0 ? Object.keys(rows[0]) : [];
} catch (err) {
queryError = isBudgetError(err)
? "Truy vấn này phải đọc quá nhiều dữ liệu và đã bị dừng. Hãy thêm điều kiện lọc, " +
"hoặc dùng cột đã có chỉ mục (so_bao_danh, toan, khtn, khxh, ten_cum_thi)."
: err instanceof Error
? err.message
: String(err);
queryError = err instanceof Error ? err.message : String(err);
} finally {
running = false;
}
@@ -125,13 +118,6 @@
{#if execTime !== null}
<span class="text-sm text-ink-muted">{rows.length} kết quả · {execTime}ms</span>
{/if}
{#if db?.lastCost}
<!-- What this query cost, then what the session has cost so far. -->
<span class="text-sm text-ink-subtle">
Truy vấn này: {db.lastCost.requests} yêu cầu · {formatBytes(db.lastCost.bytes)}
· Phiên: {db.requests} yêu cầu · {formatBytes(db.bytesRead)}
</span>
{/if}
</div>
</form>
@@ -0,0 +1,68 @@
<script>
import { formatBytes } from "$lib/database.svelte";
/**
* The gate a visitor passes before the dataset page works.
*
* The whole database is downloaded and queried in the browser, so there is
* nothing to show until it arrives. Saying so plainly — with both numbers
* that matter, the transfer and the memory — is fairer than a spinner that
* silently spends 30 MB of someone's mobile data.
*
* It has no dismiss: without the database the page has no answers to give.
*/
let { db, transferBytes = null, onDownload } = $props();
const percent = $derived(Math.round(db.progress * 100));
</script>
<div
class="fixed inset-0 z-50 flex items-center justify-center bg-black/60 p-4"
role="dialog"
aria-modal="true"
aria-labelledby="gate-title"
>
<div class="w-full max-w-[480px] rounded-xl bg-surface p-6 shadow-xl">
<h2 id="gate-title" class="mb-2 text-xl font-semibold">Tải cơ sở dữ liệu</h2>
<p class="mb-4 text-[0.95rem] text-ink-muted">
Trang này tra cứu hoàn toàn trong trình duyệt, không có máy chủ. Bạn cần tải toàn bộ cơ sở dữ
liệu về máy một lần để bắt đầu.
</p>
<dl class="mb-5 grid grid-cols-2 gap-3 text-sm">
<div class="rounded-lg bg-surface-alt p-3">
<dt class="text-ink-muted">Dung lượng tải</dt>
<dd class="text-lg font-semibold">
{transferBytes ? formatBytes(transferBytes) : "~31 MB"}
</dd>
</div>
<div class="rounded-lg bg-surface-alt p-3">
<dt class="text-ink-muted">Bộ nhớ (RAM) cần</dt>
<dd class="text-lg font-semibold">{formatBytes(db.expectedBytes)}</dd>
</div>
</dl>
<p class="mb-5 text-sm text-ink-subtle">
Dữ liệu được giữ trong bộ nhớ của tab và sẽ mất khi bạn đóng trang, nên lần truy cập sau cần
tải lại. Không nên dùng trên thiết bị có ít bộ nhớ.
</p>
{#if db.error}
<p class="notice mb-4 bg-error-bg text-error-ink">Tải thất bại: {db.error}</p>
{/if}
{#if db.loading}
<div class="mb-2 h-2 w-full overflow-hidden rounded-full bg-surface-alt">
<div class="h-full bg-primary transition-[width]" style="width: {percent}%"></div>
</div>
<p class="text-center text-sm text-ink-muted" aria-live="polite">
Đang tải… {formatBytes(db.received)} / {formatBytes(db.expectedBytes)} ({percent}%)
</p>
{:else}
<button type="button" class="btn-primary w-full" onclick={onDownload}>
{db.error ? "Thử lại" : "Tải xuống và bắt đầu"}
</button>
{/if}
</div>
</div>
+154
View File
@@ -0,0 +1,154 @@
import initSqlJs from "sql.js";
import wasmUrl from "sql.js/dist/sql-wasm.wasm?url";
/**
* The database is downloaded once and then queried in memory.
*
* The previous design read it where it lay, a page at a time over HTTP range
* requests. That works, but every page read is a separate round trip and they
* are serial, so a search that touched 390 pages waited 17 seconds for 608 KB
* — and the first visitor after a deploy waited another 20 for the CDN to fill
* its cache with a 288 MB object. Downloading trades one large transfer, which
* a browser and a CDN are both good at, for queries that cost nothing.
*
* The cost of that trade is memory: SQLite runs inside WebAssembly, so the
* whole file sits in the tab's address space for as long as the page is open.
* That is what the gate on the dataset page tells the visitor before they
* commit to it.
*/
/** Rows to return before a query is considered too broad to render. */
export const MAX_ROWS = 1000;
let sqlPromise = null;
/** The SQLite WASM module, initialised once per tab. */
function sqlEngine() {
sqlPromise ??= initSqlJs({ locateFile: () => wasmUrl });
return sqlPromise;
}
/**
* One dataset's database: its download, and the queries that follow.
*
* `expectedBytes` is the uncompressed size from the registry. The server sends
* the file gzipped, so `Content-Length` counts compressed bytes while the
* stream yields decompressed ones; progress is measured against the registry
* figure instead, and the transfer size is reported separately by `weigh()`.
*/
export class LocalDatabase {
/** Ready to answer queries. */
ready = $state(false);
/** Downloading, or opening the file once it has arrived. */
loading = $state(false);
error = $state(null);
/** Uncompressed bytes received so far. */
received = $state(0);
/** How long the download and open took, once done. */
ms = $state(0);
#db = null;
constructor(url, expectedBytes) {
this.url = url;
this.expectedBytes = expectedBytes;
}
/** Fraction downloaded, 0 to 1, for the progress bar. */
get progress() {
if (!this.expectedBytes) return 0;
return Math.min(1, this.received / this.expectedBytes);
}
/**
* Fetch the database and open it.
*
* Read as a stream rather than with response.arrayBuffer() so the visitor
* sees movement: this is tens of megabytes on a phone connection, and a
* progress bar is the difference between waiting and giving up.
*/
async open() {
if (this.ready || this.loading) return;
this.loading = true;
this.error = null;
this.received = 0;
const started = performance.now();
try {
const [SQL, response] = await Promise.all([sqlEngine(), fetch(this.url)]);
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);
const reader = response.body.getReader();
const parts = [];
for (;;) {
const { done, value } = await reader.read();
if (done) break;
parts.push(value);
this.received += value.length;
}
const bytes = new Uint8Array(this.received);
let at = 0;
for (const part of parts) {
bytes.set(part, at);
at += part.length;
}
this.#db = new SQL.Database(bytes);
this.ready = true;
this.ms = performance.now() - started;
} catch (err) {
this.error = err instanceof Error ? err.message : String(err);
} finally {
this.loading = false;
}
}
/**
* Run a query and return its rows as objects.
*
* Synchronous: the database is local, and SQLite answers from memory. A
* scan of every row takes a few hundred milliseconds, which is why the
* schema carries no secondary indexes.
*/
query(sql, params = []) {
if (!this.#db) throw new Error("Cơ sở dữ liệu chưa sẵn sàng");
const statement = this.#db.prepare(sql);
try {
if (params.length > 0) statement.bind(params);
const rows = [];
while (statement.step() && rows.length < MAX_ROWS) rows.push(statement.getAsObject());
return rows;
} finally {
statement.free();
}
}
close() {
this.#db?.close();
this.#db = null;
this.ready = false;
}
}
/**
* What the download will cost on the wire, from the server's own headers.
*
* The browser advertises gzip, and GitHub Pages compresses this file, so this
* reports the compressed size — the number that matters to someone on mobile
* data. Returns null when the server does not say.
*/
export async function weigh(url) {
try {
const response = await fetch(url, { method: "HEAD" });
const length = Number(response.headers.get("Content-Length"));
return Number.isFinite(length) && length > 0 ? length : null;
} catch {
return null;
}
}
export function formatBytes(n) {
return n < 1024 * 1024 ? `${Math.round(n / 1024)} KB` : `${(n / 1048576).toFixed(0)} MB`;
}
+5 -31
View File
@@ -99,38 +99,12 @@ export function pathOf(dataset, base) {
}
/**
* The chunk the whole database lives in. One chunk covers the file, so this is
* the only index the library ever asks for, and the assembler publishes the
* database under a name ending in it.
*/
const CHUNK_INDEX = "0";
/**
* Everything a database URL has except the chunk index, e.g.
* dbPrefixOf(d, "/thptqg") → "/thptqg/db/2017.sqlite3".
* The published database, e.g. dbOf(d, "/thptqg") → "/thptqg/db/2017.sqlite3".
*
* sql.js-httpvfs reads the file in chunked mode and builds each URL as
* prefix + chunk index. One chunk holds the whole database, so the index is
* always 0 and the published file is "<id>.sqlite30".
*/
export function dbPrefixOf(dataset, base) {
return `${base}/db/${dataset.id}.sqlite3`;
}
/**
* The published database file, e.g. dbOf(d, "/thptqg") → "/thptqg/db/2017.sqlite30".
*
* Uncompressed on purpose: the browser reads byte ranges of it, and a range of
* a gzip stream is not a range of the database.
* The browser downloads this whole file and queries it in memory. The server
* compresses it on the wire, which is why the transfer is a third of the size
* the registry records.
*/
export function dbOf(dataset, base) {
return `${dbPrefixOf(dataset, base)}${CHUNK_INDEX}`;
}
/**
* Both forms of the database's location, for RemoteDatabase: the file to read,
* and the prefix the library appends the chunk index to.
*/
export function dbSourceOf(dataset, base) {
return { url: dbOf(dataset, base), urlPrefix: dbPrefixOf(dataset, base) };
return `${base}/db/${dataset.id}.sqlite3`;
}
+6 -20
View File
@@ -1,28 +1,14 @@
import { describe, expect, it } from "vitest";
import { DATASETS, dbOf, dbPrefixOf, dbSourceOf } from "./datasets.js";
import { DATASETS, dbOf, pathOf } from "./datasets.js";
const [dataset] = DATASETS;
/**
* sql.js-httpvfs builds every request URL as urlPrefix + chunk index, and one
* chunk holds the whole database, so the index is always 0. If these two ever
* stop agreeing, the library asks for a file the assembler never published and
* every query 404s — which no other test would catch.
*/
describe("database location", () => {
it("puts the chunk index where the published file name ends", () => {
expect(dbOf(dataset, "/thptqg")).toBe(`${dbPrefixOf(dataset, "/thptqg")}0`);
describe("dataset locations", () => {
it("derives the site path from the id", () => {
expect(pathOf(dataset, "/thptqg")).toBe(`/thptqg/${dataset.id}/`);
});
it("names the file the assembler publishes", () => {
expect(dbOf(dataset, "/thptqg")).toBe(`/thptqg/db/${dataset.id}.sqlite30`);
});
it("hands RemoteDatabase both forms of the same location", () => {
const source = dbSourceOf(dataset, "/thptqg");
expect(source).toEqual({
url: dbOf(dataset, "/thptqg"),
urlPrefix: dbPrefixOf(dataset, "/thptqg"),
});
it("names the database file the assembler publishes", () => {
expect(dbOf(dataset, "/thptqg")).toBe(`/thptqg/db/${dataset.id}.sqlite3`);
});
});
-88
View File
@@ -1,88 +0,0 @@
/**
* How long the database is, asked in the one way that survives a CDN.
*
* `sql.js-httpvfs` sizes a file with a HEAD request. That request carries no
* Range header, so the browser advertises gzip, and GitHub Pages answers with
* `Content-Encoding: gzip` and the length of the *compressed* body — 64 MB for
* a 288 MB database. The library rightly refuses to believe it and gives up
* with "Length of the file not known. It must either be supplied in the config
* or given by the HTTP server."
*
* Range requests do not have that problem: the Fetch standard requires
* `Accept-Encoding: identity` on any request carrying a Range header, so the
* page reads that do the actual work always come back uncompressed. Asking for
* the first hundred bytes therefore yields both a trustworthy total, from
* Content-Range, and the file header itself to check it against.
*/
/** A SQLite file opens with these characters and then a NUL byte. */
const MAGIC = "SQLite format 3";
/** Enough for the whole SQLite header. */
const PROBE_BYTES = 100;
/** Page size lives at offset 16, big-endian; the value 1 encodes 65536. */
const PAGE_SIZE_OFFSET = 16;
/** Total size of the representation, from `bytes <from>-<to>/<total>`. */
export function parseTotalBytes(contentRange) {
const match = /\/\s*(\d+)\s*$/.exec(contentRange ?? "");
if (!match) {
throw new Error(
`the server did not say how large the database is (Content-Range: ${contentRange ?? "absent"})`,
);
}
return Number(match[1]);
}
/** Page size the file was written with, from its header. */
export function readPageSize(header) {
const raw = (header[PAGE_SIZE_OFFSET] << 8) | header[PAGE_SIZE_OFFSET + 1];
return raw === 1 ? 65536 : raw;
}
/** True when these bytes begin a SQLite database. */
export function looksLikeSqlite(header) {
const text = Array.from(MAGIC).every((ch, i) => header[i] === ch.charCodeAt(0));
return text && header[MAGIC.length] === 0;
}
/**
* Read the file header over a range request and return the database's length.
*
* Doubles as the check that the host is serving raw database bytes: a body that
* does not start with the SQLite magic means something rewrote it in transit —
* compression being the way that happens — and every later page read would be
* reading the wrong bytes.
*/
export async function probeDatabase(url, expectedPageSize, fetchImpl = fetch) {
const response = await fetchImpl(url, { headers: { Range: `bytes=0-${PROBE_BYTES - 1}` } });
if (response.status !== 206) {
throw new Error(
`${url}: expected 206 for a range request, got ${response.status}. ` +
"The host must serve byte ranges of the database.",
);
}
const total = parseTotalBytes(response.headers.get("Content-Range"));
const header = new Uint8Array(await response.arrayBuffer());
if (!looksLikeSqlite(header)) {
throw new Error(
`${url}: the first bytes are not a SQLite header, so the host is not ` +
"serving the database as stored — check for Content-Encoding on ranged responses.",
);
}
const pageSize = readPageSize(header);
if (expectedPageSize && pageSize !== expectedPageSize) {
// Not fatal: it still reads correctly, just at more requests per page.
console.warn(
`[httpvfs] ${url} has page size ${pageSize}, but requests are ${expectedPageSize} bytes. ` +
"Every page read now spans more than one request.",
);
}
return total;
}
-90
View File
@@ -1,90 +0,0 @@
import { describe, expect, it, vi } from "vitest";
import { looksLikeSqlite, parseTotalBytes, probeDatabase, readPageSize } from "./db-probe.js";
/** The first bytes of a real database: magic, NUL, then the page size at 16. */
function header(pageSize = 1024, magic = "SQLite format 3") {
const bytes = new Uint8Array(100);
for (let i = 0; i < magic.length; i += 1) bytes[i] = magic.charCodeAt(i);
bytes[16] = pageSize >> 8;
bytes[17] = pageSize & 0xff;
return bytes;
}
function respond(bytes, { status = 206, contentRange = `bytes 0-99/${317096960}` } = {}) {
return {
status,
headers: { get: (name) => (name.toLowerCase() === "content-range" ? contentRange : null) },
arrayBuffer: async () => bytes.buffer,
};
}
describe("parseTotalBytes", () => {
it("takes the total from a Content-Range", () => {
expect(parseTotalBytes("bytes 0-99/317096960")).toBe(317096960);
});
it("refuses an unknown total", () => {
expect(() => parseTotalBytes("bytes 0-99/*")).toThrow(/how large/);
expect(() => parseTotalBytes(null)).toThrow(/absent/);
});
});
describe("readPageSize", () => {
it("reads the two big-endian bytes at offset 16", () => {
expect(readPageSize(header(1024))).toBe(1024);
expect(readPageSize(header(4096))).toBe(4096);
});
it("treats 1 as 65536, as the file format does", () => {
expect(readPageSize(header(1))).toBe(65536);
});
});
describe("looksLikeSqlite", () => {
it("accepts a real header", () => {
expect(looksLikeSqlite(header())).toBe(true);
});
it("rejects a gzip stream, which is what a compressing host returns", () => {
const gzip = new Uint8Array([0x1f, 0x8b, 0x08, 0x00, 0x00, 0x00, 0x00, 0x00]);
expect(looksLikeSqlite(gzip)).toBe(false);
});
it("requires the NUL that terminates the magic", () => {
expect(looksLikeSqlite(header(1024, "SQLite format 3x"))).toBe(false);
});
});
describe("probeDatabase", () => {
it("asks for the header by range and returns the total length", async () => {
const fetchImpl = vi.fn(async () => respond(header()));
const total = await probeDatabase("/db/2016.sqlite30", 1024, fetchImpl);
expect(total).toBe(317096960);
expect(fetchImpl).toHaveBeenCalledWith("/db/2016.sqlite30", {
headers: { Range: "bytes=0-99" },
});
});
it("fails when the host ignores the range", async () => {
const fetchImpl = async () => respond(header(), { status: 200 });
await expect(probeDatabase("/db/2016.sqlite30", 1024, fetchImpl)).rejects.toThrow(/expected 206/);
});
it("fails when the bytes are not a database", async () => {
const gzip = new Uint8Array([0x1f, 0x8b, 0x08]);
const fetchImpl = async () => respond(gzip);
await expect(probeDatabase("/db/2016.sqlite30", 1024, fetchImpl)).rejects.toThrow(
/not a SQLite header/,
);
});
it("warns, but continues, when the page size is not the request size", async () => {
const warn = vi.spyOn(console, "warn").mockImplementation(() => {});
const fetchImpl = async () => respond(header(4096));
await expect(probeDatabase("/db/2016.sqlite30", 1024, fetchImpl)).resolves.toBe(317096960);
expect(warn).toHaveBeenCalledWith(expect.stringMatching(/page size 4096/));
warn.mockRestore();
});
});
+18 -70
View File
@@ -4,92 +4,40 @@ import { toAscii } from "./to-ascii";
export const MAX_RESULTS = 100;
/**
* Name search over the name_word table.
* Name search over the downloaded database.
*
* A query is matched word by word, each word as a prefix, in any order — so
* "buu loc" finds "Nguyễn Bửu Lộc". The work is arranged so that only one word
* is ever seeked on and the rest are filtered inside the same b-tree:
* Each word of the query must appear at the start of a word in the candidate's
* ASCII name, in any order, so "buu loc" finds "Nguyễn Bửu Lộc". Prefixing the
* stored name with a space lets one pattern match the first word too.
*
* 1. ask name_word_freq how many entries each word prefix covers (the whole
* vocabulary is ~4,400 rows, so this is a couple of pages);
* 2. seek on the rarest one — for a real name that is a few hundred to a few
* thousand entries rather than the 300,000 a word like "thi" would walk;
* 3. filter the other words against the ho_ten_ascii copy carried in
* name_word, so nothing is read from student until a row has matched;
* 4. join to student for the rows that survive, at most MAX_RESULTS of them.
*
* Every step is an index seek. A search costs a few hundred KB.
* This scans every row. That is the deliberate trade of downloading the file:
* a scan of 877,460 rows in memory takes a few hundred milliseconds, and
* paying for it means the published database carries no name index — which was
* more than half its size.
*/
export async function searchByName(db, query) {
export function searchByName(db, query) {
const words = tokenise(query);
if (words.length === 0) return [];
const seek = await rarest(db, words);
const others = words.filter((w) => w !== seek);
// A word matches at a word boundary: the leading space makes the first word
// reachable by the same pattern as the rest.
const filters = others.map(() => `(' ' || w.ho_ten_ascii) LIKE ? ESCAPE '\\'`);
const sql = `
SELECT s.* FROM name_word w
JOIN student s ON s.so_bao_danh = w.so_bao_danh
WHERE w.word >= ? AND w.word < ?${filters.length ? " AND " + filters.join(" AND ") : ""}
LIMIT ${MAX_RESULTS}`;
const filters = words.map(() => `(' ' || ho_ten_ascii) LIKE ? ESCAPE '\\'`).join(" AND ");
return db.query(
sql,
[seek, upperBound(seek), ...others.map((w) => `% ${escapeLike(w)}%`)],
`name ${JSON.stringify(words.join(" "))} seeking ${JSON.stringify(seek)}`,
`SELECT * FROM student WHERE ${filters} LIMIT ${MAX_RESULTS}`,
words.map((w) => `% ${escapeLike(w)}%`),
);
}
/** Exact lookup by exam number: a primary-key seek, a few pages. */
export async function lookupExamId(db, id) {
return db.query(
"SELECT * FROM student WHERE so_bao_danh = ? LIMIT ?",
[normaliseExamId(id), MAX_RESULTS],
`exam id ${JSON.stringify(normaliseExamId(id))}`,
);
/** Exact lookup by exam number, straight down the primary key. */
export function lookupExamId(db, id) {
return db.query(`SELECT * FROM student WHERE so_bao_danh = ? LIMIT ${MAX_RESULTS}`, [
normaliseExamId(id),
]);
}
/** Fold to ASCII and split into words, the same shape name_word was built in. */
/** Fold to ASCII and split into words, the same shape ho_ten_ascii is stored in. */
export function tokenise(query) {
return toAscii(query).split(/\s+/).filter(Boolean);
}
/**
* The word whose prefix covers the fewest entries, which is the one worth
* seeking on. One round trip for all of them.
*/
async function rarest(db, words) {
if (words.length === 1) return words[0];
const sql = words
.map(() => "SELECT ? AS word, COALESCE(SUM(n), 0) AS n FROM name_word_freq WHERE word >= ? AND word < ?")
.join(" UNION ALL ");
const params = words.flatMap((w) => [w, w, upperBound(w)]);
const counts = await db.query(sql, params, "word frequencies");
let best = words[0];
let bestN = Infinity;
for (const { word, n } of counts) {
if (n < bestN) {
best = word;
bestN = n;
}
}
return best;
}
/**
* The exclusive upper bound of a prefix range. U+FFFF sorts above any character
* that can follow the prefix, which is what turns "starts with" into a range
* the index can seek.
*/
function upperBound(prefix) {
return prefix + "￿";
}
/** Escape the LIKE wildcards so a user typing % or _ searches for them. */
function escapeLike(s) {
return s.replace(/[\\%_]/g, (c) => `\\${c}`);
-199
View File
@@ -1,199 +0,0 @@
import { createDbWorker } from "sql.js-httpvfs";
import workerUrl from "sql.js-httpvfs/dist/sqlite.worker.js?url";
import wasmUrl from "sql.js-httpvfs/dist/sql-wasm.wasm?url";
import { probeDatabase } from "./db-probe.js";
/**
* The database is read where it lies. SQLite asks for pages, the virtual file
* system turns each into an HTTP range request, and only the pages a query
* touches ever cross the network — a few hundred KB for a lookup, against the
* 45 MB the whole file used to cost before the first query.
*
* That only holds while every query is index-driven. The schema exists for it:
* name_word serves name search, and the score indexes serve the SQL presets.
* An unindexed query walks the table and pulls all 100+ MB of it, which is what
* the byte budget below is for.
*/
// Must equal the page size the parser writes (PRAGMA page_size in
// parser/internal/writer/writer.go), so one request is exactly one page. A
// mismatch makes every logical page read span two requests.
//
// Reads are serial — the worker uses synchronous XHR — so the count of
// requests, not their size, is what a search waits on.
const CHUNK_BYTES = 4096;
/** Generous for indexed work: a name search costs well under 1 MB. */
export const SEARCH_BUDGET_BYTES = 25 * 1024 * 1024;
/** What the SQL tab gets once the user has accepted the cost of a scan. */
export const PLAYGROUND_BUDGET_BYTES = 250 * 1024 * 1024;
/**
* What one query cost over the network: `{ requests, bytes, ms }`.
*/
/**
* One remotely-paged database, with the load state the UI needs.
*
* `budgetBytes` is a hard ceiling for the worker's lifetime: past it a query
* fails instead of quietly downloading the file. Raising it means a new worker,
* which costs only the header pages.
*/
export class RemoteDatabase {
ready = $state(false);
error = $state(null);
/** Bytes fetched by this database so far, prefetch included. */
bytesRead = $state(0);
/** HTTP range requests issued so far. */
requests = $state(0);
/** What the most recent query cost on its own. */
lastCost = $state(null);
#worker = null;
#opening;
#closed = false;
// Previous cumulative reading, so a query's own cost is a subtraction.
#seen = { requests: 0, bytes: 0 };
/**
* `source` carries both forms of the same location: `url` is the file, and
* `urlPrefix` is what the library appends the chunk index to. They must
* agree — see dbOf/dbPrefixOf, which derive one from the other.
*/
constructor(source, budgetBytes = SEARCH_BUDGET_BYTES) {
this.url = source.url;
this.urlPrefix = source.urlPrefix;
this.budgetBytes = budgetBytes;
this.#opening = this.#open();
}
async #open() {
const opened = performance.now();
try {
// Chunked mode over a single chunk, which looks odd but is the only way
// to tell this library how long the file is: the worker reads
// databaseLengthBytes in chunked mode and hardcodes the length to
// undefined in full mode. Left to itself it sizes the file with a HEAD
// request, and GitHub Pages answers that with the gzipped length, which
// it then refuses to use.
//
// One chunk covers the whole database, so the chunk index is always 0
// and every request goes to urlPrefix + "0" — the file the assembler
// publishes as <id>.sqlite30.
const databaseLengthBytes = await probeDatabase(this.url, CHUNK_BYTES);
const worker = await createDbWorker(
[
{
from: "inline",
config: {
serverMode: "chunked",
urlPrefix: this.urlPrefix,
serverChunkSize: databaseLengthBytes,
databaseLengthBytes,
suffixLength: 1,
requestChunkSize: CHUNK_BYTES,
},
},
],
workerUrl,
wasmUrl,
this.budgetBytes,
);
if (this.#closed) throw new Error("closed");
this.#worker = worker;
this.ready = true;
// Opening is not free either: the header and schema pages are read before
// any query runs, and that shows up in every later session total.
await this.#account(worker, `open ${this.url}`, performance.now() - opened);
return worker;
} catch (err) {
if (!this.#closed) {
this.error = message(err);
this.ready = false;
}
throw err;
}
}
/**
* Run a query and return its rows as objects.
*
* `label` names the query in the console trace — the only way to see what a
* search actually costs, since the byte count depends on how much the read
* heads prefetched, not just on the pages the plan needed.
*/
async query(sql, params = [], label) {
const worker = this.#worker ?? (await this.#opening);
const started = performance.now();
try {
// One array argument, never spread. The worker forwards its arguments to
// sql.js's exec(sql, params), which takes the whole parameter list as its
// second argument and ignores anything that is not an array — leaving
// every ? unbound, which reads as NULL. The library's own type says
// `...params`, and it is wrong.
return await worker.db.query(sql, params);
} finally {
await this.#account(worker, label ?? firstLine(sql), performance.now() - started);
}
}
/**
* Read the cumulative counters and report the delta.
*
* getStats() rather than the worker's `bytesRead`: that one is the budget
* counter and resets itself to zero when a query trips the ceiling, so it
* would under-report exactly when the number matters most.
*/
async #account(worker, label, ms) {
const stats = await worker.worker.getStats().catch(() => null);
if (!stats) return;
const cost = {
requests: stats.totalRequests - this.#seen.requests,
bytes: stats.totalFetchedBytes - this.#seen.bytes,
ms,
};
this.#seen = { requests: stats.totalRequests, bytes: stats.totalFetchedBytes };
this.requests = stats.totalRequests;
this.bytesRead = stats.totalFetchedBytes;
this.lastCost = cost;
console.info(
`[httpvfs] ${label} — ${cost.requests} request(s), ${formatBytes(cost.bytes)}, ${ms.toFixed(0)} ms` +
` · session ${stats.totalRequests} request(s), ${formatBytes(stats.totalFetchedBytes)}` +
` of ${formatBytes(stats.totalBytes)}`,
);
}
/**
* Drop this database. createDbWorker owns the Worker and exposes no handle to
* it, so the thread outlives this call; a page creates at most one per
* dataset and one more if the SQL budget is raised, which is why that is
* tolerable rather than a leak worth working around.
*/
close() {
this.#closed = true;
this.#worker = null;
this.ready = false;
}
}
/** True when a query failed because it would have exceeded the byte budget. */
export function isBudgetError(err) {
return /maxBytesToRead|too much data|exceeded/i.test(message(err));
}
function message(err) {
return err instanceof Error ? err.message : String(err);
}
/** Enough of a query to recognise it in the console. */
function firstLine(sql) {
const line = sql.trim().split("\n")[0];
return line.length > 70 ? `${line.slice(0, 70)}…` : line;
}
export function formatBytes(n) {
return n < 1024 * 1024 ? `${Math.round(n / 1024)} KB` : `${(n / 1048576).toFixed(1)} MB`;
}
+28 -114
View File
@@ -3,18 +3,14 @@
import { base, resolve } from "$app/paths";
import { page } from "$app/state";
import CustomQuery from "$lib/components/custom-query.svelte";
import DownloadGate from "$lib/components/download-gate.svelte";
import ScoreTable from "$lib/components/score-table.svelte";
import SearchForm from "$lib/components/search-form.svelte";
import StudentDetail from "$lib/components/student-detail.svelte";
import { dbSourceOf } from "$lib/datasets";
import { LocalDatabase, weigh } from "$lib/database.svelte";
import { dbOf } from "$lib/datasets";
import { isExamId } from "$lib/query-mode";
import { MAX_RESULTS, lookupExamId, searchByName } from "$lib/search";
import {
PLAYGROUND_BUDGET_BYTES,
RemoteDatabase,
SEARCH_BUDGET_BYTES,
formatBytes,
} from "$lib/sqlite.svelte";
let { data } = $props();
const dataset = $derived(data.dataset);
@@ -23,27 +19,25 @@
let results = $state(null);
let searchError = $state(null);
let searching = $state(false);
let elapsedMs = $state(0);
/** What the last search cost over the network, for the line under the results. */
let cost = $state(null);
/** How long the last search took, in ms. */
let searchMs = $state(null);
let activeTab = $state("search");
let sqlWarningOpen = $state(false);
// Raised once the user has accepted that a hand-written query may fetch a lot.
let budget = $state(SEARCH_BUDGET_BYTES);
/** Compressed size of the download, as the server reports it. */
let transferBytes = $state(null);
// Owned here, not in SearchForm, so it can be bound to the URL both ways.
// The query string is unreadable while prerendering — there is no request —
// so a deep link is picked up on the client only.
let query = $state(browser ? (page.url.searchParams.get("q") ?? "") : "");
const opening = $derived(db !== null && !db.ready && db.error === null);
const loadError = $derived(db?.error ?? null);
const busy = $derived(!db?.ready);
// Opened in the browser only: $effect does not run while prerendering. A new
// budget means a new worker, which costs only the header pages.
// Created in the browser only: $effect does not run while prerendering. The
// download itself waits for the visitor to accept it.
$effect(() => {
const opened = new RemoteDatabase(dbSourceOf(dataset, base), budget);
const url = dbOf(dataset, base);
const opened = new LocalDatabase(url, dataset.dbSizeMb * 1024 * 1024);
db = opened;
void weigh(url).then((bytes) => (transferBytes = bytes));
return () => {
opened.close();
db = null;
@@ -78,76 +72,34 @@
query = q;
writeUrlQuery(q);
// Counted across the whole search rather than per query: a name search runs
// two, one for the word frequencies and one for the rows.
const before = { requests: source.requests, bytes: source.bytesRead };
const startedAt = performance.now();
// A name search scans every row, which is a few hundred milliseconds of
// blocked main thread. Yielding a frame first lets the spinner paint.
searching = true;
cost = null;
// The worker reads pages with synchronous XHR, so it cannot answer while a
// query runs and there is no request count to show until it finishes.
// Elapsed time is the one honest live signal.
elapsedMs = 0;
const ticking = setInterval(() => (elapsedMs = performance.now() - startedAt), 100);
searchMs = null;
const startedAt = performance.now();
await new Promise(requestAnimationFrame);
try {
results = isExamId(q) ? await lookupExamId(source, q) : await searchByName(source, q);
results = isExamId(q) ? lookupExamId(source, q) : searchByName(source, q);
} catch (err) {
searchError = err instanceof Error ? err.message : String(err);
} finally {
clearInterval(ticking);
searchMs = performance.now() - startedAt;
searching = false;
cost = {
requests: source.requests - before.requests,
bytes: source.bytesRead - before.bytes,
ms: performance.now() - startedAt,
sessionRequests: source.requests,
sessionBytes: source.bytesRead,
};
}
}
function clearSearch() {
results = null;
query = "";
cost = null;
searchMs = null;
searchError = null;
writeUrlQuery("");
}
/** "16,9 giây", or "0,4 giây" — one decimal is enough to compare searches. */
function seconds(ms) {
return `${(ms / 1000).toLocaleString("vi-VN", { minimumFractionDigits: 1, maximumFractionDigits: 1 })} giây`;
}
function openSqlTab() {
if (budget >= PLAYGROUND_BUDGET_BYTES) {
activeTab = "sql";
return;
}
sqlWarningOpen = true;
}
function acceptSqlWarning() {
sqlWarningOpen = false;
// Reopening with the larger budget throws away the current worker, and with
// it the pages it had cached — a few hundred KB, refetched on demand.
budget = PLAYGROUND_BUDGET_BYTES;
activeTab = "sql";
}
function declineSqlWarning() {
sqlWarningOpen = false;
activeTab = "search";
}
// Global shortcuts: Ctrl+Enter submits the SQL query, "/" focuses the search
// box unless the user is already typing somewhere.
function onKeydown(event) {
if (event.key === "Escape" && sqlWarningOpen) {
declineSqlWarning();
return;
}
if (event.ctrlKey && event.key === "Enter" && activeTab === "sql") {
document.querySelector(".query-form")?.requestSubmit();
return;
@@ -183,13 +135,8 @@
</header>
<main>
{#if opening}
<p class="my-4 text-center text-ink-muted">Đang mở cơ sở dữ liệu…</p>
{/if}
{#if loadError}
<p class="notice bg-error-bg text-error-ink">Lỗi: {loadError}</p>
{/if}
<!-- Load state belongs to the gate: until the database is here there is
nothing behind it to report on. -->
<div class="mx-auto mb-6 flex max-w-[600px] border-b-2 border-line">
<button
@@ -212,7 +159,7 @@
class:border-primary={activeTab === "sql"}
class:text-primary={activeTab === "sql"}
class:font-semibold={activeTab === "sql"}
onclick={openSqlTab}
onclick={() => (activeTab = "sql")}
>
Truy vấn SQL
</button>
@@ -234,7 +181,7 @@
border-t-primary"
aria-hidden="true"
></span>
Đang tra cứu… {seconds(elapsedMs)}
Đang tra cứu…
</p>
{/if}
@@ -254,13 +201,9 @@
</p>
{/if}
{#if cost && !searching}
{#if searchMs !== null && !searching}
<p class="mt-2 text-center text-xs text-ink-subtle">
Mạng: {cost.requests.toLocaleString("vi-VN")} yêu cầu · {formatBytes(cost.bytes)} · {seconds(
cost.ms,
)} · Cả phiên: {cost.sessionRequests.toLocaleString("vi-VN")} yêu cầu · {formatBytes(
cost.sessionBytes,
)}
Truy vấn tại chỗ trong {Math.round(searchMs)} ms
</p>
{/if}
{:else}
@@ -281,35 +224,6 @@
</footer>
</div>
{#if sqlWarningOpen}
<!--
The SQL tab is the one place a user can write a query that reads the whole
database. Everything else here is a seek; this is not, so it is opt-in.
-->
<div
class="fixed inset-0 z-10 flex items-center justify-center bg-black/50 p-4"
role="dialog"
aria-modal="true"
aria-labelledby="sql-warning-title"
>
<div class="max-w-[520px] rounded-xl border border-line bg-surface p-6 shadow-card">
<h2 id="sql-warning-title" class="mb-3 text-lg font-semibold">Truy vấn SQL tốn dữ liệu mạng</h2>
<p class="mb-3 text-sm text-ink-muted">
Tra cứu thường chỉ tải vài trăm KB. Truy vấn SQL tự viết có thể quét toàn bộ bảng và tải tới
hàng trăm MB — tốn dữ liệu di động và có thể rất chậm.
</p>
<p class="mb-5 text-sm text-ink-muted">
Cơ sở dữ liệu này nặng {dataset.dbSizeMb} MB. Số byte đã tải sẽ hiển thị bên cạnh thời gian
chạy để bạn theo dõi.
</p>
<div class="flex flex-wrap justify-end gap-3">
<button type="button" class="btn-chip" onclick={declineSqlWarning}>
Quay lại tra cứu
</button>
<button type="button" class="btn-primary" onclick={acceptSqlWarning}>
Tôi hiểu, tiếp tục
</button>
</div>
</div>
</div>
{#if db && !db.ready}
<DownloadGate {db} {transferBytes} onDownload={() => db.open()} />
{/if}