mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
The browser fetches this file one page per HTTP request, so the page size is the granularity of every read. At SQLite's 4 KiB default a row reached by an index seek dragged 4 KB across the network; at 1 KiB it drags 1 KB. A name search returns up to 100 scattered rows, so its row fetches fall from about 400 KB to about 100 KB. Measured on the rebuilt 2016 file: 6.3 rows share a page where 27 did. The index walks are sequential and unaffected in bytes — the library's read-ahead already collapses those into few requests. Cost is 4% file size: 2016 288.6 -> 302.4 MB, 2017 237.7 -> 247.3 MB, the site 528 -> 552 MB against the 1 GB GitHub Pages limit. Both sql.js-httpvfs and sqlite-wasm-http recommend this page size. The PRAGMA has to run before the DDL, since a page size is fixed once a table exists, and requestChunkSize on the client has to match or every page read spans two requests. Row counts unchanged and through the assembler guards; query plans re-checked and still index-driven on the rebuilt files.
3.4 KiB
3.4 KiB
Brainstorm: which httpvfs best practices to adopt
2026-08-14 13:17. Follows
the research report.
Branch feat/httpvfs-range-queries.
Problem
Research surfaced three upstream practices we do not follow. Decide which are worth the change, knowing nothing can be verified in a browser on this machine.
Codebase context (scout)
| Concern | Touch points |
|---|---|
| Page size | parser/internal/writer/writer.go:41-47 (PRAGMA must precede the DDL — page size is fixed once a table exists), :185 (VACUUM already applies it), web/src/lib/sqlite.svelte.ts:18,53 |
| Chunked mode | databases.go:41 Extension, size guard, Clean() deletes anything not <id>.sqlite3 — would eat every chunk, site.go:135-138,153, datasets.ts:91-97, datasets.json, ~15 tests |
| Library swap | sqlite.svelte.ts:1-3,50-58,76,82, package.json:15; all consumers go through RemoteDatabase, so the blast radius is one file |
Options evaluated
A. page_size 1024 + requestChunkSize 1024 — ADOPTED
- Upstream consensus: phiresky and mmomtchev both recommend 1024.
- Honest sizing for our pattern: row fetches 400 KB → 100 KB per search; the index walk is sequential so bytes are unchanged and only the request count rises, which prefetch read-heads collapse. Net ≈ 300 KB saved per search — bandwidth, not latency.
- Cost: ~10 lines; file size +5% (528 → ~555 MB total, still under the 1 GB Pages limit); both databases rebuilt.
B. serverMode chunked — REJECTED
- Only benefit is CDN cache efficiency. GitHub Pages serves everything with
Cache-Control: max-age=600, and each deploy relays out SQLite pages anyway, so cross-deploy caching is zero either way. - Cost: split step, config JSON,
Clean()/guards/dbOf()/datasets.jsonrework, ~15 tests. - Complexity buying a benefit the host cancels. Revisit only behind a CDN with long TTLs.
C. swap to sqlite-wasm-http — DEFERRED
- For: maintained (Dec 2025), official SQLite WASM instead of a 2022 fork; matches this repo's posture on stale dependencies. Swap is one file.
- Against: its differentiator (shared cache) needs COOP/COEP headers GitHub Pages cannot send, so we would get the synchronous fallback and the maintenance benefit only. And the current integration has never run in a browser — swapping now means two unverified variables and no way to tell which broke.
- Revisit after the current build is verified live.
Decision
Adopt A only.
Implementation notes
PRAGMA page_size = 1024inwriter.OpenDB, betweensql.Openanddb.Exec(schema.DDL). The existing VACUUM inFinishapplies it.CHUNK_BYTES = 1024inweb/src/lib/sqlite.svelte.ts; it feedsrequestChunkSizeand must equal the page size.- Rebuild both databases, update
dbSizeMbindatasets.jsonto the new sizes, re-run the assembler guards.
Risks
- The benefit is arithmetic plus upstream authority, not measurement. The byte counter in the SQL tab is the check, once deployed.
- Prefetch deliberately overfetches ahead of the cursor, so the 25 MB search budget may trip earlier than a strict page count suggests. Tune after a real measurement, not before.
Unresolved questions
- Does 1024 actually beat 4096 for our queries in a browser?
- Is Fastly's caching of ranges over a 300 MB object good enough that chunked mode stays unnecessary?