mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-11 03:13:48 +00:00
Merge pull request #13 from tiennm99/fix/httpvfs-file-length
fix(web): take the database length from a range probe, not the host
This commit is contained in:
8 files changed
+256
-14
No files matched your search
@@ -121,3 +121,30 @@ jobs:
|
||||
steps:
|
||||
- id: deployment
|
||||
uses: actions/deploy-pages@v5
|
||||
|
||||
- uses: actions/checkout@v7
|
||||
|
||||
# The site reads the databases a page at a time over range requests, so
|
||||
# what matters is not that the host leaves the file alone in general — it
|
||||
# gzips the un-ranged response, and browsers work around that by sending
|
||||
# Accept-Encoding: identity whenever a request carries a Range header —
|
||||
# but that a ranged read returns raw database bytes. Checking headers is
|
||||
# what missed this before: a bare `curl -sI` advertises no encoding and
|
||||
# so passes whatever the host does. Check the bytes instead.
|
||||
- name: Verify ranged reads return database bytes
|
||||
env:
|
||||
PAGE_URL: ${{ steps.deployment.outputs.page_url }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
for id in $(jq -r '.datasets[].id' datasets.json); do
|
||||
url="${PAGE_URL%/}/db/${id}.sqlite3"
|
||||
if ! magic=$(curl -sf -r 0-14 -H 'Accept-Encoding: identity;q=1, *;q=0' "$url"); then
|
||||
echo "::error::$url is not fetchable"
|
||||
exit 1
|
||||
fi
|
||||
if [ "$magic" != "SQLite format 3" ]; then
|
||||
echo "::error::$url did not return database bytes over a range request"
|
||||
exit 1
|
||||
fi
|
||||
echo "$url: SQLite format 3"
|
||||
done
|
||||
@@ -61,8 +61,15 @@ hashes every real input file. That is the point of it; do not skip it.
|
||||
assembler refuses to publish an artifact that falls below a ratio of it.
|
||||
- **The databases ship uncompressed, as `<id>.sqlite3`.** The browser reads
|
||||
byte ranges of them, and a range of a gzip stream is not a range of the
|
||||
database. The host must not apply `Content-Encoding` either — check with
|
||||
`curl -sI` after a deploy.
|
||||
database.
|
||||
- **The file length comes from a range request, not from the host's HEAD.**
|
||||
GitHub Pages gzips `application/octet-stream`, so a HEAD reports the
|
||||
compressed size and `sql.js-httpvfs` refuses to open the file. Ranged reads
|
||||
are unaffected — browsers must send `Accept-Encoding: identity` whenever a
|
||||
request carries a `Range` header — so `web/src/lib/db-probe.js` reads the
|
||||
header over a range and passes `fileLength`. Verify the way a browser asks:
|
||||
`curl -sI -r 0-99 -H 'Accept-Encoding: identity' …`, never a bare `curl -sI`,
|
||||
which advertises no encoding and hides the problem.
|
||||
- **Every query the site runs must be index-driven.** Over range requests an
|
||||
unindexed query fetches the whole table. Hence no index on `ho_ten` (nothing
|
||||
can use one), `name_word` for name search, partial indexes for the score
|
||||
|
||||
@@ -87,10 +87,23 @@ artifact — one missing line away from publishing it.
|
||||
- **Total artifact is about 552 MB**, inside the 1 GB site limit but with less
|
||||
headroom than before: a third dataset of this size would not fit. The fallback
|
||||
is `sql.js-httpvfs`'s chunked mode, which splits a database into parts.
|
||||
- **The server must not compress the databases.** Ranges of a compressed body
|
||||
address the wrong bytes, and the library refuses to open a file whose HEAD
|
||||
carries a `Content-Encoding`. `.sqlite3` is an unknown type to Pages, so it is
|
||||
served as `application/octet-stream` and left alone — verify after a deploy.
|
||||
- **Pages does compress the databases, and that is survivable.** `.sqlite3` is
|
||||
unknown to Pages, so it is served as `application/octet-stream`, which is
|
||||
marked compressible in `mime-db` and gzipped: a plain request returns
|
||||
`Content-Encoding: gzip` and the compressed length. Ranged reads are not
|
||||
affected, because the Fetch standard makes browsers send
|
||||
`Accept-Encoding: identity` on any request carrying a `Range` header. Only
|
||||
the length probe breaks, and the site supplies `fileLength` itself instead of
|
||||
trusting HEAD — see `web/src/lib/db-probe.js`.
|
||||
- **Verify the way a browser asks.** A bare `curl -sI` advertises no encoding
|
||||
and so reports success whatever the host does; it is what let this reach
|
||||
production. Check ranged reads instead, and check the bytes, not the headers:
|
||||
|
||||
```bash
|
||||
curl -s -r 0-15 -H 'Accept-Encoding: identity;q=1, *;q=0' \
|
||||
https://<user>.github.io/thptqg/db/2016.sqlite3 | head -c 16
|
||||
# must print: SQLite format 3
|
||||
```
|
||||
|
||||
## Rollback
|
||||
|
||||
@@ -104,7 +117,8 @@ run rebuilds the older state. There is no data to migrate.
|
||||
| Blank page, 404 on assets | `paths.base` in `svelte.config.js` does not match the repo name |
|
||||
| `Failed to fetch database: 404` | Dataset id in `datasets.json` does not match the file in `db/` |
|
||||
| A route 404s | The site step did not run, or the id is missing from `datasets.json` |
|
||||
| Database fails to open | The host compressed it. `curl -sI …/db/<id>.sqlite3` must show no `content-encoding`; ranges of a compressed body are unusable |
|
||||
| `Length of the file not known` | The host gzipped the un-ranged response, so HEAD reports the compressed size. The site supplies `fileLength` from a range probe; if this returns, that probe failed |
|
||||
| Database fails to open | A ranged read did not return raw database bytes. The range check above must print `SQLite format 3` |
|
||||
| Every query is slow or huge | It is not using an index. `EXPLAIN QUERY PLAN` it: a `SCAN` means the browser is fetching the whole table |
|
||||
| Deploy fails on assembly | An uncompressed database artefact reached the output; the error names the files |
|
||||
| Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints |
|
||||
@@ -204,10 +204,15 @@ total descending.
|
||||
- **Unindexed queries are expensive.** The SQL tab can express a query that
|
||||
walks the table, which over range requests means fetching 100+ MB. A byte
|
||||
budget stops one before it gets that far, and the tab warns before it opens.
|
||||
- **`Content-Encoding` breaks everything.** If the host ever compresses
|
||||
`<id>.sqlite3` on the wire, ranges address compressed bytes and
|
||||
`sql.js-httpvfs` refuses to open the file. Verify after a deploy:
|
||||
`curl -sI …/db/2016.sqlite3` must show no `content-encoding`.
|
||||
- **`Content-Encoding` on a ranged response would break everything.** A range
|
||||
of a compressed body addresses the wrong bytes. In practice browsers prevent
|
||||
it: the Fetch standard requires `Accept-Encoding: identity` on any request
|
||||
carrying a `Range` header. GitHub Pages *does* gzip the un-ranged response —
|
||||
`application/octet-stream` is compressible in `mime-db` — which is why the
|
||||
file length is probed with a range request and passed as `fileLength` rather
|
||||
than left to the library's HEAD. `db-probe.js` checks the returned bytes
|
||||
start with the SQLite magic, so a host that ever compresses a ranged response
|
||||
fails loudly instead of returning nonsense.
|
||||
- **`sql.js-httpvfs` is unmaintained** (0.8.12, September 2022) and ships its
|
||||
own SQLite WASM. `sqlite-wasm-http`, on the official build, is the fallback.
|
||||
- **Hosted size.** 552 MB for both datasets against the 1 GB GitHub Pages
|
||||
|
||||
@@ -50,8 +50,9 @@ func OpenDB(dbPath string) (*sql.DB, error) {
|
||||
// file a page at a time over HTTP: a row fetched by index seek costs one
|
||||
// page, so a search that returns 100 scattered rows transfers 100 KB
|
||||
// instead of 400 KB. It costs about 5% file size, and both sql.js-httpvfs
|
||||
// and sqlite-wasm-http recommend it. web/src/lib/sqlite.svelte.ts must
|
||||
// request the same size.
|
||||
// and sqlite-wasm-http recommend it. web/src/lib/sqlite.svelte.js must
|
||||
// request the same size, and web/src/lib/db-probe.js warns when a published
|
||||
// database disagrees with it.
|
||||
if _, err := db.Exec("PRAGMA page_size = 1024"); err != nil {
|
||||
db.Close()
|
||||
return nil, fmt.Errorf("set page size: %w", err)
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
/**
|
||||
* How long the database is, asked in the one way that survives a CDN.
|
||||
*
|
||||
* `sql.js-httpvfs` sizes a file with a HEAD request. That request carries no
|
||||
* Range header, so the browser advertises gzip, and GitHub Pages answers with
|
||||
* `Content-Encoding: gzip` and the length of the *compressed* body — 66 MB for
|
||||
* a 302 MB database. The library rightly refuses to believe it and gives up
|
||||
* with "Length of the file not known. It must either be supplied in the config
|
||||
* or given by the HTTP server."
|
||||
*
|
||||
* Range requests do not have that problem: the Fetch standard requires
|
||||
* `Accept-Encoding: identity` on any request carrying a Range header, so the
|
||||
* page reads that do the actual work always come back uncompressed. Asking for
|
||||
* the first hundred bytes therefore yields both a trustworthy total, from
|
||||
* Content-Range, and the file header itself to check it against.
|
||||
*/
|
||||
|
||||
/** A SQLite file opens with these characters and then a NUL byte. */
|
||||
const MAGIC = "SQLite format 3";
|
||||
|
||||
/** Enough for the whole SQLite header. */
|
||||
const PROBE_BYTES = 100;
|
||||
|
||||
/** Page size lives at offset 16, big-endian; the value 1 encodes 65536. */
|
||||
const PAGE_SIZE_OFFSET = 16;
|
||||
|
||||
/** Total size of the representation, from `bytes <from>-<to>/<total>`. */
|
||||
export function parseTotalBytes(contentRange) {
|
||||
const match = /\/\s*(\d+)\s*$/.exec(contentRange ?? "");
|
||||
if (!match) {
|
||||
throw new Error(
|
||||
`the server did not say how large the database is (Content-Range: ${contentRange ?? "absent"})`,
|
||||
);
|
||||
}
|
||||
return Number(match[1]);
|
||||
}
|
||||
|
||||
/** Page size the file was written with, from its header. */
|
||||
export function readPageSize(header) {
|
||||
const raw = (header[PAGE_SIZE_OFFSET] << 8) | header[PAGE_SIZE_OFFSET + 1];
|
||||
return raw === 1 ? 65536 : raw;
|
||||
}
|
||||
|
||||
/** True when these bytes begin a SQLite database. */
|
||||
export function looksLikeSqlite(header) {
|
||||
const text = Array.from(MAGIC).every((ch, i) => header[i] === ch.charCodeAt(0));
|
||||
return text && header[MAGIC.length] === 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* Read the file header over a range request and return the database's length.
|
||||
*
|
||||
* Doubles as the check that the host is serving raw database bytes: a body that
|
||||
* does not start with the SQLite magic means something rewrote it in transit —
|
||||
* compression being the way that happens — and every later page read would be
|
||||
* reading the wrong bytes.
|
||||
*/
|
||||
export async function probeDatabase(url, expectedPageSize, fetchImpl = fetch) {
|
||||
const response = await fetchImpl(url, { headers: { Range: `bytes=0-${PROBE_BYTES - 1}` } });
|
||||
|
||||
if (response.status !== 206) {
|
||||
throw new Error(
|
||||
`${url}: expected 206 for a range request, got ${response.status}. ` +
|
||||
"The host must serve byte ranges of the database.",
|
||||
);
|
||||
}
|
||||
|
||||
const total = parseTotalBytes(response.headers.get("Content-Range"));
|
||||
const header = new Uint8Array(await response.arrayBuffer());
|
||||
|
||||
if (!looksLikeSqlite(header)) {
|
||||
throw new Error(
|
||||
`${url}: the first bytes are not a SQLite header, so the host is not ` +
|
||||
"serving the database as stored — check for Content-Encoding on ranged responses.",
|
||||
);
|
||||
}
|
||||
|
||||
const pageSize = readPageSize(header);
|
||||
if (expectedPageSize && pageSize !== expectedPageSize) {
|
||||
// Not fatal: it still reads correctly, just at more requests per page.
|
||||
console.warn(
|
||||
`[httpvfs] ${url} has page size ${pageSize}, but requests are ${expectedPageSize} bytes. ` +
|
||||
"Every page read now spans more than one request.",
|
||||
);
|
||||
}
|
||||
|
||||
return total;
|
||||
}
|
||||
@@ -0,0 +1,90 @@
|
||||
import { describe, expect, it, vi } from "vitest";
|
||||
import { looksLikeSqlite, parseTotalBytes, probeDatabase, readPageSize } from "./db-probe.js";
|
||||
|
||||
/** The first bytes of a real database: magic, NUL, then the page size at 16. */
|
||||
function header(pageSize = 1024, magic = "SQLite format 3") {
|
||||
const bytes = new Uint8Array(100);
|
||||
for (let i = 0; i < magic.length; i += 1) bytes[i] = magic.charCodeAt(i);
|
||||
bytes[16] = pageSize >> 8;
|
||||
bytes[17] = pageSize & 0xff;
|
||||
return bytes;
|
||||
}
|
||||
|
||||
function respond(bytes, { status = 206, contentRange = `bytes 0-99/${317096960}` } = {}) {
|
||||
return {
|
||||
status,
|
||||
headers: { get: (name) => (name.toLowerCase() === "content-range" ? contentRange : null) },
|
||||
arrayBuffer: async () => bytes.buffer,
|
||||
};
|
||||
}
|
||||
|
||||
describe("parseTotalBytes", () => {
|
||||
it("takes the total from a Content-Range", () => {
|
||||
expect(parseTotalBytes("bytes 0-99/317096960")).toBe(317096960);
|
||||
});
|
||||
|
||||
it("refuses an unknown total", () => {
|
||||
expect(() => parseTotalBytes("bytes 0-99/*")).toThrow(/how large/);
|
||||
expect(() => parseTotalBytes(null)).toThrow(/absent/);
|
||||
});
|
||||
});
|
||||
|
||||
describe("readPageSize", () => {
|
||||
it("reads the two big-endian bytes at offset 16", () => {
|
||||
expect(readPageSize(header(1024))).toBe(1024);
|
||||
expect(readPageSize(header(4096))).toBe(4096);
|
||||
});
|
||||
|
||||
it("treats 1 as 65536, as the file format does", () => {
|
||||
expect(readPageSize(header(1))).toBe(65536);
|
||||
});
|
||||
});
|
||||
|
||||
describe("looksLikeSqlite", () => {
|
||||
it("accepts a real header", () => {
|
||||
expect(looksLikeSqlite(header())).toBe(true);
|
||||
});
|
||||
|
||||
it("rejects a gzip stream, which is what a compressing host returns", () => {
|
||||
const gzip = new Uint8Array([0x1f, 0x8b, 0x08, 0x00, 0x00, 0x00, 0x00, 0x00]);
|
||||
expect(looksLikeSqlite(gzip)).toBe(false);
|
||||
});
|
||||
|
||||
it("requires the NUL that terminates the magic", () => {
|
||||
expect(looksLikeSqlite(header(1024, "SQLite format 3x"))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("probeDatabase", () => {
|
||||
it("asks for the header by range and returns the total length", async () => {
|
||||
const fetchImpl = vi.fn(async () => respond(header()));
|
||||
const total = await probeDatabase("/db/2016.sqlite3", 1024, fetchImpl);
|
||||
|
||||
expect(total).toBe(317096960);
|
||||
expect(fetchImpl).toHaveBeenCalledWith("/db/2016.sqlite3", {
|
||||
headers: { Range: "bytes=0-99" },
|
||||
});
|
||||
});
|
||||
|
||||
it("fails when the host ignores the range", async () => {
|
||||
const fetchImpl = async () => respond(header(), { status: 200 });
|
||||
await expect(probeDatabase("/db/2016.sqlite3", 1024, fetchImpl)).rejects.toThrow(/expected 206/);
|
||||
});
|
||||
|
||||
it("fails when the bytes are not a database", async () => {
|
||||
const gzip = new Uint8Array([0x1f, 0x8b, 0x08]);
|
||||
const fetchImpl = async () => respond(gzip);
|
||||
await expect(probeDatabase("/db/2016.sqlite3", 1024, fetchImpl)).rejects.toThrow(
|
||||
/not a SQLite header/,
|
||||
);
|
||||
});
|
||||
|
||||
it("warns, but continues, when the page size is not the request size", async () => {
|
||||
const warn = vi.spyOn(console, "warn").mockImplementation(() => {});
|
||||
const fetchImpl = async () => respond(header(4096));
|
||||
|
||||
await expect(probeDatabase("/db/2016.sqlite3", 1024, fetchImpl)).resolves.toBe(317096960);
|
||||
expect(warn).toHaveBeenCalledWith(expect.stringMatching(/page size 4096/));
|
||||
warn.mockRestore();
|
||||
});
|
||||
});
|
||||
@@ -1,6 +1,7 @@
|
||||
import { createDbWorker } from "sql.js-httpvfs";
|
||||
import workerUrl from "sql.js-httpvfs/dist/sqlite.worker.js?url";
|
||||
import wasmUrl from "sql.js-httpvfs/dist/sql-wasm.wasm?url";
|
||||
import { probeDatabase } from "./db-probe.js";
|
||||
|
||||
/**
|
||||
* The database is read where it lies. SQLite asks for pages, the virtual file
|
||||
@@ -61,8 +62,17 @@ export class RemoteDatabase {
|
||||
async #open() {
|
||||
const opened = performance.now();
|
||||
try {
|
||||
// Supplied rather than left to the library, which would size the file
|
||||
// with a HEAD request. GitHub Pages compresses that response and reports
|
||||
// the compressed length, which the library refuses to use. See db-probe.
|
||||
const fileLength = await probeDatabase(this.url, CHUNK_BYTES);
|
||||
const worker = await createDbWorker(
|
||||
[{ from: "inline", config: { serverMode: "full", url: this.url, requestChunkSize: CHUNK_BYTES } }],
|
||||
[
|
||||
{
|
||||
from: "inline",
|
||||
config: { serverMode: "full", url: this.url, requestChunkSize: CHUNK_BYTES, fileLength },
|
||||
},
|
||||
],
|
||||
workerUrl,
|
||||
wasmUrl,
|
||||
this.budgetBytes,
|
||||
|
||||
Reference in new issue
Block a user