Fixes from code review of canvas-on-do migration (commit c3f7c02):
- worker.js /api/ws: rewrite request URL to '/ws' so the DO pathname
switch dispatches correctly. The original c.req.raw kept '/api/ws'
which the DO never matched → 404 on every WS upgrade.
- migrate-from-upstash.js pickSampleOffsets: use TOTAL_PIXELS - 1 for the
last byte instead of CANVAS_WIDTH * CANVAS_WIDTH (only correct when
the canvas is square; constants explicitly invite non-square).
- chunk-storage.js writePixels: clarify atomicity comment — the loop is
atomic *because it has no awaits*, not because of any implicit DO
transaction. Added guidance for future maintainers.
- cooldown-store.js tryAcquire: GC sweep wrapped in try/catch so a
transient failure can't drop the user's allowed: true response.
Docs:
- README.md: drop Upstash from tech stack, redraw architecture,
document new project layout (durable-objects/lib, admin/), add
CHUNK_BYTES to configuration table.
- docs/system-architecture.md: full rewrite for DO-storage data flow,
document SQLite schema, race-safe rate-limit pattern, free-tier table.
- docs/deployment-guide.md: drop Upstash setup, add optional one-shot
migration runbook, update free-tier table to actual May 2026 limits.
Tests: 112 pass, 6 skipped (pending Phase 4 rewrite via
@cloudflare/vitest-pool-workers). Bundle dry-run clean.
Local wrangler dev smoke test was attempted but the sandboxed env
hangs HTTP requests at the workerd layer (TCP connects, no response).
Routing fix verified by code inspection; user must verify in their
own dev or production.
6.5 KiB
System Architecture
Overview
rplace is a collaborative pixel canvas deployed as a single Cloudflare Worker.
The frontend (Svelte SPA) is served as static assets, and a single
Durable Object (CanvasRoom, idFromName('main')) owns canvas state,
rate-limit cooldowns, and the WebSocket broadcast hub. The Worker is a
thin validation/routing proxy.
Component Map
Browser (Svelte SPA + WebSocket)
| GET /api/canvas → 16 MB binary, edge-cached 10s
| POST /api/place → batch pixel placement (validated at edge)
| WS /api/ws → real-time pixel deltas
v
Cloudflare Worker (Hono — thin proxy)
└─▶ CanvasRoom Durable Object (single instance, 'main')
├── canvas_chunks SQLite BLOB rows (256 × 64 KB = 16 MB)
├── cooldowns SQLite TTL rows (1s rate-limit)
└── WebSocket hub Hibernation API broadcasts pixel deltas
Data Flow
Pixel Placement
1. User draws on canvas → optimistic render into a pending buffer.
2. User hits Submit → POST /api/place { pixels: [{x, y, color}, ...] }.
3. Worker validates input (bounds, types, batch ≤ 2048, body ≤ 128 KB).
4. Worker resolves userId from CF-Connecting-IP and forwards to DO /place.
5. DO atomically:
a. cooldowns.tryAcquire(userId, 1s) — UPDATE expired or INSERT new row.
On conflict (active claim), responds 429.
b. chunk_storage.writePixels — group pixels by chunk_id, read each
touched chunk's BLOB, modify in memory, INSERT OR REPLACE.
c. Broadcast `{type:'pixels', pixels}` to all hibernating WebSockets.
6. Response: { ok: true } (or { error, retryAfter } on rate limit).
The DO is single-threaded; the entire 5a-c sequence runs without preemption, so cooldown check + write + broadcast are effectively atomic without explicit transactions.
Canvas Loading
1. Client fetches GET /api/canvas (10s edge-cache; CDN serves 99% of hits).
2. Worker forwards to DO /canvas on miss.
3. DO chunk_storage.readAllChunks: SELECT all canvas_chunks rows, copy each
BLOB into a single 16 MB Uint8Array offset by chunk_id × CHUNK_BYTES.
4. Client receives raw bytes (CF auto-gzip), maps each byte → COLORS_RGBA.
5. Renders via OffscreenCanvas + ImageData.
Real-time Updates
1. Client connects WS /api/ws.
2. Worker forwards upgrade to DO /ws (URL rewritten to /ws so the DO
pathname dispatch matches).
3. DO state.acceptWebSocket(server) — Hibernation API; idle sockets
survive DO eviction at zero CPU cost.
4. On placement → DO iterates state.getWebSockets() and sends JSON.
5. Client merges received pixels into local ImageData; re-renders dirty
region.
6. On disconnect → exponential-backoff reconnect (1 s → 30 s).
Storage
canvas_chunks (canvas pixels)
| Column | Type | Notes |
|---|---|---|
chunk_id |
INTEGER PRIMARY KEY | 0 .. CHUNK_COUNT − 1 |
bytes |
BLOB NOT NULL | Exactly CHUNK_BYTES bytes (last chunk may be short) |
- Linear byte layout:
offset = y * CANVAS_WIDTH + x chunk_id = floor(offset / CHUNK_BYTES)CHUNK_BYTES = 65536,CHUNK_COUNT = ceil(TOTAL_PIXELS / CHUNK_BYTES)- Lazy initialization: missing rows read as zero-filled buffers. Resizing the canvas is just a constants change — new chunks materialize on first read. No migration script needed.
- Hard caps: CF DO BLOB row size 2 MB → 64 KB chunks have 32× headroom.
cooldowns (rate-limit windows)
| Column | Type | Notes |
|---|---|---|
user_id |
TEXT PRIMARY KEY | Hashed CF-Connecting-IP |
expires_at |
INTEGER NOT NULL | ms epoch; row becomes stale past this point |
idx_cooldowns_expireskeeps lazy GC sweeps cheap.- GC runs at 1% sample rate inside
tryAcquire, wrapped intry/catch— best-effort; never blocks the rate-limit decision.
Rate Limiting
Single-row per user, 1 s window. Race-safe inside the DO via:
UPDATE cooldowns SET expires_at = ? WHERE user_id = ? AND expires_at <= ?
→ if rowsWritten > 0: claim acquired (existing row was expired).
→ else: INSERT INTO cooldowns ...
if INSERT throws on PK conflict → claim denied (active row exists).
Hardness comes from the DO single-threaded model; no SQL-level locking needed.
Migration Endpoint (transitional)
POST /admin/migrate-from-upstash (token-gated): one-shot importer that
pulls the legacy Upstash canvas via lib/canvas-storage.js (4 chunked
GETRANGE calls) and posts the raw 16 MB bytes to DO /import. The DO
splits into CHUNK_COUNT BLOB rows in a single sync transaction, then
the worker round-trips a sample-byte verification.
Removed in Phase 4 of plans/260509-2309-canvas-on-do-storage along with
@upstash/redis and ioredis dependencies.
Free-tier Footprint (CF, 2026)
| Resource | Quota | rplace usage at hobby scale | Headroom |
|---|---|---|---|
| Workers requests | 100K/day | ~100/day (50 users) | 1000× |
| DO storage / object | 10 GB | 16 MB | 600× |
| DO storage / account | 5 GB | 16 MB | 300× |
| BLOB row size | 2 MB | 64 KB chunks | 32× |
| Per-DO request rate | 1,000 / s soft cap | <1 / s | 1000× |
| WS connections / DO | tens of thousands | ~50 | huge |
/api/canvas carries Cache-Control: public, s-maxage=10, stale-while-revalidate=30.
Verify edge HIT in production via cf-cache-status: HIT; if absent, wrap
the worker handler with the Cache API to enforce caching.
Security
- Rate limiting: race-safe at the actor (single-threaded DO).
- Identity: CF-Connecting-IP hashed to userId — unspoofable.
- Input validation: strict bounds + type checks at the worker edge, re-validated at the DO trust boundary.
- Body cap: ~128 KB per
/api/place(2048 pixels × ~64 B JSON each). - Migration endpoint: Bearer-token gated. Token-comparison is direct string equality — adequate at hobby scale; consider constant-time comparison if exposure profile changes.
- DO isolation:
/canvas,/place,/import,/wsare intra-DO paths only — not internet-reachable except via the worker.
Operational Notes
- Single-region. DO is anchored to one Cloudflare colo; latency for far-away users mirrors the previous Upstash Redis topology — no regression.
- DO eviction mid-place is safe: the SQLite write commits before
broadcast. If the DO evicts before broadcast fires, client WS receives
no message but auto-reconnects and re-fetches
/api/canvas. No data loss; minor latency tail. - Storage billing. Per-account 5 GB free cap on Free plan as of Jan 7 2026 — current 16 MB is far under. Monitor on resize.