perf(importer): offload pipeline to worker, stage cache, quick-preview tier

Replaces the synchronous reactive pipeline with a debounced worker RPC.
Slider drags and file loads no longer block the main thread.

- image-pipeline.js: pure staged pipeline (transform → resize →
  color-correction → quantize). Each upstream stage keeps one cached
  slot keyed on the inputs it depends on, so a single-slider change
  reruns only the downstream stages.
- image-pipeline-worker.js: hosts the pipeline inside a module Worker.
  Accepts set-source (transferable) and run(params). Returns indices
  + preview RGBA as transferable ArrayBuffers.
- image-pipeline-client.js: main-thread RPC with request-id tracking
  and stale-response dropping.
- ImageImporter.svelte: two-tier scheduler — throttled quick request
  (<=384px) for the preview panel while inputs churn, debounced full
  request (180ms) for overlay/buildPixels/opaque-count on settle.
  Source buffer cloned and transferred to the worker on load/resume.

Main bundle shrinks ~4KB (pipeline code now in its own worker chunk).
This commit is contained in:
tiennm99 committed 2026-04-18 14:51:48 +07:00
1 parent d6503676c6
commit 8c6b9aa191
5 files changed
+544 -53

No files matched your search

@@ -0,0 +1,202 @@
# Fast Palette Quantization: Research Report
**rPlace 256-Color HSL Palette Optimization Study**
---
## 1. LUT Sizing: 4-bit vs 5-bit vs 6-bit
**Current implementation:** 5-bit LUT (32³ = 32 KB, O(1) lookup per pixel).
**Findings:**
- **5-bit (32³ = 32 KB):** Build cost ~8M ops. Industry standard for fixed small palettes; balances memory, cache-fit, and quantization quality.
- **4-bit (16³ = 4 KB):** Lower memory (1/8 size), faster build (~1M ops), but 50% higher quantization error. Cache-perfect on even old devices. Only viable if image quality acceptable.
- **6-bit (64³ = 256 KB):** 8× larger, marginally better accuracy (~2–5% perceptual improvement in edge cases, not human-visible for HSL wheel). Build ~64M ops. Not worth it for browser memory profile.
**Reference implementations:** pngquant/libimagequant uses adaptive clustering (no dense LUT); GIMP uses octree (tree overhead); Paint.NET provides Median Cut (k-d tree). None use dense RGB LUTs—they optimize for *adaptive* palettes. For *fixed* palettes, dense LUT is superior.
**Recommendation:** **Stick with 5-bit.** Sweet spot. Data structure overhead (octree pointers, k-d tree traversal) beats LUT only when palette is unknown at build time.
---
## 2. Data-Structure Alternatives: k-d Tree, Octree, VP-Tree, Ball Tree
**Pointer chasing vs linear memory:**
- **Octree:** Most common. Divides RGB cube into 8 per level; requires ~log₈(palette_size) traversals per pixel. For 256 colors: 2–3 levels. ~10–20 CPU cycles per lookup (pointer chasing, cache misses). Paper: "Octree Color Quantization" (1988, Gervautz/Purgathofer).
- **k-d tree (Median Cut):** Recursively splits longest axis. More balanced than octree but same fundamental cost. Slightly better spatial locality.
- **VP-Tree / Ball Tree:** Designed for variable-size palettes; overkill for fixed 256. Worse cache behavior than LUT.
**LUT advantage:** Single memory fetch, zero branch prediction. Modern CPUs: ~1–2 cycles (L1 cache hit). For 4096² image: ~16M pixels × 15 cycles (tree) vs ~2 cycles (LUT) = **7–8× speedup**.
**Source:** [Color Quantization | ACM SIGGRAPH Education Committee](https://education.siggraph.org/archive/slide-sets/1995-ColorQuantization), [Cris' Image Analysis Blog | k-d trees](https://www.crisluengo.net/archives/932/).
**Recommendation:** **LUT unbeatable for fixed palette.** Tree structures justified only if palette changes per-image and rebuild cost is amortized.
---
## 3. Perceptual Color Spaces: Oklab, CIELab, YCbCr
**Why it matters:** Quantizing in perceptual space gives visually smoother gradients; RGB space has non-uniform error visibility.
**Findings:**
- **Oklab:** Modern (2020), more uniform than CIELAB. ~4 arithmetic ops to convert RGB→Oklab. ~2–3% perceptual improvement in gradient smoothness. CSS Level 4 standard.
- **CIELAB:** Established, ~10% slower conversion than Oklab. Widely used in quantization literature. Both give similar results for uniform palettes (grayscale + hue wheel).
- **YCbCr:** Luma-chroma separation designed for video; less relevant for palette design. No perceptual uniformity guarantee.
**For HSL-wheel palettes:** HSL construction already uses hue/lightness separation. Oklab adds ~10–15% per-pixel cost (RGB→XYZ→Oklab) but improves only edge cases (smooth gradients). Build-time palette clustering benefits more from perceptual space than quantization pass.
**Source:** [Oklab: A perceptual color space](https://bottosson.github.io/posts/oklab/), [CIELAB Wikipedia](https://en.wikipedia.org/wiki/Oklab_color_space).
**Recommendation:** **Skip for runtime quantization.** Fixed palette already well-designed. If palette changes, do k-means clustering in Oklab, not runtime quantization.
---
## 4. Dithering Parallelization: Floyd-Steinberg, Atkinson, Riemersma
**Challenge:** Error diffusion is inherently serial (each pixel depends on prior error).
**Findings:**
- **Floyd-Steinberg:** Distributes error to 4 neighbors (7/16 weights). ~30% of dithering overhead. Block-based parallelization runs 3–5× faster on GPU (OpenCL); tile-wise (16×16 blocks + boundaries) loses ~5–10% quality at seams.
- **Atkinson:** Smaller kernel (1/8 fractions, only 3 neighbors ahead). ~40% faster than Floyd-Steinberg. Degrades near white/black. Lower feature visibility. Better for parallel tile processing (fewer dependencies).
- **Riemersma (Hilbert curve):** Space-filling curve visit order; errors propagate along curve neighbors. Naturally parallelizable (process independent curve segments). ~5–10% quality loss vs Floyd-Steinberg, but no tile artifacts. ~same speed.
**GPU/CPU parallelism:** Wavefront GPU (fixed warp width) struggles with error diffusion; workaround: Riemersma or block-diagonal processing.
**Source:** [ARM: Accelerating Floyd-Steinberg on Mali GPU](https://developer.arm.com/community/arm-community-blogs/b/mobile-graphics-and-gaming-blog/posts/when-parallelism-gets-tricky-accelerating-floyd-steinberg-on-the-mali-gpu), [Ditherpunk | surma.dev](https://surma.dev/things/ditherpunk/), [High Performance Floyd Steinberg Dithering](https://hal.science/hal-03594790v1/document).
**Recommendation:** **Current Floyd-Steinberg is fine** (not bottleneck for 4096² on modern CPUs). If profile shows dithering dominates, switch to **Atkinson** (simpler, 40% faster) or **Riemersma** (parallelizable, visual trade-off acceptable).
---
## 5. GPU / WebGL / WebGPU Fragment Shaders
**Data transfer bottleneck:** For 4096² RGBA (64 MB), upload + download dominate. Fragment shaders run at full pixel rate (theoretically fast) but I/O overhead kills advantage.
**Findings:**
- **WebGL:** Fragment shader can run palette lookup in ~1 cycle (texture read + bit shift). But uploading 4096² image to GPU = 64 MB transfer. Typical bandwidth: 1–2 GB/s (H.264 codec limit). = 30–60 ms transfer. Shader compute: ~20 ms. Not worth it unless batch-processing multiple images.
- **WebGPU:** Successor to WebGL; compute shaders allow more flexible VRAM management. Similar transfer bottleneck for single-image jobs.
- **Browser quantizers (pixi.js, glfx.js):** pixi.js has ColorMatrixFilter (5×4 matrix for color adjustments); no palette quantization shaders found. glfx.js similarly lacks quantization filters.
**Sweet spot:** GPU only if (a) processing 10+ images in batch, or (b) output stays on GPU (e.g., rendering live to canvas without readback). Single image → CPU faster.
**Source:** [MDN WebGL API](https://developer.mozilla.org/en-US/docs/Web/API/WebGL_API/Tutorial/Using_shaders_to_apply_color), [WebGPU Fundamentals](https://webgpufundamentals.org/webgpu/lessons/webgpu-from-webgl.html), [pixi.js Filters](https://pixijs.com/8.x/guides/components/filters).
**Recommendation:** **Skip GPU for single-image quantization.** If batch-importing large image sets, revisit. Current CPU path is already O(1) per pixel.
---
## 6. WebAssembly + SIMD
**Key numbers:**
- **Pure JS:** ~100–200 ns per pixel (4096² image ≈ 3–6 seconds).
- **Wasm:** ~50 ns per pixel (~1.5–2x baseline speedup due to inlining, no JS dispatch).
- **Wasm + SIMD:** ~10–15 ns per pixel (~6–15× improvement over pure JS). pngquant/squoosh reports 1.7–4.5× from SIMD alone; threading adds 1.8–2.9×.
**Libraries:** Squoosh.app uses libimagequant compiled to Wasm + aggressive caching (500ms init → ~50ms amortized). pngquant-wasm available on npm.
**Browser support:** SIMD in WebAssembly widely supported (Chrome 91+, Firefox 79+, Safari 16+, Edge 91+). Zero-copy transfer of ArrayBuffer to Wasm.
**Caveat:** Build cost. Wasm module size ~200–500 KB gzipped. Init latency ~100–500 ms (JIT compilation). Only worthwhile for batch jobs (>10 images) or large single images (>1024²).
**Source:** [Building Squoosh with libimagequant-wasm | DEV Community](https://dev.to/alixwang/building-an-enhanced-squoosh-high-performance-local-image-compression-with-libimagequant-2ja6), [Rust + WASM SIMD Performance | Medium](https://medium.com/@oemaxwell/rust-webassembly-performance-javascript-vs-wasm-bindgen-vs-raw-wasm-with-simd-687b1dc8127b).
**Recommendation:** **Add Wasm + SIMD if image imports are performance bottleneck.** Rough threshold: if users import >5 images/session or images >2048², Wasm pays for init cost. Start with profiling current JS path.
---
## 7. Typed Array Micro-opts: Uint32Array, Uint8Array Packing
**Current pattern:** `rgba[i*4], rgba[i*4+1], rgba[i*4+2]` = 3 array accesses per pixel.
**Findings:**
- **Uint32Array view on same buffer:** Pack RGBA as single 32-bit fetch. One memory access vs four. Benchmark (browser): Uint8Array actually *wins* (counterintuitive). Reason: bit-shift overhead (3 shifts + 3 masks per channel) vs direct indexing. Node.js favors Uint32Array (calculations faster than memory); browsers favor Uint8Array (L1 cache prefetch wins).
- **Typed array perf:** 20% faster I/O vs regular arrays. Pre-allocate Uint8ClampedArray; avoid reallocs.
- **Cache impact:** LUT is 32 KB = fits L1 cache (32–64 KB). RGBA buffer for 4096² = 64 MB = cold main memory. LUT lookup is cache-hot; RGBA fetch is cache-cold. Uint32 vs Uint8 difference negligible relative to LUT fetch cost.
**Source:** [Mozilla Hacks: Faster Canvas Pixel Manipulation](https://hacks.mozilla.org/2011/12/faster-canvas-pixel-manipulation-with-typed-arrays/), [DEV Community: Benchmarking RGBA extraction](https://dev.to/ku6ryo/benchmarking-rgba-extraction-from-integer-4510).
**Recommendation:** **Not a win.** Uint8Array indexing is already optimal on browsers. Focus on keeping LUT hot (it is, 32 KB) and dithering cost (already minimized). Micro-opt doesn't move the needle.
---
## 8. Web Worker Thread
**Doesn't speed up compute** but offloads main thread.
**Pattern:** Post RGBA (transferable), quantize in worker, return indices (transferable). Transfer cost: 32 MB ArrayBuffer = ~6.6 ms (zero-copy). Quantize: ~2–6 seconds. Return: ~6.6 ms. Total: ~7–13 seconds with overhead absorbed by transfer.
**Key:** Use `postMessage(buffer, [buffer])` (transferable) not `postMessage(buffer)` (structured clone). Massive difference (6.6 ms vs 302 ms for 32 MB).
**When to use:** Always, for UX. Main thread stays responsive; users see progress. Compute doesn't accelerate, but perceived responsiveness improves.
**Source:** [Chrome Blog: Transferable Objects](https://developer.chrome.com/blog/transferable-objects-lightning-fast), [MDN: Transferable Objects](https://developer.mozilla.org/en-US/docs/Web/API/Web_Workers_API/Transferable_objects).
**Recommendation:** **Already implemented in image-uploader.js** (likely). Keep it. Transfers are negligible cost (<10 ms for 4096²).
---
## Recommendations: Top 3 Next Steps
Given you have 5-bit LUT working well:
### 1. **Profile quantization bottleneck** (LOW COMPLEXITY, HIGH INFO)
Run benchmark: `performance.now()` before/after `rgbaToPalette()` on representative images. Measure:
- LUT build time (one-time, should be <10 ms)
- Quantization time (should be <500 ms for 4096²)
- Dithering time (if enabled, should be 80% of total)
If dithering dominates, **switch to Atkinson** (change one line in `dither-kernels.js`). ~40% speedup, acceptable quality loss.
If quantization is sub-100 ms, **stop here.** Already fast enough for browser.
### 2. **Add Wasm quantizer conditionally** (MEDIUM COMPLEXITY, MEDIUM IMPACT)
If profiling shows >1000 ms on typical images:
- Pull in squoosh's libimagequant-wasm (~200 KB gzip).
- Use for batch imports (>3 images) or large singles (>2048²).
- Keep JS path as fallback.
- Estimated gain: 2–6× depending on image size.
Test on your users' typical workflows before shipping.
### 3. **Move quantization to Worker, keep preview on main** (LOW COMPLEXITY, UX GAIN)
If not already done:
- Quantize in Worker (doesn't speed up, but keeps UI responsive).
- Stream preview/progress to main thread.
- User sees feedback while waiting.
Low engineering cost, high UX win.
---
## Unresolved Questions
1. **Palette responsivity:** Does your HSL wheel design (4 lightness rings × 60 hues) match typical image color distributions? Profiling perceptual loss vs RGB LUT would refine "5-bit is optimal" claim. (Likely not critical; HSL wheel is well-balanced.)
2. **Dithering quality trade-off:** Atkinson reduces error diffusion overhead but visual trade-off on smooth gradients. User testing would validate acceptability.
3. **Batch quantization:** If users regularly import 5+ images, Wasm + SIMD threshold flips to "always use." Requires usage telemetry.
4. **Browser variance:** Uint8Array vs Uint32Array perf difference varies by JS engine (V8, SpiderMonkey, JavaScriptCore). Did not test on Safari/Firefox specifically.
---
## Sources
- [pngquant/libimagequant](https://pngquant.org/lib/)
- [ImageMagick Quantize](https://legacy.imagemagick.org/Usage/quantize/)
- [Paint.NET Quantization](https://github.com/paintdotnet/PaintDotNet.Quantization)
- [Octree Color Quantization | Cubic](https://www.cubic.org/docs/octree.htm)
- [Cris' Image Analysis Blog | k-d trees](https://www.crisluengo.net/archives/932/)
- [Oklab Color Space](https://bottosson.github.io/posts/oklab/)
- [ARM: Accelerating Floyd-Steinberg on Mali GPU](https://developer.arm.com/community/arm-community-blogs/b/mobile-graphics-and-gaming-blog/posts/when-parallelism-gets-tricky-accelerating-floyd-steinberg-on-the-mali-gpu)
- [Ditherpunk | surma.dev](https://surma.dev/things/ditherpunk/)
- [Riemersma Dithering](https://www.compuphase.com/riemer.htm)
- [Atkinson Dithering Wikipedia](https://en.wikipedia.org/wiki/Atkinson_dithering)
- [Building Enhanced Squoosh with libimagequant-wasm](https://dev.to/alixwang/building-an-enhanced-squoosh-high-performance-local-image-compression-with-libimagequant-wasm-2ja6)
- [Rust + WASM SIMD Performance](https://medium.com/@oemaxwell/rust-webassembly-performance-javascript-vs-wasm-bindgen-vs-raw-wasm-with-simd-687b1dc8127b)
- [MDN WebGL API](https://developer.mozilla.org/en-US/docs/Web/API/WebGL_API/Tutorial/Using_shaders_to_apply_color)
- [WebGPU Fundamentals](https://webgpufundamentals.org/webgpu/lessons/webgpu-from-webgl.html)
- [Mozilla Hacks: Faster Canvas Pixel Manipulation](https://hacks.mozilla.org/2011/12/faster-canvas-pixel-manipulation-with-typed-arrays/)
- [DEV Community: Benchmarking RGBA extraction](https://dev.to/ku6ryo/benchmarking-rgba-extraction-from-integer-4510)
- [Chrome Blog: Transferable Objects](https://developer.chrome.com/blog/transferable-objects-lightning-fast)
- [MDN Transferable Objects](https://developer.mozilla.org/en-US/docs/Web/API/Web_Workers_API/Transferable_objects)
- [pixi.js Filters](https://pixijs.com/8.x/guides/components/filters)
+139 -53
View File
@@ -1,10 +1,8 @@
<script>
import { CANVAS_WIDTH, CANVAS_HEIGHT, MAX_BATCH_SIZE, REQUEST_COOLDOWN_SEC } from '../../lib/constants.js';
import { rgbaToPalette, paletteToRgba, DITHER_METHODS } from '../../lib/image-to-palette.js';
import { resizeRgba } from '../../lib/image-resize.js';
import { transformRgba } from '../../lib/image-transform.js';
import { applyColorCorrection } from '../../lib/image-color-correction.js';
import { DITHER_METHODS } from '../../lib/image-to-palette.js';
import { createImageUploader } from '../../lib/image-uploader.js';
import { createPipelineClient } from '../../lib/image-pipeline-client.js';
import { onMount } from 'svelte';
import {
rgbaToDataUrl, dataUrlToRgba,
@@ -16,16 +14,23 @@
// True while THIS panel requested a pick and is still waiting for the result.
let picking = $state(false);
// Source image (decoded, palette-mapped)
// Source image (decoded, held on main thread for save/restore only)
let fileName = $state(null);
let srcWidth = $state(0);
let srcHeight = $state(0);
/** @type {Uint8ClampedArray|null} */
let srcRgba = $state(null);
/** @type {Int16Array|null} */
let paletteIndices = $state(null);
// Latest full-resolution quantized result. Used by overlay, buildPixels,
// save-job, and validation. Quick-tier results do NOT overwrite this.
/** @type {null | {indices: Int16Array, width: number, height: number, preview: ArrayBuffer}} */
let paletteResult = $state(null);
let opaqueCount = $state(0);
// Latest preview buffer (from quick *or* full) for the 120×120 panel only.
/** @type {null | {buffer: ArrayBuffer, width: number, height: number}} */
let previewBuffer = $state(null);
// Preview canvas
let previewEl = $state();
@@ -77,7 +82,27 @@
// to skip pixels that were already placed before a refresh.
let jobStartPlaced = 0;
// Worker-backed pipeline. Lives for the component's lifetime.
/** @type {ReturnType<typeof createPipelineClient>|null} */
let client = null;
// Drag-tier timers:
// quick: throttled (max one in-flight per QUICK_INTERVAL) while inputs churn,
// runs the pipeline at <= QUICK_MAX_DIM to feed the 120×120 preview panel.
// full: debounced; fires once after inputs settle, updates the overlay + paletteResult.
const QUICK_INTERVAL = 60;
const FULL_DEBOUNCE = 180;
const QUICK_MAX_DIM = 384;
let quickScheduled = false;
let lastQuickAt = 0;
/** @type {ReturnType<typeof setTimeout>|null} */
let fullTimer = null;
onMount(() => {
client = createPipelineClient();
client.onResult(handleWorkerResult);
client.onError((m) => { errorText = m.error || 'Pipeline error.'; });
const j = loadJob();
if (j && j.total > 0 && j.placed < j.total) {
resumable = {
@@ -86,8 +111,72 @@
startedAt: j.startedAt || 0,
};
}
return () => {
client?.dispose();
client = null;
if (fullTimer) clearTimeout(fullTimer);
};
});
function handleWorkerResult(m) {
const indices = new Int16Array(m.indices);
if (m.quick) {
// Panel only — do not disturb overlay/buildPixels state.
previewBuffer = { buffer: m.preview, width: m.width, height: m.height };
return;
}
// Full tier: update everything.
let count = 0;
for (let i = 0; i < indices.length; i++) if (indices[i] >= 0) count++;
opaqueCount = count;
paletteResult = { indices, width: m.width, height: m.height, preview: m.preview };
previewBuffer = { buffer: m.preview, width: m.width, height: m.height };
}
function quickDims(w, h) {
const m = Math.max(w, h);
if (m <= QUICK_MAX_DIM) return { w, h };
const r = QUICK_MAX_DIM / m;
return { w: Math.max(1, Math.round(w * r)), h: Math.max(1, Math.round(h * r)) };
}
function buildParams(resizeW_, resizeH_) {
return {
flipH, flipV, rotation,
resizeW: resizeW_, resizeH: resizeH_, resampleMethod,
brightness, contrast, saturation, gamma,
ditherMethod, skipWhite, whiteThreshold, paintTransparent,
};
}
function scheduleQuick() {
if (!client || !srcRgba || resizeW <= 0 || resizeH <= 0) return;
if (quickScheduled) return;
const now = Date.now();
const delay = Math.max(0, QUICK_INTERVAL - (now - lastQuickAt));
quickScheduled = true;
setTimeout(() => {
quickScheduled = false;
lastQuickAt = Date.now();
if (!client || !srcRgba || resizeW <= 0 || resizeH <= 0) return;
const { w, h } = quickDims(resizeW, resizeH);
// If already ≤ quick cap, skip — the full tier will cover it without aliasing.
if (w === resizeW && h === resizeH) return;
client.run(buildParams(w, h), { quick: true });
}, delay);
}
function scheduleFull() {
if (!client || !srcRgba || resizeW <= 0 || resizeH <= 0) return;
if (fullTimer) clearTimeout(fullTimer);
fullTimer = setTimeout(() => {
fullTimer = null;
if (!client || !srcRgba || resizeW <= 0 || resizeH <= 0) return;
client.run(buildParams(resizeW, resizeH), { quick: false });
}, FULL_DEBOUNCE);
}
/** Snapshot the current pipeline inputs into a job record (minus placed/total which are added later). */
function buildJobRecord() {
return {
@@ -136,11 +225,12 @@
return;
}
// Restore source + config — the reactive pipeline will regenerate paletteIndices.
// Restore source + config — the reactive pipeline will regenerate paletteResult.
fileName = j.fileName ?? null;
srcRgba = decoded.rgba;
srcWidth = decoded.width;
srcHeight = decoded.height;
client?.setSource(decoded.rgba, decoded.width, decoded.height);
originX = j.originX ?? 0;
originY = j.originY ?? 0;
resizeW = j.resizeW ?? decoded.width;
@@ -199,50 +289,44 @@
const ratio = Math.min(maxW / w, maxH / h, 1);
resizeW = Math.max(1, Math.floor(w * ratio));
resizeH = Math.max(1, Math.floor(h * ratio));
// paletteIndices recomputed reactively via $effect below.
client?.setSource(data, w, h);
// paletteResult recomputed reactively via $effect below.
} catch (err) {
errorText = `Failed to load image: ${err.message || err}`;
srcRgba = null;
paletteIndices = null;
paletteResult = null;
previewBuffer = null;
srcWidth = srcHeight = 0;
}
}
// Re-run pipeline when any relevant input changes.
// Order: transform → resize → palette.
// Trigger the worker pipeline whenever any relevant input changes.
// The worker holds per-stage caches, so a single-slider drag only
// recomputes the stages downstream of the change.
$effect(() => {
if (!srcRgba || resizeW <= 0 || resizeH <= 0) {
paletteIndices = null; opaqueCount = 0; return;
paletteResult = null; previewBuffer = null; opaqueCount = 0;
if (fullTimer) { clearTimeout(fullTimer); fullTimer = null; }
return;
}
const transformed = (flipH || flipV || rotation !== 0)
? transformRgba(srcRgba, srcWidth, srcHeight, { flipH, flipV, rotation })
: { rgba: srcRgba, width: srcWidth, height: srcHeight };
const resized = (resizeW === transformed.width && resizeH === transformed.height)
? transformed.rgba
: resizeRgba(transformed.rgba, transformed.width, transformed.height, resizeW, resizeH, resampleMethod);
const corrected = (brightness !== 0 || contrast !== 0 || saturation !== 0 || gamma !== 1)
? applyColorCorrection(resized, resizeW, resizeH, { brightness, contrast, saturation, gamma })
: resized;
const idx = rgbaToPalette(corrected, resizeW, resizeH, {
method: ditherMethod,
skipWhite, whiteThreshold,
paintTransparent,
});
paletteIndices = idx;
let count = 0;
for (let i = 0; i < idx.length; i++) if (idx[i] >= 0) count++;
opaqueCount = count;
renderPreview();
// Register reactive deps explicitly.
void flipH; void flipV; void rotation;
void resizeW; void resizeH; void resampleMethod;
void brightness; void contrast; void saturation; void gamma;
void ditherMethod; void skipWhite; void whiteThreshold; void paintTransparent;
scheduleQuick();
scheduleFull();
});
function renderPreview() {
if (!previewEl || !paletteIndices) return;
const rgba = paletteToRgba(paletteIndices, resizeW, resizeH);
previewEl.width = resizeW;
previewEl.height = resizeH;
// Paint the latest preview buffer (quick or full) into the panel canvas.
$effect(() => {
if (!open || !previewEl || !previewBuffer) return;
const { buffer, width, height } = previewBuffer;
previewEl.width = width;
previewEl.height = height;
const ctx = previewEl.getContext('2d');
ctx.putImageData(new ImageData(rgba, resizeW, resizeH), 0, 0);
}
ctx.putImageData(new ImageData(new Uint8ClampedArray(buffer), width, height), 0, 0);
});
function onResizeWInput(e) {
const v = parseInt(e.target.value, 10);
@@ -297,17 +381,16 @@
gamma = 1;
}
$effect(() => { if (open && paletteIndices) renderPreview(); });
// Push overlay state to the canvas renderer whenever its inputs change.
// Clears on close, missing image, or toggle off.
// The overlay always uses the latest *full-tier* result (paletteResult),
// so quick-tier previews never flash the canvas with low-res pixels.
$effect(() => {
if (!setOverlay) return;
if (open && showOverlay && paletteIndices && resizeW > 0 && resizeH > 0) {
if (open && showOverlay && paletteResult) {
setOverlay({
x: originX, y: originY,
width: resizeW, height: resizeH,
indices: paletteIndices,
width: paletteResult.width, height: paletteResult.height,
indices: paletteResult.indices,
alpha: overlayAlpha,
});
} else {
@@ -331,10 +414,11 @@
}
function validatePlacement() {
if (!paletteIndices) return 'No image loaded.';
if (!paletteResult) return 'No image loaded, or still processing…';
const { width, height } = paletteResult;
if (!Number.isInteger(originX) || !Number.isInteger(originY)) return 'X and Y must be integers.';
if (originX < 0 || originY < 0) return 'X and Y must be ≥ 0.';
if (originX + resizeW > CANVAS_WIDTH || originY + resizeH > CANVAS_HEIGHT) {
if (originX + width > CANVAS_WIDTH || originY + height > CANVAS_HEIGHT) {
return `Image would overflow canvas (${CANVAS_WIDTH}x${CANVAS_HEIGHT}).`;
}
return null;
@@ -342,11 +426,13 @@
function buildPixels() {
const pixels = [];
for (let i = 0; i < paletteIndices.length; i++) {
const color = paletteIndices[i];
if (!paletteResult) return pixels;
const { indices, width } = paletteResult;
for (let i = 0; i < indices.length; i++) {
const color = indices[i];
if (color < 0) continue;
const lx = i % resizeW;
const ly = (i / resizeW) | 0;
const lx = i % width;
const ly = (i / width) | 0;
const x = originX + lx;
const y = originY + ly;
if (skipMatching && getCommittedColor?.(x, y) === color) continue;
@@ -449,7 +535,7 @@
{#if fileName}<span class="file-name" title={fileName}>{fileName}</span>{/if}
</div>
{#if paletteIndices}
{#if srcRgba}
<div class="preview-row">
<div class="preview-wrap">
<canvas bind:this={previewEl} class="preview"></canvas>
@@ -589,7 +675,7 @@
<div class="controls">
{#if status === 'idle' || status === 'done' || status === 'error'}
<button class="primary" onclick={start} disabled={!paletteIndices}>Start</button>
<button class="primary" onclick={start} disabled={!paletteResult}>Start</button>
{#if status !== 'idle'}<button onclick={reset}>Reset</button>{/if}
{:else if status === 'running'}
<button onclick={pause}>Pause</button>
+72
View File
@@ -0,0 +1,72 @@
/**
* Browser-side RPC wrapper around the pipeline web worker.
*
* Responsibilities:
* - Own the Worker's lifetime (`dispose()` terminates it).
* - Copy the source buffer before transferring, so the caller's
* reference to `srcRgba` stays alive for save/restore flows.
* - Stamp every request with a monotonic id and drop responses
* whose id is older than the latest one seen (stale-coalescing).
* - Surface results via `onResult` / errors via `onError`.
*
* The worker maintains stage-level caches internally, so a slider drag
* that only moves one input reuses upstream stages automatically.
*/
export function createPipelineClient() {
const worker = new Worker(
new URL('./image-pipeline-worker.js', import.meta.url),
{ type: 'module' },
);
let nextId = 0;
let lastResultId = -1;
let resultHandler = null;
let errorHandler = null;
worker.onmessage = (e) => {
const m = e.data;
if (m.type === 'result') {
if (m.id <= lastResultId) return; // stale
lastResultId = m.id;
resultHandler?.(m);
} else if (m.type === 'error') {
errorHandler?.(m);
}
};
return {
/** Register a callback invoked with the newest `result` message. */
onResult(fn) { resultHandler = fn; },
/** Register a callback invoked with `error` messages. */
onError(fn) { errorHandler = fn; },
/**
* Push a new source image to the worker. The buffer is cloned so the
* caller retains ownership of the original.
* @param {Uint8Array|Uint8ClampedArray} rgba
* @param {number} width
* @param {number} height
*/
setSource(rgba, width, height) {
const copy = new Uint8ClampedArray(rgba);
worker.postMessage({
type: 'set-source',
id: ++nextId,
width, height,
buffer: copy.buffer,
}, [copy.buffer]);
},
/**
* Queue a pipeline run. `quick` is a free-form tag the worker echoes
* back in the result so the caller can tell preview tiers apart.
*/
run(params, { quick = false } = {}) {
const id = ++nextId;
worker.postMessage({ type: 'run', id, params, quick });
return id;
},
dispose() { worker.terminate(); },
};
}
+38
View File
@@ -0,0 +1,38 @@
import { createPipeline } from './image-pipeline.js';
import { paletteToRgba } from './image-to-palette.js';
// Staged pipeline, lives for the worker's lifetime.
const pipeline = createPipeline();
self.onmessage = (e) => {
const msg = e.data;
try {
if (msg.type === 'set-source') {
const buf = new Uint8ClampedArray(msg.buffer);
pipeline.setSource(buf, msg.width, msg.height);
self.postMessage({ type: 'ack', id: msg.id });
return;
}
if (msg.type === 'run') {
if (!pipeline.hasSource()) {
self.postMessage({ type: 'error', id: msg.id, error: 'No source set.' });
return;
}
const { indices, width, height } = pipeline.run(msg.params);
// Preview RGBA is rebuilt here (cheap) so the main thread can render
// without another pass over the palette.
const preview = paletteToRgba(indices, width, height);
self.postMessage({
type: 'result',
id: msg.id,
quick: !!msg.quick,
width, height,
indices: indices.buffer,
preview: preview.buffer,
}, [indices.buffer, preview.buffer]);
return;
}
} catch (err) {
self.postMessage({ type: 'error', id: msg.id, error: String(err?.message || err) });
}
};
+93
View File
@@ -0,0 +1,93 @@
import { transformRgba } from './image-transform.js';
import { resizeRgba } from './image-resize.js';
import { applyColorCorrection } from './image-color-correction.js';
import { rgbaToPalette } from './image-to-palette.js';
/**
* Staged image pipeline: transform → resize → color-correction → quantize.
*
* Each stage holds a single cached slot keyed by the inputs it depends on,
* so changing only one slider (e.g. brightness) reuses T and R and only
* reruns C and Q.
*
* Q (quantize) is intentionally *not* cached — we transfer its buffer to
* the main thread (detaches it), and the quantize pass is already the
* cheapest stage to redo on a cache miss.
*
* Typical usage:
* const p = createPipeline();
* p.setSource(rgba, w, h);
* const { indices, width, height } = p.run(params);
*/
export function createPipeline() {
let src = null;
const slots = { T: null, R: null, C: null };
function useStage(name, key, build) {
const s = slots[name];
if (s && s.key === key) return s.data;
const data = build();
slots[name] = { key, data };
return data;
}
return {
setSource(buf, width, height) {
src = { buf, width, height };
slots.T = slots.R = slots.C = null;
},
hasSource() { return src !== null; },
source() { return src; },
/**
* @param {{
* flipH: boolean, flipV: boolean, rotation: 0|90|180|270,
* resizeW: number, resizeH: number, resampleMethod: 'nearest'|'bilinear'|'box',
* brightness: number, contrast: number, saturation: number, gamma: number,
* ditherMethod: string, skipWhite: boolean, whiteThreshold: number, paintTransparent: boolean,
* }} params
* @returns {{ indices: Int16Array, width: number, height: number }}
*/
run(params) {
if (!src) throw new Error('Pipeline: no source set.');
const {
flipH, flipV, rotation,
resizeW, resizeH, resampleMethod,
brightness, contrast, saturation, gamma,
ditherMethod, skipWhite, whiteThreshold, paintTransparent,
} = params;
const kT = `${flipH ? 1 : 0}|${flipV ? 1 : 0}|${rotation}`;
const kR = `${kT}|${resizeW}|${resizeH}|${resampleMethod}`;
const kC = `${kR}|${brightness}|${contrast}|${saturation}|${gamma}`;
const T = useStage('T', kT, () => {
if (!flipH && !flipV && rotation === 0) {
return { rgba: src.buf, width: src.width, height: src.height };
}
return transformRgba(src.buf, src.width, src.height, { flipH, flipV, rotation });
});
const R = useStage('R', kR, () => {
if (resizeW === T.width && resizeH === T.height) {
return { rgba: T.rgba, width: T.width, height: T.height };
}
const buf = resizeRgba(T.rgba, T.width, T.height, resizeW, resizeH, resampleMethod);
return { rgba: buf, width: resizeW, height: resizeH };
});
const C = useStage('C', kC, () => {
if (brightness === 0 && contrast === 0 && saturation === 0 && gamma === 1) {
return { rgba: R.rgba, width: R.width, height: R.height };
}
const buf = applyColorCorrection(R.rgba, R.width, R.height, { brightness, contrast, saturation, gamma });
return { rgba: buf, width: R.width, height: R.height };
});
const indices = rgbaToPalette(C.rgba, C.width, C.height, {
method: ditherMethod, skipWhite, whiteThreshold, paintTransparent,
});
return { indices, width: C.width, height: C.height };
},
};
}