Skip to content

Lesson images (binary, R2 + IndexedDB) ​

Lesson images are stored as binary, keyed by their SHA-256 content hash, not as base64 inside the lesson doc. Locally they live as blobs in IndexedDB (so large drafts aren't capped by localStorage's ~5 MB quota); in the cloud they live in an R2 bucket. The lesson doc only references images by hash, as { image: { hash, mime, ext } } on an image block or on a VAKT activity that carries a picture.

On the device, the blobs sit in the images object store (keyed by hash) of the s2c-lesson-maker IndexedDB database, next to the lesson library's own stores (packages/core/src/browser/imageStore.js). The database name is a historical one and stays as it is, because renaming it would strand every saved lesson.

Worker endpoints (apps/api/src/routes/images.js, registered in apps/api/src/app.js):

  • GET /images/:hash (and HEAD): public; serves the image bytes from R2 (immutable cache), with whatever content type the stored object has.
  • PUT /images/:hash: authenticated (Supabase JWT); the body must have an image/* content type, be at most 8 MB, and hash to :hash before it is stored. Called on save/publish to upload locally drafted images (packages/core/src/imagesClient.js). On the way in, the Worker re-compresses raster images to WEBP (convertImageToWebp in apps/api/src/imageConvert.js), falling back to the original bytes for formats it can't decode or when the WEBP isn't smaller. The key is still the original content hash, so dedup and references are unaffected.

The conversion uses @jsquash's WASM codecs rather than a native dependency like sharp, which is what lets the same code run in the Workers runtime and in Node with no per-platform build step. Each codec's init() is handed an already-compiled WebAssembly.Module, so the Emscripten glue never tries to fetch its binary at runtime; the Workers sandbox forbids that, and in Node it would resolve against the wrong base.

Where that module comes from is the only per-runtime part, and it sits behind the #image-codec-wasm subpath import (apps/api/package.json, src/codecs/): wrangler's bundler resolves the .wasm imports on Workers, and Node reads and compiles them off disk, memoized and deferred so an instance that never receives an upload doesn't pay for ~700 KB of codecs at startup. It is the same conditional-import pattern apps/mcp uses for its #standards-md.

The two halves are verified in different places, because no single runner covers both. The Node half runs in the API's Node test project; the Workers half is resolved by wrangler's bundler, which CI exercises with pnpm --filter @spelling-creator/api bundle (a wrangler deploy --dry-run). vitest-pool-workers cannot resolve a .wasm out of node_modules. It uses Vite's module graph, not wrangler's, so the image tests are excluded from that project (apps/api/vitest.workers.config.js) rather than made to limp along in it.

The browser also tries to do this conversion itself before an image ever reaches the Worker: readImageFile (packages/core/src/browser/imageFile.js) re-encodes PNG/JPEG picks to WEBP on a canvas (same quality target, same keep-whichever-is-smaller rule) while it's already decoding the file to measure its dimensions. This is purely an optimization (it saves upload bandwidth and R2 conversion work for the common case), not a correctness requirement: browsers without WEBP canvas encoding, GIF/SVG (skipped client-side on purpose, same as server-side), and any client that skips or fails the step all still land on the Worker as their original raster type, which converts them exactly as described above.

Setup:

bash
# Create the R2 bucket the IMAGES binding points at (see apps/api/wrangler.jsonc).
wrangler r2 bucket create spelling-creator-images

# Secret for the one-time backfill endpoints (below).
wrangler secret put ADMIN_MIGRATE_TOKEN

The routes never touch the IMAGES binding directly. They ask the platform seam for "the image store", which on Cloudflare is that R2 bucket and on a self-hosted Node instance is the S3-compatible bucket named by S3_BUCKET_IMAGES. Without either, the image routes answer 500 and everything else keeps working.

Migrating existing lessons ​

Existing cloud lessons (with base64 images inline) are converted by a one-time, idempotent backfill. It's gated by the ADMIN_MIGRATE_TOKEN secret (sent as the X-Admin-Token header) and pages through lessons, uploading each inline image to R2 and rewriting the doc. limit defaults to 25 and is capped at 100:

bash
# Repeat, passing the returned nextCursor each time, until nextCursor is null.
curl -X POST https://<worker-host>/admin/migrate-images \
  -H "X-Admin-Token: $ADMIN_MIGRATE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"cursor": 0, "limit": 25}'

Local lessons migrate automatically on first load (the old localStorage doc moves to IndexedDB, then to the lesson library; see migrateLocalStorage and migrateToLibrary in packages/core/src/browser/storage.js). Readers tolerate legacy base64 throughout, so the backfill can run any time after deploy. Deploy order: deploy the Worker (so /images exists), then ship the web build, then run the backfill.

A second, separate backfill re-compresses images already in R2 to WEBP, for objects uploaded before the PUT handler started converting. It's gated by the same ADMIN_MIGRATE_TOKEN and pages through the bucket with R2's list cursor, overwriting each PNG/JPEG object at the same key (only when the WEBP is smaller). It's idempotent (already-WEBP and untranscodable objects are skipped). Each object is a decode and an encode, so limit defaults to 10 and is capped at 50:

bash
# Repeat, passing the returned nextCursor each time, until nextCursor is null.
curl -X POST https://<worker-host>/admin/backfill-webp \
  -H "X-Admin-Token: $ADMIN_MIGRATE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"limit": 10}'

Because GET /images/:hash is edge-cached as immutable, a colo that already cached an object keeps serving the pre-conversion bytes until its cache evicts it. The handler clears the current colo's entry as it goes; the bytes look the same either way, so this only delays the size saving.

Staying within R2's free tier ​

R2's free tier allows 10 GB-month storage, 1M class-A (write) ops/month, and 10M class-B (read) ops/month. The design keeps usage well inside these:

  • Class A (writes) is roughly the number of distinct images, not the number of saves. Images are content-addressed, so PUT /images/:hash first does a head() (class B) and only put()s when the object is missing; the client also caches which hashes it has uploaded this session (imagesClient.js), so re-saving a lesson uploads nothing new. Identical images (across all users/lessons) share one object.
  • Class B (reads) stays low because GET /images/:hash responses are cached at Cloudflare's edge (the bytes are immutable, so they're safe to cache forever). Repeat views of a popular lesson (and the og-image/prerender browser) are served from cache and don't hit R2.
  • Storage is bounded by global content-hash dedup plus an 8 MB-per-image cap (MAX_IMAGE_BYTES in apps/api/src/lib/images.js, enforced by the Worker; the editor has no size check of its own, though the MCP server checks the same 8 MB before it uploads a Commons picture, in apps/mcp/src/wikimedia.js). This is the one limit without a hard code guard, so set an R2 storage alert in the Cloudflare dashboard (Notifications) if you want a heads-up as the bucket grows. Cloudflare does not offer a hard spend cap, so monitoring is the safety net here.

Copyright © 2026 Spelling Creator.