Skip to content

Interactive mode ​

For how to use this, see Interactive mode.

Interactive mode is the lesson's Practice tab (/hub/:id/practice, see Pages & routing), opened from Start lesson on the lesson page. A section's material appears as one step, then that section's questions one after another, each with a field to type an answer into. At the end, a signed-in user's answers are filed privately to their account; an unfinished run-through is kept in the browser so it can be resumed.

The route only decides when it is open. InteractiveLesson is still a dialog, deliberately: it is a focus mode with its own bottom bars and its own idea of the viewport. LessonPractice renders it with open and navigates back to the overview on close (rather than history.back(), since a deep link has no history to go back to). A lesson whose sections are all empty isn't playable (isInteractivePlayable): LessonTabs hides the tab, LessonLayout hides the button, and a direct link to /practice redirects to the lesson.

Derived from the document ​

There is no "interactive lesson" document type and nothing to switch on when authoring. The walkthrough is derived from the lesson document (buildInteractiveSteps in packages/core/src/interactive.js), so every lesson ever published works, including ones made before this feature existed and ones written by the MCP server. Nothing is added to a lesson to make it playable, and a lesson stays exactly as printable as it was.

In the documentBecomes
A section's text, image, spelling and VAKT blocksOne content step (kind: "content"), holding them together in document order.
Each question blockOne question step (kind: "question"), after that section's content.
A section with only questionsNo content step; it opens straight on its first question.
A section with nothing in itNothing.
A lesson with no questions at allA read-through: every content step, no answer fields, nothing saved.

VAKT blocks belong to the content step rather than being steps of their own: a regulation break is something whoever runs the session does with the speller, not something the speller answers, so it must not be counted by the progress bar.

Each step has a stable key (<sectionId>:content or <sectionId>:<blockId>, falling back to positions for documents without ids). Answers are keyed by block id (answerKey), so re-ordering a lesson between sittings doesn't shuffle which answer belongs to which question.

Text blocks keep their bold, italics and underlining, but not their footnote markers: this is the screen the speller reads, and a superscript number with nowhere on the screen to lead to is clutter there. The voice reads the plain words (see Formatting, footnotes & sources).

Questions are numbered from 1 within each section, matching the editor's Q7 numbering (see Navigating large lessons).

Presentation ​

Interactive mode is full-screen and drawn in the app's own theme, light or dark, as is the lesson page below it, rather than reproducing the white sheet the DOCX/PDF export produces. The blocks are re-rendered at a scale for reading and answering over a whole session: prose at reading size, images framed in the app's border and radius and sized by the reading column rather than by the size and alignment they carry (so a picture fills the width on a phone), spelling words as large cards, and a VAKT activity set apart as a red-edged card. Only the presentation differs; the content is the same blocks.

Showing the answers ​

questionAnswer(block) flattens each question type's stored answer into one shape, { answer, answers, steps, suggested }, or null when there is nothing to reveal:

TypeWhat the reveal shows
single, backgroundanswer
numberanswer plus the working steps (either alone is still revealed)
multipleevery entry in answers, the whole accepted set
multiple_openthe same list with suggested: true, labeled as suggestions
open, paraphrase, wyr, an unknown type, or a blank answernull: "This question has no set answer."

revealedAnswers turns that into the flat list of boxes the UI draws (the working is not one of them; it stays a numbered list below), and hasRevealableAnswers(steps) decides whether the toggle is rendered at all.

The two semi-open types share a color and an answers list because the S2C guidebook treats them as one family (see the comment in packages/core/src/questions.js). That grouping is the app's default, taken from one source; other Spelling practitioners categorize questions differently. suggested is what tells the two apart, and it travels with the answers so the person looking at the reveal doesn't have to remember the question type.

The toggle state lives in InteractiveLesson and is reset to off every time the dialog opens; unlike the speech settings it is never persisted. On a question step each answer box is a button (RevealedAnswer with onUse) that replaces the field's value and focuses it with the caret at the end. On the summary onUse is absent, which is what keeps the boxes there read-only. Nothing compares the typed answer with the author's anywhere.

Progress on the device ​

packages/core/src/browser/interactiveProgress.js keeps the unfinished run-through (every answer, plus the current step's key) in localStorage under spelling-creator:interactive-progress.

Progress (unfinished)A saved run-through (finished)
Lives inthis browser (localStorage)lesson_responses, on the server
Needsnothing; signed out works tooa signed-in session
Travelsno: this device onlyyes: any device you sign in on
Kept untilit is filed, you start again or discard it, or 90 days passyou delete it
Anyone else?never sent anywhere at allonly you can read it
  • Records are keyed by owner and lesson. The owner is the signed-in user's id, or "" when signed out, so signed-out users of one browser share a record per lesson. Per-owner keys exist for the shared computer (a clinic, or a home with more than one speller), where resuming into the previous person's answers would be worse than not resuming at all. The shared signed-out record is why a resumed run-through always shows the resume notice with Start again.
  • A run-through is pinned to the owner it started as (runOwner in InteractiveLesson). If the signed-in account changes with the dialog open, writes keep going to the original record, and Finish refuses to file to the new account (summary.signedInSince).
  • MAX_PROGRESS_RECORDS is 20 (most recently touched lessons) and PROGRESS_MAX_AGE_MS is 90 days, applied by pruneProgress. Pruning is fine here in a way it isn't for saved run-throughs: this is a resume cache, not the only copy of anything someone chose to keep.
  • Every write reports whether it landed. Where storage is refused (private browsing, a full quota) or missing, the leave confirmation switches back to warning that leaving discards the answers, so it only promises what was actually kept.
  • The record is cleared as soon as a run-through is filed. A failed save leaves it, so closing and coming back is a way to retry; signed out, it stays, since it is the only copy.
  • If the remembered step key no longer exists (the step was deleted), the lesson opens at the top with the answers still restored.
  • hasInteractiveProgress is what makes LessonLayout label the button Continue lesson.

Syncing progress across devices was rejected on purpose: it would mean putting half-written answers on the server, a much bigger promise than "your tab remembers".

Privacy of saved answers ​

  • Every endpoint that touches saved answers requires a signed-in session, and the Worker scopes each query to user_id = <verified caller>; that filter is the only way a row is ever addressed, not a check layered on top of one. Deleting someone else's row matches nothing and returns 404.
  • There is no endpoint that returns another user's answers: not for the lesson's author, a moderator or an admin.
  • lesson_responses has RLS enabled with no policies, unlike lessons, comments and ratings; only the service-role Worker reads or writes it. user_id is never returned.
  • Answers are not run through the profanity filter that comments go through. There's no audience to protect.
  • The in-progress copy has no endpoint at all: nothing to scope server-side, and no way for anyone to learn that a lesson was even opened.

MyLessonAnswers renders the "Your answers" panel on the lesson's overview tab (LessonOverview.jsx), under the lesson, and renders nothing when signed out or when there are none.

Worker endpoints ​

Method & pathAuthResponse
GET /lessons/:id/responsesBearer <Supabase JWT>{ "responses": [{ id, lessonId, answers, completedAt }] }; the caller's own only, newest first
POST /lessons/:id/responsesBearer <Supabase JWT>{ "response": { id, lessonId, answers, completedAt } }
DELETE /lessons/:id/responses/:ridBearer <Supabase JWT>{ "ok": true }, the caller's own only; else 404
  • POST body is { answers }, one entry per question: { blockId, sectionId, sectionName, questionType, prompt, answer }. The Worker rebuilds every entry from known fields and drops anything else, so the stored jsonb can only hold that shape. answer must be a string of at most 5,000 characters or the request is refused with 400; the other fields are cut to their maximum length (100 for ids, 300 for the section name, 2,000 for the prompt), and an unrecognized questionType is stored as open.
  • The prompt is snapshotted alongside the answer on purpose: a saved run-through has to stay readable after the lesson is edited, re-ordered, or has that question deleted.
  • Skipped questions are stored as blank answers rather than dropped, so the set still says which questions were asked.
  • Limits (shared between browser and Worker in packages/core/src/interactive.js): MAX_RESPONSE_LENGTH 5,000 characters per answer and MAX_RESPONSES 500 answers per submission.
  • You may keep 20 saved run-throughs of any one lesson (MAX_STORED_RESPONSES). Past that a POST is rejected with 409 and a message asking you to delete an older one, rejected rather than silently pruning the oldest, for the same reason the draft cap (see Lesson hub & accounts) is: they're the user's own answers, and quietly deleting them to make room isn't ours to decide.
  • POST also checks the lesson is one the caller could have read in the first place (canPlayLesson): published and not shadowbanned, or theirs, trusted or moderated. Otherwise 404.

Supabase schema ​

sql
create table if not exists public.lesson_responses (
  id           uuid primary key default gen_random_uuid(),
  lesson_id    uuid not null references public.lessons (id) on delete cascade,
  user_id      uuid not null references auth.users (id) on delete cascade,
  answers      jsonb not null,
  completed_at timestamptz not null default now()
);

create index if not exists lesson_responses_user_lesson_idx
  on public.lesson_responses (user_id, lesson_id, completed_at desc);

-- No public read policy, unlike lessons/comments/ratings: this data is private.
alter table public.lesson_responses enable row level security;

The full schema, with the reasoning in comments, is apps/api/schema.sql.

Reading aloud ​

Speech uses the browser's Web Speech API (speechSynthesis), or a natural voice where the device can run one. Like lesson summaries, this runs entirely on the reader's own device: no Worker call, no API key, no cost, and the lesson text never leaves the machine. Speech is probed for (speechSupported) rather than assumed, from an effect so server rendering and hydration agree; where it's missing, neither the controls nor the settings section is rendered.

stepSpeechText(step) builds what a step says: the section name, then the prose, image captions (never their credits), spelling words, a VAKT activity's text without its label or links, or the question prompt. A question's answer is deliberately never included, even with the reveal on.

The on/off, voice and pace preferences live in localStorage under spelling-creator:tts-enabled, spelling-creator:tts-voice and spelling-creator:tts-rate, read and written through apps/web/src/lib/speechPrefs.js by both interactive mode and the Reading aloud section of the settings page. SPEECH_RATES is [0.7, 0.85, 1, 1.25, 1.5]; an unrecognized stored rate falls back to 1. A change made in one place reaches the other the next time it mounts, not live (no storage listener).

Three platform quirks are handled between the two files. speechPrefs.js takes the one that belongs to the voice list: voices load asynchronously, announced by voiceschanged. useSpeech.js takes the two that belong to speaking: Chromium cuts off a single utterance after about 15 seconds (so text is split into chunks of at most 180 characters, MAX_CHUNK, and queued), and cancel() isn't synchronous (so a new utterance is deferred a tick after one).

Natural voices (Kokoro) ​

The natural voices are Kokoro, an 82M-parameter voice model that runs in the page with transformers.js (@huggingface/transformers 4), like the summary and import models. The model is pinned to a commit (MODEL_REVISION in readAloudVoices.js) and loaded at fp32 on WebGPU. Eight of its English voices, the ones graded C+ or better on the model card, are listed under Natural voices in the voice picker, above the browser's own, in both the popover and the settings page (apps/web/src/components/SpeechVoiceSelect.jsx):

IdNameAccent
af_heart (default)Hearten-US
af_bellaBellaen-US
af_nicoleNicoleen-US
am_fenrirFenriren-US
am_michaelMichaelen-US
am_puckPucken-US
bf_emmaEmmaen-GB
bm_georgeGeorgeen-GB

The choice is stored in the same voice preference as kokoro:<voice id>; no browser voiceURI starts with that prefix.

How it behaves:

  • Only where it can run. The list is shown only on a device with a WebGPU adapter that isn't a phone or tablet, and isn't on a metered connection unless the model is already downloaded, since loading it from the cache uses no data (readAloudPossible in packages/core/src/browser/readAloud.js). Phones and tablets are excluded by name, because WebGPU alone doesn't rule them out: an iPad has it, and Kokoro froze and then crashed Safari on one. iPadOS calls itself a Mac, so it's recognized by having a touch screen. A stored natural voice the device can't use right now (on a metered connection before it's downloaded, say) shows as the browser default in the picker, since that is what reads; the choice itself is kept.
  • Opt-in, because of the download. The browser default stays the default. Choosing a natural voice says that the first use downloads about 330 MB (DOWNLOAD_MB), once. Nothing downloads until practice mode speaks with speech on, so choosing one on the settings page, or with speech off, downloads nothing. The device check runs again just before, so a connection that has turned metered since the page loaded doesn't start one. The weights go into the shared model cache, which Delete models on the settings page clears.
  • Never silent while it downloads. During the download the browser's voice reads, and a line under the step count shows the progress. Screen readers hear that it's downloading once, not every percent. The natural voice takes over from the next thing spoken.
  • Ready from the first step once downloaded. When the model is already in the cache (readAloudCached in readAloud.js looks for its weights), it loads as soon as practice mode opens with speech on, which takes a couple of seconds. A step spoken meanwhile waits for it, up to five seconds from when it was asked for (CACHED_WAIT_MS), instead of being read in the browser's voice, and the line under the step count says "Getting the natural voice ready...". If the load runs past that or fails, the browser's voice reads that step and later steps don't wait again: the natural voice takes over once it's loaded, and a failed load is tried again once a step, like a failed download. Web Audio is started as the step is asked for, while the click behind it still counts, so the wait doesn't cost the browser's permission to play. A voice changed during the wait is the one that reads.
  • The browser's voice is the fallback, in three ways, and the same line says which:
    • a download that fails is tried again on the next step, up to three attempts in a visit (MAX_LOAD_ATTEMPTS; files that finished are cached, so a retry picks up where a dropped connection left off);
    • a chunk the model fails to make costs only that step: what was already queued plays out, then the browser's voice reads from the failed chunk. Two steps in a row like that (MAX_READ_FAILURES) and the natural voice is dropped for the visit;
    • if the browser won't start Web Audio (no recent click, a strict autoplay rule), that step is read by the browser's voice rather than queued in silence.
  • Steps play as one stream. Each chunk is made, then queued on a Web Audio timeline straight after the one before, so playback runs on while the next chunk is made.
  • The next step is made ahead. While a step plays, the first two chunks of the next step are made too (prepare in useSpeech.js, PREPARED_CHUNKS), so pressing Next starts speaking at once. This waits until the current step has been made: the model does one thing at a time, so running it alongside would hold up the chunks being listened to. Made clips are kept (up to 24, MAX_CACHED_CLIPS), so replaying a step or a spelling word doesn't make it again.
  • English only. Like choosing an English browser voice, picking one for a lesson in another language reads it with English pronunciation.

Kokoro reads phonemes, not text. packages/core/src/browser/readAloudEngine.js spells out numbers and abbreviations, turns the words into IPA with Spellophone (our WebAssembly build of espeak-ng, the phonemizer Kokoro was trained against, installed as @spelling-creator/spellophone), and passes that to the model. The text clean-up follows kokoro.js, the reference JavaScript port. That package isn't used itself because it pins transformers.js 3, which would put a second ONNX runtime in the bundle. Only Spellophone's English data is bundled (about 830 KB), as hashed assets of the build, not fetched from a CDN. readAloudEngine.js is only ever reached through a dynamic import() in readAloud.js, so none of this weight lands in the main bundle.

Timing it ​

To time it on a device, run pnpm dev:web and open /bench/read-aloud.html (apps/web/bench/read-aloud.html). The page reads the first six steps of a real hub lesson, chunked as useSpeech.js chunks them, and reports:

  • first audio: from pressing play to sound, once the model is loaded;
  • RTF (real-time factor): time to make a chunk over how long it speaks. Under 1 keeps ahead of playback;
  • stalls: silence mid-step while the next chunk is still being made.

The page is served by the dev server only; the production build's one entry is index.html. WebGPU needs a secure context, so a phone or tablet has to reach it over HTTPS.

Measured on a Mac, 110 s of speech, af_heart:

Browser and backendFirst audioMedian RTFStalls
Chrome 154, WebGPU, fp320.56 s0.17none
WebKit 26.6, WebGPU, fp320.85 s0.250.4 s
Chrome 154, WASM, fp322.4 s1.1639 s
Chrome 154, WASM, q83.0 s1.4870 s

What the numbers decided:

  • WebGPU or nothing. The site isn't cross-origin isolated, so the WASM backend runs on one thread, and on one thread Kokoro is slower than speech even on a Mac. q8 is slower than fp32 there, too. The device check asks for WebGPU for this reason.
  • A throwaway run while loading. The first run is slow while WebGPU compiles its shaders (1.9 s against 0.4 s for the same short line), so the engine does one as part of loading.
  • Making the next step ahead. The one WebKit stall is a short section name followed by a long chunk: the name finishes playing before the next chunk is ready. The first two chunks of every step after the first are made while the step before plays.
  • Computers only. On an iPad it freezes and then crashes (Safari, WebGPU, fp32). Kokoro is aimed at computers, which is what the spellers we know use for sessions, and phones and tablets keep the browser's voices.

Where the code lives ​

FileWhat it does
packages/core/src/interactive.jsTurns a document into steps; speech text; the answers the reveal shows; limits and validation.
packages/core/src/browser/interactiveProgress.jsThe unfinished run-through kept on the device: resume, expiry, pruning.
packages/core/src/lessonResponses.jsClient for the three endpoints above.
apps/api/src/routes/lessonResponses.jsThe endpoints, and the privacy scoping.
apps/web/src/components/InteractiveLesson.jsxThe full-screen walkthrough, the reveal, speech controls, resume and summary.
apps/web/src/pages/lesson/LessonPractice.jsxThe Practice tab route that opens it.
apps/web/src/components/MyLessonAnswers.jsxThe private "Your answers" panel on the lesson page.
apps/web/src/pages/lesson/LessonLayout.jsxStart vs. Continue lesson on the lesson page's button.
apps/web/src/lib/useSpeech.jsSpeaking: the browser's voices or a natural one, the fallback, making ahead.
apps/web/src/lib/speechPrefs.jsThe read-aloud preferences and voice lists, shared with the settings page.
apps/web/src/components/SpeechVoiceSelect.jsxThe voice picker the popover and the settings page share.
packages/core/src/browser/readAloud.jsThe device check for natural voices, the cache check, and the door to their engine.
packages/core/src/browser/readAloudEngine.jsKokoro: text clean-up, espeak-ng phonemes, the model.
packages/core/src/browser/readAloudVoices.jsThe model id and pin, and the Kokoro voices on offer, listable without loading the engine.
apps/web/bench/read-aloud.htmlThe dev-only timing page for Kokoro.

Copyright © 2026 Spelling Creator.