Interactive mode
For how to use this, see Interactive mode.
Interactive mode is the lesson's Practice tab (/hub/:id/practice, see Pages & routing), opened from Start lesson on the lesson page. A section's material appears as one step, then that section's questions one after another, each with a field to type an answer into. At the end, a signed-in user's answers are filed privately to their account; an unfinished run-through is kept in the browser so it can be resumed.
The route only decides when it is open. InteractiveLesson is still a dialog, deliberately: it is a focus mode with its own bottom bars and its own idea of the viewport. LessonPractice renders it with open and navigates back to the overview on close (rather than history.back(), since a deep link has no history to go back to). A lesson whose sections are all empty isn't playable (isInteractivePlayable): LessonTabs hides the tab, LessonLayout hides the button, and a direct link to /practice redirects to the lesson.
Derived from the document
There is no "interactive lesson" document type and nothing to switch on when authoring. The walkthrough is derived from the lesson document (buildInteractiveSteps in packages/core/src/interactive.js), so every lesson ever published works, including ones made before this feature existed and ones written by the MCP server. Nothing is added to a lesson to make it playable, and a lesson stays exactly as printable as it was.
| In the document | Becomes |
|---|---|
| A section's text, image, spelling and VAKT blocks | One content step (kind: "content"), holding them together in document order. |
| Each question block | One question step (kind: "question"), after that section's content. |
| A section with only questions | No content step; it opens straight on its first question. |
| A section with nothing in it | Nothing. |
| A lesson with no questions at all | A read-through: every content step, no answer fields, nothing saved. |
VAKT blocks belong to the content step rather than being steps of their own: a regulation break is something whoever runs the session does with the speller, not something the speller answers, so it must not be counted by the progress bar.
Each step has a stable key (<sectionId>:content or <sectionId>:<blockId>, falling back to positions for documents without ids). Answers are keyed by block id (answerKey), so re-ordering a lesson between sittings doesn't shuffle which answer belongs to which question.
Text blocks keep their bold, italics and underlining, but not their footnote markers: this is the screen the speller reads, and a superscript number with nowhere on the screen to lead to is clutter there. The voice reads the plain words (see Formatting, footnotes & sources).
Questions are numbered from 1 within each section, matching the editor's Q7 numbering (see Navigating large lessons).
Presentation
Interactive mode is full-screen and drawn in the app's own theme, light or dark, as is the lesson page below it, rather than reproducing the white sheet the DOCX/PDF export produces. The blocks are re-rendered at a scale for reading and answering over a whole session: prose at reading size, images framed in the app's border and radius and sized by the reading column rather than by the size and alignment they carry (so a picture fills the width on a phone), spelling words as large cards, and a VAKT activity set apart as a red-edged card. Only the presentation differs; the content is the same blocks.
Showing the answers
questionAnswer(block) flattens each question type's stored answer into one shape, { answer, answers, steps, suggested }, or null when there is nothing to reveal:
| Type | What the reveal shows |
|---|---|
single, background | answer |
number | answer plus the working steps (either alone is still revealed) |
multiple | every entry in answers, the whole accepted set |
multiple_open | the same list with suggested: true, labeled as suggestions |
open, paraphrase, wyr, an unknown type, or a blank answer | null: "This question has no set answer." |
revealedAnswers turns that into the flat list of boxes the UI draws (the working is not one of them; it stays a numbered list below), and hasRevealableAnswers(steps) decides whether the toggle is rendered at all.
The two semi-open types share a color and an answers list because the S2C guidebook treats them as one family (see the comment in packages/core/src/questions.js). That grouping is the app's default, taken from one source; other Spelling practitioners categorize questions differently. suggested is what tells the two apart, and it travels with the answers so the person looking at the reveal doesn't have to remember the question type.
The toggle state lives in InteractiveLesson and is reset to off every time the dialog opens; unlike the speech settings it is never persisted. On a question step each answer box is a button (RevealedAnswer with onUse) that replaces the field's value and focuses it with the caret at the end. On the summary onUse is absent, which is what keeps the boxes there read-only. Nothing compares the typed answer with the author's anywhere.
Progress on the device
packages/core/src/browser/interactiveProgress.js keeps the unfinished run-through (every answer, plus the current step's key) in localStorage under spelling-creator:interactive-progress.
| Progress (unfinished) | A saved run-through (finished) | |
|---|---|---|
| Lives in | this browser (localStorage) | lesson_responses, on the server |
| Needs | nothing; signed out works too | a signed-in session |
| Travels | no: this device only | yes: any device you sign in on |
| Kept until | it is filed, you start again or discard it, or 90 days pass | you delete it |
| Anyone else? | never sent anywhere at all | only you can read it |
- Records are keyed by owner and lesson. The owner is the signed-in user's id, or
""when signed out, so signed-out users of one browser share a record per lesson. Per-owner keys exist for the shared computer (a clinic, or a home with more than one speller), where resuming into the previous person's answers would be worse than not resuming at all. The shared signed-out record is why a resumed run-through always shows the resume notice with Start again. - A run-through is pinned to the owner it started as (
runOwnerinInteractiveLesson). If the signed-in account changes with the dialog open, writes keep going to the original record, and Finish refuses to file to the new account (summary.signedInSince). MAX_PROGRESS_RECORDSis 20 (most recently touched lessons) andPROGRESS_MAX_AGE_MSis 90 days, applied bypruneProgress. Pruning is fine here in a way it isn't for saved run-throughs: this is a resume cache, not the only copy of anything someone chose to keep.- Every write reports whether it landed. Where storage is refused (private browsing, a full quota) or missing, the leave confirmation switches back to warning that leaving discards the answers, so it only promises what was actually kept.
- The record is cleared as soon as a run-through is filed. A failed save leaves it, so closing and coming back is a way to retry; signed out, it stays, since it is the only copy.
- If the remembered step key no longer exists (the step was deleted), the lesson opens at the top with the answers still restored.
hasInteractiveProgressis what makesLessonLayoutlabel the button Continue lesson.
Syncing progress across devices was rejected on purpose: it would mean putting half-written answers on the server, a much bigger promise than "your tab remembers".
Privacy of saved answers
- Every endpoint that touches saved answers requires a signed-in session, and the Worker scopes each query to
user_id = <verified caller>; that filter is the only way a row is ever addressed, not a check layered on top of one. Deleting someone else's row matches nothing and returns 404. - There is no endpoint that returns another user's answers: not for the lesson's author, a moderator or an admin.
lesson_responseshas RLS enabled with no policies, unlikelessons,commentsandratings; only the service-role Worker reads or writes it.user_idis never returned.- Answers are not run through the profanity filter that comments go through. There's no audience to protect.
- The in-progress copy has no endpoint at all: nothing to scope server-side, and no way for anyone to learn that a lesson was even opened.
MyLessonAnswers renders the "Your answers" panel on the lesson's overview tab (LessonOverview.jsx), under the lesson, and renders nothing when signed out or when there are none.
Worker endpoints
| Method & path | Auth | Response |
|---|---|---|
GET /lessons/:id/responses | Bearer <Supabase JWT> | { "responses": [{ id, lessonId, answers, completedAt }] }; the caller's own only, newest first |
POST /lessons/:id/responses | Bearer <Supabase JWT> | { "response": { id, lessonId, answers, completedAt } } |
DELETE /lessons/:id/responses/:rid | Bearer <Supabase JWT> | { "ok": true }, the caller's own only; else 404 |
POSTbody is{ answers }, one entry per question:{ blockId, sectionId, sectionName, questionType, prompt, answer }. The Worker rebuilds every entry from known fields and drops anything else, so the storedjsonbcan only hold that shape.answermust be a string of at most 5,000 characters or the request is refused with400; the other fields are cut to their maximum length (100 for ids, 300 for the section name, 2,000 for the prompt), and an unrecognizedquestionTypeis stored asopen.- The prompt is snapshotted alongside the answer on purpose: a saved run-through has to stay readable after the lesson is edited, re-ordered, or has that question deleted.
- Skipped questions are stored as blank answers rather than dropped, so the set still says which questions were asked.
- Limits (shared between browser and Worker in
packages/core/src/interactive.js):MAX_RESPONSE_LENGTH5,000 characters per answer andMAX_RESPONSES500 answers per submission. - You may keep 20 saved run-throughs of any one lesson (
MAX_STORED_RESPONSES). Past that aPOSTis rejected with409and a message asking you to delete an older one, rejected rather than silently pruning the oldest, for the same reason the draft cap (see Lesson hub & accounts) is: they're the user's own answers, and quietly deleting them to make room isn't ours to decide. POSTalso checks the lesson is one the caller could have read in the first place (canPlayLesson): published and not shadowbanned, or theirs, trusted or moderated. Otherwise404.
Supabase schema
create table if not exists public.lesson_responses (
id uuid primary key default gen_random_uuid(),
lesson_id uuid not null references public.lessons (id) on delete cascade,
user_id uuid not null references auth.users (id) on delete cascade,
answers jsonb not null,
completed_at timestamptz not null default now()
);
create index if not exists lesson_responses_user_lesson_idx
on public.lesson_responses (user_id, lesson_id, completed_at desc);
-- No public read policy, unlike lessons/comments/ratings: this data is private.
alter table public.lesson_responses enable row level security;The full schema, with the reasoning in comments, is apps/api/schema.sql.
Reading aloud
Speech uses the browser's Web Speech API (speechSynthesis), or a natural voice where the device can run one. Like lesson summaries, this runs entirely on the reader's own device: no Worker call, no API key, no cost, and the lesson text never leaves the machine. Speech is probed for (speechSupported) rather than assumed, from an effect so server rendering and hydration agree; where it's missing, neither the controls nor the settings section is rendered.
stepSpeechText(step) builds what a step says: the section name, then the prose, image captions (never their credits), spelling words, a VAKT activity's text without its label or links, or the question prompt. A question's answer is deliberately never included, even with the reveal on.
The on/off, voice and pace preferences live in localStorage under spelling-creator:tts-enabled, spelling-creator:tts-voice and spelling-creator:tts-rate, read and written through apps/web/src/lib/speechPrefs.js by both interactive mode and the Reading aloud section of the settings page. SPEECH_RATES is [0.7, 0.85, 1, 1.25, 1.5]; an unrecognized stored rate falls back to 1. A change made in one place reaches the other the next time it mounts, not live (no storage listener).
Three platform quirks are handled between the two files. speechPrefs.js takes the one that belongs to the voice list: voices load asynchronously, announced by voiceschanged. useSpeech.js takes the two that belong to speaking: Chromium cuts off a single utterance after about 15 seconds (so text is split into chunks of at most 180 characters, MAX_CHUNK, and queued), and cancel() isn't synchronous (so a new utterance is deferred a tick after one).
Natural voices (Kokoro)
The natural voices are Kokoro, an 82M-parameter voice model that runs in the page with transformers.js (@huggingface/transformers 4), like the summary and import models. The model is pinned to a commit (MODEL_REVISION in readAloudVoices.js) and loaded at fp32 on WebGPU. Eight of its English voices, the ones graded C+ or better on the model card, are listed under Natural voices in the voice picker, above the browser's own, in both the popover and the settings page (apps/web/src/components/SpeechVoiceSelect.jsx):
| Id | Name | Accent |
|---|---|---|
af_heart (default) | Heart | en-US |
af_bella | Bella | en-US |
af_nicole | Nicole | en-US |
am_fenrir | Fenrir | en-US |
am_michael | Michael | en-US |
am_puck | Puck | en-US |
bf_emma | Emma | en-GB |
bm_george | George | en-GB |
The choice is stored in the same voice preference as kokoro:<voice id>; no browser voiceURI starts with that prefix.
How it behaves:
- Only where it can run. The list is shown only on a device with a WebGPU adapter that isn't a phone or tablet, and isn't on a metered connection unless the model is already downloaded, since loading it from the cache uses no data (
readAloudPossibleinpackages/core/src/browser/readAloud.js). Phones and tablets are excluded by name, because WebGPU alone doesn't rule them out: an iPad has it, and Kokoro froze and then crashed Safari on one. iPadOS calls itself a Mac, so it's recognized by having a touch screen. A stored natural voice the device can't use right now (on a metered connection before it's downloaded, say) shows as the browser default in the picker, since that is what reads; the choice itself is kept. - Opt-in, because of the download. The browser default stays the default. Choosing a natural voice says that the first use downloads about 330 MB (
DOWNLOAD_MB), once. Nothing downloads until practice mode speaks with speech on, so choosing one on the settings page, or with speech off, downloads nothing. The device check runs again just before, so a connection that has turned metered since the page loaded doesn't start one. The weights go into the shared model cache, which Delete models on the settings page clears. - Never silent while it downloads. During the download the browser's voice reads, and a line under the step count shows the progress. Screen readers hear that it's downloading once, not every percent. The natural voice takes over from the next thing spoken.
- Ready from the first step once downloaded. When the model is already in the cache (
readAloudCachedinreadAloud.jslooks for its weights), it loads as soon as practice mode opens with speech on, which takes a couple of seconds. A step spoken meanwhile waits for it, up to five seconds from when it was asked for (CACHED_WAIT_MS), instead of being read in the browser's voice, and the line under the step count says "Getting the natural voice ready...". If the load runs past that or fails, the browser's voice reads that step and later steps don't wait again: the natural voice takes over once it's loaded, and a failed load is tried again once a step, like a failed download. Web Audio is started as the step is asked for, while the click behind it still counts, so the wait doesn't cost the browser's permission to play. A voice changed during the wait is the one that reads. - The browser's voice is the fallback, in three ways, and the same line says which:
- a download that fails is tried again on the next step, up to three attempts in a visit (
MAX_LOAD_ATTEMPTS; files that finished are cached, so a retry picks up where a dropped connection left off); - a chunk the model fails to make costs only that step: what was already queued plays out, then the browser's voice reads from the failed chunk. Two steps in a row like that (
MAX_READ_FAILURES) and the natural voice is dropped for the visit; - if the browser won't start Web Audio (no recent click, a strict autoplay rule), that step is read by the browser's voice rather than queued in silence.
- a download that fails is tried again on the next step, up to three attempts in a visit (
- Steps play as one stream. Each chunk is made, then queued on a Web Audio timeline straight after the one before, so playback runs on while the next chunk is made.
- The next step is made ahead. While a step plays, the first two chunks of the next step are made too (
prepareinuseSpeech.js,PREPARED_CHUNKS), so pressing Next starts speaking at once. This waits until the current step has been made: the model does one thing at a time, so running it alongside would hold up the chunks being listened to. Made clips are kept (up to 24,MAX_CACHED_CLIPS), so replaying a step or a spelling word doesn't make it again. - English only. Like choosing an English browser voice, picking one for a lesson in another language reads it with English pronunciation.
Kokoro reads phonemes, not text. packages/core/src/browser/readAloudEngine.js spells out numbers and abbreviations, turns the words into IPA with Spellophone (our WebAssembly build of espeak-ng, the phonemizer Kokoro was trained against, installed as @spelling-creator/spellophone), and passes that to the model. The text clean-up follows kokoro.js, the reference JavaScript port. That package isn't used itself because it pins transformers.js 3, which would put a second ONNX runtime in the bundle. Only Spellophone's English data is bundled (about 830 KB), as hashed assets of the build, not fetched from a CDN. readAloudEngine.js is only ever reached through a dynamic import() in readAloud.js, so none of this weight lands in the main bundle.
Timing it
To time it on a device, run pnpm dev:web and open /bench/read-aloud.html (apps/web/bench/read-aloud.html). The page reads the first six steps of a real hub lesson, chunked as useSpeech.js chunks them, and reports:
- first audio: from pressing play to sound, once the model is loaded;
- RTF (real-time factor): time to make a chunk over how long it speaks. Under 1 keeps ahead of playback;
- stalls: silence mid-step while the next chunk is still being made.
The page is served by the dev server only; the production build's one entry is index.html. WebGPU needs a secure context, so a phone or tablet has to reach it over HTTPS.
Measured on a Mac, 110 s of speech, af_heart:
| Browser and backend | First audio | Median RTF | Stalls |
|---|---|---|---|
| Chrome 154, WebGPU, fp32 | 0.56 s | 0.17 | none |
| WebKit 26.6, WebGPU, fp32 | 0.85 s | 0.25 | 0.4 s |
| Chrome 154, WASM, fp32 | 2.4 s | 1.16 | 39 s |
| Chrome 154, WASM, q8 | 3.0 s | 1.48 | 70 s |
What the numbers decided:
- WebGPU or nothing. The site isn't cross-origin isolated, so the WASM backend runs on one thread, and on one thread Kokoro is slower than speech even on a Mac. q8 is slower than fp32 there, too. The device check asks for WebGPU for this reason.
- A throwaway run while loading. The first run is slow while WebGPU compiles its shaders (1.9 s against 0.4 s for the same short line), so the engine does one as part of loading.
- Making the next step ahead. The one WebKit stall is a short section name followed by a long chunk: the name finishes playing before the next chunk is ready. The first two chunks of every step after the first are made while the step before plays.
- Computers only. On an iPad it freezes and then crashes (Safari, WebGPU, fp32). Kokoro is aimed at computers, which is what the spellers we know use for sessions, and phones and tablets keep the browser's voices.
Where the code lives
| File | What it does |
|---|---|
packages/core/src/interactive.js | Turns a document into steps; speech text; the answers the reveal shows; limits and validation. |
packages/core/src/browser/interactiveProgress.js | The unfinished run-through kept on the device: resume, expiry, pruning. |
packages/core/src/lessonResponses.js | Client for the three endpoints above. |
apps/api/src/routes/lessonResponses.js | The endpoints, and the privacy scoping. |
apps/web/src/components/InteractiveLesson.jsx | The full-screen walkthrough, the reveal, speech controls, resume and summary. |
apps/web/src/pages/lesson/LessonPractice.jsx | The Practice tab route that opens it. |
apps/web/src/components/MyLessonAnswers.jsx | The private "Your answers" panel on the lesson page. |
apps/web/src/pages/lesson/LessonLayout.jsx | Start vs. Continue lesson on the lesson page's button. |
apps/web/src/lib/useSpeech.js | Speaking: the browser's voices or a natural one, the fallback, making ahead. |
apps/web/src/lib/speechPrefs.js | The read-aloud preferences and voice lists, shared with the settings page. |
apps/web/src/components/SpeechVoiceSelect.jsx | The voice picker the popover and the settings page share. |
packages/core/src/browser/readAloud.js | The device check for natural voices, the cache check, and the door to their engine. |
packages/core/src/browser/readAloudEngine.js | Kokoro: text clean-up, espeak-ng phonemes, the model. |
packages/core/src/browser/readAloudVoices.js | The model id and pin, and the Kokoro voices on offer, listable without loading the engine. |
apps/web/bench/read-aloud.html | The dev-only timing page for Kokoro. |