Rich text (comments & bios)
The two things users write about each other, lesson comments and profile bios, are rich text. Both are authored with a small tiptap editor (RichTextInput.jsx, built directly on @tiptap/react with a shadcn ToggleGroup toolbar) and stored as sanitized HTML.
Lesson text blocks are formatted too, but differently: they are stored as a tiptap JSON document rather than HTML, and carry footnotes rather than links. See Formatting, footnotes & sources.
Users can format text and link out. They cannot embed media: no images, video, audio or frames, whether uploaded, dragged in, pasted, or hand-written into a request. That is a deliberate product rule, and most of this page is about where it is actually enforced.
What a user can write
| Available | Not available |
|---|---|
Bold, italic, underline, strikethrough, inline code | Images, video, audio, <iframe>, <svg> |
| Bulleted and numbered lists, blockquotes | Headings, horizontal rules, code blocks, tables |
| Links (http/https/mailto) | Any attribute other than a validated href |
Headings and tables are left out on purpose: a comment is a paragraph, not a document.
The pieces
| File | Role |
|---|---|
@spelling-creator/core/richText | The policy. Allow-list, link schemes, link rel/target, HTML to text. |
packages/core/src/richText.test.js | Its tests: link safety, the allow-list, the browser-side helpers. |
apps/api/src/lib/richtext.js | The boundary. Sanitizes on write (parse5). |
apps/api/src/lib/richtext.test.js | Its tests: media, XSS, links, legacy values. |
apps/web/src/components/RichTextInput.jsx | The editor (toolbar, limits, no media affordances). |
apps/web/src/components/RichText.jsx | The renderer, for both rich text and legacy plain text. |
@spelling-creator/core/browser/sanitizeRichText | Render-time sanitizing (DOMPurify). |
The two sanitizers are deliberately separate, but not for the reason they originally were. The write-side one used to be built on HTMLRewriter, which only exists in the Workers runtime; it parses with parse5 now and would run anywhere. What keeps them apart is what they are for: the write-side pass decides what may be stored, and the render-side DOMPurify pass is defence in depth against rows written by an older server. What they agree on is shared: which tags survive, which link schemes are real links, what a surviving link is rewritten to carry, and how markup flattens to text. Those used to be two hand-maintained copies under different names (richTextToPlain on the Worker, richTextToPlainText in the browser), each with a comment asking the reader to keep them in sync.
Where "no media" is actually enforced
The editor has no image button, but a toolbar is a suggestion, not a boundary. Anyone can POST hand-written HTML straight at the Worker, so the rule is enforced in four places, only the third of which is load-bearing:
- No media button in the toolbar, so nothing offers it.
- No image/media node in the tiptap schema, so a pasted
<img>has nowhere to go and is dropped on the way in;handlePaste/handleDropadditionally refuse dropped and pasted files, which is what stops the browser from helpfully inlining a dragged-in screenshot as a giantdata:URI. - The Worker's sanitizer (
sanitizeRichText), the only one of these a hostile client cannot skip. It is an allow-list: anything not explicitly permitted is dropped, so a tag nobody thought of fails closed. Media tags are removed along with their content, and every attribute is stripped except a validatedhref. - A second sanitizing pass at render time (DOMPurify, in
@spelling-creator/core/browser/sanitizeRichText). The render path usesdangerouslySetInnerHTML, and a reader's safety shouldn't depend on assuming every row in the database was written by the current server code. Rows predating rich text, or written by some future path that forgets to sanitize, are caught here.
The sanitizer is built on parse5, a WHATWG-conformant HTML parser. Using a real parser is the non-negotiable part: it sees the same tag soup a browser would, where a regex-based "sanitizer" is the classic way to ship an XSS hole (<img/src=x onerror=…>, <scr<script>ipt>).
It was HTMLRewriter, the Workers runtime's own streaming parser, until self-hosting made that a problem: HTMLRewriter runs in one runtime, so any other host would have needed a second sanitizer, and two implementations of a security boundary that have to agree is a worse thing to own than a dependency, because the drift would be silent and it would be an XSS hole.
Swapping a streaming parser for a tree-building one brought one new concern with it. The routes cap raw input at 20,000 characters and <b> is three of them, so a caller can send a tree nearly 7,000 levels deep, enough to exhaust the stack in any tree-shaped walk. The walk is iterative, and nesting past 100 levels is flattened to its text, the same thing that happens to unrecognised markup everywhere else. The editor's deepest possible output is five.
It brought one other, sharper concern. An unknown tag is normally unwrapped (the tag goes and its words stay), but that is only safe for tags whose content was parsed as markup. HTML's raw-text elements (script, style, xmp, iframe, noembed, noframes, plaintext, noscript) hold literal text, and a serializer writes their content back unescaped because that is what a browser expects to find inside them. Unwrapping one would therefore turn its <script>alert(1)</script> body from words into markup on the way out. All of them are dropped with their content instead, and anything that survives the walk is re-parented onto the node it is actually written under, so escaping is never decided by an element that was removed.
Links
Links are the one thing a user may embed. Only http:, https: and mailto: targets survive (javascript: executes, and data: is an HTML/media smuggling channel), and the check is run against the URL with control characters stripped, so java\tscript: doesn't sneak past it. Every surviving link is rewritten to carry target="_blank" and rel="nofollow ugc noopener noreferrer": user-generated links open away from the app, earn spammers no SEO value, and can't reach back through window.opener. A link whose target fails the check keeps its words and loses the link.
The editor applies the same rule up front (tiptap's isAllowedUri hook runs isSafeLinkCandidate from the shared policy), so a user is never shown a link, typed or autolinked, that the server would only strip on save. The "candidate" variant exists because autolink validates the matched text before a scheme is attached: example.com is judged as the http://example.com it will be stored as, while anything that already names a scheme is judged as-is. As everywhere else, this is the honest-user half; the Worker's check is the boundary.
Text, not markup
Everything downstream of storage wants text, not markup, and gets it from richTextToPlain in @spelling-creator/core/richText:
- The profanity filter scans words, not tag names and URLs.
- Length limits (2000 characters for a comment, 500 for a bio) count what the user wrote, so wrapping a sentence in
<strong>never costs them their budget. The editor's live counter calls the same function the server's limit does, so the number a user sees is the number the server enforces. This used to be two mirrored implementations that had to be changed together; it is now one. - The Atom feed's
<summary>, notification bodies, the profile meta/OG description, and the one-line bio in the followers list are all plain-text contexts. Markup rendered into them would appear as literal escaped tags in feed readers, search snippets and link previews.
A comment or bio consisting of nothing but media sanitizes down to nothing, and is rejected (or, for a bio, stored as empty) rather than saved as blank markup.
Editing
An author may edit their own comment after posting (PATCH /lessons/:id/comments/:commentId). Ownership is decided by comparing the stored author_id against the verified JWT, never by anything the request claims. The new body runs through the identical sanitize, length and profanity pipeline as a fresh post, so editing is not a way to launder content past the rules that applied when it was written.
A successful edit stamps comments.edited_at, which the thread renders as an "edited" marker; a comment never changes silently under someone who already read or replied to it. Moderators cannot edit a comment, only delete it (Moderation): rewriting someone's words under their own name is a power worth not having.
Bios have always been editable, and are simply saved again through POST /profile/bio.
Values written before rich text
Every comment and bio predating this feature is a bare plain-text string. They are detected (isRichTextHtml) and rendered as text, with white-space: pre-wrap so their line breaks survive, exactly as they always looked. They are never reinterpreted as markup years after the fact. The flattening helpers pass them through unchanged, so a plain-text value with no tags is its own plain text.