Skip to content

Document import experiment ​

Status (checked October 2026): decided, and it shipped. The rule-based parser from the second pass is the editor's Import from text (packages/core/src/documentImport.js). The second fine-tune is its optional on-device step for sections the rules can't read: the editor offers "Read with the on-device model" only on a device with WebGPU that can hold the weights, and loads the q4f16 file (about 640 MB) from playforgecoding/LFM2-1.2B-Extract-lesson-ONNX, pinned to commit 6dbcb51 (packages/core/src/browser/documentModelEngine.js). There is no CPU path in the browser and no hosted fallback. The findings below are kept as measured.

Can a small on-device model turn a lesson that exists only as a document (a Word file, a typed page) into lesson JSON? This page records a measured answer, taken in October 2026 with the script in packages/core/scripts/extract-eval/, so the question is settled by numbers rather than by instinct before any import feature is built on a local model.

The idea under test: the app already runs on-device models through transformers.js for translation and summaries, so an extraction model such as LFM2 Extract could take the same route, keep the document on the author's machine, and cost nothing per request.

Method ​

  1. Test set. The newest six-section lessons on the public hub (four at the time: The History of Domestic Cats, Pompeii, Volcanoes, An Introduction to Quantum Physics), fetched from GET /lessons/:id and cached.
  2. Documents. Each lesson is rendered two ways. docx is the app's own Word export (buildDocument), read back as raw text with mammoth, which is what a generic importer sees: no section headings, question types carried only by color and so invisible, answers after a gap on the question line. plain is a hand-typed style: section headings, a "Spelling words:" line, numbered questions with the answer in brackets.
  3. Splitting. A document is cut into sections before any model sees it, by structure rather than headings: a passage paragraph that follows questions starts a new section, a chunk with no questions is a closing paragraph and stays with its section, and whatever trails the last section (sources, footnote bodies) is dropped. This found 6 of 6 sections in all 8 documents. A small model does far better on one section (about 1,000 tokens) than on a whole lesson.
  4. Extraction. Each model runs on the CPU through transformers.js at q4, greedy decoding, one section at a time, with the lesson shape given as a JSON Schema in the system prompt (the LFM2 Extract model card's format). Keys a model drifts to (section_title, spelling_words, a singular answer) are mapped back before scoring, as an import would do.
  5. Scoring. Each extraction is compared with the section it was rendered from: share of passage words recovered, paragraph-level similarity, spelling words (F1), question prompts found (F1), question type, and answers (exact set match, ignoring case and punctuation). Then the extractions are rebuilt into a lesson through normalizeLessonFile and run through the lesson checks.

The question type is scored twice. "Types (model)" is the model's own guess. "Types (derived)" is a deterministic rule applied afterwards: "Would you rather" is wyr, "in your own words" is paraphrase, no answer is open, several answers are multiple (or multiple_open when one is not in the passage), a numeric answer is number, and a lone answer is single when it is in the passage and background when it is not. That is how an import would assign types, and the composite score uses it.

Two sections per lesson, so 8 sections per model and style, on an Apple M4 with 16 GB.

Results ​

modelstyleparsedpassage wordsspellingprompts F1types (model)types (derived)answerscompositetok/ss/section
LFM2-1.2B-Extractdocx100%32%45%44%20%23%24%34%3028
LFM2-1.2B-Extractplain100%74%100%97%26%49%50%74%3031
LFM2-350M-Extractdocx100%48%71%40%0%37%33%46%5410
LFM2-350M-Extractplain88%52%81%54%0%34%46%53%5712
LFM2-1.2B (chat)docx88%65%18%58%30%31%26%40%1584
LFM2-1.2B (chat)plain100%89%56%88%45%42%41%63%1583
Qwen3-0.6B (chat)docx63%59%50%49%27%35%36%46%1197
Qwen3-0.6B (chat)plain13%11%13%10%8%8%10%10%1179

Lesson checks on the rebuilt two-section lessons, averaged over the four lessons (the originals average 27 errors and 19 warnings, because three of the four predate the current checks):

modelstyleerrorswarnings
LFM2-1.2B-Extractplain12.010.5
LFM2-350M-Extractdocx3.35.8
LFM2-350M-Extractplain5.55.3
LFM2-1.2B (chat)plain12.06.8

The low error counts for the 350M model are not good news: it drops most of the questions, and a section with few questions has little for the checks to object to.

What the models got wrong ​

  • LFM2-1.2B-Extract on the Word export. On half the sections it ignored the schema and produced flat keys of its own (paragraph_1, question_1), which no alias mapping recovers. On one section it fell into a repetition loop ("word": "BOAT" to the token cap). Where it did follow the schema it split paragraphs into sentences and invented answers for open questions ("Wolf", "Fox", "Snake" for "Name a wild animal that hunts"), which is the Extract family's instinct to fill every field.
  • LFM2-1.2B-Extract on the typed style is the one result that looks like an import: 97% of prompts and 100% of spelling words found, 74% of passage words. But answers were right half the time (merged, renamed, or made up), and a type derived from a wrong answer is wrong too.
  • LFM2-350M-Extract is fast, keeps to the schema, and finds only a fifth to a third of the questions.
  • The two chat models are three to six times slower than the Extract models at the same size, and no more accurate. Qwen3-0.6B answered the open questions itself and broke its own JSON on most sections.

Both Extract models guessed the type no better than chance. Deriving it from the answers afterwards does better everywhere, so an import should never ask a model for it.

Conclusions ​

  1. Stock, none of these models is import-ready. The best case, the 1.2B Extract model on a tidy typed document, reaches about three quarters on the composite and half on answers. An import that gets half the answers wrong makes more work than typing the lesson in.
  2. The structure is not the hard part. Section splitting is deterministic and found every section. For the two regular formats tested, a parser (the way docxImport already reads the app's own export) would beat all four models on every column. A model earns its place only on genuinely messy documents.
  3. Fine-tuning is the next step if the local route is wanted, and the data is nearly free. The renderer in this script turns any hub lesson into training pairs (document text, lesson JSON) in both styles, and more layouts are a few lines each. Liquid publishes SFT notebooks for the Extract family. The targets to move are answers (copy, never invent; empty when none is printed), paragraphs (whole, not sentences), and schema adherence on the Word-export style.
  4. The type should always be derived, not extracted, and the lesson checks should run on the result either way, exactly as they do on a hand-written lesson.
  5. Speed is fine for a desktop import. About 30 seconds per section at 30 tokens per second on a CPU; WebGPU would be faster. The chat models are too slow to be worth their accuracy.

Second pass: parser first, model last ​

The first pass said the structure is not the hard part. The second pass tested that by building the import the other way round (parse.mjs): the line classifier that already finds the sections also says which lines are passage, spelling, question and working-out, so the parser takes those directly and asks a model one thing only, where a question line's prompt ends and its answer begins, and only for a line no rule can split. With no model, a heuristic (a trailing run of capitals, a short tail after the last colon) stands in.

To give the parser something to be unsure about, five more typed-up layouts were added (layouts.mjs): question and answer on separate lines (qa), answers as a bare run of capitals with no headings at all (caps), bulleted questions with answers in square brackets (bullets), numbered questions with the answer after a colon (colon), and a number question's working-out on a line of its own (worked). The splitter found 6 of 6 sections in all 28 documents. All six sections of all four lessons, 24 per layout:

strategystylepassage wordsspellingprompts F1types (derived)answerscompositemodel callss/section
parser + heuristicdocx100%100%100%89%100%98%00.0
parser + heuristicplain100%100%100%89%100%98%00.0
parser + heuristicqa100%100%100%89%100%98%00.0
parser + heuristiccaps100%100%99%86%96%96%00.0
parser + heuristicbullets100%100%100%89%100%98%00.0
parser + heuristiccolon100%100%100%89%100%98%00.0
parser + heuristicworked100%100%100%89%100%98%00.0
parser + LFM2-1.2B-Extractcaps100%100%83%27%32%68%11.625.3
parser + LFM2-1.2B-Extractcolon100%100%90%24%30%69%11.925.7

The parser with a heuristic is right about every passage, every spelling word and every prompt, and about every answer except a handful on the capitals layout. Handing the ambiguous lines to the model made things worse, not better: given one line and a two-field schema, the 1.2B Extract model still wrote its own answers ("Domestication is the slow change of a wild animal into a tame one", four variants of it) and returned objects where strings were asked for. On these layouts it should not be asked.

The 11 percent the derived type misses is not the parser's. Of the 371 questions, 21 are "In your own words, explain..." questions that three older lessons typed as open; the current standard and the newest lesson call them paraphrase, and the rule follows the standard. The rest are a handful of single/background swaps where "is the answer in the passage" is too blunt a test ("THE ROMANS" against a passage that says "Roman town"). Using the lesson checks' own grounding logic for that call would close most of it.

Training pairs for the fine-tune come from the same renderers. make-dataset.mjs takes every published hub lesson (13 at the time), renders each in every layout, splits them with the import's own splitter, and pairs every section with its lesson JSON in chat format (system schema, user document, assistant JSON). The first run, on the seven regular layouts, made 412 training examples and 84 held out from the two newest lessons. Lessons written by a hosted model to the authoring standard would add volume without any hand labelling.

Layouts the rules cannot read. That first dataset had a blind spot: the import only calls the model for what the rules cannot read, and the rules read all seven layouts, so nothing the model was trained on could ever reach it. Two more layouts fill that gap, both modelled on how a real document loses its structure:

  • nomarks: questions typed with no question marks and no numbers, the answer tacked on after a space, and the spelling line unlabelled ("Words to learn ...").
  • runon: a numbered list that lost its line breaks, as when it is copied out of a web page or a PDF, so a section's questions are all on one line.

For these two (HARD_LAYOUTS), a section becomes a training example only if sectionNeedsModel would send it to the model, and its target is what the document says: a prompt with no question mark where the layout dropped it, and no section name where the document shows no heading. The parser on them, same four lessons:

stylesections foundpassage wordsspellingprompts F1answerscompositeoffered to the model
nomarks24/24100%0%98%61%66%24/24
runon24/24100%100%0%0%40%24/24

Getting the sections right took two changes to the splitter, both of which leave the seven regular layouts exactly where they were (and none of their 168 sections offered to the model): a line with no closing punctuation is never a passage however long it is, and a numbered line that runs into the next number is questions. Before that, a long question with no question mark started a section of its own, and each run-on list did too. With the two layouts, the dataset is 529 training examples and 108 held out, 70 and 71 of them from nomarks and runon.

Where this leaves the local model ​

  1. Ship the parser. For any document with a recognizable layout, which covers the app's own export and every typed-up style tried here, the rules get 96 to 98 percent with no model, no download and no wait. The type is derived, and the lesson checks run on the result. This is now the editor's Import from text; the parser lives in packages/core/src/documentImport.js and the scripts here call it.
  2. Keep the model out of the regular path. Stock, it loses to a regular expression on the one job left for it. A fine-tuned LFM2 Extract is still the right tool for a document the parser cannot read at all (a scan, a page with no consistent layout), and the dataset for that is generated. Measure it with the same script on the held-out lessons before it goes anywhere near the import.
  3. Fall back to the hosted Worker only when neither the parser nor the model produces a lesson that passes the checks, mirroring the translator's built-in-first chain.

Where that stands now: the first two shipped (the fine-tuned model below is what the editor offers for sections the rules can't read). The third was not built. The API has no import route, and a section neither the rules nor the model can read stays marked as unread in the import dialog.

Converting a fine-tuned LFM2 to ONNX for transformers.js ​

The stock LFM2 files the app loads come from onnx-community, and transformers.js v4 no longer ships the conversion script that made them, so the route for a fine-tuned checkpoint had to be worked out. The onnx-community graphs carry the fingerprints of Microsoft's onnxruntime-genai model builder (its node names, GroupQueryAttention and MatMulNBits), and that builder is public, supports Lfm2ForCausalLM, and takes a local checkpoint folder:

bash
pip install onnxruntime-genai onnx onnx_ir onnxscript
python -m onnxruntime_genai.models.builder -i merged -o build-cpu -p int4 -e cpu \
    --extra_options shared_embeddings=false
python -m onnxruntime_genai.models.builder -i merged -o build-webgpu -p int4 -e webgpu
python relayout-onnx.py build-cpu onnx-repo q4 merged
python relayout-onnx.py build-webgpu onnx-repo q4f16 merged

Neither optimum-onnx (no LFM2 support) nor Liquid's own LiquidONNX wrapper (its output targets onnxruntime-genai, and its README says it is not loadable by transformers.js) does this on its own. The builder's raw output is not loadable either, for three small reasons that relayout-onnx.py fixes:

  • the convolution caches are named past.N.conv, where transformers.js feeds past_conv.N and maps present_conv.N back to it;
  • the key/value cache's head dimension is left symbolic, and transformers.js sizes the first empty cache from the declared shape, so it allocated a zero-width cache;
  • the chat template is a separate chat_template.jinja file, and the config lacks the transformers.js_config block that tells the browser to fetch the external weights file. The files also go under onnx/model_<dtype>.onnx.

One more for the CPU build: for a model that ties its input and output embeddings, which LFM2 does, the builder emits an int8 embedding lookup (GatherBlockQuantized) that onnxruntime-web's wasm backend has no kernel for. shared_embeddings=false keeps the embedding as a plain table, as the onnx-community files have it. The WebGPU backend runs the quantized one, so the q4f16 build keeps the smaller default.

Verified on the stock 350M Extract model with the same transformers.js version the app uses: through onnxruntime-node (38 tokens in 0.3 seconds), and in headless Chromium on both backends, WebGPU with q4f16 (0.8 seconds) and wasm with q4 (15.7 seconds, single-threaded). All three produced the same correct JSON. The notebook's last cells run exactly this, and the result is a repo run.mjs --models and the app's own loader can take.

One caveat found on the fine-tuned checkpoint itself: its WebGPU q4f16 export copies a real section faithfully, but its CPU q4 export paraphrases the passage instead of copying it, with or without the embedding option and with full-precision matmul compute, while the stock model's CPU export is fine. The cause was not found. publish-colab.ipynb builds an int8 CPU export instead and checks it before uploading, and that one is faithful.

The fine-tuned model through the same scorer ​

The published int8 export, scored with run.mjs --dtype int8 on the two lessons held out of training (two sections each, all seven layouts):

layoutparsedpassage wordsspellingprompts F1types (model)types (derived)answerscomposite
docx100%100%100%100%99%88%100%98%
plain100%100%100%100%96%88%100%98%
qa100%100%81%100%93%88%98%94%
caps75%75%75%75%73%63%75%73%
bullets100%100%100%100%96%88%100%98%
colon100%100%100%98%94%89%96%97%
worked100%100%100%100%93%88%100%98%

Set beside the stock model's 74 percent composite on the typed layout and 34 percent on the Word export, and the rules' 96 to 98 percent, the fine-tune has caught up with the parser on every regular layout. The one miss in the capitals layout is a single long section whose reply hit the 1,500-token cap before closing its JSON. Two things the sample in the training notebook did not show: the fine-tuned model's own type guesses are now right 93 to 99 percent of the time, above the derived rule, and the answers are exact on every layout but the capitals one. About 30 seconds a section on an M4's CPU through onnxruntime-node.

The second fine-tune, with the layouts the rules cannot read ​

Retrained on the dataset with nomarks and runon (521 training examples), and scored the same way but on every section of both held-out lessons, 12 per layout rather than 4, so the two tables are not on the same sample:

layoutparsedpassage wordsspellingprompts F1types (model)types (derived)answerscomposite
docx100%99%98%97%93%87%99%96%
plain100%100%100%98%96%87%100%97%
qa100%99%85%99%96%87%100%94%
caps100%99%98%98%97%87%99%96%
bullets92%91%92%90%87%80%92%89%
colon100%99%100%99%96%85%96%96%
worked92%89%92%89%88%78%92%88%
nomarks100%100%98%98%94%87%97%96%
runon92%92%92%91%88%80%91%89%

The two rows that matter are the last two, since those are the sections the import actually hands to the model: 96 and 89 percent, where the rules get 66 and 40. Every row below 94 is one section out of twelve whose reply is not valid JSON, and none of them hit the token cap: one run-on section wrote the spellingWords key inside the paragraph list, one worked section dropped an answer in as a bare string where a question belonged, and one bulleted section repeated itself. The import shows such a section as unread rather than guessing. On the sections that parse, every layout is 96 to 100 percent, and the capitals layout's long section now fits. The app pins this export, and loads only its q4f16 file, on WebGPU. The int8 file that runs well on a CPU is about 2.5 GB, too large to ask a browser to download, so a device without a capable WebGPU adapter is never offered the model at all (packages/core/src/browser/documentModel.js).

The model cards for both repos live in scripts/extract-eval/model-cards/ and are pushed with the Hub CLI (the README there has the commands). The published exports: LFM2-1.2B-Extract-lesson (merged weights) and LFM2-1.2B-Extract-lesson-ONNX (q4f16 for WebGPU and int8 for a CPU; the first fine-tune's q4 file was removed when the second replaced it, since its weights no longer matched). The training dataset is public too, under CC BY 4.0, as playforgecoding/spelling-creator-document-import; its card is scripts/extract-eval/dataset-card.md, and the README there has the upload command.

Running it again ​

bash
cd packages/core
node scripts/extract-eval/run.mjs --dry-run --styles all            # render + split only
node scripts/extract-eval/run.mjs --lessons 4 --sections 2          # first pass, models
node scripts/extract-eval/run.mjs --strategy rules --styles all     # the parser, no model
node scripts/extract-eval/rescore.mjs                               # re-score saved outputs
node scripts/extract-eval/make-dataset.mjs                          # fine-tuning pairs

The fine-tune is scripts/extract-eval/finetune-colab.ipynb: open it in Google Colab on a GPU runtime, upload the two JSONL files, set your Hugging Face name in the first cell, and run it top to bottom. It trains a LoRA adapter on the base Extract model, scores the held-out sections, merges the adapter, converts the result to ONNX with the onnxruntime-genai builder and relayout-onnx.py (q4f16 for WebGPU, int8 for a CPU; see above), and pushes a repo in the onnx-community layout, which run.mjs --models then measures against the same held-out lessons. A new upload only reaches the app once MODEL_REVISION in packages/core/src/browser/documentModelEngine.js points at its commit.

--models takes any transformers.js-compatible causal LM on the Hub; a model id containing "Extract" gets the model card's schema prompt, anything else the same schema inside an instruction. --local-models <dir> reads them from a folder laid out the Hub way instead (<dir>/<name>/onnx/model_<dtype>.onnx), for an export that has not been uploaded yet. Outputs, rendered documents and the summary table land in scripts/extract-eval/out/ (or the --out folder); models are cached in scripts/extract-eval/.cache/ (about 3.5 GB for the four above). All of it is gitignored.

Copyright © 2026 Spelling Creator.