You're learning the vitamins for an exam. At the stop for Vitamin C, you say "Vitamin D". The app turns the stop green. You just practised the wrong answer.
That app is mine. It's called Loci, and I found this while writing this post. Type "Vitamin D" at that stop and Loci marks it wrong, as it should. Say it out loud and the stop lights up.
Loci's answer checker has one rule: forgive the typo, never the wrong answer. This post is how that rule was built and tested, and how the voice path got around it.
- Live: https://loci.edycu.dev
- Code: https://github.com/edycutjong/loci (MIT)
- Demo video (2:30): https://youtu.be/I5sx5C28xR0
What Loci does
Loci is the memory palace, with your own room as the palace. Take a photo of your room and paste a list of 3 to 12 items, in order. An AI finds the objects in the photo. Plain code picks one object per item and joins them into one route. A second AI call writes a short, strange scene for each stop.
Then the lights go out. At each stop you say or type the item. A right answer turns the pin green and lights that part of the room again.

The real app, recorded from the production build. The room is AI-generated, and a script types the answers.
The AI proposes. Plain code decides. Whether you remembered is decided by shared/score.ts, plain TypeScript, never by the AI.
The first rule: one typo budget for the whole answer
The first rule in my planning notes was simple. Allow typos up to 20% of the answer's length, at most 2, anywhere in it. The repo keeps that rule in a regression test:
// tests/regressions.test.ts: the planning notes' first rule
const wholeAnswerRule = (said: string, target: string) => distance(normalize(said), normalize(target)) <= Math.min(2, Math.floor(0.2 * normalize(target).length));
"vitamin c" is 9 characters with the space. 20% of 9, rounded down, is 1 typo. "vitamin d" is exactly 1 typo away, so it passes. "Henry VII" passes for "Henry VIII" the same way.
That rule never shipped. Before any checker code existed, the plan listed pairs that must be wrong, and these two broke the rule on paper. The plan was made with the Devpost Learn skill pack, and it's in the repo's devpost/ folder.
The fix: the same 20%, per word
The fix kept the 20%. It moved it from the answer to each word:
// shared/score.ts
/** Typos forgiven in one word: 0 for words of up to 4 letters, 1 for 5–9 letters, 2 for 10 or more. */
export const allowance = (wordLength: number) => Math.min(2, Math.floor(0.2 * wordLength));
/** At most this many typos across a whole answer. */
export const MAX_TOTAL_TYPOS = 3;
/** Is `answer` close enough to one target text? */
export function closeEnough(answer: string, target: string): boolean {
const a = normalize(answer);
const t = normalize(target);
if (!a || !t) return false; // an empty answer is never right
if (a === t) return true;
const aWords = a.split(" ");
const tWords = t.split(" ");
if (aWords.length !== tWords.length) return a.replaceAll(" ", "") === t.replaceAll(" ", ""); // "oatmilk" = "oat milk"
let total = 0;
for (let i = 0; i < tWords.length; i++) {
const typos = distance(aWords[i], tWords[i]);
if (typos > allowance(tWords[i].length)) return false; // "vitamin c" ≠ "vitamin d": short words must be exact
total += typos;
}
return total <= MAX_TOTAL_TYPOS;
}
| Word | Typos forgiven |
|---|---|
| up to 4 letters | 0 |
| 5 to 9 letters | 1 |
| 10 letters or more | 2 |
| the whole answer | 3 at most |
"Vitamin" may have one typo. "C" may have none. So "occulomotor" counts for Oculomotor, and "Vitamin D" never counts for "Vitamin C".
Before comparing, normalize lowercases both texts and drops accents, punctuation and numbering like "1.". distance counts two swapped letters as one typo.
60,000 answers with known verdicts
Unit tests pin the edges: a word at its allowance passes, and one more typo fails. On top of that, fast-check generates 60,000 answers per run, across 6 properties.
Each answer is built so its verdict is known without asking the checker. Change, drop or add one letter in a word of up to 4 letters, and the answer must be wrong:
// tests/score.property.test.ts (itemWith, said and replace are small helpers at the top of the file)
it("a word of up to 4 letters must be exact: one letter changed, missing or added is never right ('Vitamin D' for 'Vitamin C')", () => {
const edit = fc.record({ kind: fc.constantFrom("change", "drop", "add"), at: fc.nat(), shift: fc.integer({ min: 1, max: 25 }), letter: fc.constantFrom(...ALPHA) });
fc.assert(
fc.property(itemWith([1, 4]), edit, ({ words: ws, i }, { kind, at, shift, letter }) => {
const w = ws[i];
const p = at % w.length;
const offByOne =
kind === "change" ? w.slice(0, p) + ALPHA[(ALPHA.indexOf(w[p]) + shift) % 26] + w.slice(p + 1)
: kind === "drop" ? w.slice(0, p) + w.slice(p + 1)
: w.slice(0, p) + letter + w.slice(p);
return !closeEnough(said(replace(ws, i, offByOne)), said(ws));
}),
{ numRuns: RUNS },
);
});
Every one of those 60,000 answers goes through closeEnough. Keep that in mind.
Voice needs more ways to say yes
Speech is messier than typing, in three ways:
- Split words. Speech-to-text can break a long word in two: "vestibular cochlear" for Vestibulocochlear.
- Sound-alikes. Loci's scenes use them on purpose, as memory hooks: "trochlear" sounds like "truck-lear".
- Echoes. When Chrome's on-device speech model is installed, Loci passes it your list as hints. While recording the demo (recorded audio, not a live microphone), its final answers echoed them about once in a hundred words: "Cambrian Jurassic Quaternary" when only "Quaternary" was said.

A scene with its sound-alike. The room is AI-generated.
So for each stretch of speech, the voice path tries the typing checker first, then two more rules:
// shared/voice.ts
/** Is this stretch of speech the item (or one of its accepted answers)? */
export function spokenMatch(spoken: string, item: Item): boolean {
for (const target of [item.text, ...item.accepts]) {
if (closeEnough(spoken, target)) return true;
const a = normalize(spoken).replaceAll(" ", "");
const b = normalize(target).replaceAll(" ", "");
if (b.length >= 5 && distance(a, b) <= allowance(b.length)) return true; // "vestibular cochlear"
if (b.length >= 5 && soundsAlike(a, b)) return true; // "truck lear"
}
return false;
}
The sound-alike rule compares consonant skeletons. It drops the vowels, h, w and y, and merges letters that sound close, so "truck lear" and "trochlear" both become trklr. Keys shorter than 4 letters never count. For echoes, a second reading skips leading words that name stops you already answered.
Voice is forgiving in one more, safe way. A phrase that matches nothing in your list shows "didn't catch that", and it costs nothing.
The old rule came back
Look at the second rule again. Join the words, then allow typos by the length of the whole thing. That's the first rule, the one the plan threw out. It came back in the voice slice of the build, the same day, under a comment about "vestibular cochlear".
// what the two functions return at commit d0e5df9
import { closeEnough } from "./shared/score";
import { spokenMatch } from "./shared/voice";
closeEnough("Vitamin D", "Vitamin C"); // false: typed, the stop stays dark
spokenMatch("vitamin d", { text: "Vitamin C", accepts: [] }); // true: said, the stop turns green
"vitamind" against "vitaminc" is 8 letters, so 1 typo is allowed, and 1 is used. The full voice interpreter agrees: at the Vitamin C stop, hearing "vitamin d" returns { kind: "right", stop: 0 }. "Henry VII" and "Henry VIII" pass for each other the same way, in both directions. So would two numbered items one digit apart, like "Type 1" and "Type 2". The sound-alike key drops digits too.
Step 7 of the diagram above says "short words exact" and names two files. Only shared/score.ts keeps that promise.
Why didn't the tests catch it?
-
The 60,000 answers only test
closeEnough.spokenMatchcalls it first, then has two more ways to say yes. No property runs through those. - The voice tests use the example lists. They check that no two items in the same list match each other. That passes, because no example list has two items that close.
- The echo fix's regression test even uses Vitamin C and Vitamin D. It checks that "vitamin d" still counts at the Vitamin D stop. It never asks what happens at the Vitamin C stop.
A list of pairs that must fail only protects the functions you run it through. Every path that can say yes needs the whole list.
Update (2026-10-05): fixed in commit 60f6fd4. The two extra voice rules now run only when every short word of the item was heard exactly, and the must-fail pairs run through spokenMatch and the interpreter in their own test. The price: voice is now as strict as typing on short words. "World War 2" no longer counts for "World War II" unless you add it as an accepted answer.
The AI half, briefly
The AI does two jobs: find the objects in your photo (Gemini first, with fallbacks) and write the scenes (DeepSeek first). Code checks both answers before anything uses them. Malformed boxes and empty scenes are dropped, never repaired.
One prompt fix from the build: with only a rule telling it to spell the item exactly, DeepSeek dropped the item's exact spelling in 8 of 12 scenes that used a sound-alike. One worked example in the prompt fixed it:
- Write the item itself, spelled exactly as given, in every scene, even when the image rests on a sound-alike. Example for the item "Trochlear" on a wall clock: "Your wall clock turns into a little truck-lear, honking 'Trochlear!' as it circles the room."
On the live site, a script built 9 palaces the way a person would: 3 AI-generated rooms, each uploaded as if it were your own photo, times 3 lists. Nothing was stubbed.
- 9 / 9 palaces built
- 102 / 102 scenes spell their item
- median wait 14.45 s from "Build my palace" to the first scene, slowest 30.02 s
- $0.0115 billed for all nine ($0.1151 at list prices; the Gemini calls ran on a free-tier key)
The full record, with every model's answer, is in DEMO.md.
Limits
- One sitting only. Loci shows a first-try score, learning time and recall time. It makes no claim about next week and has no spaced repetition.
- Voice is Chrome and Edge only, in US English. It's checked with a stand-in recognizer and with recorded audio, never with a person at a microphone.
- The example rooms are AI-generated. Real rooms are messier.
- Your photo goes to an AI once (Gemini, or DeepSeek if Gemini is busy). A free-tier Gemini key lets Google use what it receives, so leave people and papers out.
-
Typing is strict about extra words. "optic nerve" stays dark for "Optic". Add it as an accepted answer:
Optic / Optic nerve.
Try it
Open https://loci.edycu.dev and tap Try it: 12 cranial nerves. No photo, no sign-up, no AI call. Tap Next stop a few times, then Lights out, and type what you remember. In Chrome or Edge, Say it instead switches to voice.
- Code: https://github.com/edycutjong/loci
- The 30-second path for judges: https://loci.edycu.dev/judge/
- Story page: https://loci.edycu.dev/story/
- Devpost: https://devpost.com/software/loci-4isou6
Loci was built for Build With AI: Basics on Devpost.
If you write a fuzzy matcher, write the pairs that must fail before you write the rule. Then run them through every function that can say yes, not just the first one.
I wrote this with help from an AI assistant, working from the repo's code and docs. The rooms in the images are AI-generated, and the demo video uses an AI voice.

Top comments (1)
The update's stricter voice boundary is a useful correction. I'd also run the must-fail pairs through the interpreter at different positions in the route, not only through spokenMatch: before answering Vitamin D, after answering it, and while Vitamin C is current. That would pin the interaction with the rule that strips already-answered stops from an echoed transcript.
Explicit accepted answers seem worth testing at the list level too. If someone gives two stops the same alias, does the interpreter reject the ambiguity or let the current stop decide? A shared alias could bring back a false green even with the short-word guard intact.