DEV Community

Cover image for My first-choice AI model said no to 5 of 6 room photos. The app still answered each one in about 4 seconds.
Edy Cu
Edy Cu

Posted on

My first-choice AI model said no to 5 of 6 room photos. The app still answered each one in about 4 seconds.

After a stroke, some people with aphasia understand everything but can't find the words.

Hospital picture cards (EAT, DRINK, TOILET) look nothing like their life. Another kind of board is made from a photo of the person's own surroundings. In AAC (tools that speak for people who can't), it's called a visual scene display. Someone has to build it by hand, object by object.

Room to Speak does that part with AI. A caregiver takes one photo of the room where a parent with aphasia spends the day. A vision model finds up to 8 things he'd want to talk about and writes a short first-person phrase for each. Rings appear on the photo. He taps the kettle, and the phone speaks for him, with a phrase like "Aku mau kopi" (I'd like coffee).

Room to Speak in use: an example bedroom is picked, rings appear on its objects, and tapping the walker lights it up while the phone speaks

Real AI output on an AI-generated example bedroom, answered by deepseek-flash (step 3 of the ladder below). The wait is trimmed and the GIF has no sound.

Try it at roomspeak.edycu.dev (tap an example room, no login), read the code on GitHub (MIT), or watch the 1:48 demo. All room photos in this post are AI-generated.

Everything he does every day runs on the device. Only setup needs the internet, and setup is one AI call. On 2026-10-04 I sent six room photos to the live function. My first-choice model said no to five of them. All six still came back with spots, with a median wait of 4.03 s, for $0.00.

This post is about the code behind that, the bug that taught me to write it, and what happens to the answer once it arrives.

One call does all the AI work

How Room to Speak works: setup runs once and is the only step that needs the internet (photo, shrink, POST /api/detect, a three-model ladder with a 55 s budget, validate, box to ring, caregiver review, save on the device). Every day, he taps, the nearest ring wins, and the device voice speaks.

The browser shrinks the photo to 1600 px at most and posts it to one Vercel function, api/detect.ts. It's the only server code, and the API keys live there, so they never reach the browser. (A test builds the app with canary keys and checks that no built file contains them.)

The prompt in shared/prompt.ts asks for structure, not prose:

Pick up to 8 objects he would most plausibly want to talk about for daily needs, comfort,
people, or activities (e.g. kettle/cup -> a drink, TV/radio -> turn it on, window -> fresh air,
bed/chair -> rest, family photo -> a person, medicine, fan/AC, door, phone).
Skip structural or trivial things: walls, floor, ceiling, ceiling lamps, outlets, cables, generic clutter.
Only include objects that are clearly visible; fewer good spots beats 8 weak ones.

For each object return:
- label: short object name in ${name}
- phrase: what HE says, first person, natural everyday spoken ${name} (casual, not formal), max 6 words
- box_2d: [ymin, xmin, ymax, xmax] normalized to 0-1000
Enter fullscreen mode Exit fullscreen mode

Gemini also gets a JSON response schema, so the answer comes back as an array. Each box is [ymin, xmin, ymax, xmax] on a 0–1000 grid, whatever the photo size.

The line I'd keep is fewer good spots beats 8 weak ones. Here a wrong spot isn't noise. It's a button that says the wrong thing.

Three models, one 55-second budget

The function tries gemini-3.8-flash, then gemini-3.1-flash-lite, then deepseek-flash on DeepSeek's own API:

async function findSpots(image: string, lang: Lang) {
  const deadline = Date.now() + TOTAL_BUDGET_MS;
  const tried: string[] = [];
  for (const a of ATTEMPTS) {
    const key = a.key();
    if (!key) continue;
    const left = deadline - Date.now();
    if (left < 4_000) break;
    try {
      const raw = await a.run(image, lang, key, AbortSignal.timeout(Math.min(a.ms, left)));
      return { spots: validateSpots(raw), model: a.name };
    } catch (err) {
      tried.push(`${a.name}: ${(err as Error).name === "Error" ? (err as Error).message : (err as Error).name}`);
    }
  }
  throw new Error(tried.length ? tried.join("; ") : "no provider key configured");
}
Enter fullscreen mode Exit fullscreen mode

The rules in there:

  • Each step has its own time limit (a.ms): 12, 10 and 25 seconds.
  • All steps share one 55 s budget, which keeps the function under the browser's 70 s wait.
  • With less than 4 s left, it stops instead of starting a call it can't finish.
  • A provider without a key is skipped, so the same code runs with one key or two.

The timeout that skipped the fallback

The first version only fell back on error responses, like the 503 "high demand" the very first real call got. Then a real run timed out on the main model.

A timeout is not a response. AbortSignal.timeout() makes fetch throw a TimeoutError, and my loop only looked at res.status. So the error flew past the fallback and failed the whole request. The next model was never asked.

The fix was a try/catch around each step (863498a). A regression test now fakes that exact error and checks that the next model still answers:

// What fetch throws when AbortSignal.timeout() fires.
const timedOut: Answer = () => Promise.reject(new DOMException("The operation was aborted due to timeout", "TimeoutError"));
Enter fullscreen mode Exit fullscreen mode

The receipt, and how to check it

The run from the top of the post: six AI-generated room photos, sent to the production function on 2026-10-04, encoded the way the app does it. Nothing mocked.

Step Model Answered
1 gemini-3.8-flash 1 of 6
2 gemini-3.1-flash-lite 5 of 6
3 deepseek-flash 0 of 6

6/6 rooms answered · 43 spots · p50 4.03 s · max 7.38 s · $0.00 (Gemini API free tier)

The five rooms step 1 turned down still finished in 3.38 to 4.28 s, end to end. Step 1 said no fast, and step 2 picked them up. The one room step 1 did answer was the slowest of the six, at 7.38 s.

You can make one of these calls yourself, with no key and no clone. This sends the app's example bedroom to the live function:

{ printf '{"lang":"en","image":"'; curl -s https://roomspeak.edycu.dev/samples/bedroom.jpg | base64 | tr -d '\n'; printf '"}'; } \
  | curl -s https://roomspeak.edycu.dev/api/detect -H 'Content-Type: application/json' --data-binary @- -w '\n%{time_total} s\n'
Enter fullscreen mode Exit fullscreen mode

On 2026-10-04 it returned 8 spots (bed, walker, radio, water glass, medicine, window, prayer rug, thermos) from gemini-3.1-flash-lite in 6.9 s. The full six-room run, every spot and phrase, is in DEMO.md.

An answer is not a safe answer yet

The answer becomes buttons, so validateSpots checks every item first. Anything malformed is dropped, never repaired. The checks, from shared/validate.ts:

for (const item of raw) {
  if (out.length >= MAX_AI_SPOTS) break;
  if (!item || typeof item !== "object") continue;
  const { label, phrase, box_2d } = item as Record<string, unknown>;
  if (typeof phrase !== "string" || !phrase.trim()) continue;
  if (!Array.isArray(box_2d) || box_2d.length !== 4) continue;
  if (!box_2d.every((v) => typeof v === "number" && Number.isFinite(v) && v >= 0 && v <= 1000)) continue;
  const [ymin, xmin, ymax, xmax] = box_2d as number[];
  if (ymin >= ymax || xmin >= xmax) continue;
Enter fullscreen mode Exit fullscreen mode

Why not repair? A swapped box could be swapped back, but that's a guess about what the model meant. The caregiver can add a missed spot by hand. A wrong spot speaks.

To test it, I didn't hand-write answers. A fast-check property test builds them from spots that are valid or invalid by construction. Each bad spot breaks exactly one rule: a value outside 0–1000, min and max swapped, zero height, the wrong number of values, a blank phrase, no box. So the right result is known without re-implementing the validator.

Over 100,000 generated answers, the validator never returns a box outside 0–1000, a box with min ≥ max, a blank phrase or more than 8 spots, and never drops a valid spot before the cap. The test also asserts that the answers really do go over the 8-spot cap, so that branch is exercised, not just present. I wrote about why that check matters in an earlier post.

Rings a hand can hit

Each box becomes a ring at its center, sized to the object but kept within calm bounds:

const r = Math.min(0.08, Math.max(0.03, 0.3 * Math.min(boxW, boxHInWidthUnits)));
Enter fullscreen mode Exit fullscreen mode

Overlapping rings get nudged apart a little. In Speak mode, no ring is smaller than 64 px across.

Then the tap. It counts when the finger lifts, and only inside the same ring it started in. A second finger cancels it. So a resting palm, or a slide across the photo, never speaks:

function down(e: PointerEvent<HTMLDivElement>) {
  if (!interactive) return;
  if (press.current) {
    press.current.cancelled = true;
    return;
  }
  press.current = { id: e.pointerId, hit: locate(e).hit?.id ?? null, cancelled: false };
}

function up(e: PointerEvent<HTMLDivElement>) {
  const p = press.current;
  if (!p || p.id !== e.pointerId) return;
  press.current = null;
  if (p.cancelled) return;
  const { fx, fy, hit } = locate(e);
  if (onPoint) onPoint(fx, fy);
  else if (hit && hit.id === p.hit) onActivate?.(hit);
}
Enter fullscreen mode Exit fullscreen mode

When rings overlap, locate picks the one whose center is nearest. The rings are still real <button>s labelled with their phrase, so a keyboard and a screen reader reach them too. A click from Enter or Space has detail === 0, so the button's own handler only fires for the keyboard:

onClick={(e) => e.detail === 0 && onActivate?.(s)}
Enter fullscreen mode Exit fullscreen mode

The voice is the browser's own speechSynthesis. Each new phrase calls cancel() before speak(), so a second tap never waits behind the first. With no voice for the chosen language, the phrase shows as large text instead.

One more wall: fine everywhere except Vercel

import { validateSpots } from "../shared/validate" worked in Vite, Vitest and tsc. Vercel runs api/*.ts as Node ESM, which needs the .js extension, so only the deploy broke (cf62a94). A test now walks every relative import reachable from api/ and fails on any without .js.

What I'd carry over

  • A timeout is a thrown error, not a status code. A fallback that only reads res.status doesn't cover it.
  • Give each step its own time limit and one shared budget, so a slow model can't eat the next one's time.
  • Treat a model's JSON like user input. Check it, drop what fails, never repair it.
  • Build test inputs whose right answer you know by construction, and check that the edge you care about actually gets hit.

Honest limits

  • No private photos on the free tier. Google's terms for unpaid use let it use submitted images to improve its products, and let human reviewers read them. Fine for the AI-generated example rooms; a family's real room needs the paid tier first. If you try the app, use the example rooms.
  • Measured on AI-generated rooms only. There is no benchmark on real family photos yet.
  • A proof of concept, not a medical device. No clinical claims. One room, one device, the device's own voice. It still needs testing with real families and speech therapists.
  • The ladder doesn't say why. The function reports which step answered, not why the earlier ones said no.

Try it

App: roomspeak.edycu.dev (tap an example room, no login) · Code: github.com/edycutjong/roomspeak (MIT) · 30-second path: roomspeak.edycu.dev/judge · Story: roomspeak.edycu.dev/story

Planned and built with the Devpost Learn skill pack for Build With AI: Basics. The planning docs are in the repo's devpost/ folder.

If you work in AAC, or care for someone after a stroke, I'd like to hear what this gets wrong.

Top comments (0)