DEV Community

王冲
王冲

Posted on

One HTML File, Zero Requests: Building a Read-Along Player With Word-Level Highlighting

My favorite artifact from the app I'm building isn't in the app. It's the file the app exports: a single HTML page that plays a chapter of audio while every spoken word lights up in sync — karaoke-style — with zero network requests, zero dependencies, zero accounts. Open it on a phone from 2015 with airplane mode on; it works. Email it to a grandparent; it works.

This is the heart of Elmren Voice, a local-first text-to-speech app for Mac I've covered here before — how I picked the TTS engine, then how I packaged the Python sidecar. This post is the last leg: how synthesized audio becomes a self-contained read-along player, and the alignment trick that makes the highlighting exact.

Why a file instead of a page on our server

The product's promise is that a child's reading material and a family's voice recordings never leave the machine. A hosted player — even a nice one — would break that promise the moment playback needed a server. A plain file keeps it: generate locally, open anywhere, and the privacy story is structural instead of contractual. It also means the thing you made in 2026 still works in 2036 regardless of anyone's business model, including ours.

The constraint drove the format: one HTML file, audio embedded as a data URI, the word-timestamp table inline, and a few hundred lines of dependency-free vanilla JS for the highlighter.

The pipeline: text → audio → word timestamps

Synthesis runs sentence-by-sentence (that's a cloning-quality requirement, not a design preference). Then the generated audio goes through whisper.cpp for alignment — not transcription quality, alignment. The numbers from my reference machine: 58.8 seconds of audio → complete word-level timestamps in 2.87 seconds.

Then the part nobody warns you about.

Whisper mishears words. The highlighter must not care.

The timestamp says when a word was spoken, but the transcript whisper produces says what it heard. Feed the transcript's words straight into the UI and you get the classic failure mode of naive karaoke: the highlight drifts to what was heard while the page shows what was written, the kid's eyes track the wrong word, and the whole point of read-along — eyes locked on the correct print — collapses.

The fix: the highlighter renders the source text, never the transcript. The source is ground truth (it's literally what we sent to the TTS engine); the transcript is only a carrier of timing. A SequenceMatcher pass maps each transcribed word to its counterpart in the source — on my 166-word reference passage it covers 166/166 words — and any word whisper dropped or mangled simply inherits timing from its neighbors. Mishearings degrade to slightly looser sync, never to highlighting the wrong word.

If you build anything with forced alignment, this inversion — source text for rendering, transcript for timing only — is the whole ballgame.

The runtime: boring on purpose

The player itself is deliberately uninteresting technology, which is the compliment: an <audio> element, a word table, and a timeupdate-driven binary search for the current word. A <span> per word gets a class toggle. Pages chunk by paragraph so a long chapter never renders a wall of ten thousand spans at once. There's a progress memory in localStorage, a dark mode that respects prefers-color-scheme, and — because some kids share devices with siblings — everything also works with the keyboard alone.

No framework. The file has to outlive every framework.

The same timestamps, two more products

Once word timing exists, it composes:

  • M4B audiobook export — chapters and bookmarks straight from the sentence boundaries the pipeline already knows.
  • Karaoke MP4 — the word-highlight render as video. The shipped ffmpeg build carries no subtitle filters, so it's rendered per-word as PNG frames and concatenated — a constraint I've written about before, and frame-exact highlighting came out of it.

One alignment pass, three deliverables.

What it looks like

There's a live demo player up — it's a real export from the product (a chapter of The Secret Garden), not a mockup: play it and watch the words light up. The app itself is $39.99 once for macOS, with a free trial (120 synthesis minutes in your first week), and voices are either one of the 64 built-in or a clone of your own — learned and spoken entirely on your Mac.

One more thing: Elmren Voice launches on Product Hunt tomorrow, Tuesday Oct 6. If read-along tooling for kids who struggle with reading is your corner of the world, I'd genuinely value a look — and criticism; the comment threads are where the good corrections come from.

As before, thanks to the stack that makes a one-person product possible: whisper.cpp, Kokoro, Qwen3-TTS and mlx-audio.

Top comments (0)