In my last post I covered how I picked the TTS engines for Elmren Voice — a Mac app that clones your own voice locally and turns any text into a read-along player with word-by-word highlighting. Everything runs on-device: no cloud, no API keys, no account.
That post was the fun part. This one is the part nobody warns you about: getting a Python ML stack to live inside a signed, notarized, sandbox-friendly Mac app — and keeping the download at ~122MB while the models themselves weigh gigabytes.
The stack: Tauri 2 shell (Rust + React/TypeScript), Python sidecar for the engine (mlx-audio, Kokoro, whisper.cpp, ffmpeg). Nine traps, all paid for in real debugging hours.
1. Your sidecar "launcher" silently swallows arguments
The packaged launcher must start with no arguments. My first build passed rw-engine --serve style args through a wrapper that resolved to python --serve — Python happily starts an interactive session and your bridge never handshakes. Symptom: app launches, engine "runs", zero responses. The fix is boring: the launcher takes nothing; the serve mode is the default; mode switches ride the JSONL protocol, not argv.
2. Version your wire protocol… as strings
Our bridge is newline-delimited JSON over stdin/stdout. Every message is enveloped, and the envelope is {"v":"1"} — v is a string. I shipped one call site with a numeric 1 once. Strict parsers see two different protocols. Pick string versions on day one; JSON numbers will bite you across languages (Rust serde and Python json disagree on int-vs-float edge cases).
3. Control commands must never queue behind data commands
The nastiest bug I shipped: user clicks Cancel on a synthesis job — nothing happens. Why? One mutex guarded the whole stdio bridge, and a 10-minute chapter synthesis held it on every roundtrip. The cancel command was delivered, then sat in the write queue behind the job it was supposed to kill.
The fix, shipped in v1.9.1: split the hub into a data plane (long roundtrips) and a control plane (short commands: cancel, pause, settings) with an independent write-side lock, and make control messages fire-and-forget. Cancels now land even mid-roundtrip. If your GUI freezes "sometimes" during long native calls, audit your lock topology before anything else.
4. PEP 668 means zero runtime pip. Ever.
Modern macOS Pythons are externally managed; runtime pip install was never going to survive anyway (no compiler, no network assumptions, Gatekeeper). The rule that falls out: the entire Python environment ships inside the app as a vendored sidecar-site — including things you'd normally pip install at build time, like en_core_web_sm for spaCy. If a wheel can't be vendored, that dependency doesn't exist.
5. brew's whisper-cli dies on every machine but yours
whisper.cpp does our word-level alignment. The Homebrew build hardcodes its ggml backend search path to the Cellar absolute path — on my dev Mac, perfect; on a clean Mac, Abort trap 6. The only reliable answer is building from source statically: v1.9.2, BUILD_SHARED_LIBS=OFF, Metal embedded. One self-contained binary, seeded on first launch.
6. Hardened runtime rejects your relocated dylibs
Relocating Python extension modules and native tools means library paths that no longer match their signatures. Under Hardened Runtime (required for notarized distribution) that's a hard kill at load. You need disable-library-validation in the signing entitlements for those binaries — plus re-signing them after every copy/move. Also: keep native binaries in their own native-tools/ tree, or your Rust build will try to scan Mach-O files sitting inside the Python site-packages and waste minutes per build.
7. The bundled ffmpeg has no subtitle filters. Karaoke videos anyway.
The redistributable ffmpeg build carries no ass/drawtext filters, which kills the obvious "render word-highlighting onto video" path. The workaround that shipped: render each word as a PNG, concat them per-frame, and pad the video to audio length. Less elegant, zero extra dependencies, and frame-exact highlighting came out of it. Constraint beat cleverness.
8. A symlink in your data root is a jailbreak
Users will symlink ~/Application Support/.../models to an external drive. Tempting to allow — until path resolution follows the link out of the sandbox root and your escape checks pass things they shouldn't. Our PathService.resolve follows links and rejects any path that escapes the data root; the supported path for moving data is cp -cR (clonefile, near-free on APFS). Related: Hugging Face downloads behind some proxies hang forever on the new Xet channel — HF_HUB_DISABLE_XET=1 forces the classic HTTP path.
9. Ship thin, hydrate on first launch
The final architecture: the DMG is ~122MB — app shell, UI, a sample book, and no models. On first run the app downloads engine packs on demand: ~700MB for the base stack (54 Kokoro preset voices + alignment + media), ~1.7GB adds the multilingual set, ~2.3GB all-in with the Qwen3-TTS cloning engine (64 voices total). Downloads are cancelable, disk space is pre-checked, and a background reconciliation loop repairs interrupted packs. First launch went from "download 2.3GB before you see anything" to "usable in a minute, clone-ready when you are."
The takeaway
A local-first AI product is two products: the ML pipeline, and the distribution engineering around it. The second one has no benchmark blogs. If you're heading down this road: vendor everything, sign everything, keep your control plane off your data plane's locks, and treat every binary that worked on your machine as untested.
I'm building Elmren Voice ($39.99 one-time, macOS 13+) — there's a live demo player exported from the real product if you want to hear the word-highlighting in action. Engine selection backstory in part 1.
Top comments (0)