Built for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
My mother and father ask me to read WhatsApp messages, pictures, labels and documents. Both have weak eyesight and use Redmi A4 phones. I am in another city right now, so that request can mean waiting for me to be available.
I built Suniye, or рд╕реБрдирд┐рдП, an Android reading aid for those moments. It reads small print aloud in Hindi, with large controls for рд░реЛрдХрд┐рдП (Stop), рдлрд┐рд░ рд╕реБрдирд┐рдП (Repeat) and рдзреАрд░реЗ рд╕реБрдирд┐рдП (Slow). The pilot supports Android 11 and later.
Hear the historical pre-0.3 image-and-PDF demo (24 seconds). This footage predates the 0.3 parent-UX revision. The preview above is silent. The video shows the source image, then рдирдорд╕реНрддреЗ read aloud; the PDF pages read рдмрд┐рд▓ and рдзрдиреНрдпрд╡рд╛рдж. These simple synthetic samples establish the demonstrated path. Dense documents need separate checks.
An amount needed more care than sending printed text to a voice API. тВ╣1,250 becomes тАЬрдПрдХ рд╣рдЬрд╝рд╛рд░ рджреЛ рд╕реМ рдкрдЪрд╛рд╕ рд░реБрдкрдпреЗтАЭ in the speech input. The original remains visible, with a separate amount hint. Dates keep their written field order.
рдХрд╛рдЧрдЬрд╝ рдкрдврд╝рд┐рдП opens the camera, рдлрд╝реЛрдЯреЛ рдкрдврд╝рд┐рдП selects a photo, and рдлрд╝рд╛рдЗрд▓ рдкрдврд╝рд┐рдП opens a PDF. Android Share accepts text, images and PDFs. After caregiver setup, a floating рд╕реБрдирд┐рдП control is the intended WhatsApp route. The daily controls use large Hindi labels; settings and optional spoken commands stay behind caregiver setup. Stop remains below the scrolling content, where a long reading cannot push it away.
The pilot has bounded Android 11 and 15 emulator checks. The parent-focused revision passed internal layout and bundled-Raju playback assertions at 160% fonts, but a System UI dialog covered the capture and later system crashes blocked the full touch walkthrough. That check remains open. I have not observed my parents using it on their Redmi A4s. Parent check and limits.
Demo
Watch the historical v0.1 walkthrough ┬╖ Download the parent-UX APK
The linked APK is the released 0.3.0 parent-UX revision. The films show earlier demonstrated flows; the newer revision's internal assertions do not establish a completed touch walkthrough. The cancellation/privacy safeguards and release-link corrections are published in the source repository. The project page is live on the existing Free Render service, with the current APK link and historical screenshots clearly labelled.
The walkthrough shows a synthetic electricity bill, Hindi amount wording, Repeat and Stop, followed by a photo and two PDF pages. It displays the source files before their recognized text. The WhatsApp section is a labelled setup walkthrough; actual WhatsApp reading remains untested.
The inputs use real Android content URIs, bundled OCR and real Raju v4 API audio. Android screenrecord did not capture device sound, so I mixed the provider MP3s separately at the observed playback rate. Source previews are inserted stills; native waiting stays in the film. The Gemma reply is a separate saved backend check. Recording notes.
For a quick trial without configuring a backend, download the prepared PDF reader, open the HTML and tap рд╕реБрдирд┐рдП. It displays the original and Hindi amount wording, and plays its embedded Raju recordings without provider requests. This caregiver export is separate from the Android app.
Code
Source repository ┬╖ Current 0.3 usability scope and limits ┬╖ Caregiver setup for the current release
I began the scaffold on October 2 during the challenge. The code is MIT licensed. The debug APK includes synthetic fixtures; production source excludes them. Provider keys and family messages stay out of the public repository. Gemma weights have separate terms, and Google's bundled OCR SDK is proprietary. Third-party notes.
How I Built It
The phone reads Android's accessible text or runs bundled Devanagari and Latin OCR. When OCR recognizes text, the image stays on the phone. Online narration sends that text through the authenticated backend to ElevenLabs. An image with no recognized text can be sent for a separate Gemma description. Gemma is prohibited from supplying guessed image text when OCR misses it.
New narration uses ElevenLabs Raju exclusively. Fixed help prompts are bundled Raju recordings; successful readings cache their audio for offline Repeat and Slow. A new reading needs a connection. If audio is unavailable, the original remains visible. The app never silently switches to Android or browser TTS.
Stop must stay stopped
Mastra runs the online reading workflow: validate the source, prepare the original or an optional explanation, then request speech if needed. Each request has its own cancellation signal. Stop invalidates the phone's operation and discards late replies. A deliberately delayed-response test checked that a stopped reading did not restart. The voice-only Android 15 check also covered cached replay, Slow, Stop, missing or corrupt audio, wrong-provider audio and late replies. Its host audio was disabled, so it verifies media state rather than another listening judgment. Android receipt.
Keep the original when the model gets it wrong
Gemma 3 4B runs through Ollama on my laptop for simpler Hindi explanations and non-text picture descriptions. A warm check changed тАЬрдкреНрд░рд╡реЗрд╢ рд╕реЗ рдкреВрд░реНрд╡ рдЕрдкрдиреЗ рдЬреВрддреЗ рдЙрддрд╛рд░рдирд╛ рдЕрдирд┐рд╡рд╛рд░реНрдп рд╣реИредтАЭ to тАЬрдкреНрд░рд╡реЗрд╢ рдХрд░рдиреЗ рд╕реЗ рдкрд╣рд▓реЗ рдЕрдкрдиреЗ рдЬреВрддреЗ рдирд┐рдХрд╛рд▓реЗрдВредтАЭ Both tell the reader to remove their shoes before entering; the Mastra response took 2.66 seconds. Saved reply.
Other replies changed a date, replaced тАЬattachтАЭ with тАЬsubmitтАЭ, or added context absent from the source. The quantity and negation guard caught some errors and missed others. Six synthetic Backboard calls comparing Gemma 3 4B and Qwen 2.5 72B also passed literal amount/date checks while exposing added-context errors. That small comparison supports keeping explanations experimental and separate from the original; it does not establish a general model ranking. Model results ┬╖ Comparison.
Bundled OCR misread a dense bill too. I added a retake below .85 confidence. An early PDF invitation triggered that guard, so the final demo uses simpler text. This is a recovery heuristic, not an accuracy guarantee. Amounts, dates and medicine labels still need a family member's check.
I investigated TabPFN for choosing a retake or caregiver review. I prepared 100 CC BY 4.0 printed Hindi training-word crops from Mozhi-Hindi and built an isolated batch harness using Suniye's unchanged ML Kit OCR pipeline. Both emulator attempts stopped at the host's disk-space check before boot. There are zero actual OCR observations, no model inference and no measured benefit. The confidence rule remains unchanged, and TabPFN is not a twelfth category claim.
Make the wait explainable
Sentry showed one local Gemma call taking 32.64 seconds, including 24.96 seconds loading and 2.73 seconds generating. That gives me a specific delay to investigate. The outgoing trace omitted reading text and kept timing and token metadata. Trace evidence.
Render's Free Singapore service passed an authenticated original-reading request with real Raju audio and rejected an unauthenticated request with 401. The newer hosted voice-only check returned unchanged bill text and 27,420 bytes of Raju audio. The free service can sleep; hosted Gemma and Render-to-Atlas remain unconfigured. Local Gemma evidence should not be mistaken for a hosted model. Hosted check.
I tried to use the $50 hackathon credit for Render's $7/month always-on compute plan. Render required payment information on file even with the credit, so the service remains Free. Credit-only upgrade check.
Why Does Open Innovation Matter?
I can inspect Mastra's workflow, change its stages and test cancellation. The model and speech providers are separate modules, so the reading policy is visible in code. The original-reading path does not depend on a generative rewrite. The current speech policy accepts only Raju; replacing it would require an explicit policy change.
Local Gemma lets me repeat failures on synthetic Hindi text and inspect the replies. It runs on my laptop, not inside the APK. Remote explanations and picture descriptions still need an authenticated model backend. The open pieces make those decisions and their limits available to another builder; the OCR and speech services have their own terms.
While preparing the post, I read Samajh and ReadAloud. Their work focuses on source-linked explanations and following spoken words. My focus is my parents' small-print problem, Hindi amounts and reachable Android controls.
My Agent Session
I used Entire to import the development session locally. A checkpoint lookup recovered my instruction: тАЬthey ask me to do it so ux and ui needs to be specifically for themтАЭ. That explains the large Hindi controls and separate caregiver setup. Curated provenance links those instructions to the code; full private transcripts stay private.
Claude reviewed the PRD, specification and selected source. I fixed findings and ran separate runtime checks. Source review helped find defects; it did not tell me how the app feels to my parents.
Prize Categories
I am entering eleven categories. Their jobs and evidence are different:
| Category | Why I used it | What actually ran and its limit |
|---|---|---|
| ElevenLabs | My parents need to hear the original in Hindi and replay it. | Raju v4 API audio played in Android; cached replay and bundled help prompts were checked. Redmi listening is pending. Attribution: elevenlabs.io. |
| Mastra | Stop needs to cancel work through the model and speech stages. | The reading workflow validates input, separates explanation and handles cancellation; delayed-reply tests ran. Redmi network behavior remains untested. |
| Gemma | Formal Hindi may need a simpler, separate explanation. | Local Hindi explanation and non-text picture trials ran. Meaning errors remain; hosted Gemma is not connected. |
| Render | The phones need a reachable, authenticated endpoint while I am elsewhere. | Free HTTPS original reading and Raju speech passed. The service can sleep; hosted Gemma and Atlas remain unconfigured. |
| Sentry Agent Tracing | I needed to find where a long Gemma wait was spent. | A local trace exposed model loading as most of one request's wait. Outgoing traces omitted reading text; this is developer diagnosis. |
| Backboard | I wanted to compare how faithfully two models explain Hindi. | Six synthetic Gemma/Qwen calls exposed added context. Three cases do not rank models generally; memory/search/tools were off and the temporary key was revoked. |
| Entire | I wanted to trace interface choices back to my family brief. | Eighteen imported checkpoints and lookups linked requirements to decisions. Curated provenance is public; full sessions stay private. |
| SerpApi | Caregiver setup needed official Android references. | October 2 searches supplied two help articles for a Gemma summary. A later search returned no approved pages and stopped; exact Redmi menus remain unverified. |
| MongoDB Atlas | Caregiver preferences need to survive a backend restart. | Local preference saving, restoration and restricted routes passed. Render-to-Atlas and Android sync remain untested. |
| Temporal | Caregiver PDF preparation should recover after a worker fails. | A local job recovered, retried an unfinished page and rejected late writes after cancellation. Preparation makes no speech/model calls; no Temporal Cloud deployment. |
| Tiger Data / pgvector | Setup searches should retrieve relevant, approved instructions. | Local PGlite/pgvector plus keyword retrieval ran over three references. English queries were tested; Hindi search and Tiger Cloud were not. |
The sponsor ledger links the dated receipts. Setup, comparison and tracing tools stay outside the daily parent screen. SerpApi searches use public help questions, and the Backboard comparison used synthetic text.
The next family check is specific: can each parent read a WhatsApp message and a paper label, then stop and replay without my help? I have not observed that yet. Emulator checks establish the tested app behavior; their Redmi phones and their own use need a separate check.

Top comments (0)