This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
TL;DR. My friend Zee has ADHD. I asked him what actually goes wrong, and he answered in two lines: things he can't see disappear, and the one thing he locks onto eats the whole day. So I built Actually. You brain-dump in your own messy words. A Qwen3.5-4B I fine-tuned on Tinker turns that into tasks, and TabPFN reads your own history to tell you how long each one will actually take you. Nothing you dump disappears until you tick it off. Everything except the parser runs on my laptop, and the whole build cost $2.81.
What I Built
I had the idea before I asked him. That was mistake one.
My plan was a time-blindness tool, a thing that tells you "this'll take 15 minutes" is a lie. It was a fine idea. Then I messaged Zee and asked what his actual problem was, and he sent back this, verbatim:
"can't remember. if infront remember or else gone from brain. person or task"
"addiction pro max. if i addict i addict too much. idk these 2 main."
Two problems, and only half of one was the thing I'd planned to build.
The first is out of sight, out of mind. A task or a person that isn't physically in front of him stops existing. The second is hyperfocus. When something grabs him, it takes the afternoon, and everything else silently falls off the end of the day.
So I rebuilt the plan around his words, not mine.
Actually is a single page. You type everything on your mind the way you'd text it: "email sir abt the attendance thing (5 min), finish dbms assignment should take like 30 min, laundry maybe, call nani back i keep forgetting." Three things happen:
- It splits your rant into tasks, and every one of them goes on a "Still on your plate" list that stays there until you press Done. People get an amber dot, because "call nani back" is exactly the kind of thing that vanishes.
- It tells you how long each task will take you. Not how long tasks take in general. How long your admin tasks take relative to your guesses, from your own history.
- It fits the day to your free time and warns you when a single task could swallow it.
The brief, in his words. The left side is Zee's description. The right side is the whole product.
Demo
Demo: docs/actually-demo.mp4.
Code and setup: the repo runs locally in four commands (below). There is no hosted link yet, and I'll explain why in the concessions.
This is the result screen for the brain-dump above, on the demo profile:
The verdict. 35 minutes of stated guesses, 3 hours 14 minutes of predicted reality.
Every task gets a short lavender bar for your guess, a cobalt bar for the prediction, and a faint band for the 10–90% range. Tasks that don't fit go under "Move to tomorrow" in amber. They don't get deleted, and they don't fall off the list.
Moved, not forgotten. Laundry doesn't fit in three hours, so it waits on the plate for tomorrow.
Code
ArqamWaheed
/
actually
How long will it actually take? A time-blindness helper for a friend with ADHD — fine-tuned open model (Tinker) + TabPFN, runs on one box.
Actually
How long will it actually take you? A planning aid built for one friend with ADHD.
My friend described his two problems like this:
"if infront remember or else gone from brain. person or task" "if i addict i addict too much"
Out of sight is out of mind, and when he does lock in, one thing eats the day. So Actually does two things:
- Nothing you dump disappears. Type a messy brain-dump. A small open model, fine-tuned on Tinker, splits it into tasks. Every task and every person you mention stays on a "Still on your plate" list until you tick it off.
- It shows the real cost of each task. TabPFN, a tabular foundation model, reads your own history of guessed vs real durations and predicts how long each task will take you, with a range, the chance you finish it today, and a…
MIT licensed. eval/ reproduces every number in this post.
Interesting files:
-
app/parser.py: the brain-dump to JSON parser. One interface, two backends (Tinker sampling or any OpenAI-compatible local server). -
app/forecast.py: TabPFN regressor and classifier over the user's own task history. -
train/sft.py: the whole LoRA fine-tune in 65 lines of Tinker SDK. -
eval/parser_eval.pyandeval/tabpfn_eval.py: the comparisons in the tables below.
uv pip install --python .venv/bin/python -r requirements.txt
cp .env.example .env # TINKER_API_KEY, TABPFN_TOKEN, TINKER_SAMPLER_PATH
python scripts/make_demo_data.py
PARSER_BACKEND=tinker uvicorn app.main:app --port 8000
How I Built It
The pipeline. Two open models, each doing the one job it's actually good at.
Why two models, not one. I could have asked one large model to read the rant and guess durations. It would have answered confidently and been wrong in a generic way, because it has never seen how long Zee's tasks take. Reading messy language is a language model's job. Predicting a number from a small table of personal history is a tabular model's job. Splitting them made each one honest.
The parser: a 4B model that reads a brain-dump. Brain-dumps are hard for small models. They merge two tasks into one, miss the "maybe" that makes a task optional, and lose deadlines buried mid-sentence. I needed strict JSON with seven fields per task, from text like "ugh exams are killing me rn... reply to dad about fee slip. also need batteries for the remote?? maybe."
The data: every model in the pipeline is open-weight. I didn't have hundreds of real brain-dumps, and I wasn't going to fake Zee's. So Qwen3.6-35B-A3B wrote 450 messy brain-dumps in a Pakistani-student register, and Qwen3.5-397B-A17B labelled every task in them. Both ran on Tinker's sampling API. 448 survived validation. I held out 48 and trained on 400.
The fine-tune: LoRA rank 32, 3 epochs, 75 steps. Tinker gives you a training loop, not a black box. train/sft.py renders each example with the chat template, masks the loss to the assistant's JSON tokens only, and runs forward_backward plus optim_step with a linear decay. Loss went from 335.5 to 43.3.
The training loop. The amber note is the caveat, and it stays in the picture on purpose.
Did it beat the baseline? Same 48 held-out dumps, same prompt, temperature 0:
| Parser | valid JSON | task F1 | kind | deadline | p50 latency |
|---|---|---|---|---|---|
| Qwen3.5-4B zero-shot | 98% | 95% | 78% | 69% | 3.19 s |
| Qwen3.5-4B 5-shot | 96% | 93% | 82% | 55% | 3.28 s |
| Qwen3.5-9B zero-shot | 100% | 94% | 88% | 61% | 4.53 s |
| Qwen3.5-4B + LoRA | 100% | 97% | 93% | 76% | 3.28 s |
The fine-tuned 4B beats a model more than twice its size on every column that matters, and it's 1.4x faster than that 9B. Few-shot prompting made deadlines worse (69% to 55%). Five examples in the prompt taught the base model the format, not the judgement.
Two honest caveats. First, the test set is synthetic, and its labels come from the 397B teacher, so this measures agreement with a much larger open model, not accuracy against a human. Second, "all fields exactly right on every task in a dump" is still only 8% for the fine-tune. Seven fields times six tasks is a brutal bar, and the misses are mostly dreaded and deadline, which are judgement calls.
The forecaster: TabPFN, because the table is tiny. A person's task history is maybe 40 to 100 rows. You can't train a model on that, and you shouldn't try. TabPFN doesn't train. It reads the whole table in-context and predicts. I run TabPFNRegressor on log-duration and ask for the 10th, 50th and 90th quantiles, so every estimate comes with a range. A TabPFNClassifier predicts whether the task gets done the day it was planned. The features are boring on purpose: kind, his guess, start hour, weekday, dreaded, leaves the house, has a deadline.
In-context means every finished task counts immediately. When you press Done and log how long something took, that row joins the history and the next plan uses it. No retraining step exists, because TabPFN has nothing to retrain. For a tool about a single person, that's the whole bet.
The carry-over list is the boring feature that matters most. db.remember() writes every parsed task to SQLite and dedupes near-identical titles, so saying "call nani back" three days running doesn't create three entries. The list renders above the input box, so it's the first thing on screen, which is the literal fix for "if infront remember."
Out of sight is the bug. This list is the fix: it stays in front of you until you clear it.
The hyperfocus warning is one line of logic. If a task's 90th-percentile estimate is at least half your free time, the card says "This one can swallow your day. Set a timer before you start." It isn't clever. It's the sentence Zee needs to see before he starts, not after.
Sentry found the slowest part, and it was not the language model. Every plan request is one agent run: invoke_agent actually, then a gen_ai.chat span for the fine-tuned model with the model name and token counts, then execute_tool tabpfn_forecast. OpenAIIntegration(include_prompts=False) and send_default_pii=False keep the brain-dumps out of Sentry entirely.
I assumed the model call would dominate. It didn't. The TabPFN tool span was taking 2.4 to 3.5 seconds per request, because predict() was re-attending over every history row on every call.
The fix, as Sentry saw it. Same tool, same history, 3.25 seconds down to 268 milliseconds.
TabPFN has a fit_mode="fit_with_cache" that computes the history's attention state once at fit time. Fitting gets slower (2.50 s to 3.43 s, once per history change). Predicting got 9x faster, from 2.43 s to 0.27 s. End to end, a plan went from 5.6 s to 3.5 s.
The slow part was re-reading 70 rows of history on every request, not the model.
The trace also showed what's left. The gen_ai.chat span is now 3.2 s of a 3.5 s request, and most of it is two HTTP round-trips to Tinker with a gap between them, not token generation. That's the clearest argument I have for the next step: serve the parser locally.
Where the time goes now. The remaining cost is network, which is a hosting decision, not a model one.
Things I refused to do:
- Do not invent Zee's data. The live demo runs on a profile labelled "Demo person (made-up data)".
- Do not claim the parser is local while it's served from Tinker. The footer of the app changes depending on which backend is running.
- Do not send brain-dumps to the observability tool. Sentry gets durations and token counts, never prompts.
- Do not use a Shapley chart where a plain sentence works. The "why" chips are per-kind overrun ratios from your own history: "admin tasks take you ~2.9x your guess."
Why Does Open Innovation Matter?
Three reasons, all specific to Zee.
I could teach a model his way of writing. The fine-tune is the whole trick: task type 78% to 93% on held-out dumps for $2.81 of compute. You can't fine-tune a closed model this way, and you can't download what you trained. I downloaded mine. The LoRA adapter is a 291 MB file on my disk, and Tinker's cookbook can convert it to a standard PEFT adapter or merge it into the base model.
His history never leaves the machine. An ADHD diagnosis, what he avoids, how long he spends on things: that's health data. TabPFN-3.5 runs locally on CPU with open weights, so the part that reads his life is on a laptop, not behind an API. Sentry only ever sees timings.
The tools are inspectable all the way down. When the forecast was slow, I could read TabPFN's source, find fit_mode, and fix it in one line. The parser's chat template let me switch off the model's thinking mode, so it answers in JSON instead of reasoning first. A closed API gives you neither lever.
Where closed would have been better, honestly: a frontier API would probably parse these brain-dumps about as well as my fine-tune with zero training, and it would be faster to call than Tinker's beta sampling endpoint. I traded that for owning the weights and keeping his data local. For a tool about one person's head, that trade is the point.
My Agent Session
I built this with Claude Code, and every session is captured with Entire. Entire stores the prompts, reasoning and tool calls next to the commits they produced, with secrets redacted before anything leaves the machine. I added explicit redaction rules for every key this project touches, then grepped the checkpoint for each one before pushing. Zero hits.
- Checkpoint
01M4494DVYMY4YYYJF1ZKAK7S3: the full session, from reading the challenge page and measuring every prize pool, to picking the stack, training the parser and the Sentry fix. 17 turns. - Inspect it with
entire checkpoint explain 01M4494DVYMY4YYYJF1ZKAK7S3in a clone of the repo.
The checkpoint's generated summary is the best changelog this project has. One line from it: "Tinker's raw LoRA adapter export through tinker-cookbook's PEFT builder downloads the full 8GB base model. That won't fit on a low-disk machine, so convert the adapter directly instead." That decision, why the parser isn't self-hosted yet, now lives next to the code it explains.
What I Learned
- Ask the person before you design. I had a polished time-blindness pitch. Zee's two-line answer turned it into a carry-over list plus a hyperfocus warning, and those are the parts he'll actually use.
- A small model plus a fine-tune beats a bigger model plus a prompt. The 4B LoRA beat the 9B on every column, and few-shot prompting made the hardest field worse.
- Profile before you optimise, and don't trust your guess about where the time goes. I'd have bet on the LLM. It was a tabular model re-reading 70 rows.
- Concessions belong in the product, too. The demo profile says "made-up data" in the dropdown. The footer says where the parser runs. A tool for someone who's hard on himself should never lie to him.
Prize Categories
- Best Use of Tinker: LoRA fine-tune of Qwen3.5-4B (rank 32, 400 examples). On 48 held-out brain-dumps it beats the base 4B, base 4B 5-shot, and the 9B: task type 93% vs 78/82/88%, deadline 76% vs 69/55/61%, 100% valid JSON, and 1.4x faster than the 9B. The dataset was also generated and labelled on Tinker with open models. Total spend $2.81. Adapter downloaded.
-
Best Use of TabPFN: local TabPFN-3.5 regressor (log-duration, 10/50/90% quantiles) and classifier (done the same day?) over the user's own small history. In-context, so every logged task updates the next plan with no retraining.
fit_with_cachecut predict latency from 2.43 s to 0.27 s. -
Best Use of Sentry Agent Tracing:
invoke_agent→gen_ai.chat(model, input and output tokens) →execute_tool tabpfn_forecast, prompts excluded. Traces found the real bottleneck (TabPFN predict, not the LLM), I fixed it, and the traces confirmed it: 3.25 s to 268 ms on the tool span, 5.6 s to 3.5 s end to end. - Best Use of Entire: the full build session captured as a redaction-checked checkpoint linked to its commit, and quoted above to explain why the code is the way it is.









Top comments (1)
The strongest engineering lesson here is that the “best model” wasn't the main optimization target. The system worked because each component had a narrow responsibility: the language model handled messy text, TabPFN handled the small personal dataset, and the persistent task list handled the actual product constraint.
The TabPFN profiling result is a great example of why end-to-end measurement matters too. It would have been easy to optimize the LLM because that is the obvious AI component, while the real bottleneck was repeatedly processing 70 rows of history. The 2.43s → 0.27s improvement came from changing the execution model, not changing the model.
I also really like the decision not to treat synthetic evaluation as proof of real-world accuracy. With a personal assistant like this, the hardest validation isn't whether the parser matches a teacher model; it's whether the predictions remain useful for the actual person using it.
That combination narrow model responsibilities, profiling the whole pipeline, and being explicit about what the evaluation does and doesn't prove is probably more reusable than the particular models in the stack.