DEV Community

Cover image for A dinner planner for a friend who shouldn't have to check the AI's work
Pierre-Laurent Medori
Pierre-Laurent Medori

Posted on

A dinner planner for a friend who shouldn't have to check the AI's work

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

A friend of mine spends too much time deciding what to cook. Not cooking: deciding. He has a wife, a kid, a mollusc allergy, and the same question waiting for him at the end of every day. He also doesn't see what AI would do for him in daily life, and I understand why. Most of what gets said about AI has nothing to do with dinner.

So I didn't argue. I built him the most ordinary thing I could think of.

Weeknight plans the week's dinners on Sunday evening and hands over the shopping list. You tell it how many people are eating, the allergies, the foods someone won't touch, what the family likes, and which nights are busy. You can show it a photo of the fridge. Gemma, Google's open-weight model, suggests the dinners from the Mac itself, through Ollama. Then plain code checks every dish and adds up the list, aisle by aisle.

When I showed him a prototype, before my last round of fixes, he told me he was glad to get the time back, and glad to stop carrying the question around in his head.

The second half of his sentence is the real specification. Someone who uses a tool to stop thinking about dinner is not going to proofread it. He won't scan Tuesday's stir-fry for oyster sauce. So the app has to.

Later I sent him a shopping list from Weeknight, and he seemed happy with it. One suggestion from that first demo stuck, too: tacos. He never used to make them at home. Now he does, and he tells me he and his son are very happy about it.

Demo

The demo uses my friend's numbers, three people and a mollusc allergy, with tastes made up for the occasion.

First, Gemma reads a photo of the fridge and lists what it sees. It saw lettuce where there was spinach, so I fix that one word.

Gemma reads a photo of the fridge and lists what it sees, then lettuce is corrected to spinach

Then I plan the week. The code sends a Saturday dish back because it had olives, which this household won't eat, and the week arrives with its shopping list. The list sets apart what is probably already in the fridge.

Weeknight sends back a dish with olives, then shows the week's dinners and the shopping list

The whole run, in one minute:

Fridge photo: Alabama Extension, CC0, via Wikimedia Commons.

Code

GitHub logo sintineddi / weeknight

Plan the family's dinners on Sunday evening with a local open model (Gemma via Ollama): allergen checks, shopping list, A2A-ready.

Weeknight is a Python server built on the standard library alone, and one page of plain JavaScript. There is no pip install and no npm install. With Ollama running:

ollama pull gemma4:e4b
python3 server.py
Enter fullscreen mode Exit fullscreen mode

How I Built It

Gemma suggests, the code checks

Gemma Plain code
Ideas the dishes, ingredients, amounts and recipe steps
Shape answers in JSON constrained by a schema
Fridge reads the photo of the fridge
Allergens checks every dish name and ingredient; a dish that fails goes back to the model, twice at most
Amounts flags quantities that can't be right ("4000 g of pasta for 4") and asks for a fix once
Variety sends back a third dinner built on the same protein, and any dish the family rated 👎
Shopping list adds up the quantities, converts the units and sorts the aisles

The model is good at ideas and bad at arithmetic, so it only does the ideas.

Gemma 4 runs through Ollama, and its answers are constrained by a JSON schema, so the code never has to parse prose. One setting mattered more than any prompt. Gemma 4 thinks before it answers by default, and with thinking on, e4b took 74 to 84 seconds for a week while 12b ran past ten minutes and got cut off. With "think": false, e4b plans a week in 30 to 50 seconds on my M3 Max, depending on how many dishes the checks send back. The fridge photo goes to the same model, which reads it in about 7 seconds. The photo is never written to disk. Weeknight also speaks A2A, so another agent can ask it for the shopping list, or receive it.

The allergen check is a keyword list over the 14 allergens that EU food labels must declare. It knows the usual hidden sources: pesto for nuts, soy sauce for wheat, Worcestershire sauce for fish, oyster sauce for molluscs. It reads "peanut-free" as "read the label", never as a guarantee, and it knows that lactose-free milk is still milk. It trusts the ingredients over the dish's name, because the ingredients are what gets bought. A dish that still fails after two tries stays on the plan, in red. Nothing is hidden.

I tried to make it fail

A checker you have never seen fail is a checker you haven't tested. So I built households to tempt the model. One is allergic to nuts and loves satay and pesto. One can't have milk or eggs and loves carbonara. One is coeliac and loves pizza. One is allergic to sesame and fish and loves sushi and hummus. The last one has my friend's allergy: molluscs, in a family that loves paella, calamari and oyster-sauce stir-fries. Three Gemma 4 sizes planned three weeks for each household, and I counted the problems in the model's first answer, before any check.

First answer: dinners with an allergen coeliac dairy-egg nuts sesame-fish molluscs
gemma4:e2b 3 of 18 0 of 18 0 of 18 5 of 18 0 of 18
gemma4:e4b 9 of 18 0 of 18 0 of 18 0 of 18 0 of 18
gemma4:12b 0 of 18 0 of 18 0 of 18 0 of 18 0 of 18

Three weeks per model and household is enough to see a pattern, not to measure a rate.

I expected the model to forget allergies. It doesn't forget them. It gives in. e4b knew the household was coeliac, and it still put plain pasta, buns or tortillas in half of its dinners: the very dishes that household had asked for. Twice a dish was even called "(Gluten-Free)" while its ingredient list said plain bread or spaghetti. The shopping list follows the ingredients, so that is what would have been bought.

My friend's allergy turned out to be the easy one. The mollusc household asked for paella, calamari, seafood pasta and oyster-sauce stir-fries, and in 54 dinners no model served it a single mollusc. The two small models mostly ignored those tastes and planned sausages, tacos and quesadillas. The big one adapted them: a Spanish-style rice with chicken and peppers, a "seafood pasta" made with cod, a beef and broccoli stir-fry with no oyster sauce in its ingredients. My best explanation is that squid and mussels are foods you name, with obvious stand-ins, while gluten hides inside the staples a family eats every week. A model thinks "pasta", not "wheat". Nine weeks can't prove that, but it changed what I worry about. The check matters most where the allergen doesn't look like one.

The biggest model made no allergen slip in 18 weeks. It reached for gluten-free pasta, tamari and coconut aminos. It failed somewhere else. In 17 of its 18 weeks, the very first quantity of its answer, the first ingredient of the first dinner, was impossible, usually ten times too big: "5000 g of chicken breast" for four. The two smaller models never got their first number wrong. My guess is something in how the first number gets decoded under the schema, but it is only a guess. Asked to fix them, 12b rarely did. That, and the doubled waiting time, is why e4b stays the default: its mistakes are the kind the code catches.

One last run removed the prompt's single sentence about hidden allergens. The slip rate barely moved, but the kind of slip did: soy sauce showed up four times, and mayonnaise once. The checks sent them back. Twice they gave up after two tries, and those two dinners stayed in red.

The checker grades its own homework

The numbers after the checks are measured by the same keyword list that did the cleaning, so they can't show what it misses. To find out, every final dinner of the allergic households was read one by one, and listed in the repo so anyone can check: 342 dinners. Claude Code, which I built Weeknight with, did that reading. No allergen got through unflagged.

The reading found bugs in the checker itself, though. It rejected a "Chicken pasta" made with gluten-free pasta, because a dish's name alone was enough to send it back. It didn't understand "no-added-wheat" or "tahini-free". It let bread pass for milk and stock pass for gluten, where it should have asked to read the label. A test week after all that found one more gap the reading had missed: taco shells. All of them are fixed, and 43 tests keep them fixed.

It is still a keyword list, not a doctor, and it reads English only: a dish written in French would walk right past it. Under every plan, the app ends with the advice you would get without it: read the labels.

Why Does Open Innovation Matter?

I didn't compare Weeknight against a closed model, so I won't claim a closed one would plan worse dinners. What I can describe is what running an open model on the Mac changed while I was building it.

The family's details stay at home. My friend's allergy, his family's busy nights and the inside of his fridge never leave the Mac it runs on. The fridge photo isn't even saved. The shopping list leaves only when he decides to send it somewhere.

Testing costs nothing, so I tested a lot. The benchmark is 66 weeks of deliberately hostile households, about 400 dinners, plus every rerun after a checker fix. On a metered API, I would have rationed those runs. Running them is how I found that my own checker was throwing away good dishes, and that 12b fumbles its first number.

The model is a setting, not a contract. e2b, e4b and 12b are the same family at three sizes, so the trade-off between speed, allergen slips and arithmetic is something I measured, not something I guessed. If a better open model shows up next month, trying it is one flag, and the benchmark tells me whether to keep it.

My friend doesn't need any of that, and he shouldn't have to. What he needs is a Tuesday dinner with no squid in it, and the list ready on Sunday night. The tacos were a bonus.

My Agent Session

I built Weeknight with Claude Code, in a day.

Prize Categories

Best Use of Gemma.

Top comments (0)