DEV Community

Cover image for We Let Vibe Coders Merge to Production. We Kept the Core.

We Let Vibe Coders Merge to Production. We Kept the Core.

Olexandr Uvarov on September 28, 2026

How non-engineers ship their own AI tools to production from a chat — and the one zone the engineer kept. Part 2 of A Prompt Is a Wish. A Tool Is a...
Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

"A prompt is a wish. A tool is a law." is going on my wall. The part I liked most is that the platform looked for existing tools before writing anything. That's the step agents skip most in code, too: they'd rather write a new helper than find the one a colleague built last month. In Semitexa, the PHP framework I'm building, I went the same way: the agent queries a map of what already exists before it writes, instead of relying on how well it searches. How do you keep the tool registry from filling up with near-duplicates as more non-engineers add to it?

Collapse
 
ouvarov profile image
Olexandr Uvarov •

Agreed on the search step, agents in code skip it all the time.

For us the search part is a law, not a wish. The builder can't get to the plan until it has listed the registry, the tool just refuses. And the plan has a separate "reused" part, so the author sees what's picked up and what's new before any code exists.

Then structure does the rest. Every tool sits in a domain folder with a domain prefix in the name, and the registry is append only. So a near-duplicate would land right next to the original, where both the builder and the review bot see it. The bot has "reuse, don't duplicate" as a high rule too.

What's still up to the model is the final call, "is this the same thing". Two tools can look alike by name and still be split on purpose, and that's kinda hard to put into a rule. If it ever gets messy, that's where I'd add a law.

How does your map in Semitexa decide two things are the same, by signature or by description?

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

The map doesn't decide "same" for us either, and I'd be wary of anything that claims to. What it does is make the structural duplicates hard to miss: if two handlers end up handling the same payload type, a usages query shows two handles edges where you expected one, and that's a fact, not a judgment. Duplicates by meaning stay with the model and the reviewer, much like your "is this the same thing" call. The separate "reused" section in the plan is a nice touch, the author sees reuse before any code exists. Do you ever see the builder reuse something that was split on purpose?

Thread Thread
 
ouvarov profile image
Olexandr Uvarov •

"two handles edges where you expected one, that's a fact not a judgment" this is well said. for us the domain prefix plays kind of the same role, a duplicate just lands next to the original

about your question, the builder reusing something that was split on purpose, I didn't catch that. I think because the reason for the split lives right next to the tool

example. we have three tools that look the same from outside, upload an image to a presigned url. first takes base64, second runs it through a paid model and re-encodes, third just copies bytes from url to url. the ADR of the third one has a separate section why we didn't reuse the first two: base64 blows up the context, and the model changes the bytes and costs money

and when the builder looks at the registry, it gets a description and an ADR link for every tool. so "why it's separate" it sees in the same moment as "what already exists"

but catching this is harder than a duplicate. a duplicate you see in the diff, a wrong reuse from outside looks like good behavior))

Thread Thread
 
syntaxwanderer_26 profile image
Taras Hanych •

The ADR link next to every registry entry is the part I'd copy: "why it's separate" arriving at the same moment as "what already exists" is exactly when the builder needs it. Your three upload tools are a good example of why sameness can't come from the signature: same input and output, three different costs. And you're right that wrong reuse is the harder failure, because it looks like discipline. One check that might catch some of it: if a tool's ADR lists reasons not to reuse X, a new caller that picks X for the same job gets flagged for a human look. Has the builder ever argued with an ADR, or does it treat them as final?

Collapse
 
carbonlayer profile image
CarbonLayer •

The important design choice here isn’t self-merge; it’s bounding what self-merge can touch. The protected core, server-rebuilt registry, commit-pinned checks, and reversible changes make that authority limited rather than blind. And the caveat matters: this fits reversible internal tools, not payments or other hard-to-undo actions.

As adoption grows, what signal would tell you the self-merge boundary needs to tighten?

Collapse
 
ouvarov profile image
Olexandr Uvarov •

yeah, bounding it is the whole point, self-merge alone is just a button))

funny thing, so far we went the other way. first version of the boundary was stricter and only ~20% of PRs passed it, so people would just keep coming to devs to merge, same queue as before. so we loosened it to 81%.

when to tighten — for me it's about blast radius. while a mistake stays inside one author's tool, the zone is ok. if something self-merged breaks someone else's tools, that's the signal. we actually had one case like this, even before self-merge: a registry regression, other people's tools just disappeared from the list. and we didn't fix it by putting a human back, we took the registry away from the model, now the server rebuilds it.

so for me tighten means move the next risky thing out of the zone or add a rule for the bot, not more approvals. probably next is shared helpers, one edit there hits every tool that imports it.

and the practical trigger: fixes landing after a merge the bot already approved. that would mean the review lets bugs through and something has to change. right now we don't see that.

Collapse
 
carbonlayer profile image
CarbonLayer •

That’s a useful distinction: tightening the boundary doesn’t have to mean adding approvals; it can mean moving a risky capability out of the model’s reach. The registry example makes the blast-radius test concrete, and the jump from 20% to 81% shows why a control that just sends most PRs back to the old queue isn’t much of a win. Shared helpers sound like the next interesting boundary—one change can affect a lot of tools at once.

Thread Thread
 
ouvarov profile image
Olexandr Uvarov •

yep, the 20% version taught us that. a gate nobody passes is just the old queue with extra steps))

helpers are partly covered already. every PR comes with an ADR, and the review bot checks the code against it. so if a helper changes and the ADR doesn't say it and why, no approval. it catches the accidental edits, which is most of the risk. for now that's enough, we're watching. if problems start we'll do something, but right now there's nothing to fix

Thread Thread
 
carbonlayer profile image
CarbonLayer •

That makes sense—the 20% version is a good reminder that a control only helps if it catches real risk without rebuilding the old queue in disguise. Having each PR come with an ADR, then checking the code against it, sounds like a practical way to catch accidental drift while keeping reviews moving.

And if accidental edits are the main risk you’re seeing, “watch and add more only if the evidence calls for it” feels like the right level of process. What would tell you it’s time to tighten the gate—repeated ADR mismatches, regressions, or something else?

Thread Thread
 
ouvarov profile image
Olexandr Uvarov •

neither ADR mismatches nor regressions are a signal by themselves. a bug after merge happens to experienced devs too, it's normal))

so far we didn't have a single case where someone broke something after self-merge. I think it's because of how the flow is built

  1. the code is not written how the AI wants, it goes through a pipeline. first clarifying questions, then a plan where you see what's reused and why, an ADR, and code by the repo rules

  2. two layers of review. first one on the client side, before the PR even exists. second is the bot on GitHub with a clean context, it never saw how the code was written and looks only at the diff and the ADR. together they catch almost everything

so a vibe coder with this flow works not worse than a dev. and tightening makes sense if that stops being true, if their stuff starts breaking more often than ours

Collapse
 
mudassirworks profile image
Mudassir Khan •

the clarifying questions gate refusing to skip is the part that actually scales. most internal tool builders I've seen skip it entirely and then spend a week debugging why the tool does the wrong thing in edge cases.

the reuse of existing components (the review fetcher and the date filter) before writing new code is the version of this that compounds. context length is the tax you pay for not knowing what already exists.

what does the merge gate reject most often? curious whether it's schema mismatches or unexpected side effects on shared tools.

Collapse
 
ouvarov profile image
Olexandr Uvarov •

good question, I didn't want to guess so I went and counted))

no merge without the bot's approve, so whatever it finds the builder fixes before merge. on the builder PRs it reviewed there were 62 findings. and none of your two options is the top one

about half are just bugs inside the author's own tool. a regex with \w that silently doesn't match Cyrillic. a field that is read but never copied into the config. placeholder arrays shipped empty, so the lookup never finds anything. on the scenario the author checked everything works

second most common is runtime limits. we run on Workers, and there a cron stops fitting the time budget after someone adds retries. or a whole file gets loaded into memory with no limit. vibe coders don't know these limits exist, and the model doesn't care until you remind it

third is when the ADR doesn't match the code. ADR says the whole catalog was moved, the code has empty stubs. that's the check I wrote about above

schemas, there was one real case, a payload in the new shape went out before the external service could accept it. side effects on other people's tools, also basically one, the model dropped other people's entries from the registry. after that we took the registry away from the model completely, so it can't break this way anymore

so the zone boundary works. what's left lives inside the author's own tool, and breaking things there is cheap

Collapse
 
elijahbrown profile image
Elijah Brown •

Self-merge for vibe-built tools is a brave gate. One check I'd keep on the core: any new route that writes email or phone should parse the phone for a country and look up the email domain's MX, because generated forms often stop at client-side format checks.

Collapse
 
ouvarov profile image
Olexandr Uvarov •

thanks) but I think it's not really our case, maybe I didn't explain it clear enough in the article

it's an internal platform, there are no public forms or signup routes at all. only our own people log in, or a call comes with api key. so nobody from outside types a phone or email there

and the core is exactly the part vibe coders can't touch, it's in "Who merges": a vibe coder can't change the core without an engineer. any change to the core goes through a developer anyway

skills can't even make network calls, keys exist only in actions and are injected on the server. that's in "Where the platform stops". and the last point there is what it's not for: money and anything you can't roll back

so the risk for us is not bad input from users, it's what generated code can touch. that's why the whole article is about the zone boundary

Collapse
 
elijahbrown profile image
Elijah Brown •

That makes sense. In that setup the boundary around generated code and server-side actions matters more than public input handling.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

That makes sense. In that setup the boundary around generated code and server-side actions matters more than public input handling.

Some comments have been hidden by the post's author - find out more