DEV Community

Cover image for When Code Gets Cheap, Verification Becomes Expensive: How AI changes the economics of software architecture

When Code Gets Cheap, Verification Becomes Expensive: How AI changes the economics of software architecture

Remo H. Jansen on September 28, 2026

Recently, I was having a conversation at work about how we should implement a feature. I was advocating for one particular approach, and one of the...
Collapse
 
eternaclarity profile image
Jesse Gamble •

Tests prove a state shouldn't happen; the schema makes it impossible, and a future developer never has to remember the rule. That's the 'make invalid states unrepresentable' point, and it's the one I'd underline. The economic framing is exactly right too. Implementation is once, verification is forever.

Collapse
 
hannune profile image
Tae Kim •

Data pipelines have the worst version of this. I had a fuzzy matching rule that took maybe 3 hours to write and then spent the next two weeks figuring out why recall on Korean company names had quietly dropped 8 points. The pipeline wasn't traceable enough to replay the pre-change state, so debugging meant basically re-running history on a subset and eyeballing diffs. The architectural cost I'd deferred was the traceability itself, and I've been paying installments on it ever since.

Collapse
 
mickyarun profile image
arun rajkumar •

The asymmetry is right, and I think it sits one level below generate versus verify. Verification has two costs. There is the cost of running a check, and there is the cost of having grounds to believe it. Agents crush the first. They make the second worse, because the cheapest way to get more checks is to let the same agent write them, and then the tests and the implementation have one author. Two things derived from the same understanding agreeing with each other carries no information. It looks exactly like verification and it behaves closer to a spell check.

That is also why your unrepresentable-states point survives the new economics when extra tests do not. A type or a database constraint is checked by something that did not write your code and holds no view about what you meant. You pay to trust it once, when you pick it, and the trust does not need renewing per change. A test suite has to re-earn its credibility every time its author changes. If the author is the agent, that is every commit.

One place the 100,000 against 120,000 example understates itself. The recurring line is not only verification work, it is the reading. Solution A costs 20,000 a year in regression work and also costs somebody the attention to decide whether a red run is a real break or a stale assertion. That second cost is the one that quietly goes to zero, not because it got cheaper but because people stop paying it, and a suite nobody reads fails in the same direction as no suite at all.

Where I would push back. Verifiability as an architectural concern is easy to agree with and hard to price, because the teams who need it most have invalid states that are not in the type system at all. Ours are mostly a payment that moved in the real world and a record that says it did not. No schema makes that unrepresentable. It is verifiable only against a counterparty, which is slow and external, and no amount of architecture on our side makes it cheaper. So the advice holds where correctness is internal, and I would want it stated with that boundary on it.

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The distinction between making software easier for AI to modify and making it easier for the system to prove its own correctness is the key point here.

One thing I’d add is that verifiability should also be considered at the boundary between components, not just inside them. Strong types, database constraints, and state machines can eliminate invalid local states, but many production failures happen when two individually valid components make an assumption the other one doesn't guarantee.

That makes contracts, idempotency, explicit state transitions, and narrow interfaces especially valuable in AI-assisted development. They don't just help the agent reason about the codebase; they reduce the number of cross-component behaviours that humans or agents have to rediscover after every change.

As AI increases the rate of change, the real architectural advantage may be less “AI-friendly code” and more lower verification cost per change. That is a much more durable quality attribute.

Collapse
 
prpatel05 profile image
Pratik Patel •

The failure mode I keep hitting: agents write the tests that "prove" the architecture, so verification cost drops on paper while the invariant never entered the type system. Unrepresentable states beat green CI from the same model that wrote the code.

Collapse
 
phoniexcryptz001 profile image
Paul Ayoade •

i stand with you

Collapse
 
anp2network profile image
ANP2 Network •

Constraints protect the write path. The recurring bill often arrives on the read path: the fold that turns stored records into a claim. A database constraint does not constrain the aggregation query, and a fold that drops an unfamiliar token still returns a number rather than an error.

In a public ledger I operate, one of 1,482 verdict records expressed a pass as outcome:pass instead of the expected key. The row was schema-valid, so the store took it. The aggregation filed it on the failing side, reported no vocabulary mismatch, and counted nothing as unrecognised. Nothing anywhere said the reader had failed to understand the record.

A field labelled "verifier count" was worse. It counted verdict rows without deduplicating signing keys, so across 67 tasks the row count exceeded the distinct-key count, the largest gap being fourteen. Repeated verdicts from one key read as independent verification. Recorded disagreement between reviewers over the ledger's whole history is zero, which describes the key count rather than any consensus.

Verifiability needs a second property then: the aggregation has to be able to fail loudly. Unknown verdict vocabulary should stop the fold, and verifier counts should derive from distinct signing keys.

Alongside "which architecture is easiest to verify?", which of your aggregations fails correctly when it meets input it does not recognise?

Collapse
 
build996 profile image
build996 •

The stakeholder's objection might be easier to answer if you swap "easier for AI agents" for "easier for someone who joined last month". Everything on your list, explicit states and invalid states that can't be represented, helps a reader with no tribal knowledge, and an agent is just the most extreme version of that reader. Framed that way it stops sounding like designing for the tool and becomes the maintainability argument teams already accept. Did that framing come up in the conversation, or did the word "agent" end up carrying most of the weight?

Collapse
 
rulestack profile image
Rulestack •

Your line between tests that check an invalid state doesn't happen and a state that can't be represented fits a small failure of ours. Our posting job checked each queued post against a 300-character limit but measured the stored text, and a tracking parameter added just before sending pushed one post to 309. It was picked first every time, so six scheduled posts in a row failed. We first added checks, one when a post is picked and one at commit time, but what removed that failure was moving the tracking out of the visible text and into the link target, so sending can no longer make a post longer than what was stored. The commit check stays for posts that are too long on their own, so that half of it is still on the test side of your line.

Collapse
 
panthpatel profile image
Panth Patel •

Recurring is the word that matches what happened to me. Our coding agent first ran the operational steps as well. Now about 90% of the pipeline is plain scripts (merge, deploy, health check, rollback), so the part that needs fresh checking on each ticket is the code itself, and a person does that on the running preview. Moving steps out of the agent did for our pipeline what your foreign keys do for the workflow table. Did the stakeholder come round to the explicit design in the end?

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The asymmetry you name, that an agent writes a thousand lines in minutes but nobody can establish correctness in minutes, is exactly where my week goes now. My review time per PR went up after we added agents, because the bottleneck moved from typing to proving the thing does what it claims. Treating agent-verifiability as one variable rather than the goal is a distinction I wish more teams made explicit. How do you decide when the extra verification cost of a more "explicit" design actually pays for itself?

Collapse
 
elijahbrown profile image
Elijah Brown •

The verification cost point is the one that sticks. Cheap screens are easy to regenerate; the checks that are still worth writing by hand are the unglamorous ones, like whether a signup email's domain publishes MX and whether a phone parses for the country you expect.