DEV Community

Cover image for The ADR That Rots Isn't the Decision, It's the Assumption Under It
Andrii B.
Andrii B.

Posted on AI-assisted

The ADR That Rots Isn't the Decision, It's the Assumption Under It

The best thing about architecture decision records is that nobody's allowed to edit them. It's also the reason they quietly lie to you.

Here's a concrete case. On July 28, 2023, AWS announced that starting February 1, 2024, every public IPv4 address would cost $0.005 per hour, attached or not. That's about $3.65 a month per address. Now imagine an ADR your team accepted in 2022: "Services get public IPs directly; we skip NAT gateways to keep the network simple and cheap." Every word of that record is still accurate. It describes exactly what you decided and why. And the "cheap" part stopped being true on a specific date that was announced six months in advance, in a blog post nobody on your team connected to ADR-014.

The decision didn't rot. The assumption under it did. And your ADR process, the one everyone agreed was a best practice, had no mechanism to notice.

The record is immutable. The world it describes isn't.

Immutability is the whole design. When Michael Nygard wrote the post that popularized ADRs back in November 2011, he put it plainly: "If a decision is reversed, we will keep the old one around, but mark it as superseded." The major cloud guidance repeats it almost word for word. AWS's prescriptive guidance says "When the team accepts an ADR, it becomes immutable." Microsoft's Well-Architected Framework calls the log "an append-only log" and tells you "Don't go back and edit accepted records."

That's the right call, and I'm not going to argue against it. A decision log you can rewrite is a decision log you can't trust. The value is in being able to read what the team believed at the time, not a cleaned-up version that makes everyone look prescient.

But look at what the model assumes. It assumes a reversal is an event: someone notices the decision is wrong, writes a new ADR, and marks the old one superseded. Supersession is reactive. It only fires when a human spots the problem.

And the nasty thing about assumption rot is that the decision usually keeps working. Your services still run on public IPs. Nothing pages. The only signal is a line on the bill that grew, and the person reading the bill isn't the person who remembers why the network looks the way it does. Nygard named this exact problem in the same 2011 post: "One of the hardest things to track during the life of a project is the motivation behind certain decisions." ADRs fixed the tracking. They didn't fix the checking.

Where the rot actually lives

Every architecture decision rests on a handful of beliefs about the world at the moment you made it. Some are about cost. Some are about licensing, scale, vendor roadmaps, regulation, or the team you had in the room. Those beliefs have a shelf life. The decision text doesn't.

Here's the part I find interesting: the earliest well-known decision template treated this as a first-class field. Jeff Tyree and Art Akerman's 2005 IEEE Software paper, "Architecture Decisions: Demystifying Architecture", had an explicit Assumptions section, described as: "Clearly describe the underlying assumptions in the environment in which you're making the decision", and the paper lists cost, schedule, and technology as examples.

Then the lightweight templates won, and for good reasons. Nygard's format folds everything into Context, which "describes the forces at play, including technological, political, social, and project local." MADR 4.0, released in September 2024, gives you Context and Problem Statement, Decision Drivers, Considered Options, Decision Outcome, Consequences, and a Confirmation section, but no dedicated slot for assumptions. So in most repos today, the beliefs a decision depends on live as prose inside a Context paragraph. Readable, yes. Checkable, no.

You don't have to look far for assumptions that broke under decisions that were still "accepted" in somebody's repo:

The decision The assumption underneath What changed
Give services public IPs, skip NAT Public IPv4 is effectively free AWS charges $0.005/IP/hour from Feb 1, 2024
Standardize on Redis as the cache It's BSD-licensed open source Redis 7.4 moved to RSALv2/SSPL (announced March 2024); the Linux Foundation launched the Valkey fork under BSD-3 on March 28, 2024; Redis 8 added AGPLv3 as an option in May 2025
Terraform for all infrastructure It's MPL-licensed open source HashiCorp moved to the Business Source License on Aug 10, 2023; OpenTofu shipped its first stable release Jan 10, 2024

Notice that none of these make the original decision wrong. Plenty of teams are still correctly running Redis and Terraform. That's exactly what makes this a gotcha, and it cuts both ways:

  • You keep a decision for a reason that no longer exists. The choice might still be right, but nobody re-derived why. You're running on inertia with a rationale that's now fiction.
  • Someone reverses a decision by pointing at the dead assumption, without checking whether the decision still holds for other reasons. "Terraform isn't open source anymore, so we're migrating" might be a great call or a six-month distraction, and the ADR can't tell you which, because it never separated the assumption from the argument.

Write assumptions so they can fail

The fix starts at write time, and it's cheap. Pull assumptions out of the Context prose into their own list, and write each one so it can be falsified by something you can observe. An assumption without a failure signal is just a mood.

Compare these two:

Weak:   "Traffic is moderate, so a single Postgres primary is fine."
Strong: "A1: Peak write load stays under ~2k writes/sec on the primary.
         Breaks when: p95 write latency on the primary exceeds 50 ms for a week,
         or we sign a customer that needs regional data residency."
Enter fullscreen mode Exit fullscreen mode

The strong version names a number, a signal, and an event. The numbers there are illustrative, not a recommendation, but the shape is the point: someone reading it in 18 months can check it in five minutes without having been in the room.

In practice I'd add a small section to whatever template you already use. It doesn't fight MADR or Nygard; it just gives the assumptions somewhere to live that isn't a paragraph:

## Assumptions

| ID | Assumption | Breaks when | Confidence |
|----|------------|-------------|------------|
| A1 | Public IPv4 costs are negligible at our scale | Any per-IP charge appears on the bill | High |
| A2 | We stay single-region for 2 years | A customer contract requires EU data residency | Medium |
| A3 | The team owning this has 2+ people who know it | Either of them leaves or moves teams | Low |

Review by: 2027-03-01
Enter fullscreen mode Exit fullscreen mode

That Confidence column isn't my invention. Microsoft's ADR guidance tells you to "record the confidence level of the decision. Sometimes an architecturally significant decision is made with relatively low confidence. Documenting that low confidence status could prove useful for future reconsideration decisions." I'd push it one level down, onto each assumption, because a decision is usually high-confidence on some legs and a guess on others. The low-confidence assumptions are the ones you review first.

The Review by date matters more than it looks. It turns "we should revisit this sometime" into something a script can find. If your ADRs live in the repo (and they should; Thoughtworks put lightweight ADRs in the Adopt ring back in November 2017 with a strong preference for source control over wikis), listing the overdue ones is a few lines:

import pathlib, re, datetime

today = datetime.date.today()
for adr in sorted(pathlib.Path("docs/adr").glob("*.md")):
    text = adr.read_text()
    if "Status: Superseded" in text:
        continue
    m = re.search(r"^Review by:\s*(\d{4}-\d{2}-\d{2})", text, re.M)
    if m and datetime.date.fromisoformat(m.group(1)) <= today:
        print(f"overdue: {adr.name} (review by {m.group(1)})")
Enter fullscreen mode Exit fullscreen mode

Run that in CI on a schedule, or just before your review meeting. It won't tell you whether an assumption broke. It tells you which ones nobody has looked at in a while, which is the part humans reliably forget.

Fitness functions catch the wrong drift

The obvious objection: isn't this what fitness functions are for? Neal Ford and Rebecca Parsons define an architectural fitness function as something that "provides an objective integrity assessment of some architectural characteristic(s)." MADR 4.0 even has a Confirmation section (renamed from Validation) for "how the implementation of/compliance with the ADR can/will be confirmed", with ArchUnit-style tests as the example.

Those are great, and you should have them. But they check a different thing. There are two kinds of drift, and they need different tools:

  • Implementation drift is when the code leaves the decision. Someone adds a direct HTTP call between services that were supposed to talk over the message bus. A fitness function catches that in the build, every time, with no human involved.
  • Premise drift is when the world leaves the decision. The code matches the ADR perfectly. The pricing page changed, or the license did, or your traffic grew 10x, or the two people who understood the system left. No test in your repo can fail on that, because nothing in your repo changed.

Fitness functions guard the decision against the code. Nothing guards the decision against reality except a person who goes and looks. That's not a tooling gap you can close with a better linter. It's a job, and at principal level it's usually yours.

Two kinds of drift: implementation drift, where code diverges from the ADR and a fitness function catches it, versus premise drift, where code still matches but the assumptions under the ADR have rotted and a decision-log review catches it

The decision-log review ritual

Here's the ritual I'd run. It's small on purpose, because an architecture ceremony people skip is worse than none.

Cadence: quarterly, plus tripwires. A calendar review catches slow rot. Tripwires catch fast rot. Hold an unscheduled review when any of these happen:

  • A vendor you depend on announces a pricing, licensing, or support-lifecycle change.
  • A load or cost metric crosses a threshold written into an assumption.
  • A team reorg, or someone who owns a key system leaves.
  • An incident postmortem names an architectural cause.

Review the assumptions, not the decisions. This is the core move, and it's what keeps the meeting short. Re-litigating decisions is a bikeshed with no bottom. Checking assumptions is mechanical: for each open ADR due for review, go down the assumptions table and ask "is this still true, and how do we know?" Low-confidence ones first.

Every assumption gets one of three outcomes:

  1. Holds. Note the date and the evidence. Move on.
  2. Weakened. It's still roughly true, but trending the wrong way. Note it, tighten the "breaks when" threshold, and pull the review date in.
  3. Broken. Now, and only now, you re-examine the decision itself. Either it still stands for other reasons, in which case you write a new ADR that re-affirms it with the current rationale and supersedes the old one, or it doesn't, in which case you write the ADR that changes course.

The point of option 3's first branch is that "we're keeping it" is also a decision, and it deserves a record with honest reasons. Otherwise the next person to read ADR-014 still sees the dead rationale and draws the wrong conclusion.

Where the review notes go. This is where people get stuck on immutability, so take a position: don't edit the accepted ADR, not even to add a "reviewed on" line. Keep a separate, append-only review log next to the records:

docs/adr/
  0014-public-ips-skip-nat.md        # never edited after acceptance
  0031-supersedes-0014-nat-gateway.md
  reviews/
    2026-q3.md                       # one file per review, append-only
Enter fullscreen mode Exit fullscreen mode

The ADR stays a faithful snapshot of what you believed. The review log is the history of whether you still believe it. Both are immutable in their own way, and together they answer the question a new engineer actually asks: "Is this still why we do it this way?"

Who's in the room. The ADR owner for each record under review, and whoever owns the thing the assumption is about (the person who reads the cloud bill, the person who tracks the vendor contract). AWS's guidance is explicit that every ADR should have an owner who actively maintains it; this ritual is what "actively maintains" means after acceptance. If an ADR's owner has left, reassigning ownership is the first action item, before anyone reads a single assumption.

The decision-log review flow: quarterly and tripwire triggers feed a review of assumptions, low-confidence first, with outcomes Holds, Weakened, or Broken leading to a new ADR that re-affirms or changes course and marks the old one Superseded

Why this sits with the principal

You could hand this to whoever maintains the ADR folder, and in a healthy team the owners do the actual checking. But the reason the job drifts up to the principal or staff engineer is that you're usually the only person who holds enough of the decision graph in your head to see that one broken assumption props up four decisions, not one.

It's also one of the cheapest ways to steer architecture without managing anyone. Thoughtworks put the architecture advice process in the Trial ring in April 2025, describing a setup where "anyone can make any architectural decision, provided they seek advice from those affected and those with relevant expertise," with ADRs as part of what keeps those decisions informed. Decentralized decisions scale well. They also scatter their assumptions across a dozen teams' ADR folders, and nobody below the principal seat is looking at all of them at once. The review log is how you notice that three teams each quietly assumed the same vendor price.

This pairs naturally with deliberately temporary systems. If you've built something as sacrificial architecture, its ADR has the most important assumption of all baked in: "we'll replace this before it matters." That's an assumption with an expiry date, and it's the one that most reliably gets forgotten until the temporary thing is load-bearing.

The honest version

Most of your assumptions will hold. Most quarterly reviews will be twenty minutes of "still true, still true, weakened, still true." That's fine. That's what a smoke detector is like most days, too.

The ritual isn't there to find something every time. It's there so that when the assumption that actually matters breaks, the price change or the license change or the one engineer who understood the billing service handing in notice, you find out in a meeting you scheduled, with the ADR open in front of you, instead of in an invoice, a legal review, or an incident channel at 2am.

Keep the records immutable. Just stop treating immutable as the same thing as still true.


Originally published at andriiboyko.com.

Top comments (0)