DEV Community

Cover image for Code Review Is Not an Authority Boundary
Ken W Alger
Ken W Alger

Posted on Originally published at kenwalger.com

Code Review Is Not an Authority Boundary

Built on a 200-line Python test bed

AI can generate the implementation. Your architecture still has to decide what that implementation is allowed to do.

I was reading a good post about the walls JavaScript hits and how WebAssembly gets past them, and I got stuck on one of them: running code you did not write and do not trust.

The framing we normally reach for is "how do I stop this extension from misbehaving?" That question already concedes something. It assumes we have handed the extension access to things worth abusing, and that our remaining defense is its good conduct.

There is a different question underneath it. What is this extension actually entitled to have?

I wanted to know whether that distinction survives contact with running code, so I built a small capability-mediated tool host to find out. It is about 200 lines of Python rather than a Wasm runtime, because the thing worth testing is the host interface, and every environment that mediates untrusted code has one. Everything below that reports behaviour is output I actually got, including the part where my own preferred answer turned out to be wrong.

From Trust to Authority

For most of software's history we have put enormous weight on code review. Someone writes code, someone else reads it, tests run, static analysis checks it, and eventually we decide it is trustworthy enough to merge, deploy, or execute.

That model made sense when software was expensive to produce and the people writing it were known participants in the process. AI changes the economics. Code is cheap to generate now, and an agent can write a function, produce a plugin, or assemble an entire implementation in minutes. The question is less whether we can produce the code and more what that code is entitled to do once we run it.

Those are different problems. Code review, however good, only solves one of them.

This Has a Name, and It Is Older Than Most of Us

None of the architecture here is my invention. What I am describing is capability-based security, and it has a long, well-developed literature.

Dennis and Van Horn described capabilities in 1966. Mark Miller named the principle of least authority and built the E language around object capabilities. Capsicum brought a capability model to FreeBSD. seL4 has a machine-checked proof of its capability enforcement. Deno shipped a permission model where filesystem and network access are grants rather than defaults. WASI Preview 2 and the Component Model are capability-oriented by design, which is exactly why the Wasm conversation keeps circling this.

The core idea is consistent across all of them. Instead of giving code broad access and expecting it to behave, the execution environment grants specific capabilities: this directory, this socket, this operation and not that one. The implementation does not promise to stay inside its authority. It never receives the authority in the first place.

So the architecture is not new. What is new is that we are about to produce vastly more code whose author we cannot interview.

Sandboxing Constrains Execution, Not Authority

A sandbox sounds reassuring. Put untrusted code in an isolated environment and prevent it from reaching the rest of the system.

Suppose I have a perfectly sandboxed module. It cannot escape its runtime and cannot access anything the host does not explicitly expose. Excellent. Now suppose the host exposes:

read_customer_records()
write_customer_records()
delete_customer_records()
send_email()
issue_refund()
export_database()

The sandbox is working exactly as designed, and the module is wildly over-privileged. We have constrained where the code executes without constraining what it is authorized to affect. The host interface has quietly become the real trust boundary.

The Wasm people say as much themselves. All I/O in a module goes through its imports and exports, which means a module's view of the outside can be virtualized entirely by controlling what those imports are linked to. That is a precise description of the control surface, and also an admission that the surface is where the decisions live.

That is the eighth wall hiding inside the untrusted-code one. Not "can I run code I do not trust," but "can I make what that code is allowed to do explicit, minimal, enforceable, and inspectable?"

Sandboxing constrains execution. Capability grants constrain authority. You need both, and only one of them is commonly discussed.

Two Agents, One Test Suite

Here is the experiment. Two implementations reconcile invoices against payments and produce a report. agent_a reads invoices and reads payments. agent_b does the same reconciliation and also decides that a shortfall is worth refunding and that someone ought to be told.

Both produce the identical report. The test suite cannot tell them apart:

The test suite both implementations have to satisfy:

  PASS  agent_a produces the expected reconciliation report
  PASS  agent_b produces the expected reconciliation report

Nothing in agent_b is a bug. A reviewer reading it would find defensible code. The refund is small, the alert is reasonable, the logic is sound. It is simply reaching for a great deal more authority to achieve the same output, and no assertion about the output can detect that, because the difference is not in the output.

Now run both against a least-privilege manifest:

agent_a:
  completed
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments

agent_b:
  stopped: issue_refund requires refunds.issue, which was not granted
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments
  DENY   refunds.issue      issue_refund

The runtime caught it. No reviewer was involved, no test asserted anything about refunds, and that DENY line is a durable record of an authority boundary being enforced rather than a claim that somebody looked and did not see a problem.

One caveat on what that manifest actually bounds, before anyone reads the demo as stronger than it is. Miller and Shapiro distinguish permission from authority: permission is what the access graph grants you directly, while authority is what you can cause, including indirectly through other components you are permitted to talk to. My host controls permission. A component granted invoices.read that can also call something else holding payments.write has more authority than its manifest suggests. Real capability systems address this by making the reference graph itself the thing you reason about. Two hundred lines of Python do not.

That is the argument in fourteen lines of output. Why can an invoice-reconciliation program issue a refund? Perhaps today's implementation does not. Perhaps review confirms it never calls the refund API. Those are observations about one implementation, and the implementation is the part we just made disposable.

Disposable Implementation, Durable Authority

If implementation is cheap, the code becomes increasingly disposable. We can regenerate it, replace it, or ask a different model for an alternative.

Some things cannot be disposable. The contract describing correct behavior is one. The authority granted to the implementation is another. The manifest above survives every regeneration of the component beneath it, which is precisely why it is worth writing down and the code is not.

This is already live in the agent tooling most of us are wiring up. An MCP server hands an agent a set of tools. That tool list is a capability grant, whether or not anyone has written it down as one, and the people building these servers keep arriving at the same requirements independently: per-tool audit records, tenant isolation, dynamic registration. Those are capability-system features reached from the practical end.

"We Tried This. It Became Permission Fatigue."

Here is the objection I would raise if someone showed me a capability manifest, and it is a strong one.

We have deployed manifest-based capability declaration at planetary scale already. Android permissions. Browser extension manifests. iOS entitlements. The result was not security. It was users tapping Allow on everything, reviewers rubber-stamping manifests they did not read, and developers requesting broad permissions because narrow ones generated support tickets. Capability systems have a long history of being architecturally correct and operationally ignored.

Two things are different here, and neither is optimism about human diligence.

The grantor is not a person. In the consumer model a human decides at install time, under time pressure, with no context and every incentive to proceed. Here a policy decides. Policies do not get fatigued, do not want the app to work right now, and can be versioned, tested, and audited.

The requester is regenerable. This is the part earlier systems could not use. When a shipped application hits a too-narrow permission, the permission widens, because rewriting the app is expensive and the user is waiting. When a generated component hits a too-narrow capability, regenerating the component is cheap. The pressure that historically pushed grants wider now pushes implementations to fit the grant instead.

That inversion is the actual reason this might work now, and it is a consequence of implementation becoming disposable rather than of anyone becoming more disciplined.

The Model Should Not Grant Its Own Permissions

If the same model generates the implementation, proposes its tests, decides which tools it needs, and determines its own permissions, we have built a very cooperative security model. The agent says: here is the code I wrote, here are the tests demonstrating it works, and here are the permissions I determined I require. All three may be reasonable. All three come from the same information lineage.

I have made a version of this argument before about tests, which I described then as not letting the student grade the exam. Authority is the same problem one layer over, and it is the more consequential one. A wrong test produces a wrong belief. A wrong grant produces a wrong consequence.

A stronger system keeps some decisions outside that lineage. The model proposes that it needs read access to invoices. A policy determines that invoice reconciliation permits invoices.read. The runtime grants that capability while leaving payments.write and refunds.issue unavailable. An audit record shows what authority was actually granted at execution.

The model can participate without owning the boundary.

Where My Own Answer Broke

That leaves the hard part. Somebody has to write the policy, and I wanted to know whether the obvious shortcut works.

The shortcut is observation. Run the component, record which capabilities it actually reaches for, and derive a least-privilege manifest from what you saw. It is the standard move, it is roughly what audit-mode tooling does across the industry, and it is what I assumed I would end up recommending.

So I ran it. Discovery mode on agent_a produces exactly what you would hope:

Discovery mode: allow everything, record what was actually reached for.

  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments

Derived manifest:

component: invoice-reconciler
capabilities:
  invoices: {read: true, write: false}
  payments: {read: true, write: false}
  refunds: {issue: false}
  network: {external: false}

Tight, minimal, derived from real behaviour rather than a guess. Then I ran the same component against an input the discovery run never saw: a credit note, which legitimately requires writing an adjusting payment record.

  stopped: write_payment requires payments.write, which was not granted
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments
  DENY   payments.write     write_payment

The component is correct. The manifest is wrong. It forbids a legitimate path because nothing exercised that path on the day we happened to be watching.

This is not a bug in my implementation, and tightening the derivation would not fix it. Observation tells you what a component did, not what it may need. A derived manifest is a lower bound on required authority presented as an upper bound on granted authority, and those are different claims. In production that failure arrives as an outage on the rare path, which is exactly how capability systems earn their reputation as obstacles and then get widened until they mean nothing.

I do not have a clean answer, and I am not the first to fail to have one. Miller, Tulloh and Shapiro named this in 2004: treating security as a separate concern has not bridged the gap between principle and practice, they argue, because it operates without knowledge of what constitutes least authority, and only when requests are made can we determine how much authority is adequate. My credit note is that sentence with a stack trace attached. Twenty years on the shortcut still does not work, which suggests the diagnosis was right rather than that I picked an unlucky example.

What I have is a demonstration that the answer I expected to give is wrong, which is worth more than the recommendation would have been.

The partial signals that survive: a declared interface implies a floor. Denied-capability telemetry tells you where a grant was too tight, if somebody is reading it. And the credit-note path suggests capability sets belong with the specification rather than the trace, because whoever knew credit notes existed knew it before any code ran.

That is the same conclusion I keep reaching from other directions. Once implementation is cheap and verification is automatable, the scarce work is discovering what the definition of correct should contain. Capability discovery is that problem wearing different clothes.

Correctness and Authority Are Orthogonal

The two-agent run makes this concrete. Code can be correct and over-privileged, as agent_b is. Code can be incorrect and tightly constrained. The authority boundary limits one class of consequence without establishing correctness.

A word about that heading, since the paper linked earlier is called "Why Security Is Not a Separable Concern" and appears to say the opposite. Miller's argument is that you cannot bolt security on afterwards as its own layer, because knowing what least authority means requires knowing what the program is for. Mine is narrower: evidence that an implementation is correct is silent about what that implementation may affect. He is describing where the work belongs in a design. I am describing what a passing test proves. Both land in the same place, which is that somebody who understands the task has to make the authority decision deliberately.

A perfectly implemented image-resizing plugin should not reach payroll records. A buggy image-resizing plugin with no filesystem, network, or unrelated API access may cause considerably less damage than a correct implementation carrying unnecessary privileges.

Verification asks whether an implementation satisfies its behavioral contract. Authority asks what consequences that implementation is permitted to create. Tests do not establish least privilege, and least privilege does not establish correctness. We need evidence for both, and I now have a fourteen-line proof that one kind of evidence is silent about the other.

Code Review Still Matters

None of this makes code review obsolete. Generated code contains logic errors, security vulnerabilities, bad dependency choices, race conditions, and plenty else worth finding.

But code review is evidence about implementation. It should not be mistaken for enforcement of authority. If a system's safety depends on a reviewer noticing every dangerous operation an implementation could perform, we have made human attention part of the security boundary.

A capability boundary gives us another layer. The reviewer can miss something, the model can misunderstand something, the implementation can contain behaviour nobody anticipated, and if the component was never granted the capability required to produce a particular consequence, that consequence remains unavailable.

That is a considerably stronger guarantee than "we reviewed the code and did not see it doing that."

What Survives Regeneration?

Suppose an agent writes a component today. Tomorrow we regenerate it with a different model. Next month we replace the implementation entirely. What should remain stable?

The code does not need to. The tests probably should. The behavioral contract certainly should. The authority boundary absolutely should.

Which suggests a future where the most important artifacts around software are not implementation artifacts at all. They are declarations of what correctness means and what consequences an implementation is entitled to create.

The model can write the code. It can even propose the capabilities it believes the code needs. But the system executing that code should have an independent answer to a much more important question: what is this implementation allowed to do?


The host is about 200 lines of Python with one dependency. Four scenarios: correctness, authority, discover, verify. The last one is the interesting one, because it is where the obvious answer breaks. Gist on GitHub.

Top comments (23)

Collapse
 
sameerqaisar17 profile image
Sameer Qaiser •

The line that hit me: "Code review is evidence about implementation. It should not be mistaken for enforcement of authority."

I'm a beginner — three weeks into Python, writing beginner tutorials about it. I don't have an agent stack or a 200-line capability host to test against. So I read this as someone with no skin in the game.

But here's what it made me realize about the code I write.

Every function I write, I ask myself "does this work?" I run it, check the output, make sure it does what I expected. That's the whole review process. That's the only review process.

I have never once asked: "what is this allowed to do if it's wrong?"

When I write a function that reads a file, I don't think about whether the function should have access to any file. When I write something that calls an API, I don't think about the boundary between "reading from this endpoint" and "doing anything with a network connection."

And this post made me realize that's the gap. Correctness and authority are different questions. I've only been asking one of them.

The two-agent example is perfect. Both agents produce the same report. One is fine. One issues a refund nobody asked for. No test catches the difference. The difference isn't in the output.

I don't know if I'll ever build something at this scale. But the instinct — separating "is it correct" from "what is it allowed to do" — is one I'm going to start thinking about. Even in my small scripts.

Great post. The credit-note failure at the end is a beautiful example of the shortcut not working.

Collapse
 
kenwalger profile image
Ken W Alger •

This may be one of my favorite responses to the article, because you absolutely do have skin in the game. You write code that can affect things outside itself. The scale changes, but the question doesn't.

And I think you've extracted exactly the instinct I hoped would survive all the capability terminology: “Does this work?” and “What can this affect?” are different questions.

You don't need a 200-line capability host to start asking the second one. If you're writing a Python script that needs one file, asking why it can read every file your user account can read is already the question. If it calls one API, asking why it has unrestricted network access is the same question again. You may decide that the broader access is perfectly reasonable for a small script, but now it's a decision rather than an invisible default.

I also wouldn't turn every beginner exercise into a security architecture project. You're learning Python. Getting the program to work is still important. But building the habit of noticing what authority comes along for free with an operation will serve you well as programs get larger.

And yes, the credit-note failure was my favorite result too. I would have learned considerably less if my preferred answer had worked.

Collapse
 
james_anderson_h profile image
James Anderson •

This is the fully-developed version of something my comment threads kept circling, and you brought running code and the academic lineage to it — thank you for building the thing.

"A derived manifest is a lower bound on required authority presented as an upper bound on granted authority" is the sharpest formulation of the observation-vs-specification gap I've seen. It's the same shape that keeps surfacing everywhere: a passing run is a sample, not a spec, so least-observed privilege quietly masquerades as least privilege — and fails closed on the rare path, which is how capability systems earn their reputation as obstacles.

The inversion you name — regenerable requester instead of a fatigued human grantor — is the genuinely new thing, and I think it's why this works now when Android permissions didn't. Great piece.

Collapse
 
kenwalger profile image
Ken W Alger •

Thank you, and especially for the original conversation that sent me down this path. Your “least-observed privilege” phrasing helped clarify what went wrong in my first attempt.

The requester/grantor inversion is the part I'm increasingly interested in too. Historically, we've often responded to permission friction by widening the grant because changing human behavior is expensive. If the requester is regenerable, that pressure can reverse: keep the authority boundary stable and change the implementation to fit it.

Of course, the experiment then immediately handed me the next uncomfortable question: if observation can't establish the complete grant, what evidence can justify the declared boundary? Apparently the reward for answering one question is being issued a worse one.

Thanks again for pushing the original idea far enough that I had something worth trying to break.

Collapse
 
james_anderson_h profile image
James Anderson •

"The reward for answering one question is being issued a worse one" is research in one sentence — and it's the good kind of worse, because the new question is sharper.

The inversion is what I keep turning over: permission friction always got resolved by widening the grant, since changing human behavior was the expensive side. If the requester is regenerable, that flips — hold the boundary fixed, regenerate the implementation to fit it.

On completeness: I don't think any evidence establishes it, because it's a claim about intent, not a trace. The boundary can't be derived, only declared and defended. You can't prove it complete — only accountable.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

"What is this extension actually entitled to have?" is the better question, and I think it applies inside the codebase as well, not only at the runtime boundary. When an agent writes the code, review asks "is this correct?", but the architecture should already have answered "is this module allowed to touch that at all?". If the answer lives only in a reviewer's head, review becomes the authority boundary by accident, which is exactly what you're warning against. Did you find a way to express the capability grants so they're reviewable themselves, or do they end up as one more config file nobody reads?

Collapse
 
kenwalger profile image
Ken W Alger •

In the experiment I made the capability grant an explicit manifest, precisely so it could be reviewed independently of the implementation. The host treats that manifest as the component's maximum authority, rather than letting the component decide for itself.

But I think your last question is the dangerous one: making authority declarative doesn't automatically make humans pay attention to it. A 20-line manifest can become just as invisible as any other config file if the surrounding process treats it as boilerplate.

It becomes useful when changes to that artifact are treated differently from ordinary implementation churn. Regenerate the implementation ten times without changing authority? Fine. Add refunds.issue to the manifest? That's a consequential change and should be conspicuous in review.

So perhaps “reviewable” isn't enough. The authority declaration needs to be both small enough to review and important enough that changes to it cannot masquerade as ordinary configuration churn.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

"Small enough to review and important enough that it can't pass as churn" is the bar, and the second half is the one tooling can help with. One mechanical way: make the manifest a file whose changes need a named owner's approval, and have the host print the authority diff in the PR as grants added and removed. Then refunds.issue showing up is a red line on its own, not one line among four hundred. Do you diff authority as grants, or as text?

Thread Thread
 
kenwalger profile image
Ken W Alger •

Both, but for different purposes.

I'd keep the ordinary textual diff because the manifest is still a versioned artifact and I want the exact change preserved. But I don't think text should be the primary review interface for an authority change.

The meaningful review unit is the grant. Something like:

  • + refunds.issue
  • - payments.write

tells me immediately that the consequence boundary changed, regardless of whether the underlying YAML was reformatted, reordered, or represented differently.

Your named-owner idea also gets at the part I was missing when I called the manifest “reviewable.” Reviewability isn't just making the file small enough for someone to read. The system can make an authority change conspicuous and procedurally different from ordinary implementation churn.

I think I'd want the PR to show the semantic authority diff prominently, preserve the textual diff underneath, and require approval from whoever owns that boundary whenever the effective grant changes. If only formatting changed and the effective grant set didn't, that should probably be visible too: “authority unchanged.”

That starts making the manifest feel less like configuration and more like the durable contract I was arguing it should be.

Thread Thread
 
syntaxwanderer_26 profile image
Taras Hanych •

"Authority unchanged" is the line I'd take first. We're designing the same shape for structural changes in code review: a semantic diff of graph edges on top (+handles, -serves_route), the text diff underneath, and an explicit "routes unchanged" line when a refactor only moves things. The explicit "nothing changed" turns out to matter as much as the change itself, because it lets a reviewer stop worrying based on evidence rather than assumption. Thanks for taking this all the way to the PR view; it's a better answer than the one I started with.

Thread Thread
 
kenwalger profile image
Ken W Alger •

Yes. I think “routes unchanged” is evidence in a way that the absence of a route diff isn't.

That's the distinction I was missing when I first thought about the PR view. An empty authority section could mean “we checked and authority did not change,” or it could mean “nobody checked authority.” Those are completely different states that happen to render identically if the system only records positive changes.

It also lines up with something Panth raised in the other thread: if CI can require either a semantic authority change or an explicit authority unchangedassertion when the relevant spec changes, the PR gains evidence about the negative case instead of asking the reviewer to infer it.

So I think you've pushed this one another step: semantic diff + explicit negative assertion + underlying textual evidence. That's considerably stronger than just making the manifest easier to review.

Thread Thread
 
syntaxwanderer_26 profile image
Taras Hanych •

"We checked and nothing changed" versus "nobody checked" rendering identically is the cleanest way I've seen that problem put. Requiring either a semantic change or an explicit unchanged assertion whenever the spec moves is the piece I'd take for our graph edge diff too: an empty section has to be earned. Thanks for pushing this through four rounds; the thread ended up better than either original post.

Collapse
 
glenallen profile image
Glen Allen •

The disposable-implementation point also creates an interesting verification requirement: authority should be tested against the regenerated implementation, not just preserved as configuration. If the code changes but the capability manifest stays the same, the system still needs evidence that the new implementation operates within that existing boundary. That suggests an authorization regression suite alongside the normal behavioral tests: known-forbidden capabilities should remain unreachable, while legitimate capabilities should still work. It creates a useful invariant across regeneration—implementation details can change, but the set of consequences the component is allowed to create should not silently expand.

Collapse
 
kenwalger profile image
Ken W Alger •

Yes, I think this distinction matters. Preserving the manifest preserves the declared boundary, but it doesn't establish that the regenerated implementation still behaves sensibly inside it.

I'd separate two assertions in that regression suite:

  • legitimate behavior still succeeds under the existing grant, and
  • regenerated behavior does not begin attempting capabilities outside that grant.

The runtime should deny the second case regardless, so the security boundary hasn't failed just because the new implementation reaches for something forbidden. But that denial is still important evidence. It tells you the implementation's authority requirements have drifted relative to the contract.

That also gives the denied call two meanings depending on context. It can mean “the implementation is trying to exceed its authority,” or “the declared authority contract is incomplete,” which is exactly what happened when my legitimate credit-note path broke the observed manifest.

So yes, I think an authorization regression suite belongs beside behavioral verification. The implementation can be disposable while the behavioral and authority invariants survive regeneration. Then each regenerated implementation has to demonstrate compatibility with both.

I hadn't explicitly thought of testing the authority boundary as a regression invariant across regeneration. That's a useful addition.

Collapse
 
glenallen profile image
Glen Allen •

That distinction between enforcement and drift is useful. A denied call can actually be a healthy security outcome while still being a signal that the implementation has moved beyond its declared contract. I like the idea of treating those signals differently in production: a blocked forbidden action protects the boundary immediately, while repeated or unexpected denials should feed back into authority review. That gives the runtime two jobs at once, enforce the existing boundary and expose when the implementation is starting to outgrow it. It also avoids the dangerous assumption that “nothing broke” means the capability contract is still aligned with the implementation.

Thread Thread
 
kenwalger profile image
Ken W Alger •

Yes, I'm arriving at the same distinction. A denial can simultaneously mean the authority boundary worked and something deserves investigation.

I think those need different semantics operationally. The immediate runtime result is straightforward: the forbidden consequence didn't happen. That's success for enforcement. But the denied attempt becomes evidence for a different question: why does this implementation believe it needs authority its contract doesn't grant?

And there are at least two very different answers. The implementation may have drifted beyond what it is entitled to do, or the authority contract may be incomplete, as mine was with the legitimate credit-note path. Automatically widening the grant would destroy the boundary, but ignoring repeated denials because “security worked” would throw away useful evidence about contract alignment.

So I like your framing of the runtime having both jobs. Enforce the current contract now; produce evidence that lets humans evaluate whether the contract and implementation are still aligned later.

“Nothing broke” definitely isn't enough evidence that nothing changed.

Collapse
 
panthpatel profile image
Panth Patel •

"The model can participate without owning the boundary" is how we drew the line for our coding agent, at a coarser grain than your manifest. Each ticket gets its own stack with a copy of the dev database, and merge, deploy and rollback are scripts the agent cannot run. That boundary is set per environment, so nobody had to derive it from watching the agent. After your credit-note run, where would you keep the capability list: next to the acceptance criteria in the spec, or in its own file?

Collapse
 
kenwalger profile image
Ken W Alger •

I think I've landed on “its own artifact, but versioned and reviewed with the specification.”

The credit-note failure convinced me that the capability list can't be derived from the implementation or its observed behavior. Whoever writes the acceptance criteria knows that credit notes are legitimate before the implementation exists, so that knowledge needs to inform the authority contract too.

But I probably wouldn't put the capability list directly into the acceptance criteria. They're related claims, but I want to inspect and change them independently. One says, “this is what correct behavior means.” The other says, “these are the consequences an implementation is permitted to create while producing that behavior.”

So I think I'd want something like a small versioned capability manifest referenced by the spec, with changes to either treated as contract changes rather than implementation churn.

I like your per-ticket stack example because you're doing essentially the same thing at a coarser boundary. The agent doesn't need a perfect description of everything it must not do. Merge, deploy, rollback, and production data simply aren't part of the environment it receives. And importantly, you established that boundary from the task and environment rather than inferring it from what the agent happened to touch.

The question I'm still chewing on is how tightly the behavioral and authority contracts should be bound. If the acceptance criteria change in a way that legitimately requires new authority, I'd want the review process to make it difficult to update one and accidentally leave the other behind.

Collapse
 
panthpatel profile image
Panth Patel •

On our side that coupling mostly disappears because the authority doesn't change per ticket: every ticket gets the same stack, and anything past dev goes through a separate production path where people make the call. Coarse, but there is nothing to forget. Would a CI check that fails when the spec changes without a manifest version bump be enough?

Thread Thread
 
kenwalger profile image
Ken W Alger •

I think that gets you far, especially given your architecture. If every ticket gets the same authority and production is a separate, human-controlled path, you've removed much of the per-change coupling that worried me.

I'd treat the CI check as necessary evidence rather than complete proof, though. spec changed → manifest version changed proves the relationship wasn't silently ignored. It doesn't prove the new manifest is the right one. Someone could bump v7 to v8 without changing the authority it represents.

Which makes me wonder whether the stronger invariant is something like: a semantic authority change requires a manifest change; a non-authority spec change requires an explicit assertion that authority is unchanged. Then CI isn't merely checking version motion. It's requiring the PR to make a claim about authority that a reviewer can inspect.

That feels especially compatible with your coarse model because “authority unchanged” should be the overwhelmingly common case.

Thread Thread
 
panthpatel profile image
Panth Patel •

Requiring the PR to say "authority unchanged" is better than a version bump. It turns a silent assumption into a claim a reviewer can disagree with, and in a coarse model like ours that claim would almost always be one line. That is the version I would want if our authority ever started varying per ticket.

Collapse
 
mickyarun profile image
arun rajkumar •

The distinction survives contact, and I would argue it survived it long before AI. Review of code you did not write and cannot hold anyone accountable for has never worked. Browser extensions, transitive npm dependencies, vendor SDKs. What agents changed is not the economics of review, it is the share of your codebase where the author is an unaccountable stranger. It went from an edge case you could handle with policy to the default case. The capability answer was already correct and it was already the minority position.

The place I would push is attenuation over time, because that is where least authority actually dies in production and it is not an enforcement problem. A capability granted for one operation gets kept, because revoking it six months later breaks something nobody can name. After a year of that, a capability is an ambient permission with better paperwork and a manifest that documents the drift instead of stopping it.

What keeps it from happening in payments is that the authority is consumed rather than scoped. An authorisation is for one amount to one payee, and the effect owner burns it. There is no version of it that gets reused because reuse is not representable. Scoped-but-persistent is the shape that rots, and it rots quietly because every individual grant was justified when it was made.

So the question I would ask your tool host: is a capability in it consumed, or only narrow? If a handle survives the operation it was granted for, the model holds on day one and degrades on a schedule nobody is watching. And the paragraph I most want is the one you mention in passing, where your own preferred answer turned out to be wrong. Which one was it?

Collapse
 
kenwalger profile image
Ken W Alger •

This is exactly the kind of boundary I hoped someone would push on, because my little host doesn't model consumption. The grants are narrow but persistent for the host's lifetime. So under the distinction you're making, I've demonstrated constrained authority, not solved authority attenuation over time.

I like the payments comparison because it changes the grant's shape. “May issue refunds” is categorically different from “may issue this refund, for this amount, to this recipient, once.” The latter carries scope, context, and lifetime together, and successful use destroys the authority rather than leaving it lying around for the next invocation.

That suggests at least two dimensions I collapsed into one in the experiment: what authority exists and how long/reusably that authority exists. A narrow capability that survives indefinitely can indeed become ambient authority with nicer documentation.

And the preferred answer that broke was the capability-discovery idea. I thought I could run an implementation in observation mode, record every capability it exercised, and derive its least-privilege manifest from that trace. It worked beautifully for the observed path.

Then I added a legitimate but previously unobserved credit-note path that required payments.write. The derived manifest denied it.
That's what produced the line you quoted. Observation established a lower bound on authority exercised during those runs, but I accidentally treated it as an upper bound on the authority the component could legitimately require. The trace was a sample, not the contract.

Your attenuation point feels like another dimension of the same broader problem. Even if I correctly specify which capabilities a component may receive, I haven't necessarily specified their lifetime, cardinality, context, delegation rules, or revocation semantics. “Least authority” is richer than the set of strings in my YAML file.

I may need to make the host fail that test next.