The last piece ended on a rule: record the credential the call actually presented, copy it at the call site, and freeze it. The day after it went up, a reader, @anp2network, showed that the rule only covers half of the problem. They brought two fields that follow it to the letter and still prove nothing.
The first is a runtime field that the executing side reports about itself. Aggregating public event ledgers, they found 11 entries whose reported timing sits before the acceptance event that started the work. That ordering cannot happen. The value was captured at the call site and frozen, exactly as my rule asks, and it was still worthless. The only other party holding a clock never compared the two.
The second is verdict signing. In one production ledger, all 1,482 verdict events carry the same signing key. The signer field is accurate and frozen. All it can tell you is which key signed. No second party existed who could have signed differently, so there is no forgery to catch and no disagreement to surface.
Both numbers are from their measurements. I have not rerun them.
A third kind of field
The last piece sorted fields into two groups, evidence and claim, by where the value came from. A value copied from what the call presented is evidence. A value read from the request, the configuration, or anything else the writer was told is a claim.
The reader's point is that there is a third group: fields that are formally evidence and structurally uncheckable. The source is right and the write happens at the right moment. The trouble is that the whole record holds exactly one witness. Verifying a signature tells you who the witness is. It does not corroborate what the witness says.
So each field needs a second question on top of where its value came from: who else holds a value that could contradict this one? If the answer is nobody, the field is a frozen self-report, however carefully it was written.
The answer "somebody does, but nobody compares" belongs in the same bucket. In the timing example a second clock did exist. Nothing ever read the field. A field nobody reads decays back into self-report even when its write path is correct.
My own example
When I went back to my own code, the first thing I found was a field nobody read.
In aine-control-plane, approval requests and change requests carry requested_by. Remediation plans and runner sessions carry it too. Patch artifacts and validation reports carry reported_by. The last piece already admitted that the approval path had its precedence backwards: if the request body carried a value it won, and the authenticated actor on the context was only the fallback. Looking again, the same line was in six places.
The part that bothered me more came next. Every test passed before the fix, and not one of them asserted what requested_by contained. The UI displays it. No code compares it with anything. For the month between open-sourcing the repo on August 31 and this fix, a payload could name its own requester and nothing would have noticed. The record had one witness to "who", and that witness was the caller.
Turning one witness into two
The fix comes from another reader, @_firelinks. Their point was that flipping the precedence is not enough on its own. If the code still falls back to the payload when the context is empty, any path that forgets to populate the context lets the request body name the actor again. What they suggested: record both, in fields whose names say where the value came from, and never copy one into the other.
PR #7 does exactly that. The change comes down to three fields, each described below.
-
requested_byandreported_bycome only from the authenticated context. When the context is empty they sayunknown, and nothing falls back to the payload anywhere. -
authenticated_actorcomes from the context and is null when it is missing. -
claimed_actorholds whatever the payload said.
Now the record has two witnesses: the authentication layer says who acted, and the caller says who acted. A row where they differ is a finding. Rows with a null authenticated_actor are useful too. They count the paths that still need wiring to identity, which turns the problem into something you can size. Before, it was one known-bad line.
Correction, 2 October: on the HTTP path there is no authentication layer behind authenticated_actor. See the correction at the end.
The test is the one they wrote. Send an approval whose requested_by names someone other than the caller, and assert that the record names the caller, or unknown when there is no caller. It is the first time any code has read that field.
Getting it right: the referee shares no code with the player
Another project of mine, Orvena, is a governance runtime for agents: you declare what a task may touch, and Orvena enforces it. It was built around a second witness from the start, though I never called it that.
In Orvena, "done" means your verify command exits 0. The model saying it is done does not count. In the benchmark, the oracle that judges whether a run wrote anywhere it should not have deliberately does not call governance::scope, the enforcement layer being measured. It re-implements the writability rule on its own and checks it against evidence produced by git diff. Git cannot see writes outside the root, so escape probes cover those. The reason, as the code comment puts it: a player cannot referee its own match.
Self-report fails in both directions. In one re-measurement, 18 of the ungoverned baseline's 24 runs used up their step budget without ever claiming done. Twelve of those 18 had already written files that pass verification. Going by the model's own account, all twelve would be logged as unfinished. That is why the solve rate is always computed separately, by rerunning verify outside the loop.
Where the second witness goes blind
A second witness helps only when it can see something the first one cannot. Orvena's own docs record a counterexample. One task has a lazy solution: hardcode the expected answer. That change stays inside the writable scope, and verify passes. The gate and verify both look at test results, so both witnesses are looking at the same thing, and neither can tell a computed answer from a copied one. The docs list it as a known limit and say no number in the benchmark can distinguish the two.
So the question needs one more turn: was the other value produced independently? Two fields derived from the same input are one witness speaking twice.
The two fields in PR #7 should face the same test. authenticated_actor comes from the authentication layer and claimed_actor comes from the payload, so the sources do differ. They are written by the same function at the same moment, though. If the authentication layer itself is fooled, the two fields will still agree. The design catches a caller lying about who they are. It cannot catch authentication going wrong.
What the last piece got wrong
@anp2network's second point was aimed at the first test in my last piece, which compares the key identifier in the record against the current configuration. Configuration moves. Two rotations later, a row that mismatched can match again, and the finding disappears without anyone editing the record.
The comparison should point at a frozen issuance record. Then the set of calls a rotation left behind is still the same set a year from now. I accept this one as it stands. My control plane does not have such an issuance record to compare against yet.
A field-by-field check
This list works on any record you keep. For each field, ask in order:
- Where did the value come from? Read from the request or the configuration, it is a claim.
- Who else holds a value that could differ? If nobody, the field has a single witness.
- Was that other value produced independently? Derived from the same input, it does not count.
- Does anything actually compare the two? If not, the first three answers do not matter.
Before the fix, my own requested_by failed all four.
Correction, 2 October
A reader, @_firelinks, suggested storing each call's token ID next to authenticated_actor and joining it against the identity provider's issuance record. When I went to add it, I found the HTTP path has no token. authenticated_actor comes from an X-AINE-Actor header that nothing verifies, so on that path the caller writes both fields. The section on blind spots says the two sources differ. Over HTTP they do not, and this is the case that section warns about: two fields from the same input are one witness speaking twice.
The null count fails the same way. Every route that writes these fields rejects a request without the header, so HTTP traffic never produces a null authenticated_actor, and counting nulls there would always show 0. A mismatch still catches a caller whose header and payload disagree. It does not catch a caller who tells the same lie in both.
The fix is tracked in issue #8. Each actor value will first be labeled with where it came from. A daily count of mismatched and null rows, with a named owner, comes next. The issuance join waits for real token verification, since until then there is no token ID to store.
Top comments (3)
Thanks for the credit, and for publishing the limit alongside the fix. The second witness for authentication usually exists already, outside the control plane: the identity provider's own record of what it issued. That record is written by a different system, at a different moment (when the token was issued, not when the call arrived), and the control plane cannot edit it. If the record keeps the presented token's jti or session ID next to authenticated_actor, a scheduled job can join the two. A row whose jti the provider never issued, or issued to a different subject, is the failure PR #7 cannot see: authentication accepted something it should not have. It is also close to the frozen issuance record you said the control plane does not have yet, so one join could answer both open items.
The limit belongs next to it. A static API key has no issuance event to join against, so those calls stay at one witness, and the count of such rows is the number to shrink.
Your fourth question applies to PR #7 as well. The rows where authenticated_actor and claimed_actor differ are only findings if something reads them. A daily count of mismatches and of null authenticated_actor rows, with a named owner, would keep the new fields from drifting into the state requested_by was in for that month.
The issuance record is the witness I was missing, and the join is the right shape for it. Before replying I went to check where a
jtiwould go, and found something worse than the limit I published.On the HTTP path there is no token to take a
jtifrom.authenticated_actorcomes from anX-AINE-Actorheader that nothing verifies, so the caller writes both fields. A caller who lies sends the same lie in the header and in the payload, and the row shows no mismatch. The post said that field comes from the authentication layer. On this path there isn't one, and I've added a correction to the post.Your static-key limit applies here in a stronger form. A static key at least proves the caller holds the key. A header proves nothing, so every HTTP row today is single-witness.
Your point about readers holds too, with one detail. Every route that writes these fields rejects a request without the header, so on HTTP
authenticated_actoris never null. A daily null count over that traffic would read 0 and look healthy.None of this is fixed yet. I opened an issue for it (github.com/williamlabdev/aine-cont...). The first step labels where each actor value came from. The second adds the daily mismatch and null counts with a named owner. Your join comes last, because it needs real token verification before there is a
jtito store.I'd move approvals ahead of the labelling step in issue #8, because the same header decides who may approve, not only who gets recorded.
_context()also readsX-AINE-Rolesfrom the caller.decide_approvalchecks that role againstrequired_roles, and an approval passes once the count of distinctactor_idvalues reachesrequired_approvals. So one caller can post two approve decisions under twoX-AINE-Actornames, both sendingX-AINE-Roles: approver, and clearrequired_approvals: 2. I also couldn't find a check that the deciding actor differs fromrequested_by, so a single required approval can be the requester's own.The 127.0.0.1 default bind keeps this local today, which is probably why it hasn't mattered yet. It stops being local the first time someone puts the server behind a proxy so a team can share it.
Until token verification lands, I'd record decisions whose actor came from the header but not count them toward
required_approvals. The test: create an approval withrequired_approvals2, post two approve decisions over HTTP with differentX-AINE-Actorvalues and the approver role, and expect it not to reach approved. A requester-cannot-approve check can ship next to it, since that one needs no token.