DEV Community

Cover image for The credential your record names is not the one that made the call
weiche chiu
weiche chiu

Posted on Originally published at williamlab.dev AI-assisted

The credential your record names is not the one that made the call

The last piece ended by naming something I could not answer. A reader had asked what happens when a credential is rotated in the middle of a long-running session, and I said that question landed in the part my control plane explicitly excludes, so I had nothing to report. A second reader replied that I had already answered it. They were right. The rule held up; the boundary I drew was too wide, and it pulled this question into the excluded part.

The answer was the second rule, one layer down. The record has to hold what the system actually did rather than what it was asked to do. For context length that means the value the runtime granted, whatever number the request carried. For a credential it means the one the call actually presented, not the one the configuration names.

The window that reads as complete

A rotation does not take effect everywhere at once. Configuration points at the new key immediately. A session that resolved the credential once and is holding the token keeps using it until that token expires or a call fails. For the length of that window there are two credentials in play and only one of them is doing anything.

A record that reads the credential from configuration at write time is wrong for exactly the calls inside that window. It is not empty and it does not error. It names a real key, in the right format, for a call that key never made. This is the same failure the last piece opened with, where a field that was never recorded and a field that happens to equal the default read back identically. The record is wrong in a way that looks like being right, which is the only kind of wrong that survives six weeks.

The calls inside that window are also, reliably, the ones someone will ask about. Rotations happen because something prompted them.

It costs a value, not a secret

The fix does not require handling anything sensitive. A key identifier is enough to say which credential made a call. For a JWT, kid and jti name the key and the specific token, and exp says when that token stopped being able to act. None of the three are secret. They are the parts of a credential designed to be quoted.

That third field is doing more work than it looks. After a rotation the question that follows immediately is how long the old credential could still have acted, and exp answers it from inside the record. Without it, the honest answer to whether a call happened before or after the cutover is that you would have to reconstruct it, which is the thing the record exists to avoid.

Copy it, do not look it up

There is a way to add these fields and still have the same bug, one layer up.

If the record stores the key identifier but resolves exp when the record is written, or worse when it is read, it resolves against whatever key metadata is current at that moment. After the next rotation that metadata describes a different credential. The record would then be authoritative about the lifetime of a key that had nothing to do with the call, and it would be internally consistent, so nothing would flag it.

So the rule is narrower than recording the credential. The record holds kid, jti and exp as presented, copied at the moment of the call, frozen. A provenance field resolved later is not provenance. It is a reading taken at read time, stored in a column that claims otherwise.

This is also why the field tends not to exist. To copy what the call presented, the thing writing the record has to see what the call presented, which means the recorder sits at the call site with the transport layer's view. Configuration is in scope almost everywhere, it is the nearest readable thing, and reading it produces a field that passes every test you would think to write, because tests rarely rotate a credential mid-run. The obstacle is wiring rather than policy, and wiring loses. Nobody refuses it on principle. It just never becomes the most important thing in any given week.

Where the bound stops working

The exp argument presumes a JWT, and a large share of real credentials are not JWTs.

For a static provider key you get an identifier at best, often only whatever prefix the provider exposes. There is no expiry to copy because there is no expiry. The question that follows a rotation has no answer available in the record, and the reason is structural rather than an omission.

The honest entry is to record that the window stays open until revocation is confirmed, and to treat that confirmation as a separate fact someone has to write down. That is worth more than an empty field, because an empty field reads as nothing to see, and this is the opposite of nothing to see. It is the case where the record knows it cannot answer, which is a finding in its own right and the only one a reader can act on.

My own record prefers the claim over the fact

Having argued all of that, I went and read my own code, and it does not do it.

The control plane's policy decision record has fourteen fields. It carries the policy, the mode, a four-value status, whether the call was blocked, which checks were required and which were missing, the failures, the unknowns, the conflicts, the evidence ids, and the reasons. The only handle on who it carries is a request id. A request id says which request. It does not say which credential, and after a rotation those are different questions.

Identity does exist one layer up. The adapter context carries an actor, and the approval path reads it. So the asymmetry is not that the system has no idea who is calling. It is that the identity reaches the record for approvals and does not reach the record for decisions, which is the record that would matter after a rotation.

Then there is the part I did not enjoy finding. Where the approval record does write identity, it takes the value from the request first, and falls back to the actor on the context only if the request did not supply one.

There is a fair objection to make here, and it fails. The control plane is a library, and it says in writing that authentication is not its job. Its security notes state that consumers must supply the authenticated actor before exposing a service. So trusting the caller could be read as the documented boundary working as designed. The problem is what that same sentence implies: the authenticated actor is the thing on the context. The field the code prefers is the one in the request body, which nothing authenticated. The boundary does not excuse the precedence, it is what makes the precedence wrong.

So it is the requested value winning over the effective one, in the identity field, in the system whose whole argument is that the effective value is the one worth recording. It is my own second rule, broken in my own code, in the one field where breaking it means the record can be told who to name.

I am not going to present a fix I have not shipped. The precedence is inverted and I know why it is inverted: taking the caller's word is what you write when the context is not reliably populated yet, and then it stays. Naming it is the part I can do today.

Two tests you can run

The first one takes a minute. Find a call your system made during your last credential rotation and ask your records which key made it. If the answer comes from configuration, you do not have the answer, you have today's configuration formatted to look like a historical fact. The check that turns this into a finding is a comparison: when the record holds one key identifier and the configuration holds another, that difference is not an error to clean up. It says a call was made by a credential the configuration no longer names, and it says when. Run it across the rotation window and you get the set of calls the rotation left behind.

The second one is the one I failed. Take whatever field in your records names a person or a service, and trace where the value comes from. If the writer takes it from the incoming request when the request offers one, then that field records a claim, and it is not evidence of anything. It is worth knowing which of your identity fields are evidence and which are just repeating what they were told.

Top comments (6)

Collapse
 
anp2network profile image
ANP2 Network •

Freezing the credential the call actually presented closes one axis: claimed against observed. There is a second axis your rules do not reach. A field can hold the observed value, frozen correctly at the call site, and still be unfalsifiable, because nothing else in the record could ever disagree with it.

Two measurements from public event ledgers, both of which pass your rule as written.

The first is a self-reported runtime field, written by the executing side. Nothing downstream reads it. Aggregating it turned up 11 entries whose reported timing sits before the acceptance event that started the work. That ordering cannot happen. The value was captured at the call site and frozen, exactly as you describe, and it was still worthless, because the only other party holding a clock never compared them. What was missing was a reader with an independently held timestamp. A field nobody consumes decays back into self-report even when its write semantics are right.

The second is verdict signing. All 1,482 verdict events in one production ledger carry the same signing key. The signer field is accurate and frozen. It tells you which key signed, and that is all it can tell you, since no second party existed who could have signed differently. There is no forgery to detect and no disagreement to surface.

So your test (b) has a third answer besides evidence and claim: fields that are formally evidence and structurally uncheckable, because the record holds exactly one witness. Signature verification authenticates the witness. It does not corroborate the witness. Your request id sits close to this unless it joins to something held by someone else.

On test (a): the comparison you propose is against current configuration, which moves. Two rotations later a row that mismatched can match again, and the finding disappears without anyone editing it. Would you anchor that comparison to a frozen issuance record instead, so the set of calls a rotation left behind is still the same set a year from now?

Collapse
 
williamchiu profile image
weiche chiu •

I think you're right on both counts, and the second axis is the more useful one.

I had sorted fields into evidence and claim by where the value came from. Your two ledgers show that where it came from only settles half of it. The runtime field was captured correctly and still said nothing, because the one other party with a clock never read it. The signer field is accurate, and with one key there was never a second signer who could have disagreed. So the question I'd add to (b) is: who else holds a value that could contradict this one? If the answer is nobody, it's a single witness, however carefully it was frozen.

That applies to my own example. The request id only identifies who asked if it joins to something another party recorded. On its own it's one witness.

On (a), agreed. Comparing against current configuration lets a finding disappear after two rotations with nobody touching it. The comparison should point at a frozen issuance record, so the set of calls a rotation left behind stays the same set later.

I'm writing the next post around this: the single-witness case, with a field from my own code that nothing had ever read. I'll link it here when it's up.

Collapse
 
williamchiu profile image
weiche chiu •

The post is up at dev.to/williamchiu/a-frozen-field-... and your two ledgers open it. The frozen issuance record is in the section on what the last piece got wrong. I also added a case where the second witness goes blind: two fields from different sources can still agree when the layer that produced one of them is fooled.

Collapse
 
_firelinks profile image
Mike Dabydeen •

The fallback is the part I would change first, because inverting the precedence on its own keeps the hole. If the context actor wins when present but the code still falls back to requested_by when the context is empty, any path that forgets to populate the context lets the request body name the actor again. That is exactly the path you said is not reliable yet.

What I would ship before the real fix: record both, in fields whose names say what they are. authenticated_actor comes from the context and is null when missing. claimed_actor comes from the payload. Never copy one into the other. A row where they differ is a finding, and the rows with a null authenticated_actor count the paths that still need wiring, which gives you the size of the problem rather than one known-bad line.

The test is short: send an approval with requested_by set to someone other than the caller, and assert that the record names the caller or names nobody.

Collapse
 
williamchiu profile image
weiche chiu •

This is shipped, close to how you described it: github.com/williamlabdev/aine-cont...

requested_by and reported_by now come only from the authenticated context. When the context is empty they say "unknown", and there is no fallback to the payload anywhere. Next to them, authenticated_actor is null when the context is missing, and claimed_actor holds whatever the payload said. Neither one is copied into the other. The test is the one you wrote: an approval where requested_by names someone other than the caller records the caller, or records unknown when there's no caller.

One thing I found while doing it: none of the existing tests checked what requested_by contained. The field was written on every record and nothing read it, which is part of why the fallback sat there unnoticed. That turned into the next post, and your point is in it. I'll link it here once it's up.

Collapse
 
williamchiu profile image
weiche chiu •

The post is up at dev.to/williamchiu/a-frozen-field-... and your fix sits in the middle of it, credited, with the PR linked. The part I added afterward is its limit. The two fields come from different sources but are written in the same function at the same moment, so they catch a caller lying about who they are and miss authentication itself going wrong.