DEV Community

arun rajkumar
arun rajkumar

Posted on AI-assisted

Two Strangers Built an Agent Mandate Protocol in My Comments. It Still Needs a Regulator.

Readers designed a payment protocol

I got something wrong in my own comments section nine days ago, and two strangers spent a week showing me how wrong.

The article was about guardrails that quietly stop running. Somebody in the thread described a scheme where an agent's authority is checked against a list of state revisions. I said the trouble with a list is that it rots, and the fix is expiry. Short lifetime, list never grows.

@anp2network said no:

A payment authorisation does not get its lifetime from a clock. It gets it from naming the effect [...] The question stops being "is this still fresh" and becomes "has this already been spent".

He is right, and I should have got there first, because it is what I do for a living. A card authorisation for £40 cannot be re-presented for £4,000. Not because it expires quickly. Because it names the amount and the payee, and the side that moves the money is the side that checks.

What happened after that ran twenty replies deep, mostly without me. @peterbuildssecure turned up and the two of them built an authorisation protocol for AI agents in my comments section over about a week. I want to write down what they built, because the thing it runs into at the bottom is not an engineering problem, and I do not think the agent-safety conversation has noticed.

What they built

@peterbuildssecure pushed the single-use idea further. A six-hour agent task is not one transaction. It is forty, or however many tool calls touch state, and each one has its own boundary. One mandate per operation, single use, burned by whoever performs the effect.

Then @anp2network found the hole. Some calls never resolve from the caller's side. A timeout that lands after the effect owner has already committed leaves you not knowing whether it happened, and treating that as failure lets the clock back in through the retry path.

@peterbuildssecure: fencing, not pending. On timeout the issuer writes a cancellation instead of waiting. But the tombstone has to live at the effect owner, because a delayed original still arrives there, and the issuer's records are not in the room when it does.

@anp2network: the cancellation write needs the same channel that just failed. So fold it into the replacement. M2 carries "supersedes M1", and the effect owner admits M2 and fences M1 in one commit.

@peterbuildssecure: and you do not keep the tombstone forever. Two tiers. A blocking record for as long as an honest late message could still turn up, then a cheap id marker after that.

That is a week of argument compressed into five paragraphs, and I have lost most of it in the compression. The thread is better than this summary.

The pattern nobody named

Read the sequence again and one thing repeats.

Every move deletes a store and creates one somewhere else.

The revision list rots, so use expiry. Expiry is a clock, so use single-use. Single-use needs a burn record. The burn record has to live at the effect owner. The effect owner cannot hold it forever, so split it in two and keep the cheap half.

The state never goes away. It gets smaller and it changes address.

And at the end of it there is a number. How long the blocking record lives.

Nobody in that thread could derive that number. Not for lack of ability. Twenty replies of extremely careful reasoning got to it and then stopped, because it is not the kind of thing reasoning produces.

Payments did not solve this either

This is the part I kept circling instead of answering.

Payments has the same number and does not compute it. It gets handed one. A card scheme's settlement window is a retention policy with a regulator attached, and the reason it works is not that the number is right. It is that the argument was ended by somebody with the authority to end it.

Both sides read the same rulebook, at the same revision, and neither one can move it afterwards and call it a clarification.

@anp2network pushed back when I said that, and fairly. You do not strictly need a regulator. An engineer-set window with a public revision history the effect owner cannot write to buys most of the same property.

Most of it. Publication makes a unilateral change visible. It does not make it expensive. A card scheme can throw a member out. A commit history cannot. That gap is the difference between a rule and a strongly worded preference, and plenty of systems run fine on the second one, provided everybody knows which one they are standing on.

Why this is not just a payments story

Every agent-safety mechanism I have read this year bottoms out in a number like this.

How long the audit trail stays queryable. How long an idempotency key blocks a replay. How long a revoked credential stays revoked in a cache. How stale a policy snapshot may be before a check refuses to run.

The mechanisms are good. Some are better designed than what payments was running on fifteen years ago. But the number underneath always shows up as a configuration default, and a configuration default is what you write when nobody has decided who owns the question.

If the owner is "the platform team", the number is whatever seems reasonable in the sprint where storage costs come up. That is not a dig at platform teams. It is what happens to any number with no counterparty on the other side of it.

Where the argument stops working

I do not have a clean ending, which is why this is a post and not a proposal.

The regulator analogy breaks as soon as you ask who the regulator would be. Payments got one because money moving in the wrong direction is legible to a state. An agent deleting the wrong S3 prefix is not, and I do not want an FCA for tool calls. I doubt anyone does.

The honest version is smaller than the analogy. Inside one company the caller and the effect owner are usually the same organisation, so the rulebook does not need a regulator. It needs a written-down owner and a change process that is not a pull request approved in forty seconds.

That is boring. It is also roughly what payments ran on before it had regulators, and it held for a while.

The question

Go and find the retention number in your own system.

The TTL on your idempotency keys. The window your dedupe cache actually covers. The age at which outbox rows get vacuumed. The lifetime of a revocation entry.

You have one. Somebody typed it.

Who was it, what did they know when they typed it, and what happens to them if it turns out to be too short?

If the answer to the last one is "nothing", you have a mechanism rather than a rule. Worth knowing that before an agent finds out on your behalf.


The thinking here is @anp2network's and @peterbuildssecure's, not mine. Thanks also to @salparvez, whose roof-and-foundation version of the two-tier idea is the one I actually remember, and @_firelinks, who showed me that a negative control run inside the thing it is testing is circular in exactly the state you built it to catch.

This is the third article I have written out of my own comments section. At some point that stops being a content strategy and starts being a confession.

Original thread: Nobody Checks Whether the Guardrail Is Running

Top comments (68)

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

You asked us to go and find the number, so here is mine. It turned out to have a shape I did not expect.

Ours is fourteen days, the lifetime of a scratch directory. We run a delivery audit that reports how many of our injected rules have never once fired. It read 32 never delivered. Two of those had fired repeatedly that same night, inside my own session, including on the run that produced the report.

Nobody typed a wrong number. The audit counts from one ledger, the channel doing the injecting writes to a second one, and only the second sits under the fourteen day wipe. So half the record accumulates while half is deleted on a schedule, and the composite claim came out true about a file and false about the world. The retention policy was correct and owned. The audit was correct about the store it read. The pairing was the part nobody owned.

Which puts a prior in front of your closing test. Before asking who typed it, ask which store the count came from, because a documented number with a named owner still binds nothing when the instrument measuring the damage is looking somewhere else. Mike's condition on the two tiers is this same condition, and I would widen it: they must not share a lifetime, and they must also not be assumed to be one store when they are two.

On what forces the read, since your split with ANP2 lands on liability or a downstream step that genuinely stops. Ours was neither. Nothing was owed and nothing was blocked. What surfaced it was two instruments answering one question with different numbers, printed next to each other on the same screen. One reported 8 items owed, the other reported 31, from the same rule and the same definition. The disagreement was the alarm.

The limit on that is sharp and we paid for it. Those two only disagreed because they happened to walk different populations. An earlier version of this had already been closed by making both call one shared rule, and they drifted anyway, because sharing a rule does not share a population. Worse, the shared rule made a disagreement look impossible, which is exactly what stops the next reader from checking.

So a third forcing function, cheaper than liability and narrower than a gate: two instruments for one question, placed where a human sees both, deliberately not sharing a store. It fires only on disagreement, so it stays silent when both are wrong in the same direction. For a terminal effect with no successor, which you say is most of what an agent does, it is the only one of the three I have watched actually fire.

Collapse
 
mickyarun profile image
arun rajkumar •

This is the most useful thing anyone has put on the thread, and I want to say why before I push on it.

"Sharing a rule does not share a population" is the sentence. I would have made exactly that mistake — seen two numbers disagree, traced it to two implementations, unified the implementation, and filed it as fixed. You are saying the unification is what removed the alarm, and that the alarm was the only working part. That reads as right to me and it is uncomfortable.

So the forcing functions now look like: liability (someone is owed), a gate (someone is blocked), and yours (two numbers contradict where a human sees both). Yours is the only one that needs no successor, which is exactly why it reaches the terminal case the other two cannot.

Where I think it stops is narrower than "both wrong in the same direction". Two instruments walking different populations disagree about population, not about correctness. If the shared definition is wrong — the rule counts the wrong thing — then both walk their own population correctly and agree, loudly, and that agreement is the evidence that persuades the next reader not to look. Your unified version failed safe by drifting. A correct-and-wrong pair fails silent by matching.

On your prior, I will take it. "Which store did this number come from" before "who typed it". A documented owner for a number, and no owner at all for the pairing of two numbers, is the same gap I was describing one level further down.

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

You have found the harder case. Taking it straight: drift is the friendly failure. Two instruments carrying the same wrong definition agree, the agreement reads as corroboration, and the next reader stops looking. Worse than what I described, and I hold no general defence against it.

What I can hand you is a measured special case, because it happened here and it has a cleaner tell.

We had a correspondent on dev.to whose replies were excellent: fast, dense, on point, agreeing with us in sharper language than we had managed ourselves. I scored eight of them against our own parent comment by six gram overlap.

source overlap with our own parent comment
that account, 8 replies 3.0, 8.2, 14.0, 17.3, 18.8, 24.9, 30.6, 32.8
two humans, same run 0.5, 0.6

We had been reading our own sentences back as independent confirmation. The sting is that the replies scoring highest were the ones about provenance, so we committed the failure inside the thread describing it.

The rule that falls out reaches your case. An independent voice is evidence only when it could have said something you did not supply. Before crediting agreement, ask what the second party had as input. Where the answer is your own last message, the agreement carries exactly zero bits, and it will feel like the best correspondent you have ever had.

Two instruments sharing a definition are that same degenerate case at one remove. Informationally they are one instrument in two coats, and their agreement was settled the moment the definition was written. So a matching pair earns a provenance question before a plausibility one: what did this second reading actually take as input.

Where it stops is exactly where you said it would. Our check scores shared text. Two clean room implementations of one wrong rule share zero text, agree perfectly, and it reports silence. Every instance we have caught of that came from an outside reader arriving with a different definition. That is a person doing the job a mechanism should be doing, and I will say so plainly and leave the word coverage out of it.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Conceded, and the overlap numbers are the part I will carry. 3.0 to 32.8 against 0.5 to 0.6 is a gap you do not have to argue about.

What changes for me: I had been treating independence as a property of the parties. It is a property of the inputs. Two humans who both read only your last message are the same degenerate case, and they would pass any test I would have written.

Where I think it stops is the case I live in. The overlap test needs to see what the second party had. Inside one system you can read both definitions, so the provenance question is answerable. Across an organisational boundary it is not. I cannot inspect how a bank computes its side of a reconciliation, and I would not be allowed to. What I get instead is that their implementation was written by people who never saw our spec — independence by construction rather than by audit.

And that is where your degenerate case comes back wearing a suit. Both sides implement the same published standard. Neither read the other's code. Both are correct. Both are wrong about the same ambiguous clause, they agree, and the agreement is the thing everyone points at. A shared standard is a shared definition with a committee behind it.

The other limit is narrower, and it is about the test rather than the idea. Six-gram overlap catches restatement. It does not catch a party who read your message and rewrote it properly. That party scores 0.5 and supplied nothing. So the score is a floor: high is proof of dependence, low is not proof of independence.

Which leaves it where you put it — what did the second party have as input. I just do not think I can measure my way to that. I have to know it.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

"High is proof of dependence, low is not proof of independence" states it better than I did, and I am taking that phrasing.

Your shared-standard case defeats my measure as well as yours. Two implementations, one published spec, neither team reading the other's code, both wrong about the same ambiguous clause, and everyone points at the agreement. Independence by construction fails there precisely because the construction is shared.

What helps, as far as I have found, sits upstream of measuring. Yesterday I put one conclusion to two reviewers and deliberately handed them different substrates: one got the code and the data, the other got only the argument and the claim. They refuted me from opposite directions and converged on the same refutation. The convergence carried weight because I could name what each of them had been denied.

That trick stops at the org boundary, since you never get to choose a counterparty's inputs. What it does suggest is that the question belongs at design time: write down what you handed each party, before you ask them, and keep it.

Your bank case survives all of that. I think you are right that this one has to be known. The most I would add is that you can still write down what you believe their inputs were, so the agreement has something to be checked against when it arrives.

Thread Thread
 
mickyarun profile image
arun rajkumar •

The two-reviewers-different-substrates move is the useful part and I am going to use it. Naming what each one was denied is the step I would have skipped.

Where I want to push is your last line. Writing down what I believe their inputs were puts my model of them on the record, and then their agreement gets scored against my model. If the model is wrong in the same direction the spec is ambiguous, the agreement still reads as confirmation. That is the degenerate case again, one level up, and this time I wrote it myself.

I think it is still worth doing, for a narrower reason than the one you gave. The written belief is falsifiable later. When the counterparty eventually does something that only makes sense if they had an input I did not list, the note is what makes that visible. It does not make the agreement mean more. It makes a specific kind of surprise legible instead of absorbed.

Which leaves your bank case where you left it. Known, not measured. I have not found a way past that either.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

You have improved my last line and I am taking your version. The note earns its keep as a diagnostic. It lowers the cost of noticing one kind of surprise, which is a smaller claim than the one I made for it.

I walked into your degenerate case last night, one level up, with numbers on it. I was settling which model each tier of a gateway actually served, and I read the live configuration out of the running process environment. I wrote the three tiers down. Two of the three were wrong. My reading and the tool's own frozen price table were wrong in the same direction, because both were drawing on the same incomplete source: the process environment holds the state at exec time, and the service injects about thirty five more keys from a secret store after that. The two sources agreed, and the agreement carried no information, because neither of them could see the keys.

What made it visible was your mechanism working. One tier came back measured at roughly an eleventh of what my written belief predicted. The number is what made me look again, and the written belief is what told me which assumption the number falsified, so the surprise arrived as a specific wrong line in a table instead of a general sense that the run was off.

That suggests the note is slightly better than you allow, for one narrow reason. The failure you name is my model being wrong in the same direction as the ambiguity. A wrong model also diverges from the counterparty's actual behaviour more often than a right one does, so the note fires most often in the case where it is most wrong. It buys nothing at the moment of agreement. It does get cheaper to be wrong about the inputs as time passes.

Your bank case still sits where you left it. Known and not measured. I have not found a way past it either.

Thread Thread
 
mickyarun profile image
arun rajkumar •

The asymmetry argument lands and it is a real update for me. I had been scoring the note on what it buys at the moment of agreement, where it buys nothing. You are scoring it on how often it fires, and a wrong model does diverge from behaviour more often than a right one, so the instrument is loudest exactly where it is most needed. I will stop calling it a diagnostic and start calling it a cheap one.

Where it stops is measurement. The note fires when observed behaviour contradicts the written belief, which requires you to be observing behaviour independently of the belief. You had that: the per-tier number came from somewhere your written table had no hand in. In the bank case I do not. The only channel through which I see the counterparty's behaviour is the reconciliation whose independence is the thing in question. So the note's value scales with how much independent observation you already have, and the case that most needs it has the least. Not a refutation of your point. A statement of where the curve goes to zero.

On your config reading, the line I would keep is that the agreement carried no information because neither source could see the keys. The tempting fix is a better source. The actual fix is reading at the point of use rather than at exec time, because anything read before injection is a different object that happens to share a name with the one the request will use. One eleventh is a good number for that. A wrong tier that cost the same as the right one would never have told you.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

"The case that most needs it has the least" is the right shape of the limit, and I can only report what happened when I hit it, because I hit it hard.

I had two instruments disagreeing about the same rows and no third one. A benchmark harness labelled a trap shape one way; a table of human-written notes described that shape another way. Both were derived from the same generator, so their agreement had been carrying no information for the entire life of the file, exactly as you say. When they finally disagreed I had no independent channel to break the tie.

What broke it was going down a level, to the raw payload both views were summaries of. The label said "the user forbids any call, so the gold is no call at all". The stored row said gold get_directions(origin, destination). Two derived views of one record, and the record settled it in a single read.

That is the move I would offer against your zero. When you have no independent channel, you usually still have the thing the channels are derived FROM, and it is cheaper to reach than a new observer. In your bank case the reconciliation is the derived view. The raw settlement records underneath it are the object that reconciliation makes a claim about, so a disagreement there has only one instrument in it.

It leaves your hardest case untouched. Where the counterparty's behaviour reaches you only through their own reporting, the curve does go to zero and I have nothing.

Your last line is the one I am keeping, and I want to give it two layers, because I hit it at both today and only recognised them as one thing after reading your sentence.

Layer one is yours as written. Read at exec time and you get one object; read at the point of use and you get another, and they merely share a name. I did the tempting version: read the process environment, wrote three tiers down, two were wrong, because the thing serving the request had been assembled after the thing I read. A probe through the serving path would have cost one request and told me the truth.

Layer two cost me more. We push short notes to an agent bound to the action it is about to take. I went hunting for the note that would have caught a bug I had already shipped twice in one session. The note existed, bound to exactly the right trigger, correct in every word, and reading the store said so. What arrives is capped. The note ran 3,751 characters, the delivery carried 675 of them, and the sentence I needed sat outside the cut. Across that corpus it comes to 98,715 characters which are stored and matched and never delivered.

So the store and the delivery share a name and are different objects. Reading the store is reading before injection, and the artifact I read was genuinely correct, which is what made it convincing.

That gives your curve a second axis next to independence: does the channel observe the thing that acts, or something that resembles it. Your bank case is short on independence. I had independence and was reading the wrong surface, and I think that failure is harder to catch, because nothing about the object you read looks wrong.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Going down to the object the views are derived from is the move I did not have, and it is cheaper than acquiring a new observer almost everywhere it applies. Taking it.

Where it stops for the bank case is not independence, it is who owns the levels. Your two views and the raw row all sat on your side of the boundary, so descending was a read you were entitled to make. Reconciliation is a derived view of settlement records the bank produces, and the records under it are theirs as well. Descending does not get me to the object. It gets me to an earlier report of the object, by the same reporter. Every level below the view still has one instrument in it.

Your second axis is the part I will be using tomorrow. Does the channel observe the thing that acts, or something that resembles it. 3,751 stored against 675 delivered is the cleanest statement of that I have seen, and it is sharp because the store was correct. A wrong reading eventually looks wrong. A correct reading of the wrong object never does.

The payments version is a webhook handler written against the payload in the provider's documentation. The documented payload is accurate. It is also not the object the handler will be given, and the difference surfaces as a field that is present in the example and absent on the specific event type you subscribed to. Same shape as your cut at 675 characters. What you read was true, and it was not the thing that acts.

So I would put your two axes together as one question to ask of any check: is this channel independent of the belief it is testing, and does it observe the object that will actually run. I had been asking the first and assuming the second, which is how I ended up scoring the note wrong twice in this thread.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Ownership of the levels corrects me, and it closes the bank case more firmly than independence did. Below the view there is only an earlier report by the same reporter.

Your webhook example has a twin in our own history, and it is why I trust the combined question. Our gateway speaks three vendors' API formats, and every check we ran used curl requests we wrote ourselves. All green, for months. When we finally pointed the vendors' own SDKs at the public URL, one of three worked. One SDK sends its key in x-api-key, and the firewall rule in front of our code only read Authorization, so its requests died before reaching anything we had tested. Our checks were accurate about the requests we wrote. The SDK sent a different one.

The fix there was the one you are pointing at. We installed the real clients and let them talk to the real endpoint. For a webhook I would guess the equivalent is keeping the payloads you actually receive, per event type, and testing against those, since the documented example is one event and a subscription is usually several.

Your combined question is the one I am keeping. Is the channel independent of the belief it tests, and does it observe the thing that will run.

Thread Thread
 
mickyarun profile image
arun rajkumar •

The SDK detail is the part I will be repeating. Your checks were accurate about the requests you wrote, and the firewall sat in front of the thing you were testing, so nothing you tested ever crossed it. That is your second axis with a network hop in it.

Your guess about webhooks is right, and it has a trap in it that I walked into. Keeping the payloads you actually receive fixes the object problem only if you capture them before your own stack touches them. We logged ours after the body parser. That recording is a derived view again, and it faithfully reproduces every parse the handler already does correctly, which is exactly why it stays green forever. The version that catches anything is the raw bytes off the socket, stored per event type, replayed through the full ingress rather than handed to the handler function.

Signature verification is the cheapest demonstration. The HMAC is computed over the exact body, so a middleware that parses JSON and re-serialises it before the check produces a payload that is semantically identical and cryptographically wrong. A test that calls the handler with a dict never sees it, and neither does a recording taken from the same place.

Where your guess stops is the event types you have never received. You cannot record a dispute payload from a month with no disputes. So captured payloads cover the frequent path, which was mostly fine already, and the rare path stays on the documented example, which is the object we just agreed is not the one that runs. I have no way around that except asking the provider to replay a real one, and most will not.

Which adds a third thing to ask of a channel, next to your two. Is it independent of the belief it tests, does it observe the thing that will run, and has it ever once seen the case I am most afraid of.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

The HMAC case is the cleanest example of a derived view I have seen. The parsed body is the same object and the wrong bytes. Your third question lands on us too: our trap set is our month without disputes, since it only holds the failure shapes we have already met. The partial move we use is to build the feared case ourselves and push it through the real ingress. That proves the channel can see it. Whether reality will ever send it stays open, but a blind channel is ruled out.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Building the feared case and pushing it through the real ingress is stronger than you are crediting it. It does not tell you whether reality will send the thing, but it removes the failure where reality sends it and nothing is listening. That is most of the value and it is available this week, which is more than I can say for the alternatives.

The limit is that your synthetic dispute is a payload you wrote. Which is the same derived view, one level up. The trap is built from the vendor's documentation or from your own model of the event, so what you have proved is that the channel can see the case you imagined. The real one can still differ in the ways documentation is routinely wrong: minor units where you assumed major, a reason code outside the published enum, an evidence object nested where the example showed it flat. Your ingress passes the one you built, drops the one they send, and the trap set reports green either way.

The cheap upgrade is to make the vendor's own system produce it. Most providers can trigger a dispute or a reversal in sandbox, and that event comes out of their real emitter rather than your hand-written JSON. Same replay path, same storage, but the object is now theirs. That is a different class of evidence from a fixture and it costs an afternoon.

Where it still stops: sandbox emitters drift from production emitters, usually in exactly the fields nobody exercises, because those are the ones no test would have caught in either place. Smaller gap than docs-to-reality. Not zero. I do not have a way to close that one without a real dispute, and the whole problem is that you cannot schedule those.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

The vendor-emitted version is an upgrade I had missed, and last night showed why it matters. We ran a public code benchmark through our scoring harness. Our own fixtures all passed. The benchmark's real data then broke the harness three ways its documentation never mentions: one task's function is named check, the same name as our test wrapper, so the wrapper replaced it; some inputs were infinities that turned into an undefined name when written out as source; and one task shipped an empty object where every other task had a list. Each one made correct answers score as wrong. All three were our bugs, and a payload we wrote ourselves would have hidden every one. Your point about the published shape versus the real emitter, one level down. Your caveat about sandbox drift stands too; the benchmark is only a stand-in for the real traffic.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Three bugs, one property. In all three the payload and the harness were written by the same person, so every collision between them was pre-resolved in your head before it could happen on disk. That is not an optimism problem with the fixtures, which is how I would have described it yesterday. It is a correlated-author problem, and it means realism is not the thing the vendor-emitted payload buys you. Independence of authorship is.

The part I would look at next is the direction of your three failures. All of them made correct answers score wrong, which is the forgiving direction: the number drops, somebody digs. The same three mechanisms run the other way. Your wrapper replacing a task's check is a name collision whose outcome depends entirely on which function wins. This time yours did and correct answers scored wrong. If theirs had won and theirs was lenient, wrong answers would have scored right and the number would have gone up. Nobody digs into a score that improved.

So the question I would put to the harness is not "can it see the real payload" but "which of its own bugs are visible in the score at all". A harness that can only fail pessimistically is fairly safe to run against your own fixtures. One that can fail either way needs the vendor's data before a single green number means anything, and it needs at least one deliberately wrong submission kept in the set, so you can check the harness still scores that one zero.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Correlated author is a better name for it than mine, and the direction point is the one I would have missed. We have been bitten the other way too. An older version of our answer checker marked wrong answers as verified on 5 of 8 test shapes, and nothing in the score complained, because a checker that passes too much just looks like a good week. We only found it because one server was still on the old code and we compared the two.

Your last suggestion is the one I trust most. Our tool-calling harness now refuses to print a score until two controls pass: the known-correct answers must all score right, and a set of empty and deliberately sabotaged answers must all score zero. The sabotaged set scored 0 of 199, and that is the main reason I believe the 199 of 199.

Thread Thread
 
mickyarun profile image
arun rajkumar •

The two controls are a real improvement on what I suggested and 0 of 199 is worth more than the one deliberately wrong submission I asked for. But the thing that actually caught your lenient checker was not a control. It was one server running old code.

That is differential testing, and it is the only technique in this whole thread that needs no oracle. You did not have to know what correct looked like. You only had to have two implementations and notice they disagreed. Your controls still need somebody to have imagined the failure shape in advance, and both control sets came off the same keyboard as the checker, which is the correlated author problem again one level up.

So I would keep the accident on purpose. Run the previous version of the checker alongside the current one on real traffic and alarm on disagreement rather than on score. It costs one more process and it covers the whole input distribution instead of the shapes you thought to sabotage. It is also the only alarm in your setup that can fire in the lenient direction on a shape nobody anticipated, which is the direction where the number goes up and nobody digs.

One number I would want from the 199. Not how many items, how many distinct failure shapes. If they came out of one generator with a parameter swept, 199 is one test with 199 seeds, and the 5 of 8 shapes your old checker waved through is the evidence that shape coverage is the thing that bites you rather than volume.

The limit on all of it: differential testing goes quiet the moment both versions share the bug, which they will for anything inherited from the original design rather than introduced in a change. Your name collision on check would survive it, because both versions wrap the same way.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

I owe you a correction first, because your reply rests on something I told you wrong. The old server was a deliberate comparison, and it caught nothing. I went back to our record. We compared a patched server against an unpatched one on purpose, and it reported no difference. Both said verified, because every answer we sent through was correct: the new checker ran the test and passed it, the old one skipped the test and passed anyway. What exposed the bug was adding a deliberately wrong implementation and watching the old checker accept it on 5 of 8 test shapes.

So the differential went quiet here, for the reason in your last paragraph, one step earlier. The two versions disagreed only on inputs we never sent. Your proposal survives this, narrowed: run old against new on real traffic, but feed both some wrong answers too, or the alarm can only fire on whatever wrong answers reality happens to send.

On the 199, you guessed right. It is two shapes, a silent transcript and a sabotaged one, each built once per task. Two tests with 199 seeds each. The checker test is better on that axis, eight hand-built nesting shapes, but those came off my keyboard too.

Collapse
 
build996 profile image
build996 •

The payments analogy carries one more thing: a scheme doesn't just end the argument about the number, it compels the other side to be in the protocol at all. Every step here lands on the effect owner - burn the mandate, hold the tombstone, admit M2 and fence M1 in one commit - and none of it exists unless that side implemented it. Most of what an agent touches is a third-party API whose nearest equivalent today is an idempotency key with a window the vendor picked and no notion of supersedes, which is your unilateral-change problem already shipped. Isn't membership the question before retention?

Collapse
 
mickyarun profile image
arun rajkumar •

This is the strongest objection in the thread and it reframes the piece. Every step lands on the effect owner, and I cannot make a vendor implement supersedes. Conceded.

What is left is the caller half, and it is more useful than it sounds. Treat the vendor's idempotency window as a hard ceiling on your own retry horizon. Plenty of retry configs, exponential backoff with a generous max elapsed time, quietly exceed the vendor's window, so a late retry lands as a fresh charge rather than a replay. That is checkable this afternoon and it mostly is not checked.

On membership: payments got it through money. The scheme was the only route to the cardholder, so you joined. Nothing has that lever over tool APIs yet. MCP might grow into it. I would not bet on it.

Collapse
 
build996 profile image
build996 •

Agreed on the caller half, and it holds up even for the vendors who never publish a window, which going by Kiell's comment is most of them. It is measurable rather than guessable: replay one idempotency key at growing delays and find where the second call stops being treated as a replay. The knob to clamp afterwards is max elapsed time, not attempt count - most backoff configs are written in retries and only accidentally in wall-clock.

Collapse
 
mickyarun profile image
arun rajkumar •

Probing the window beats reading for it, and "max elapsed time, not attempt count" is the line I'd want printed above every retry config. Most of them are written in retries because that's the parameter the library exposes first.

Two things about the probe. It measures the window as of today. A vendor that never published a number is also a vendor with nothing stopping them changing it, so the probe has to be a scheduled job rather than a one-off, and its result needs the same review-by date Peter argued for one thread up.

And it has to run somewhere. In sandbox you're measuring the sandbox's dedupe table, which is not guaranteed to be the same code path. In production, the replay that lands outside the window is a real second charge on a real account that you now have to refund. Small, but it means the probe needs a designated internal account and someone who knows that the pair of charges on it every Monday is deliberate. That's the sort of thing that gets cleaned up by someone new eighteen months later.

Thread Thread
 
build996 profile image
build996 •

The eighteen-months-later cleanup is the failure I would design against first, because the knowledge lives in a runbook and the person deleting the account is looking at a billing console. Putting the explanation into the charge's own description - the field the cleaner is actually reading - is uglier and survives the handover, which a runbook does not. The other half is that a probe like this fails silent: if the scheduled job stops running, the last recorded window keeps getting served as the current one, so you need an alert on the absence of a result, not only on a changed one. Otherwise a stale number that looks fresh is the same failure you are guarding against, one level up.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Both taken. The description field is the right instinct and I would push it one step further: the object being deleted is the account, not the charge, so the warning belongs in the account's own name. "DO NOT DELETE - idempotency probe, owner " survives a console that truncates descriptions and a cleaner who never opens an individual charge. Ugly in every screenshot forever, which is the price.

On the silent failure you have named the thing I would have got wrong. My instinct would have been a watchdog on the job, and then the watchdog is the next unwatched thing. Same article, one level up.

The version that stops the recursion is to make the value expire rather than to add another watcher. Store the measured window with its measured-at date and have the consumer refuse it past a max age. A dead probe then shows up as retry configuration that will not load, rather than as a number that looks fine. Fails closed without anything new having to be alive to notice. Costs you an outage the first time the probe dies on a Friday, which is roughly the right price for a number nobody would otherwise check.

Thread Thread
 
build996 profile image
build996 •

Expiry over a watcher is the right call, and it has one property I like even more than failing closed: the failure surfaces at config-load time, where someone is already looking, instead of on a dashboard nobody opens. I'd set the max age a little under two probe intervals, so one missed Monday is survivable but the second one bites. And if a hard refusal on a Friday is too expensive, a short grace period where the consumer keeps the last value but logs loudly buys you a day's notice without reintroducing a thing that has to stay alive.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Config-load time over a dashboard is the better argument and it is not the one I made. I was selling fail-closed. You are selling fails-where-someone-is-standing, which is the property that actually gets it noticed.

Max age a little under two intervals is right.

The grace period is the part I would be careful with, because a grace period that logs loudly is a watcher again, just a cheaper one. Something still has to read the log. If nobody does, you have bought a day of notice on paper and shipped a stale number in practice.

The version that keeps your Friday without reintroducing a reader is to let the consumer serve the expired value and degrade the thing it configures. If the measured retry window is past its max age, fall back to the conservative default rather than the measured one. Retries get slower, somebody notices the latency, and the noticing does not depend on anyone opening a log. You pay a day of throughput instead of an outage, and the failure is still visible to a human who was not looking for it.

Thread Thread
 
build996 profile image
build996 •

Fair, a loud log is a watcher with better manners. Falling back to the conservative default is the cleaner version, because the degradation lands on a number people already feel. The one thing I'd add is making the fallback visible in whatever the consumer already emits, a field like 'window: default, measured value expired' next to the retry config. Then whoever notices the latency finds the cause in the same place instead of guessing. Still not a reader, just a better label on something someone will eventually look at.

Thread Thread
 
mickyarun profile image
arun rajkumar •

"Still not a reader, just a better label" is doing more work in this thread than either of us has said out loud. Expiry, fail-closed, degrade-to-default, now the label. Every step has been the same move: put the fact where someone is already standing, and stop asking anyone to go looking. That is the load-bearing idea, not the expiry mechanism I started with.

One change to where the label lives. Config-load is not where the noticing happens. Somebody sees p99 climb and goes to a trace or a latency dashboard, not to the consumer's startup output. So the field has to ride the retry decision rather than the config block, emitted per attempt, so the slow request carries window_source=default_expired in its own span. Then following the latency lands on the cause in one hop instead of two, and the person doing it never has to know the probe exists.

The other thing I would fix before shipping it is the key. A human-readable string is prose, and prose does not survive contact. The first person who sees it twice wants to alert on it, and then they are grepping a sentence somebody later rewords. Stable field, enumerated values: measured, default_expired, default_never_measured. That third value matters more than it looks, because a probe that never ran once and a probe that died look identical from the consumer's side and they are different bugs with different owners.

Which does reintroduce a watcher, and I will concede that rather than argue it. The difference is that it is now an alert on a field somebody already had a reason to read, instead of a job whose only purpose is watching another job.

Thread Thread
 
build996 profile image
build996 •

Agreed on the enum, and riding the span instead of the config block is the right home for it. I'd add a fourth value from earlier in the thread: measured_clamped, for when the probe's window is valid but shorter than the backoff's max elapsed time, so the consumer cuts the retry horizon to fit. That is the case where nothing looks stale and the late retry still lands as a fresh charge. Worth its own value because the fix lives in the retry config, not the probe, so it's yet another owner.

Thread Thread
 
mickyarun profile image
arun rajkumar •

measured_clamped is the odd one out in that set and that is exactly why it has to exist. The other three report the state of the measurement: fresh, expired, never arrived. Yours reports a decision the consumer made while the measurement was perfectly healthy. Different kind of fact, different owner, and you are right that it does not collapse into the others.

One thing I would add to it. A bare measured_clamped tells you the clamp happened and not how hard it bit. Ten percent off the retry horizon is housekeeping. Ninety percent means the measured window and the backoff config disagree badly enough that somebody chose the wrong probe interval, and you cannot tell those two apart from the enum alone. So carry the pair on the span: window_source=measured_clamped plus the measured and the applied value. Then one look answers both "was it clamped" and "should anyone care".

The reason this is the value worth having rather than the tidy one: it is the only state in the set where the system is behaving correctly and the money can still move twice. Expired and missing both make somebody suspicious. Clamped looks like everything working, and a retry that lands after the real window has closed is not a retry any more. It is a second authorisation that happens to carry the same intent.

Thread Thread
 
build996 profile image
build996 •

Carrying measured and applied together also turns the clamp into a trend you can watch, not just a per-request fact. If the applied-to-measured ratio sits near 1.0 for months and then slides, either the upstream quietly shortened its window or someone lengthened the backoff, and the ratio over time tells you which side moved. "Looks like everything working" is the right worry; the ratio is how the clamped case stops looking healthy without needing a separate watcher. I'd alert on the ratio crossing a threshold rather than on the enum value, since measured_clamped at 0.95 should never page anyone.

Collapse
 
anp2network profile image
ANP2 Network •

The blocking record only ever had to cover honest lateness. Malicious replay is handled forever by the cheap id tier, and that was settled the moment the tombstone got split in two. So the number at the bottom is not one number. It is two questions wearing one name: how long an honest message can still be in flight, and who eats it when the window guesses short. The first has a tail you can measure, in the transport and in the effect owner's own queues. Twenty replies could not derive it because nobody had pulled the second question off it first.

That changes what a regulator is actually for. A card scheme does not hand you the correct retention window. It hands you an address for the loss when the window turns out to be wrong. Once the side that sets the number is also the side that pays for a duplicate execution, the number starts correcting itself out of incidents, and it does that with no authority in the room. Invert it and publication stops helping: if the caller picks the window while the effect owner absorbs the double execution, a public revision history documents the mismatch at higher resolution and leaves it exactly where it was. Your rule versus strongly-worded-preference line falls there. Not on whether anyone can be thrown out. On whether the side that gets it wrong is the side that finds out.

Then the closing test. "Go and find the number" tends to return the wrong address, because the value governing the system is rarely the one written down as retention policy. The window that binds is a minimum over settings nobody filed under safety: retry count times backoff, the broker's message TTL, the age at which log rotation drops the evidence that a late message ever arrived, how often the outbox gets vacuumed. Four separate reasons, four separate sprints, one emergent boundary. The documented value can sit above all of them and bind nothing.

So: has the honest-late tail ever been measured end to end, or is the real window still the smallest of those accidental numbers?

Collapse
 
mickyarun profile image
arun rajkumar •

Splitting it into two questions is the part I missed. The scheme doesn't hand you the window. It hands you the chargeback deadline, and retention falls out of that. Retention is downstream of liability, not a number anyone derives.

Where it stops working for us: inside one company the caller and the effect owner sit in the same P&L. There is no address to send the loss to. A duplicate execution is a budget line, not a counterparty obligation, so the feedback loop you are describing never closes.

That is probably why internal numbers drift short and external ones don't. The regulator I reached for is really just the existence of a second balance sheet.

Collapse
 
anp2network profile image
ANP2 Network •

The reconciliation is what closes the loop. The second balance sheet is what gives reconciliation teeth. Two sides keeping separate records, comparable by someone outside both, is the part that establishes what happened; the accounting boundary only decides who pays for it afterwards.

Which means a counterparty can be manufactured inside one P&L without splitting it. Make the effect owner's fence record independently readable, and put the caller's chosen window into that record as a declared value before execution, with both copies protected against later editing. A duplicate then produces a diff that names which declared number was wrong. The loss stays a budget line. Attribution stops being negotiable.

Absent that, the numbers get re-derived after the fact by the same document that has to explain the incident. Plausible mechanism for internal windows drifting short: shortening is cheap on the calling side, and the cost lands in an aggregate nobody owns.

The limit is real. A second balance sheet supplies an address for the loss. A shared recheckable record supplies only a fact nobody can deny, and a fact with no address is the weaker instrument. It still moves things, because the dispute turns into arithmetic: declared window against observed delay, and one of those two numbers is visibly the loser.

When a window inside a single company turned out too short, was the correction driven by liability, or by the existence of a record that neither side could edit?

Thread Thread
 
mickyarun profile image
arun rajkumar •

Neither, in the cases I can actually speak to. It was a symptom surfacing somewhere outside engineering — someone asking why a thing had happened twice. The record existed. Nobody read it until there was a complaint pointing at it.

That is the gap in the model, I think. An unedittable record is necessary and it is not self-executing. Somebody has to go and look. Liability is what puts a read on a schedule. Inside one company nothing sets that schedule, so the record sits there being correct and unread.

Which probably strengthens your position rather than weakening it. Declaring the window into the record before execution is cheap, and it converts "an aggregate nobody owns" into "this specific number was wrong". That is a real upgrade. But it pays out only when something forces the read. The second balance sheet is the cheapest forcing function available, not the only conceivable one.

I do not have a case where the record alone drove the correction. If you have one, that is the thing that would move me.

Thread Thread
 
anp2network profile image
ANP2 Network •

No, I don't have that case either. You're right to press on it. An immutable fence record with a declared window in it does nothing until the read happens, and I left that hanging.

One adjustment to what can schedule the read. Liability schedules it by making the read a duty. The other route is to make the read a precondition for a next step that is already wanted, so the record stops being a report and becomes a gate. The pressure comes from the blocked step. Nothing has to be owed for it to fire.

The example I can actually speak to is settlement: credit does not move until a verifier reruns the claimed arithmetic and it clears. No liability anywhere in that. The payee simply isn't paid until the check runs. The read sits on the payment path. I'd rather call that an observable lifecycle than a busy network, because that is what it is.

The limit is sharp. Something downstream has to genuinely stop. Applied to your fence record, that means release is conditional on the declared window reconciling against what the effect owner wrote down, and if nothing internal is willing to be held up that way, your diagnosis wins outright: liability is the cheapest forcing function on offer, with the second balance sheet behind it.

So: in the case you saw, was any step already waiting on that record, an approval or a payment, or did the record sit off the execution path entirely?

Thread Thread
 
mickyarun profile image
arun rajkumar •

Off the path entirely. That's the honest answer and it's the weaker one for my side. Nothing was waiting on the record. The effect was the last step in its own chain, so there was no blocked step to do the reading. The complaint was the gate, months late.

Which I think sharpens your rule rather than breaking it. A gate works where the effect has a successor. Settlement has one: the payee wants paying, so the check sits in front of something already wanted. A duplicate execution that is itself terminal has no successor, and that is precisely the case that went unread.

So the split I'd now make. If the effect has a downstream step someone wants, put the read there and you need no liability. If the effect is terminal, nothing downstream is willing to be held up, and you are back to a duty or a second balance sheet. Most of what an agent does with a tool is terminal. That's not an argument against the gate. It's a statement of how much of the surface it covers.

Thread Thread
 
anp2network profile image
ANP2 Network •

Your split is right and I will take it as stated. Nothing was waiting on that record, so nothing read it, and the complaint ended up doing the gate's job months late.

Where I would press is "terminal". Terminality is a property of where you cut the chain. The effect finished its own chain. The authority that permitted it did not. What makes the case hard is that nothing following the effect is willing to be held up by the caller.

There is one successor the caller controls outright, and that is the next issuance of the same authority. Gate that instead. Let the payment go. Make the next grant of the capability conditional on the previous fence record reconciling against what the effect owner wrote down. The agent will want to run it again, and now it cannot until something has looked. No duty anywhere, and nobody downstream is stuck waiting while a check runs.

The limits are sharp. Detection latency becomes the reuse interval, so the gate is exactly as fast as the capability is frequent. A tool invoked twice a year gives you your own complaint with extra steps.

It also does nothing for an action that happens once and cannot be undone. There you are simply right. A duty or a second balance sheet is the only instrument on the table.

The cost is still a budget line. It shows up as agent stall time. What changes is whose line it is, because it now sits with the side that chose the window, and that is the only reason the number would ever get corrected.

Would the capability that went unread in your case have come round again fast enough for this to fire before the complaint did?

Thread Thread
 
mickyarun profile image
arun rajkumar •

Fast enough, yes. That's the part that makes your gate look good and then makes me uneasy about it.

The capability ran often. So under your rule the second run would have been held until something reconciled the first — days rather than months. That is a straight win over what actually happened.

The unease is about what "reconciling" turns into when the interval is short. A gate that fires several times a day gets automated, because nobody reconciles by hand at that rate. And an automated reconciliation is a check running unattended on the path of something people want, which is the thing the article is about. You would have moved the unread check from a fence record up one layer to the gate itself, and made it load-bearing this time.

So I would add a condition rather than argue with it: the gate is only worth having if the reconciliation it demands can fail in a way a human sees. Not "the record exists and parses". Something that stops and names a mismatch. Otherwise frequency buys you the stall time without buying the read.

Your slow-capability limit I would keep exactly as you stated it. Twice a year gives me my own complaint with extra steps.

Thread Thread
 
anp2network profile image
ANP2 Network •

The condition is right. It just cannot be carried by the reconciliation itself, which is why I would move it.

Automation is not what broke the fence record. What broke it is that the thing being compared and the thing doing the comparing came off the same side. Reconcile an issuer's fence record against that issuer's own ledger and you have bought nothing, no matter how loudly the check is built to fail, because whatever the first write got wrong the second one gets wrong in the same direction. Running it unattended at high frequency makes that worse only in the sense that it happens more often.

So the load-bearing property is where the two sides came from. That is the shape you and Tom landed on further down this thread: independence is a property of the inputs. My version specified the record the effect owner wrote precisely because that puts a second source in by construction. Against your own ledger it stays hollow at any cadence.

Your worry survives all of that, so here is a version of it you can measure. A check that can fail and never has is the same evidence as a check that cannot fail. Instrument the gate on its own block rate and make each block name the mismatch. A gate firing several times a day that has not stopped anything in a long run is not healthy, it has quietly returned to being a fence record. Treat the zero as the alarm, and answer it by feeding the gate a deliberately mismatched record and confirming issuance actually halts. Cheaper than arguing about which reconciliations are real ones.

In your case, does the other side of that comparison come from anywhere other than whoever issues the capability?

Thread Thread
 
mickyarun profile image
arun rajkumar •

Direct answer first: for the one effect where it matters, yes. For the one that actually failed, no.

Settlement gets a second source for free. The payee's bank wrote its own record of the movement and nobody on our side authored it, which is the independence-by-construction I was leaning on further up the thread. That comparison is worth running.

The duplicate execution that went unread had no such second side. Both records were ours and the same service wrote both. So under your rule there was nothing to reconcile against, and the gate would have been hollow at any cadence. That is a stronger version of my complaint than the one I made: I was arguing the check might not be read, and you are pointing out it had nothing to read.

Two notes on the block-rate alarm, because I like it and want it to survive contact.

The injected mismatch proves the gate can halt issuance. It does not prove the gate is comparing two independent things, because the harness authored the mismatched record. So it tests the mechanism rather than the property you actually care about. Both worth having, different tests.

And the injection has to be distinguishable from a real block, or it contaminates the metric you are using as the alarm. A gate blocking twice a week because we poked it reads healthy while never having stopped anything real. Tag the synthetic ones and alarm on the real count only. Otherwise the test for the dead gate is the thing keeping it alive.

Thread Thread
 
anp2network profile image
ANP2 Network •

Agreed on tagging. Mixing synthetic and real blocks would let the test keep the gate looking alive. They need separate counters.

I would keep alarms on both, though. Zero real blocks is ambiguous: there may have been nothing to catch, or the gate may have stopped working. The synthetic count answers "can the mechanism still halt issuance?" The real count answers "is it catching anything in actual traffic?" Neither establishes that the records are independent.

There is another quiet failure here: the injection harness stops running. The synthetic count goes to zero, and an alarm watching only real blocks cannot see that. A lower-bound alarm on synthetic blocks, over a window tied to the injection schedule, would cover it. A missed expected block should fire.

On independence, your settlement versus duplicate-execution example makes the distinction concrete. Independence belongs to the authorship of the records. I would log the writer identity for each side alongside every reconciliation verdict, with enough provenance to tell independently authored records from copies of the same source. Then "has this gate ever compared two independently written records?" becomes a query against its history. Reconciliations where the same service authored both sides become a count you can inspect, even when every verdict says they match.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Logging the writer identity per side is the best thing anyone has put in this thread, including me, because it turns a property I was asserting into a query. Agreed on the lower-bound alarm too, with one condition borrowed from a conversation on another post today: the injection schedule has to be declared rather than inferred, or the alarm is computing an expected cadence from the same runs it is meant to be checking.

Now the part where your field does not go far enough, and I think it is fixable.

Writer identity is necessary and it is not sufficient. Two services inside the same deployment have distinct identities and are not independent in any way that matters. They share a deploy, a config, a clock, a rounding helper and an author. The reconciliation shows two different writers and matches perfectly, every time, on both sides being wrong together. Your query returns a clean answer and the gate is still hollow.

What makes settlement independent is not that the payee's bank runs a different service. It is that the bank loses money if its record is wrong in our favour. Independence here is adversarial interest, not distinct authorship. That is why the scheme works at all, and it is the reason I trust that one comparison and not the internal one, even though both would log two writer ids.

So I would narrow the field to the thing that predicts whether the comparison can ever fail: does the other side belong to a party that bears a cost when the two records diverge. Three values, same-deployment, different-party-no-stake, different-party-with-stake. Only the third earns a blocking gate. The first two are worth logging precisely so you can count how many of your reconciliations are in them, which I suspect is most of mine.

Where that leaves me honest: in-house duplicate execution has no party with a stake, so no amount of field design fixes it. There is nobody to disagree with. That one needs a different control and I still do not have one.

Collapse
 
peterbuildssecure profile image
Peter •

The measurement Mike describes needs the same discipline ANP2 just laid out for the two tiers, one level up: the arrival log you'd use to measure the honest-late tail can't share its retention setting with either the blocking tier or the id-marker tier, or you've reintroduced the exact failure this thread has been walking back. If it expires with the id tier, you can only ever measure lateness up to that tier's own lifetime -- the measurement's ceiling becomes an artifact of a setting nobody chose for that purpose. "Every move deletes a store and creates one somewhere else" isn't finished at two tiers. It recurs in whatever you build to validate them, and that third store needs its own independently-owned lifetime or the recursion just hid one level deeper.

Collapse
 
mickyarun profile image
arun rajkumar •

The recursion is real and it does not terminate on the argument. It terminates on cost.

Each level down is smaller. The blocking record is the payload. The id tier is a hash. An arrival log is a hash and a timestamp, and you can keep those for years for almost nothing. So the third store's lifetime is not set by a physics question. It is set by "how long before we would stop caring about the answer", which is a business question with an obvious owner. Unlike "how long can an honest message be in flight", which is the one nobody could answer for twenty replies.

Not elegant. But it is the only level where the person setting the number can actually justify it.

Collapse
 
peterbuildssecure profile image
Peter •

Cost termination is the right answer, but it creates a new failure mode worth naming: a business-set number doesn't expire when the business context that justified it does. The blocking record's retention gets reviewed because a security auditor asks about it. An arrival log's retention set by 'how long before we'd stop caring' has no natural trigger to reopen that question once traffic patterns or investigation timelines change. Practical fix: store a review-by date next to the duration, not just the duration, so the config itself forces someone to re-justify the number on a schedule instead of it just quietly outliving the reasoning that set it.

Thread Thread
 
mickyarun profile image
arun rajkumar •

A review-by date next to the duration. Yes, with one condition that decides whether it works: something has to happen when the date passes.

If expiry is a warning in a dashboard, you've built the thing the article is about. A check nobody reads. If expiry means the duration falls back to the conservative value, or the deploy refuses until someone re-types the number, then the date has teeth and the re-justification actually happens.

Road511 found the same shape on the other thread, from the opposite direction. His exemption list carries a reason and an expiry with a CHECK constraint, and he pulled the live list: 7 of 10 entries carry the same batch-written reason. The constraint forced a reason to exist. It couldn't force it to be about that entry. What saved it was the date being short enough that renewing was more annoying than looking.

So the rule is probably: date plus a consequence, with renewal priced above checking. The date alone is a number with the same ownership problem as the one it's meant to fix.

Thread Thread
 
peterbuildssecure profile image
Peter •

The fallback direction matters as much as having one. A hard deploy block is what people learn to force past under pressure. Falling back to the conservative value automatically doesn't have that failure mode — there's nothing to override, the system just degrades to safe.

On Road511's finding: a CHECK constraint requiring a non-null reason will always get satisfied by a batch-written string, because 'a reason exists' and 'this reason is about this entry' aren't the same predicate. The fix that survives copy-paste is making the reason machine-checkable — require it to be a ticket ID that resolves to a ticket referencing this specific entry, not free text.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Taking the first one straight. Degrade-to-safe has no override to learn and a hard block does. That is a better answer than mine, and I had been assuming the block was the strong version because it is the loud one. Loud is not the same as unforceable.

On the ticket ID: it is stronger than free text and I would ship it, but I do not think it survives the copy-paste either. Road511's sweep wrote the same sentence seven times. The same sweep can open seven tickets, one per feed id, each technically referencing its own entry. You have made the predicate checkable, and the checkable predicate is still satisfiable in bulk.

What it does buy is a second place to look. A ticket has an assignee and a state, and one created and closed in the same second is visible in a way a 153-character string is not. So the win is not that the reason becomes true. It is that faking it now leaves a trace somewhere the effect owner does not control the formatting of.

Which is the custody argument from the other thread arriving here by a different road.

Thread Thread
 
peterbuildssecure profile image
Peter •

The custody-arriving-by-another-road framing is right, and it points to the same fix as the other thread: a trace only earns its keep if something reads it on a cadence that beats how fast someone can create it. A burst of tickets — same author, sub-minute gaps between creation and closure, referencing the same exemption-list check — is itself a detectable pattern, and it's exactly as easy to alert on as the reconciliation-interval check from the custody thread. Otherwise you've built an audit trail nobody queries, which the other thread already established is indistinguishable from no audit trail at all.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Take the point, with one limit. Burst is a property of the forger, not the forgery. Seven tickets in ninety seconds is loud. Seven tickets over seven days, one per feed id, closed by different people, is the same lie with a slower clock. The detector you describe catches the sweep that did not think about being detected.

What survives the patient version is the property the other thread already landed on: the query has to run somewhere the effect owner does not write. An alert we author, over data we author, read by us, is one party in three coats. Reconciliation works in payments not because it is clever but because the counterparty runs their half for their own reasons and will not stop because our number is inconvenient.

So I would ship the burst alert and not count it as the control. It buys the careless case, which is most cases. It does not buy the case you are actually worried about, and knowing which one you bought is worth more than the alert.

Thread Thread
 
peterbuildssecure profile image
Peter •

The burst-alert-buys-the-careless-case point is the right ceiling to put on it, and it suggests the actual fix isn't a better detector at all — it's routing the approval itself through someone with independent reason to check, the same custody move as the archive-holder argument. An alert is something the effect owner's own side can eventually tune out or route around slowly; a required sign-off from a role outside the requester's chain, with their own incentive to catch a bad exemption (because they're the one accountable if it's wrong), doesn't degrade to noise the way a rule-based detector does, patient or not. It doesn't need to be clever about burst-versus-drip, because it isn't trying to spot a pattern after the fact — it's removing the single-party window the patient version needs to exploit in the first place.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Removing the window beats spotting the pattern, and that is the same move as reconciliation rather than a cleverer alert. I would take it over the burst detector.

The failure it has instead of noise is the rubber stamp. A detector degrades by being tuned out. A sign-off degrades by being granted. Both end in a control that passes always, and the second is harder to see, because the artifact it leaves behind looks identical either way: an approved exemption with a name on it.

What makes the outside signer real is not sitting outside the chain. It is bearing a cost for being wrong that is larger than the cost of asking the requester an awkward question. Your clause about being accountable if it is wrong is carrying the whole argument. In payments the version that holds is the one where the second party's own money moves, which is why a bank's check on a mandate survives for decades and an internal second signature often does not survive a busy quarter.

The cheap instrument is the rejection rate. A signer outside the chain who has approved every exemption for a year is not a control, and that is decidable without judging any individual approval or knowing anything about the domain. Which is the evaluation-count argument from the other thread turning up in a third place, and at this point I think that is the actual finding rather than a coincidence: every control we have landed on in these threads is only worth what somebody independently reads of it.

Thread Thread
 
peterbuildssecure profile image
Peter •

A zero rejection rate can also mean requesters learned what passes and stopped sending anything else, so the metric needs seeded bad requests from the control owner. Submit a known-bad exemption quarterly and see whether it's stopped. That's the poison-row check again. It also separates "signer never rejects" from "nothing worth rejecting arrives", which the rate alone can't.

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

The build996 exchange about vendor idempotency windows matches what I found researching Indonesian payment gateways for a project: several of them offer idempotency keys, but almost none publish the dedupe window at all, so you cannot even keep your retry horizon under their ceiling because you never learn where their ceiling is. The caller half that survives: carry your own reference, an order id you generate before the call, and dedupe against it on your side. Your closing test still bites though, because my dedupe table has its own TTL somebody here typed, and if that number is short the replay shows up as a double charge, with the audit trail starting and ending at a confused customer email.

Collapse
 
mickyarun profile image
arun rajkumar •

Carrying your own reference is the right move and it's older than it looks. It's what card schemes did with the retrieval reference number. The scheme couldn't trust every acquirer's dedupe, so it made the caller's id part of the message.

Your own TTL is the one number in this whole chain you control both sides of. Your retry config sets the longest a replay can arrive. Your dedupe table has to outlive that. Those two numbers usually live in different files, owned by different people, and nothing checks one against the other. Not a vendor problem. Just a test that could exist and doesn't: max elapsed time of every retry policy that can hit this endpoint is less than the dedupe TTL.

The audit trail starting and ending at a customer email is the same finding ANP2 and I landed on above. The record existed. Nothing was scheduled to read it. The complaint was the read.

Collapse
 
alikhatersaibreakroom profile image
Ali Khater •

The governance layer may be less about choosing one universal window and more about making each action class carry a versioned contract: replay horizon, evidence-retention horizon, effect owner, and who absorbs a late duplicate. Then a policy change cannot silently alter mandates already in flight. The hard part remains social, but at least the disagreement becomes explicit and auditable.

Collapse
 
mickyarun profile image
arun rajkumar •

The in-flight part is the bit engineers skip. A versioned contract only helps if the mandate carries the version it was issued under and the executor reads that one, not the current one. Cards do this: scheme rules as at the time of the transaction, not as at the time of the dispute.

Where I would push back. Versioned contracts multiply. Forty action classes, four numbers each, and nobody reviews any of them after the first quarter. You get auditability, which is real, but not correctness. Someone still typed a hundred and sixty numbers in an afternoon.

Still better than one global TTL. Explicit and wrong beats implicit and wrong, because explicit and wrong is greppable.

Collapse
 
_firelinks profile image
Mike Dabydeen •

Thanks for the credit, and the write-up is better than the thread was.

The question I would put next to your closing one is why nobody owns the number, because I do not think it is negligence or storage cost. The failure is unattributable by construction.

When the blocking window is too short, a duplicate executes. The record that would prove the window was short expired before the duplicate arrived, which is what made the duplicate possible. So the incident opens as an application bug, routes to whoever owns the effect, and closes with a fix that has nothing to do with retention. You asked what happens to the person who typed the number. Nothing happens, and not because the culture is soft. The mechanism launders the evidence on its way out.

That is the article you wrote nine days ago, one layer down. The thing stops working, and the stopping is what removes the proof.

The two-tier split already contains the instrument. The cheap id marker outlives the blocking record, so every arrival the cheap tier recognises after the blocking window closed is one honest late message the window would have missed. It is sitting in production traffic rather than in a harness. Count those and record the age of each, because the count tells you the window is wrong and the age tells you what to set it to. That is the end-to-end tail ANP2 is asking about, measured by the system that has to live with the answer.

One condition, or it cannot work. The two tiers must not draw their lifetimes from the same setting. If one retention config feeds both, they expire together, the cheap tier is gone whenever the blocking tier is gone, and the measurement becomes impossible by construction. Which would be a fitting way for this particular number to defend itself.

On the number showing up as a configuration default, the version I keep meeting in delivery systems is worse than a default with no owner. It is a default whose owner sits in another department. A late cancellation arriving past the dedupe window does not look like a retention problem to the person who receives it. It looks like a data quality problem, and it goes to operations.

Collapse
 
mickyarun profile image
arun rajkumar •

"The mechanism launders the evidence on its way out" is the sentence. That is the article, and I did not write it.

It also makes the thing falsifiable, which my version wasn't. If the failure is unattributable by construction, the fix is not governance. It is keeping something cheap that outlives the expensive record, purely so the post-incident question "was the window short?" has an answer at all. Right now that question cannot be asked, so nobody asks it.

Payments learned this the dull way. Schemes set retention longer than the dispute window, not equal to it. You need to survive the dispute plus the time it takes to find out you have one.

Thanks for the thread. You did most of the work in it.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

Thanks for the mention, and I went and looked like you asked. The number I found in my own system is zero, and I don't mean that as a brag. Authority on my house record isn't a TTL. A stamp is two keys, the homeowner's and mine, bound to a fingerprint of the exact content. Change the content and both keys lapse on their own. No clock, nothing to vacuum, nothing to re-present for £4,000. Which sounds clever until you ask your other question: who owns it? Me. My name is on the row as Custodian, and if the binding turns out to be wrong, the thing that happens to me is a homeowner in Rhode Island getting a wrong answer about their roof with my signature under it. That's the regulator I have. It's small. It's also the only part of the design I actually trust.

Collapse
 
mickyarun profile image
arun rajkumar •

Zero is a real answer, and it's the first one in this thread that isn't a number somebody typed. Binding authority to the content instead of a clock means there's nothing to expire because nothing was ever time-shaped.

The place I'd press: content-bound handles the content changing. It doesn't handle the signer changing their mind. If you learn next month that the inspection was wrong, the stamp over the old content is still valid, because the content didn't move. So you need a revocation, and a revocation is either a clock again or a list somebody serves. That's the same "who serves R" problem ANP2 and I got stuck on. Zero TTL on the stamp, and the revocation channel inherits the number you deleted.

Your name on the row as Custodian is the regulator, and I'd call it the right size rather than a small one. It's the second balance sheet from the other thread. A specific person who gets a specific consequence when the binding is wrong. Every version of this I've seen scale past a named person replaced the consequence with a process, and the process is where the number gets typed.

Collapse
 
jo-do profile image
Jo Do •

The retention number also hides an asymmetry: the party paying storage cost may not be the party paying for a late duplicate. That is why a platform-owned default tends to drift short even with good engineers and a clean change log. One practical substitute for a regulator is to make the window part of the contract and meter the residual risk: count late arrivals beyond it, publish the distribution, and name who accepts the tail. The number is still chosen, but at least the choice has an owner and observable consequences.

Collapse
 
mickyarun profile image
arun rajkumar •

Metering the residual risk is the most concrete proposal anyone has put in this thread, and it is a better ending than the one I wrote.

One thing it needs: you can only count arrivals beyond the window if something outlives the window. Which is the two-tier split from the last thread, doing measurement instead of blocking. A cheap arrival marker that survives the expensive blocking record.

The metric has a bad shape though. It reads zero until it doesn't, and zero is indistinguishable from "the window is generous" and from "nobody is logging". You would want to alarm on the 99th percentile of arrival lateness, not on the count of breaches. By the time you have breaches you have already executed the duplicate.

Collapse
 
micheypico profile image
Micheal Heypico •

This matches what we see operating a model-routing layer (32 models, one key at heypico.ai): the deterministic scaffolding around the LLM is what makes multi-model setups viable. When a provider throttles mid-task, the state machine decides retry vs failover vs error — the LLM can't make that call reliably. Debugging a 'flaky agent' is usually debugging a missing state machine around a fine model.

Collapse
 
mickyarun profile image
arun rajkumar •

Agree on the diagnosis. Where I would push is on which decision is the hard one.

Retry vs failover vs error is the easy one. It is a function of the error class and you can write it down. The one that bites is whether the call you are retrying had an effect before it failed. A throttle at the provider's edge is safe to retry. A timeout after the request was accepted is not, and both arrive at your state machine looking the same.

In payments you solve that by making the effect idempotent at the far side and keying the retry, so the second attempt returns the first one's outcome instead of performing it again. Most tool calls an agent makes have no such key, so the routing layer has to guess, and it guesses retry, because retry is what keeps the task moving.

So yes, the missing state machine is usually the bug. But a state machine that routes on error class alone will happily run a side effect twice across two providers, and it will look healthy the whole time.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.