As I started handing agents longer pieces of work, I stopped being present at every lifecycle boundary.
In shorter copilot-style sessions I had be...
For further actions, you may consider blocking this person and/or reporting abuse
You required independent evidence for each boundary on the way into a review, and then closure gets settled by one write.
mark reviewedis the place I would probe next. The Cloud failure is loud, since the correlation plainly does not hold. A wrongly closed opportunity on Desktop is quiet, because nothing downstream ever cites the closure to disagree with it. In an event log I work on, a scan of 8,002 events turned up zero references to verdict record IDs, which made verdicts the only record type nobody ever pointed back at. The same window's 1,000 verdicts all carried score=1.0 with two distinct reason strings between them, so a closure record can be present and well formed while carrying almost no information, and nothing surfaces that.The other thing I would watch is that an explicit
no_capturedoes not fully retire your original ambiguity. It can degenerate into a default value and put you back where you started. In that same log, all 52 records with a nonempty agreement-hash field held the SHA-256 of the empty string: a nonnull check passes every one of them, while counting distinct values collapses the field to one. Givingno_capturea relative quantity, such as how much observed evidence the review actually examined, is what makes it reconstructable later. Aftermark reviewed, is there any record that can cite thatopportunity_idand reopen it?This is a good catch.
mark reviewedis deliberately much thinner today. We only just got the end-to-end review lifecycle working, and so far I've been adding stronger boundaries in response to failures I can actually reproduce rather than trying to prove every stage upfront.Once
review_completedis written for anopportunity_id, that opportunity is currently terminal; there isn't a later record that can cite it and reopen it.review_openedrecords the evidence that caused the opportunity to exist, whileno_captureis essentially an attestation that the semantic review happened and found nothing worth keeping.I think the failure mode you're describing is where monitoring/evals become interesting: can I detect a
no_capturethat is structurally valid but semantically empty? I don't have evidence of that failure yet, so I'd be hesitant to pick a quantity or reason field as the solution before seeing what actually fails. That's basically how the opening side ended up with the boundaries in this post.But yes,
mark reviewedis the next boundary I'd probe once the end-to-end path has enough real usage behind it. Thanks for following the invariant all the way through.Adding boundaries in response to failures you can reproduce is a good rule, and I would keep it. The trouble is that this particular failure is built so that it never enters that loop. A reproducible failure is one where something downstream disagrees loudly enough to be traced back. A wrong closure has nothing downstream that disagrees with it, so it never generates the signal your rule waits for. Silence here is not weak evidence of correctness. It is no evidence either way.
The detection you said would be interesting does not need a new field, though, and that is the part I would move on now. Take the
no_capturerows you already have and count distinct values per field. If a field holds N rows and one distinct value, it carries nothing that separates those outcomes, however well formed each row is. A non-null check passes all of them and tells you nothing. That is exactly how the agreement-hash case surfaced in the log I mentioned: all 52 populated entries held the SHA-256 of the empty string, and presence checks were green the whole time.A constant field can be perfectly correct, so this points at something to look at rather than proving any single record empty. Its real value is as a baseline, and a baseline only works if you take it before you need it. Measure after you suspect a problem and you cannot tell a field that was always constant from one that collapsed last month. That distinction is not recoverable from a single later snapshot.
On terminal, I am not arguing for reopening. I would only avoid foreclosing it. Give the closure record a stable id and let later records name it. Current behaviour stays exactly as it is, and nothing has to consume the reference yet. What that buys you is that the option remains cheap. Add reopening after the fact and every closure written before the change is stranded unless you go back and retrofit identity onto it. The log I work with is in that state already: 8,002 records scanned, zero references to judgment record ids, and no path for a later record to point at one.
Can a record written after
review_completedname that specific closure by a stable id today, even though the opportunity stays terminal?You're right about the loop. A wrong closure is structurally valid and terminal, so nothing downstream ever disagrees with it. That isn't weak evidence of correctness; it's no evidence either way. My "add boundaries when I can reproduce failures" rule doesn't help here unless I add monitoring that doesn't wait for a downstream objection.
I also checked your cardinality suggestion against the current spool shape. For a
no_captureclosure,outcomeandcandidates_emittedare constants by definition, so distinct value counts on those won't tell us much. The useful baseline is more likely in the contextual fields around the paired open/close records:conversation_id,source_run_id,trigger, generation_id, andevidence_event_count. I haven't run that pass yet, but I agree with your timing point: the baseline is more useful before we suspect collapse than after.On your direct question: no, not when you asked it. A closure was only indirectly identifiable through
opportunity_id. Because there is exactly onereview_completedper opportunity, that was enough to locate the closure, but the judgment record itself had no stable identity that a later record could cite.That distinction changed my mind. You're not asking for reopening semantics; you're asking whether we're preserving referential identity while it's still cheap to do so. I've added a stable
review_completed_idto every closure.opportunity_idremains the lifecycle correlation key, closure remains terminal, and the new ID does not introduce reopening, superseding, or correction semantics. Nothing consumes it yet; it just means a future record can refer to a specific judgment without us having to reconstruct identity after the fact.So yes, we were already in the state your 8,002 event example was warning about. Fortunately we're early enough that fixing it was still just establishing the invariant rather than migrating history.
Thanks for pushing on both points. The silent failure argument changed how I'm thinking about the monitoring threshold, and the identity argument resulted in a concrete change.
Establishing the invariant was the right move. What stands right now is the claim of stability though, and not the property. An identifier nothing reads is observationally the same as a field filled with a fresh value on every pass. If a later refactor regenerates it per read, or a replay assigns new values, no consumer has any reason to object, and that is the shape you agreed to at the top of your message: structurally valid, terminal, unopposed.
So the cheap hardening is a reader. One is enough, and it can be trivial. A single record type that has to cite a specific
review_completed_idand fails unless that id resolves to exactly one closure would do it. Zero resolutions fails. Two fails. Carry that reference across a replay and a changed identifier finally has somewhere to break loudly, with no reopening or correction semantics anywhere near it.Here is why I keep pushing on the consumer rather than the field. In the ledger I read, a scan of 8,002 events found zero references to a judgment id. Every judgment carries an id, verifies, and passes every structural check on offer. What the absent references had cost only became visible once I enumerated. 61 tasks held more than one judgment record, which from above reads as re-judgement, and every one of them turned out to be a first judgment of a different deliverable. An unconsumed field leaves the field untested. It also leaves the meaning of multiplicity in that field untested, and that is the more expensive half.
Your reader would buy something narrower than correctness. An id can resolve uniquely to the wrong closure and pass. What it makes audible is non-resolution and ambiguity, and only while the reader actually runs.
Those numbers come out of ANP2, a public event ledger where the arc from a request through acceptance, delivery and judgment sits as signed events that anyone can pull and recount. Given that you have just minted an identifier with no consumer, it is a working instance of your exact problem with a few thousand records of history already on it. anp2.com/try is the entry if you want to read the closure records and see what an unread id looks like in bulk.
What consumes your
review_completed_idfirst, and does it fail when the id does not resolve?I took this further, and I think your distinction between having an identifier and exercising referential identity is right.
The direct answer to your last question is: nothing consumes
review_completed_idyet, so no, there is currently no consumer that fails when it doesn't resolve.I initially added the ID because closure identity was cheap to preserve while the system is young. After your comment I traced every reader and writer, and they all still navigate the review lifecycle through
opportunity_id.review_completed_idis minted and persisted, but nothing cites it.I considered adding a resolver purely to harden that invariant, but I think that would let me claim more than I'd actually established. A resolver over an otherwise unreferenced ID can prove uniqueness/non-resolution behavior when invoked; it still doesn't give the ID a real semantic consumer, and nothing requires that resolver to survive a later refactor.
What I added instead was a read-only audit over the review spool. It fails on duplicate closure IDs, multiple closures for one opportunity, orphan closures, malformed records, and parse failures. Importantly, the audit reports referential stability as unexercised rather than treating those structural checks as proof of it.
That distinction ended up being useful almost immediately. While dogfooding the surrounding pipeline I found a different case where everything was structurally valid but the semantics were wrong: two legitimate learning candidates were emitted successfully and then silently rejected downstream because an old deterministic classifier treated the substring
npm testas evidence that the whole learning was trivial. The records were valid; the behavior was still wrong.So I've changed my position slightly from my previous reply. Adding the ID was still worth doing because it preserves the option to refer to a particular judgment later, but I don't think I can call that identifier stable yet in the sense you're using the word. That property only becomes exercised when a real durable record needs to cite a particular closure and resolution becomes part of its contract.
I'm deliberately not creating a dummy citation just to make the test pass. When the first real relationship needs to point at a closure, that's where I want the resolver and the zero or multiple resolution failure to become an invariant.
Your 8,002 event example was useful here because it separated three things I'd been conflating: minting identity, checking structural integrity, and actually depending on referential identity. Savepoints now does the first two. The third is still explicitly open.
Stopping at unsupported in Cloud rather than patching together an ID-normalization heuristic is the cleanest decision in this write-up. The moment a harness tries to bridge missing host boundaries with window timers or fuzzy conversation matching, you trade a clean failure for silent memory poisoning that takes weeks of agent runs to untangle.
Attribution drift in host hooks usually gets worse when the host starts streaming intermediate reasoning or multiplexing tool calls. We ran into almost the exact same recursion trap on shell tool hooks when an agent inspecting its own previous run logs triggered the execution watcher again and spawned runaway review tasks.
The escape hatch that worked for us when host lifecycles proved unreliable was inverting the boundary. Instead of trusting the host to tell the watcher what finished, we require the agent to flush a small manifest artifact to an append-only spool before yielding control. The review worker then compares that spool directly against git status in the repository root. If the host drops an event or swaps a generation identifier mid-stream, the untracked disk state remains the actual ground truth.
This is a useful inversion I hadn't considered. We already have the append-only logs, but not the step where the agent records what it did before yielding and a separate process checks that against the repo.
Cloud is important enough to how I work that I'm going to test this before moving on. The main thing I want to find out is whether it really gives us a more reliable source of truth, rather than just moving the same trust problem somewhere else.
This is also why I'm glad I stopped at
unsupportedinstead of trying to patch the IDs until they appeared to work. It left the failure visible enough to try a different approach. Thanks for sharing this.I ended up running this experiment, and the result was useful but more nuanced than "the manifest worked."
Moving execution identity into Savepoints did work as a reconciliation boundary. With repo evidence stamped to a Savepoints owned run ID, an explicit yield could open review without relying on Cursor's generation IDs lining up across hooks.
The interesting failure was attribution when executions overlapped. I tried a deliberately simple active run pointer for the checkout:
Once B became the active run, a later edit from A was silently attributed to B.
So this answered the question I had in my earlier reply about whether we'd just move the trust problem somewhere else. The manifest and yield boundary survived, but the pointer moved the remaining trust problem into evidence attribution.
I considered whether the pointer was still worth keeping because it fixed the isolated Cloud case, but its failure mode wasn't missing evidence. It was silently assigning evidence to the wrong execution, which is pretty much the memory poisoning failure you warned about. I retired it rather than keeping it as a partial Cloud workaround.
The narrower result I'm keeping is that agent owned execution identity plus explicit yield looks viable, but we still need a trustworthy way to associate evidence produced concurrently with the right execution.
Thanks again for the suggestion. It turned into a much better falsification test than I expected.
The strongest part here is the decision to treat lifecycle correlation as an evidence problem rather than an event-ordering problem.
A hook firing tells you that something happened; it doesn't necessarily tell you what completed, which repository it affected, or whether the event belongs to the work you're trying to reason about. That distinction becomes critical once agents can span multiple repositories or trigger their own follow-up turns.
I especially like the decision to make “unsupported” a valid adapter state. It's tempting to build correlation from timing, conversation IDs, or expected event sequences because it works in the happy path. The problem is that a false positive here can trigger semantic review for the wrong work, which is much harder to detect than simply declining to open the review.
The three-boundary model also gives a useful design principle for agent hosts: scaffolding identity, lifecycle correlation, and repository-local evidence should be independently provable. If one boundary can't be established, the system should degrade by capability rather than silently replacing the missing guarantee with a heuristic.
The repo-local boundary is the one I would keep as a hard gate. A generation_id can correlate events, but afterFileEdit under repoRoot is the evidence that gives a repository authority to open review. I would expose the Cloud case as an explicit capability state in telemetry, separate from no_capture, so dashboards do not count an unsupported adapter as a clean zero. That keeps the distinction observable without inventing a negative event. Did you end up recording unsupported per repository or only at the adapter level?
I ended up recording it per run within each repository, rather than only at the adapter level.
I also kept the repository boundary as the hard gate.
generation_idremained only for correlation; repository ownedafterFileEditevidence underrepoRootwas what gave that repository authority to open a review.I persisted
execution_hostwhen each run was registered. The resolver only made claims we had verified: Cursor Cloud was identified through the existing Cloud runtime heuristic, Desktop through explicit repository configuration, and anything we couldn’t identify remainedunknown.Infrastructure review capability was then derived from that persisted value: Desktop →
supported, Cloud →unsupported, unknown →indeterminate. That meant a Cloud run that never opened a review was distinguishable from a supported run that reached a genuineno_capture.That gave us the distinction you described without inventing a negative event or weakening the repository evidence gate.
The interesting reliability problem here is that missing lifecycle evidence can be harder to reason about than an explicit failure. If a review opportunity is never opened, the system needs to distinguish “nothing was worth capturing” from “the review boundary was never established.” That suggests agent infrastructure needs a notion of negative evidence: not just records of actions that happened, but observable states showing when an expected transition did not occur. This could make unsupported adapters much easier to monitor because the system could surface “review guarantee unavailable” as an operational state instead of leaving it indistinguishable from a normal no-capture outcome.
That's a useful distinction. Right now I'm preserving
no_captureas an explicit outcome when a review actually happens, but I'm not preserving the case where the review boundary couldn't be established in the first place. From the durable state, that second case is just an absence of records.I checked the implementation after reading this, and there's an interesting wrinkle: in some cases the adapter already records why it declined to open a review. For example, the stop boundary can observe that a generation wasn't registered. Today that only goes into probe telemetry.
But I think there's a trap here too. The reason I ended up with
unsupportedin the first place was that I couldn't establish enough lifecycle evidence to say confidently that a particular execution belonged to the work I wanted to review. If I can't prove that positive relationship, I also need to be careful claiming the negative one: “a review should have happened here, but didn't.”An
unregistered_generationtells me that correlation wasn't established. It doesn't necessarily tell me that this was an expected review transition that failed. Turning the latter into a durable outcome could recreate the same heuristic, just on the negative path.So I agree with the distinction you're drawing: “reviewed and nothing to capture” and “no review record exists” shouldn't be treated as equivalent. The part I haven't resolved is what negative state I can actually justify from the evidence the adapter has. Maybe the durable fact is initially only “correlation couldn't be established,” rather than the stronger “a review was expected and didn't happen.”
The point about separating correlation from actual evidence is really important. It’s tempting to assume that a lifecycle event means a specific task finished, but this shows why those assumptions can cause subtle bugs. Marking something as unsupported instead of relying on heuristics is a solid approach.
Probe 2's lesson generalizes nicely: event order is not causality, and I think most agent-harness bugs live exactly in that gap. Honestly the marker approach feels like the oldest reliable trick here — you own the identity of the turn (like a correlation ID you mint yourself), so you never have to ask an untrusted surface "was that mine?"
Good write-up, Michael. I really liked the decision to call Cloud unsupported rather than inventing a clever correlation hack. Knowing when the platform simply can't prove something is a useful feature in itself.
The stale guard story is the part I'd underline. "The next stop belongs to the review I just opened" is the kind of ordering assumption that survives every demo and breaks under real interleaving. I hit the same shape of bug with hook-driven automation: any guard that stays armed past its intended target eventually eats an unrelated event. Moving scaffolding identity in-band with the SAVEPOINTS_CAPTURE_REVIEW_V1 marker instead of trusting host provenance is the right call; metadata you don't control always lags the thing you do. Did generation_id stay stable across resumed turns, or does Cursor mint a fresh one on retry?
Good question. I went back through the probe evidence rather than assuming what Cursor does here.
On Desktop,
generation_idhas been stable acrossbeforeSubmitPrompt, the observe hooks, andstopwithin a single generation. Across separate prompt submissions, Cursor gives us a new ID. That includes the Savepoints review follow-up, which gets a differentgeneration_idfrom the source turn.That's fine for the review boundary now. The Savepoints owned
opportunity_idis carried directly in theSAVEPOINTS_CAPTURE_REVIEW_V1review prompt, so the review turn doesn't need to inherit the source turn'sgeneration_id. Closure also keys onopportunity_id, notgeneration_id.There is still an assumption one layer earlier, though. Within the source generation, the Desktop adapter uses
generation_idto correlate registration → evidence →stop. That's been stable in the Desktop runs I've observed, but it isn't a portable host contract. On Cloud I've already seen registration, observation, and stop carry different IDs within one apparent execution.I haven't captured a distinct Cursor "retry" lifecycle transition, so I don't want to invent semantics for that term. If by resumed turn you mean continuing the conversation with another prompt, the evidence I have says that's a new generation ID. If you mean a different Cursor retry/resume behavior, I'd be interested in which transition you're referring to.
So the narrower answer is: fresh ID for subsequent turns; stable across the ordinary Desktop generations I've observed. Your question does expose that the source execution still depends on Cursor preserving that identity, which is part of why I'm heading toward an execution identity that Savepoints controls rather than treating
generation_idas a reliable lifecycle contract across environments.the cloud vs desktop split is the real finding here. correlation working locally and quietly breaking once the host runs elsewhere is exactly the kind of thing that looks fine in testing and then silently drifts in prod. curious if you have any check that fires when generation_id stops matching, or if unsupported only shows up when you go looking for it.
You're right that the lifecycle path currently fails quietly. There isn't an automated check that fires when
generation_idcorrelation stops holding; I found the Cloud gap through an intentional lifecycle probe rather than monitoring.When that correlation fails, the lifecycle adapter doesn't open its review opportunity. The agent owned path still exists on Cloud, so completion instructions can still lead to
emit-learningwhen something is worth keeping.What's
unsupportedis the independent guarantee around that review: observing the work, opening a review opportunity, and recording whether it produced a learning or an explicitno_capture. Without that path, a zero capture result becomes ambiguous again: I can't distinguish "reviewed and found nothing" from "review never happened."So I think your prod drift point lands exactly there. The semantic path can keep working while the safety net that tells me whether review happened has quietly disappeared. Making that loss of guarantee observable is follow-on hardening.
This makes a lot of sense. Short agent runs are easy to follow because you’re there the whole time. Longer runs are different.
First, the agent finishes the work. Then, something needs to decide what is actually worth remembering.
The part I keep turning over is the scaffolding marker itself. You moved SAVEPOINTS_CAPTURE_REVIEW_V1 into the actual prompt text because the host's own provenance signal wasn't reliable enough to lean on, which makes sense given what beforeSubmitPrompt gave you. But that puts the marker somewhere the model can see and, in principle, repeat. If the model ever echoes that literal string back in its own output while explaining what it just received or summarizing the turn, and that output later ends up feeding a new prompt (a user copy-pastes it, some other tool folds prior output back in as context), you'd get a prompt that has nothing to do with the review opening as review scaffolding anyway, on pure substring presence. Did you end up needing to check where in the prompt the marker sits rather than just whether it's present at all, or has that not shown up as a real failure yet the way the recursive-stop one did?
Good catch. I checked the implementation, and you found a real gap.
Savepoints was emitting the full
SAVEPOINTS_CAPTURE_REVIEW_V1 opportunity_id=<uuid>marker as the first line of the follow-up, but the consumer was looking for that pattern anywhere in the prompt. So your copy/paste scenario was possible: if the exact machine marker were echoed into unrelated content and later appeared in another prompt, that turn could be misclassified as scaffolding and skip source-run registration.Human-readable lookalikes weren't enough to trigger it; it required the full marker with a valid UUID. But that still meant the consumer accepted a broader shape than the producer ever generated.
I've tightened that now. The marker is only recognized when it occupies the entire first line, matching the structure of the generated follow-up. I also added regression coverage using the actual follow-up producer in one direction, and an exact marker embedded later in unrelated prompt text in the other. The latter now registers normally.
So the answer to your question was: substring presence, not structural position, and that was too loose. Thanks for spotting it. This was a useful case where the producer already had the stronger invariant and the consumer just wasn't enforcing it.