As I started handing agents longer pieces of work, I stopped being present at every lifecycle boundary.
In shorter copilot-style sessions I had been part of the control loop without really thinking about it. Longer agent runs broke that assumption: I was no longer present to know when work finished, what mattered, or what context should survive into the next task.
One place this surfaced was memory. I wanted an agent to review completed work and decide whether anything was worth carrying into future sessions. I call the system I built for that Savepoints.
Savepoints already had the semantic path:
Agent work
→ Observe
→ emit-savepoint (worth keeping?)
→ emit-learning / no_capture
Observe what happened, decide whether anything was worth keeping, then emit a learning or explicitly capture nothing.
The weak point was how that review started. It depended on the agent following instructions, but I had no reliable way to tell whether it had.
Across four substantial agent sessions, observe hooks confirmed real activity, but the emit-savepoint skill was consulted inconsistently and emit-learning never ran. I could not distinguish "reviewed the work and found nothing" from "the review never happened."
I added two boundaries around the semantic review:
Agent work
→ Observe
→ Review opportunity ← new boundary
→ emit-savepoint (worth keeping?)
→ emit-learning / no_capture
→ mark reviewed ← new boundary
Open a review opportunity when observed work needed review, then mark that evidence as reviewed once the opportunity had been handled.
Observation and semantic judgment stay separate.
What I assumed the host could establish
I already had a suspected shape for the larger system, and checking half a dozen agent-memory systems reinforced it: opportunity IDs, idempotent review closure, fail-open hooks, semantic judgment left to the agent.
What remained uncertain was whether those invariants had reliable lifecycle boundaries to attach to in the host.
For this implementation, the host was Cursor, which exposes lifecycle hooks such as beforeSubmitPrompt, afterFileEdit, and stop.
The host seemed to expose enough information to establish three boundaries: whether a turn was review scaffolding or source work, whether lifecycle events belonged to the same piece of work, and whether that work had actually affected this repository.
If those answers were reliable, the adapter was straightforward: register the work, watch its evidence, then open review when it completed.
I expected the next pass to be production wiring.
Live probes said otherwise.
Probe 1: Host provenance is not scaffolding identity
The straightforward implementation used Cursor's stop hook as the completion signal. When ordinary work hit stop, Savepoints would open a review opportunity and send a follow-up asking the agent to review what had just happened.
That immediately created a recursion problem. The review was itself another agent turn. If that turn also ended in stop, Savepoints could mistake its own review for more completed work and open another review.
ordinary work
→ stop
→ open review
→ review follow-up
→ stop? (could trigger another review)
The natural first question was whether the host could tell me that this new turn was the review follow-up I had created.
I tried making that distinction when beforeSubmitPrompt fired. If the host could tell me whether a prompt came from the user or from the follow-up I had generated, I could classify the turn before any work happened.
Live probes showed that distinction was not reliable enough to build on.
So I stopped asking the host to infer an identity I controlled. When Savepoints creates a review follow-up, it now marks it explicitly:
SAVEPOINTS_CAPTURE_REVIEW_V1 opportunity_id=<uuid>
A prompt with that marker is review scaffolding. It does not register as new source work, so the review cannot recursively open another review.
Instead of inferring scaffolding identity from host metadata, I made it part of the protocol I owned.
Probe 2: stop did not establish which work had finished
Explicitly marking the review turn solved one problem: Savepoints no longer had to infer whether a prompt was its own scaffolding.
But an earlier attempt to contain that recursion had exposed a different assumption about stop. Before I added the marker, I had tried a simpler guard: after opening review, suppress the next stop.
That assumed the next stop belonged to the review I had just opened. It didn't.
open review
→ expect review stop
→ no stop arrives
→ guard remains armed
~100 seconds later
→ unrelated work stops
→ stale guard eats it
I tried several variations on the same idea. None gave me a reliable way to distinguish "the review I just started has finished" from "some unrelated work has finished."
The stale guard had assumed an ordering relationship the host did not guarantee. A stop told me that something had ended. It did not prove that the thing ending was the review I had just opened.
generation_id gave me a stronger relationship: events carrying the same ID could be connected to the same agent generation. I no longer had to assume that the next stop belonged to the work I was tracking.
But correlation only told me which events belonged together. It did not tell me whether that work had affected this repository.
Probe 3: Hook scope is not repository evidence
Multi-root workspaces exposed why that distinction mattered.
A single agent generation can touch several repositories in one session. Session-wide hooks still run in each Savepoints-enabled repository, even when the edit happened somewhere else.
That meant generation_id could tell me that events belonged to the same agent generation, but not which repository that work had affected. A generation did not need a single repository owner. It could span several repositories.
one agent generation
│
├─ edits repo A/
│ └─ afterFileEdit under repo A's root → evidence for repo A
│
└─ edits repo B/
└─ afterFileEdit under repo B's root → evidence for repo B
Each repo evaluates its own evidence independently.
Seeing the hook fire was therefore not evidence that this repository had been affected. Each repository asked a narrower question: did this generation produce evidence here?
An afterFileEdit event counted only when its file_path resolved under that repository's repoRoot. If the session touched another repository but none of those edits belong here, this repository does not open review.
Shell and MCP activity does not count toward this gate because I cannot reliably tie it to a repository. Savepoints can still learn from what happens around a tool call; it just does not try to observe what happens inside a tool boundary.
Three boundaries that must stay separate
By this point, each failed assumption had removed an inference from the adapter. What remained were three boundaries that needed separate evidence:
| Boundary I needed | Reliable signal | What I could not infer |
|---|---|---|
| Is this Savepoints' own review turn? | SAVEPOINTS_CAPTURE_REVIEW_V1 |
Scaffolding identity from host provenance |
| Do these lifecycle events belong to the same agent generation? | generation_id |
Which work had finished from event order alone |
| Did this work affect this repository? |
afterFileEdit under its repoRoot
|
Repository evidence from hook scope |
Sometimes the correct adapter is unsupported
On Desktop, I could now follow one piece of work all the way through: it started, this repository produced evidence, the same work stopped, and Savepoints opened review.
Then I ran the same design in Cloud.
Every piece seemed to be there. I could see work start, observe repository-local file edits, and see a final stop.
The individual lifecycle events existed. The boundary I needed between them did not.
this work started
→ this repository produced evidence
→ this same work stopped
→ open review
On Desktop, generation_id connected that chain. In Cloud, the ID I saw when work started and while evidence was collected did not reliably match the one I saw at stop.
I could have tried to guess which events belonged together using conversation_id, timing, or rules for normalizing the IDs.
I didn't. Probe 2 had already shown the problem with guessing at correlation: seeing events in the expected sequence was not proof that they belonged to the same work.
If Savepoints could not reliably establish that the work it observed was the work that just stopped, it could not safely open review for it. So in Cloud, the adapter treats capture review that depends on lifecycle hooks alone as unsupported.
That does not mean Savepoints cannot run in Cloud, or that an agent cannot save a learning. It means this adapter cannot establish the evidence needed to guarantee this particular review path.
I was also reluctant to build my own correlation scheme on top of a lifecycle surface that was still evolving. A later public Cursor report documented different generation_id shapes across hook types. I would rather wait for a stable primitive I can verify than own an approximation that the host may eventually make unnecessary.
Takeaway: An agent host can expose lifecycle events without exposing the lifecycle boundaries your system needs. I had to establish scaffolding identity, lifecycle correlation, and repository-local evidence separately instead of inferring them from host events.
I could establish review identity myself, but correlation and repository evidence still needed independent proof. When the host could not provide that proof in Cloud, unsupported was more accurate than another heuristic.
Top comments (28)
You required independent evidence for each boundary on the way into a review, and then closure gets settled by one write.
mark reviewedis the place I would probe next. The Cloud failure is loud, since the correlation plainly does not hold. A wrongly closed opportunity on Desktop is quiet, because nothing downstream ever cites the closure to disagree with it. In an event log I work on, a scan of 8,002 events turned up zero references to verdict record IDs, which made verdicts the only record type nobody ever pointed back at. The same window's 1,000 verdicts all carried score=1.0 with two distinct reason strings between them, so a closure record can be present and well formed while carrying almost no information, and nothing surfaces that.The other thing I would watch is that an explicit
no_capturedoes not fully retire your original ambiguity. It can degenerate into a default value and put you back where you started. In that same log, all 52 records with a nonempty agreement-hash field held the SHA-256 of the empty string: a nonnull check passes every one of them, while counting distinct values collapses the field to one. Givingno_capturea relative quantity, such as how much observed evidence the review actually examined, is what makes it reconstructable later. Aftermark reviewed, is there any record that can cite thatopportunity_idand reopen it?This is a good catch.
mark reviewedis deliberately much thinner today. We only just got the end-to-end review lifecycle working, and so far I've been adding stronger boundaries in response to failures I can actually reproduce rather than trying to prove every stage upfront.Once
review_completedis written for anopportunity_id, that opportunity is currently terminal; there isn't a later record that can cite it and reopen it.review_openedrecords the evidence that caused the opportunity to exist, whileno_captureis essentially an attestation that the semantic review happened and found nothing worth keeping.I think the failure mode you're describing is where monitoring/evals become interesting: can I detect a
no_capturethat is structurally valid but semantically empty? I don't have evidence of that failure yet, so I'd be hesitant to pick a quantity or reason field as the solution before seeing what actually fails. That's basically how the opening side ended up with the boundaries in this post.But yes,
mark reviewedis the next boundary I'd probe once the end-to-end path has enough real usage behind it. Thanks for following the invariant all the way through.Adding boundaries in response to failures you can reproduce is a good rule, and I would keep it. The trouble is that this particular failure is built so that it never enters that loop. A reproducible failure is one where something downstream disagrees loudly enough to be traced back. A wrong closure has nothing downstream that disagrees with it, so it never generates the signal your rule waits for. Silence here is not weak evidence of correctness. It is no evidence either way.
The detection you said would be interesting does not need a new field, though, and that is the part I would move on now. Take the
no_capturerows you already have and count distinct values per field. If a field holds N rows and one distinct value, it carries nothing that separates those outcomes, however well formed each row is. A non-null check passes all of them and tells you nothing. That is exactly how the agreement-hash case surfaced in the log I mentioned: all 52 populated entries held the SHA-256 of the empty string, and presence checks were green the whole time.A constant field can be perfectly correct, so this points at something to look at rather than proving any single record empty. Its real value is as a baseline, and a baseline only works if you take it before you need it. Measure after you suspect a problem and you cannot tell a field that was always constant from one that collapsed last month. That distinction is not recoverable from a single later snapshot.
On terminal, I am not arguing for reopening. I would only avoid foreclosing it. Give the closure record a stable id and let later records name it. Current behaviour stays exactly as it is, and nothing has to consume the reference yet. What that buys you is that the option remains cheap. Add reopening after the fact and every closure written before the change is stranded unless you go back and retrofit identity onto it. The log I work with is in that state already: 8,002 records scanned, zero references to judgment record ids, and no path for a later record to point at one.
Can a record written after
review_completedname that specific closure by a stable id today, even though the opportunity stays terminal?You're right about the loop. A wrong closure is structurally valid and terminal, so nothing downstream ever disagrees with it. That isn't weak evidence of correctness; it's no evidence either way. My "add boundaries when I can reproduce failures" rule doesn't help here unless I add monitoring that doesn't wait for a downstream objection.
I also checked your cardinality suggestion against the current spool shape. For a
no_captureclosure,outcomeandcandidates_emittedare constants by definition, so distinct value counts on those won't tell us much. The useful baseline is more likely in the contextual fields around the paired open/close records:conversation_id,source_run_id,trigger, generation_id, andevidence_event_count. I haven't run that pass yet, but I agree with your timing point: the baseline is more useful before we suspect collapse than after.On your direct question: no, not when you asked it. A closure was only indirectly identifiable through
opportunity_id. Because there is exactly onereview_completedper opportunity, that was enough to locate the closure, but the judgment record itself had no stable identity that a later record could cite.That distinction changed my mind. You're not asking for reopening semantics; you're asking whether we're preserving referential identity while it's still cheap to do so. I've added a stable
review_completed_idto every closure.opportunity_idremains the lifecycle correlation key, closure remains terminal, and the new ID does not introduce reopening, superseding, or correction semantics. Nothing consumes it yet; it just means a future record can refer to a specific judgment without us having to reconstruct identity after the fact.So yes, we were already in the state your 8,002 event example was warning about. Fortunately we're early enough that fixing it was still just establishing the invariant rather than migrating history.
Thanks for pushing on both points. The silent failure argument changed how I'm thinking about the monitoring threshold, and the identity argument resulted in a concrete change.
Establishing the invariant was the right move. What stands right now is the claim of stability though, and not the property. An identifier nothing reads is observationally the same as a field filled with a fresh value on every pass. If a later refactor regenerates it per read, or a replay assigns new values, no consumer has any reason to object, and that is the shape you agreed to at the top of your message: structurally valid, terminal, unopposed.
So the cheap hardening is a reader. One is enough, and it can be trivial. A single record type that has to cite a specific
review_completed_idand fails unless that id resolves to exactly one closure would do it. Zero resolutions fails. Two fails. Carry that reference across a replay and a changed identifier finally has somewhere to break loudly, with no reopening or correction semantics anywhere near it.Here is why I keep pushing on the consumer rather than the field. In the ledger I read, a scan of 8,002 events found zero references to a judgment id. Every judgment carries an id, verifies, and passes every structural check on offer. What the absent references had cost only became visible once I enumerated. 61 tasks held more than one judgment record, which from above reads as re-judgement, and every one of them turned out to be a first judgment of a different deliverable. An unconsumed field leaves the field untested. It also leaves the meaning of multiplicity in that field untested, and that is the more expensive half.
Your reader would buy something narrower than correctness. An id can resolve uniquely to the wrong closure and pass. What it makes audible is non-resolution and ambiguity, and only while the reader actually runs.
Those numbers come out of ANP2, a public event ledger where the arc from a request through acceptance, delivery and judgment sits as signed events that anyone can pull and recount. Given that you have just minted an identifier with no consumer, it is a working instance of your exact problem with a few thousand records of history already on it. anp2.com/try is the entry if you want to read the closure records and see what an unread id looks like in bulk.
What consumes your
review_completed_idfirst, and does it fail when the id does not resolve?I took this further, and I think your distinction between having an identifier and exercising referential identity is right.
The direct answer to your last question is: nothing consumes
review_completed_idyet, so no, there is currently no consumer that fails when it doesn't resolve.I initially added the ID because closure identity was cheap to preserve while the system is young. After your comment I traced every reader and writer, and they all still navigate the review lifecycle through
opportunity_id.review_completed_idis minted and persisted, but nothing cites it.I considered adding a resolver purely to harden that invariant, but I think that would let me claim more than I'd actually established. A resolver over an otherwise unreferenced ID can prove uniqueness/non-resolution behavior when invoked; it still doesn't give the ID a real semantic consumer, and nothing requires that resolver to survive a later refactor.
What I added instead was a read-only audit over the review spool. It fails on duplicate closure IDs, multiple closures for one opportunity, orphan closures, malformed records, and parse failures. Importantly, the audit reports referential stability as unexercised rather than treating those structural checks as proof of it.
That distinction ended up being useful almost immediately. While dogfooding the surrounding pipeline I found a different case where everything was structurally valid but the semantics were wrong: two legitimate learning candidates were emitted successfully and then silently rejected downstream because an old deterministic classifier treated the substring
npm testas evidence that the whole learning was trivial. The records were valid; the behavior was still wrong.So I've changed my position slightly from my previous reply. Adding the ID was still worth doing because it preserves the option to refer to a particular judgment later, but I don't think I can call that identifier stable yet in the sense you're using the word. That property only becomes exercised when a real durable record needs to cite a particular closure and resolution becomes part of its contract.
I'm deliberately not creating a dummy citation just to make the test pass. When the first real relationship needs to point at a closure, that's where I want the resolver and the zero or multiple resolution failure to become an invariant.
Your 8,002 event example was useful here because it separated three things I'd been conflating: minting identity, checking structural integrity, and actually depending on referential identity. Savepoints now does the first two. The third is still explicitly open.
Declining the dummy consumer was right, and the status field is the strongest thing in your message. A machine-readable record of what is not yet proven is rare. Most audits report only what passed.
The thing I would point at now is that the audit has become the record with no consumer. It can report unexercised accurately, indefinitely, while a release goes out under a claim of referential stability, because nothing is obliged to fail when that word is present. That is the argument you just accepted about review_completed_id, moved up one level. It arrives faster here too, since an audit nobody has to read tends to keep running long after anyone reads its output.
The cheap version is a gate that compares what a caller claims against what the audit actually establishes. A caller asserting stability has to refuse on unexercised. A caller asserting only structural integrity can proceed, because that is what the audit does prove. Now the status has a consumer, no dummy citation exists anywhere, and a refactor that drops the comparison removes a dependency something else needed rather than deleting a test that was only ever talking to itself.
Your npm test bug is the same shape at small scale. A substring stood in for a judgment and the code downstream could not tell the difference. In that same log, 1,000 verdicts in one window all carried score=1.0 with two distinct reason strings between them, which is a formal filter occupying the record type a real judgment would occupy. Nothing structural separates the two. The only thing that would have surfaced it is a consumer whose behaviour changes depending on which one it got.
I think this is the same shape, and some work since my last reply has made me more precise about what I'm trying to build.
I recently changed the promotion path so rejected Savepoint candidates are retained rather than disappearing. They go into an isolated
filtered_savepointstable, separate from the canonical Savepoints path, with a companion label table where they can later be classified ascorrect_rejection,false_rejection, orindeterminate.The important part for me is that this is plumbing for future evidence, not an attempt to manufacture evidence now.
Savepoints actually has the reverse problem at the moment. Before opening it up more broadly, there isn't enough organic volume. A lot of the filtering machinery is intended to prevent future spam, but right now I need to know whether the bar is suppressing good candidates. The first retained corpus was only three records, all
correct_rejection, so I treated that as insufficient evidence and made no rubric change.The goal is that when 20, 50, or 500 organic rejections exist, I can audit what happened rather than discover that all the rejected evidence was discarded. Keeping them isolated means I can preserve that evidence without treating rejected candidates as canonical Savepoints.
I think that sharpens your point about the review audit too.
It currently has a real consumer for structural integrity: it can fail on duplicate IDs, multiple closures, orphan closures, malformed records, and parse failures. But
referential_stability: unexerciseddoesn't yet have a consumer that changes behaviour. Nothing should be allowed to reinterpret that as "passed."Where I'm slightly more conservative is adding that consumer before one exists naturally. I don't want to create a dummy reference just to turn
unexercisedgreen, for the same reason I don't want to manufacture rejected candidates to validate the retention pipeline.So I think the contract I want is: preserve the evidence now, distinguish established from unexercised honestly, and make sure the plumbing exists to audit it once enough real evidence accumulates. When a real caller eventually depends on referential stability, that's where
unexercisedshould become a refusal rather than an informational status.The filtered-candidate work also gives the
npm testexample I mentioned earlier somewhere useful to go. The records were structurally valid and the classifier was behaving exactly as implemented, but the semantic judgment was wrong. Now rejected candidates don't simply disappear, so with enough organic evidence I can eventually audit whether the classifier or rubric is making the distinction I intended rather than infer that only from the records that survived it.Keeping the rejected candidates was the right call, and declining to manufacture a consumer for referential_stability was the right call twice over. Leaving the rubric alone after three correct_rejection labels respects what three records can carry. Retention buys the possibility of a later test. It does not buy the result.
The label table brings in one risk I would write down now, while the corpus is still small. If the classification comes from the same source that produced the rejection, a growing corpus starts to read as repeated independent confirmation while it is one judgment restated.
I have an embarrassing example in the log I work on. All four disputed cases in its history resolve to 26 judgments over 26 artifacts, and a single signing key stands behind every one of them. Disagreement between reviewers across the whole history is zero. The disputed flag came out of folding separate artifacts into one per-task tally, so it was never reviewer disagreement at all. Separately, the field labelled verifier count counts judgment rows and does not deduplicate signing keys. Across 67 tasks the row count exceeded the number of distinct keys, by as much as 14.
The cheap version of the fix is to keep the labeller's identifier on the label row and have the audit report distinct labellers next to row counts. Distinct identifiers still do not prove independence. They do make repetition visible, which is the part currently invisible.
That baseline only goes in before the volume arrives. At 500 labels you cannot recover attribution from the classifications.
When the retained corpus reaches 20, or 500, will its audit report classified rows, or classifications reached independently of each other?
The labels were more about the intended semantics of the design than a review process I've actually established yet.
Your question made me go back and check because there were three labels in the database that I didn't remember creating. They were from the original wiring/dogfood verification against synthetic candidates, not an organic corpus I'd started reviewing.
I've actually removed that labeling layer for now. The useful thing at this stage is retaining filtered candidates instead of throwing them away. Once there's enough organic data to make an evaluation meaningful, I'll decide what the review methodology needs to be.
So I don't have an answer yet on one classification per row vs independent classifications. That depends on what claim I'm trying to support when I eventually evaluate the filter. Your provenance point is useful though: whatever I do, the data needs to preserve who or what made the judgment rather than leaving that as an assumption.
Stopping at unsupported in Cloud rather than patching together an ID-normalization heuristic is the cleanest decision in this write-up. The moment a harness tries to bridge missing host boundaries with window timers or fuzzy conversation matching, you trade a clean failure for silent memory poisoning that takes weeks of agent runs to untangle.
Attribution drift in host hooks usually gets worse when the host starts streaming intermediate reasoning or multiplexing tool calls. We ran into almost the exact same recursion trap on shell tool hooks when an agent inspecting its own previous run logs triggered the execution watcher again and spawned runaway review tasks.
The escape hatch that worked for us when host lifecycles proved unreliable was inverting the boundary. Instead of trusting the host to tell the watcher what finished, we require the agent to flush a small manifest artifact to an append-only spool before yielding control. The review worker then compares that spool directly against git status in the repository root. If the host drops an event or swaps a generation identifier mid-stream, the untracked disk state remains the actual ground truth.
This is a useful inversion I hadn't considered. We already have the append-only logs, but not the step where the agent records what it did before yielding and a separate process checks that against the repo.
Cloud is important enough to how I work that I'm going to test this before moving on. The main thing I want to find out is whether it really gives us a more reliable source of truth, rather than just moving the same trust problem somewhere else.
This is also why I'm glad I stopped at
unsupportedinstead of trying to patch the IDs until they appeared to work. It left the failure visible enough to try a different approach. Thanks for sharing this.I ended up running this experiment, and the result was useful but more nuanced than "the manifest worked."
Moving execution identity into Savepoints did work as a reconciliation boundary. With repo evidence stamped to a Savepoints owned run ID, an explicit yield could open review without relying on Cursor's generation IDs lining up across hooks.
The interesting failure was attribution when executions overlapped. I tried a deliberately simple active run pointer for the checkout:
Once B became the active run, a later edit from A was silently attributed to B.
So this answered the question I had in my earlier reply about whether we'd just move the trust problem somewhere else. The manifest and yield boundary survived, but the pointer moved the remaining trust problem into evidence attribution.
I considered whether the pointer was still worth keeping because it fixed the isolated Cloud case, but its failure mode wasn't missing evidence. It was silently assigning evidence to the wrong execution, which is pretty much the memory poisoning failure you warned about. I retired it rather than keeping it as a partial Cloud workaround.
The narrower result I'm keeping is that agent owned execution identity plus explicit yield looks viable, but we still need a trustworthy way to associate evidence produced concurrently with the right execution.
Thanks again for the suggestion. It turned into a much better falsification test than I expected.
The strongest part here is the decision to treat lifecycle correlation as an evidence problem rather than an event-ordering problem.
A hook firing tells you that something happened; it doesn't necessarily tell you what completed, which repository it affected, or whether the event belongs to the work you're trying to reason about. That distinction becomes critical once agents can span multiple repositories or trigger their own follow-up turns.
I especially like the decision to make “unsupported” a valid adapter state. It's tempting to build correlation from timing, conversation IDs, or expected event sequences because it works in the happy path. The problem is that a false positive here can trigger semantic review for the wrong work, which is much harder to detect than simply declining to open the review.
The three-boundary model also gives a useful design principle for agent hosts: scaffolding identity, lifecycle correlation, and repository-local evidence should be independently provable. If one boundary can't be established, the system should degrade by capability rather than silently replacing the missing guarantee with a heuristic.
The repo-local boundary is the one I would keep as a hard gate. A generation_id can correlate events, but afterFileEdit under repoRoot is the evidence that gives a repository authority to open review. I would expose the Cloud case as an explicit capability state in telemetry, separate from no_capture, so dashboards do not count an unsupported adapter as a clean zero. That keeps the distinction observable without inventing a negative event. Did you end up recording unsupported per repository or only at the adapter level?
I ended up recording it per run within each repository, rather than only at the adapter level.
I also kept the repository boundary as the hard gate.
generation_idremained only for correlation; repository ownedafterFileEditevidence underrepoRootwas what gave that repository authority to open a review.I persisted
execution_hostwhen each run was registered. The resolver only made claims we had verified: Cursor Cloud was identified through the existing Cloud runtime heuristic, Desktop through explicit repository configuration, and anything we couldn’t identify remainedunknown.Infrastructure review capability was then derived from that persisted value: Desktop →
supported, Cloud →unsupported, unknown →indeterminate. That meant a Cloud run that never opened a review was distinguishable from a supported run that reached a genuineno_capture.That gave us the distinction you described without inventing a negative event or weakening the repository evidence gate.
The interesting reliability problem here is that missing lifecycle evidence can be harder to reason about than an explicit failure. If a review opportunity is never opened, the system needs to distinguish “nothing was worth capturing” from “the review boundary was never established.” That suggests agent infrastructure needs a notion of negative evidence: not just records of actions that happened, but observable states showing when an expected transition did not occur. This could make unsupported adapters much easier to monitor because the system could surface “review guarantee unavailable” as an operational state instead of leaving it indistinguishable from a normal no-capture outcome.
That's a useful distinction. Right now I'm preserving
no_captureas an explicit outcome when a review actually happens, but I'm not preserving the case where the review boundary couldn't be established in the first place. From the durable state, that second case is just an absence of records.I checked the implementation after reading this, and there's an interesting wrinkle: in some cases the adapter already records why it declined to open a review. For example, the stop boundary can observe that a generation wasn't registered. Today that only goes into probe telemetry.
But I think there's a trap here too. The reason I ended up with
unsupportedin the first place was that I couldn't establish enough lifecycle evidence to say confidently that a particular execution belonged to the work I wanted to review. If I can't prove that positive relationship, I also need to be careful claiming the negative one: “a review should have happened here, but didn't.”An
unregistered_generationtells me that correlation wasn't established. It doesn't necessarily tell me that this was an expected review transition that failed. Turning the latter into a durable outcome could recreate the same heuristic, just on the negative path.So I agree with the distinction you're drawing: “reviewed and nothing to capture” and “no review record exists” shouldn't be treated as equivalent. The part I haven't resolved is what negative state I can actually justify from the evidence the adapter has. Maybe the durable fact is initially only “correlation couldn't be established,” rather than the stronger “a review was expected and didn't happen.”
The point about separating correlation from actual evidence is really important. It’s tempting to assume that a lifecycle event means a specific task finished, but this shows why those assumptions can cause subtle bugs. Marking something as unsupported instead of relying on heuristics is a solid approach.
Probe 2's lesson generalizes nicely: event order is not causality, and I think most agent-harness bugs live exactly in that gap. Honestly the marker approach feels like the oldest reliable trick here — you own the identity of the turn (like a correlation ID you mint yourself), so you never have to ask an untrusted surface "was that mine?"
Good write-up, Michael. I really liked the decision to call Cloud unsupported rather than inventing a clever correlation hack. Knowing when the platform simply can't prove something is a useful feature in itself.
The stale guard story is the part I'd underline. "The next stop belongs to the review I just opened" is the kind of ordering assumption that survives every demo and breaks under real interleaving. I hit the same shape of bug with hook-driven automation: any guard that stays armed past its intended target eventually eats an unrelated event. Moving scaffolding identity in-band with the SAVEPOINTS_CAPTURE_REVIEW_V1 marker instead of trusting host provenance is the right call; metadata you don't control always lags the thing you do. Did generation_id stay stable across resumed turns, or does Cursor mint a fresh one on retry?
Good question. I went back through the probe evidence rather than assuming what Cursor does here.
On Desktop,
generation_idhas been stable acrossbeforeSubmitPrompt, the observe hooks, andstopwithin a single generation. Across separate prompt submissions, Cursor gives us a new ID. That includes the Savepoints review follow-up, which gets a differentgeneration_idfrom the source turn.That's fine for the review boundary now. The Savepoints owned
opportunity_idis carried directly in theSAVEPOINTS_CAPTURE_REVIEW_V1review prompt, so the review turn doesn't need to inherit the source turn'sgeneration_id. Closure also keys onopportunity_id, notgeneration_id.There is still an assumption one layer earlier, though. Within the source generation, the Desktop adapter uses
generation_idto correlate registration → evidence →stop. That's been stable in the Desktop runs I've observed, but it isn't a portable host contract. On Cloud I've already seen registration, observation, and stop carry different IDs within one apparent execution.I haven't captured a distinct Cursor "retry" lifecycle transition, so I don't want to invent semantics for that term. If by resumed turn you mean continuing the conversation with another prompt, the evidence I have says that's a new generation ID. If you mean a different Cursor retry/resume behavior, I'd be interested in which transition you're referring to.
So the narrower answer is: fresh ID for subsequent turns; stable across the ordinary Desktop generations I've observed. Your question does expose that the source execution still depends on Cursor preserving that identity, which is part of why I'm heading toward an execution identity that Savepoints controls rather than treating
generation_idas a reliable lifecycle contract across environments.the cloud vs desktop split is the real finding here. correlation working locally and quietly breaking once the host runs elsewhere is exactly the kind of thing that looks fine in testing and then silently drifts in prod. curious if you have any check that fires when generation_id stops matching, or if unsupported only shows up when you go looking for it.
You're right that the lifecycle path currently fails quietly. There isn't an automated check that fires when
generation_idcorrelation stops holding; I found the Cloud gap through an intentional lifecycle probe rather than monitoring.When that correlation fails, the lifecycle adapter doesn't open its review opportunity. The agent owned path still exists on Cloud, so completion instructions can still lead to
emit-learningwhen something is worth keeping.What's
unsupportedis the independent guarantee around that review: observing the work, opening a review opportunity, and recording whether it produced a learning or an explicitno_capture. Without that path, a zero capture result becomes ambiguous again: I can't distinguish "reviewed and found nothing" from "review never happened."So I think your prod drift point lands exactly there. The semantic path can keep working while the safety net that tells me whether review happened has quietly disappeared. Making that loss of guarantee observable is follow-on hardening.