The expensive part of working with a coding agent is rarely the last step. It is all the steps before it: the exploring, the false starts, the files it wrote and rewrote until the thing finally worked. Once an agent has got somewhere good, the state it is in is worth more than any individual output, and the uncomfortable truth is that you usually cannot get back to it. Run the same prompt again and you get something else.
DigitalOcean's Managed Agents runs each session in its own microVM and lets you checkpoint that whole machine, fork it, and roll it back. I wanted to see what that looks like in practice, with something visual enough that the result is obvious from a screenshot. So I had an agent build a small web app, then branched it three ways and pointed a browser at every branch.
The app, seen through a tunnel
The agent was OpenCode running DeepSeek v4 Pro, and the brief was a status dashboard for a fictional coffee roastery: four metric cards and a table of recent batches, in a single HTML file with no external resources. It took 91.6 seconds and 29,084 tokens.
I started a plain python3 -m http.server inside the sandbox and reached it with the runtime's port forwarding:
$ doctl harness-runtime port-forward <session> 18101:8080
Forwarding 127.0.0.1:18101 -> port 8080 in session 01a10b67-...
Ready. Press Ctrl-C to stop.
That binds to localhost on my machine and nothing else. The help text is explicit that nothing inside the sandbox is exposed publicly, which is the right default for a box with an AI agent and a shell in it. Every screenshot in this post is a real browser, Chromium driven by Playwright, loading a page served from inside a microVM through one of those tunnels.
One checkpoint, three futures
With the dashboard working, I checkpointed the session and forked it three times from that checkpoint:
checkpoint create: 25.23s cp_33dcd0324ab9
fork x3 : 15.29s
The part I keep coming back to is what was already true when the three forks woke up. I had not started anything in them. I checked each one before giving it any instructions:
fork1: PID 593 | local GET -> 200 | index.html 6695 bytes
fork2: PID 593 | local GET -> 200 | index.html 6695 bytes
fork3: PID 593 | local GET -> 200 | index.html 6695 bytes
The web server was running in all three, at the same PID, serving the same file. A fork here is not a copy of a disk that you then boot. It is a copy of a running machine, including the processes that were alive inside it, and each copy carries on from the moment the checkpoint was taken.
Then each fork got a different brief, all three in parallel:
One was asked for a light, minimal Scandinavian design, one for a retro green CRT terminal, and one to keep the design and add a seven-day chart in inline SVG. They took between 25.6 and 46.7 seconds each. Throughout all of it, the original server process kept serving every edit as it landed, because it was reading the file from disk on each request.
One honest note about fork three. Its chart looks right at a glance and is wrong on inspection: the bar labelled 27 only reaches about 18 on the axis, and the one labelled 45 stops near 40. The agent got the scaling arithmetic slightly off. I left it in the screenshot because the post is about the sandbox rather than the agent, and because it is a useful reminder that a branch you like the look of still needs reading.
Why not just run the prompt again
This is the experiment that convinced me forking is about more than saving time. I started a completely fresh session and gave it the identical prompt, word for word, on the same model:
It built a different app. Light where the first was dark, a brown header bar the first did not have, different numbers throughout: 1,847 kg of beans in stock against 2,847, eight open orders against eighteen, batch #1047 where the first run had BR-1042.
The cost was different too, and in a way that undermined a claim I was about to make. The first build took 91.6 seconds and 26,656 input tokens. The second took 31.3 seconds and 2,558. I had been planning to write that a fork saves you a minute and a half of setup, and with only two samples the setup itself varied by a factor of three.
So the real argument for forking is not speed, although the fork was quick. It is that a fork gives you exactly the state you had, the app you were looking at and liked, while a rerun gives you a different app at a cost you cannot predict. If an agent spent an hour getting a repository into shape, that shape is what you want to branch from, and re-prompting is not a way to get it back.
Undo for a whole machine
The other half of the feature is rollback, so I played the part of an agent having a bad afternoon. On the parent session I killed the web server and deleted the app directory entirely. The tunnel stopped answering:
through the tunnel: connection failed
Then I rolled the session back to the checkpoint:
rollback call: 5.59s until READY: 6.04s
app dir: index.html
server : PID 593
GET : 200
The file came back, which I expected. What I did not expect was PID 593. The process I had explicitly killed was running again, restored from the memory in the checkpoint rather than restarted, and it was serving the same page.
To check it was really the same page and not something close, I compared the new screenshot with the original pixel by pixel. They differed in 88 pixels out of 792,000, every one of them inside a small status dot in the top corner. That dot pulses with a CSS animation, so the two screenshots simply caught it at different points in its cycle, and every other pixel on the page matched the original exactly.
What each step cost
Creating a session took 15.4 seconds. Checkpointing a running one took 25.2, which is the slowest of the platform operations and the one to plan around if you want to checkpoint frequently. Forking three ways from an existing checkpoint took 15.3, and rolling back took 5.6.
The whole experiment, five sessions including the three forks and the fresh rerun, cost 27 cents, from $3.26 to $2.99 on the account. That figure is mostly inference rather than sandbox time, and it is low enough that the honest advice is to try this on your own agent workflow rather than take my word for any of it.
What I got wrong
The first rollback I attempted failed. I tried to roll fork one back to the checkpoint I had forked it from, and got:
✗ Not found
checkpoint not found
Status 404 Not Found
Checkpoints belong to the session that created them. A fork inherits the state of its parent but not its checkpoint history, and listing checkpoints on the fork returned nothing at all. If you want to be able to undo inside a fork, checkpoint the fork first. It is a sensible design once you know it, and it is not obvious from the outside.
The second was a cleanup mistake. To close the tunnels I ran pkill -f "harness-runtime port-forward", which matched the command line of the very shell running it, killed that shell, and stopped the script before it reached the lines that removed the sessions. All five were still running when I checked. I have now made exactly this mistake twice, and the cure is the same both times: list the sessions afterwards rather than trusting the script to have finished.
The third is the one already described above. I had a sentence drafted about how much setup time a fork saves, built on a single build time, and the very next run of the same prompt took a third as long. The number that survived is the one about exactness.
What I would take from it
Most agent sandboxes are disposable boxes that you throw away when the job is done. What DigitalOcean has built is closer to version control for a running computer: you can checkpoint a machine with its processes alive, branch it into independent copies that carry on from that instant, and rewind a session after something goes wrong without rebuilding anything.
That changes how you would structure longer agent work. Get an agent to a good state once, checkpoint it, and explore several directions from the same starting point rather than hoping a fresh run lands somewhere similar. Put a checkpoint in front of anything risky so a bad step costs six seconds instead of the whole session. The caveats are real but manageable: it is a public preview in one region, checkpoints take around 25 seconds, and each one belongs to a single session. None of those changed what I saw, which was three different apps growing out of one running machine, and a process I had killed coming back at the same PID.



Top comments (1)
The PID surviving identically across all three forks is the detail that proves this is memory state rather than a restart, and it's worth calling out explicitly, because a reboot-and-replay would renumber and most readers won't notice the difference.
The useful follow-on is separating what a memory checkpoint buys over a disk snapshot. On disk you get files, installed dependencies, config, build cache and git state — which for a coding agent is most of the expensive part, since the 91.6 seconds and 29,084 tokens produced files, not processes. What memory adds is the running server, open sockets, warm JIT and anything already in RAM.
So the honest framing is that disk forking captures the work and memory forking captures the moment. For exploring branches of an agent's output, the work is usually what you wanted. For resuming something mid-flight, only the moment will do.
One risk the post doesn't cover that's worth sitting with: a live process holds whatever credentials were in memory when it was checkpointed. Three forks means three copies of live tokens and session keys. A disk snapshot has the same problem for anything written to a file, but a memory checkpoint also carries the things that were never meant to touch disk.
Did you check whether the forked boxes shared an identity, or did each get its own?