Most of the agent-hosting products that turned up this year solve the same first problem, which is that nobody wants a model running rm -rf against their laptop, so the model gets a container somewhere else instead. That part has become commodity. The part that has not, and the part I wanted to check properly, is what happens to a long-running agent when you stop paying attention to it halfway through its work.
DigitalOcean's Managed Agents went into public preview recently, and the documentation makes a claim that is stronger than it first looks. Pausing a session, it says, preserves the processes, the memory and the workspace filesystem, and resuming brings them back. A stopped container loses everything that was not written to a volume, so if that sentence is literally true it is a different kind of thing, and I could not find anybody who had gone and tested it. So I spent a morning and about seven cents finding out.
Designing a test the filesystem cannot fake
The obvious version of this test is worthless. If you write a file, pause, resume, and read the file back, you have proven that a disk survived, which was never in doubt. The claim about memory needs something that lives only in memory and is never read back from anywhere.
So I wrote the dumbest possible process. A bash loop holding a counter in a shell variable, incrementing once a second, appending the current value and a timestamp to a log. The log is write-only from the process's point of view. Nothing ever reads it back, so if the session were destroyed and recreated with the filesystem restored, the counter would start again from one and the old log would simply have a new sequence appended to it.
#!/bin/bash
i=0
while true; do
i=$((i+1))
echo "$i $(date -u +%H:%M:%S)" >> /tmp/tick.log
sleep 1
done
Starting it needed a bit of care, because the exec channel into the sandbox kills its children when it closes. setsid nohup /tmp/tick.sh </dev/null >/dev/null 2>&1 & disown was what survived.
Then I let it run for about a minute, paused the session, went and made coffee, and resumed.
What came back
Two consecutive lines in the log, which is the whole result:
47 08:05:42
48 08:10:10
Forty-seven, then forty-eight, on the same process at PID 590 holding the same shell variable it had before, and the only evidence that anything happened at all is the four minute and twenty-eight second hole where a one-second tick should have been.
The pause call itself returned in 0.86 seconds and the resume in 1.16. Neither of those numbers is doing much work, since the interesting quantity is the four and a half minutes in between, during which the session was not consuming anything.
I ran the same check against the agent's own context rather than a shell variable. Before pausing, I had the agent generate a random ticket identifier, HARBOUR-7742, and write it into a file. After the resume I asked it what the ticket was called, and it answered from its conversation history in eight output tokens without touching the filesystem. Both kinds of state came back, the operating system's and the agent's.
The sandbox is a real machine
Worth confirming what the process was actually running on, since "sandbox" covers everything from a chroot to a VM.
Hypervisor detected: KVM
CPU: 2 vCPU Memory: 3939 MB Disk: /dev/vda Kernel: 6.1.176
A microVM with its own kernel, its own block device and hardware virtualisation underneath it, not a namespace on a shared host. Session creation from the API call to status READY took 15.97 seconds, which is slower than a container and about what a Firecracker-class VM costs you. Sizes run from mars-1vcpu-1gb up to mars-16vcpu-32gb.
The agent inside it was driven by DigitalOcean's own inference endpoint rather than a third-party key. A single prompt through DeepSeek v4 Pro came back in 7.3 seconds, 15,793 tokens in and 91 out, having written the file I asked for. The inference and the sandbox billing arrive on one account, which is a smaller convenience than the pause thing but not nothing.
Forking, which is where it got strange
There is a fork command, and I assumed it did what checkpoint-and-restore products usually do, which is snapshot the disk and give you a second sandbox with the same files.
It turns out not to. I forked the running parent twice, took 30.98 seconds for both, and then checked the ticker process in all three sandboxes.
| ticker PID | counter shortly after | counter a minute later | |
|---|---|---|---|
| parent | 590 | 177 | 244 |
| child 1 | 590 | 165 | 231 |
| child 2 | 590 | 162 | 228 |
Every one of them had the same process at the same PID, counting, from the value it held at the moment the fork was taken. Three copies of one running program, diverging from a common ancestor. If you have ever wanted to run an agent up to a decision point and then explore four different choices from exactly that state, without replaying the work that got you there, this is the primitive that does it.
Checkpointing separately took 25.25 seconds and reported a size of roughly 111 GB, which is the sparse allocation rather than anything you are storing.
Collecting the wall-clock cost of every operation in one place, since the spread between them is the thing that would shape how you use it:
Creating a session is the expensive one at 15.97 seconds. Pausing and resuming are close enough to instant that you would not build around them, which is what makes the pause worth reaching for in the first place.
Two things I would want stated plainly
Egress from a fresh sandbox is open, and I confirmed that by reaching both example.com and api.github.com from one without configuring anything at all. Adding a single host to the manifest flips the behaviour entirely:
egress:
- api.github.com
After that, api.github.com returned 200 and everything else returned nothing at all, which is the right design, since naming one host is an unambiguous statement that you want a deny-by-default posture. But an unconfigured session has a general-purpose language model with a shell and the open internet, and the default is the permissive one.
The other thing is that HARNESS_INFERENCE_API_KEY is readable as an ordinary environment variable from inside the sandbox, so anything running in there can print it. That is unavoidable if the process is going to call the inference endpoint, and the same is true of every runtime I know of, but it is the reason the egress allowlist matters more than it looks.
And this is a public preview in a single region without an SLA. Everything above is true of what shipped and none of it is a commitment about what will ship.
What I got wrong
Twice, and both times the same shape of mistake.
The first attempt at starting the ticker used bash -s < piped through the exec channel. The command reported success, the log file appeared with a few lines in it, and then it stopped, because the process died with the channel. I spent a while reading pause documentation for an answer to a problem that was not about pausing at all.
The second one was worse, because it produced a plausible number. I checked whether the checkpoint flag existed by running checkpoint create --name, got an error, and nearly wrote that checkpoints could not be labelled. The flag is --label. Three commands in this session refused an argument I had assumed from the shape of other tools, and in each case the refusal text was the thing that told me, which is an argument for quoting error output rather than paraphrasing it.
What it cost
Four sessions, two of them forks, one checkpoint of a 111 GB sparse image, a few dozen exec calls and one inference request. The account balance went from $4.40 to $4.33, so the whole morning came to seven cents.
I mention the figure because the thing that usually stops people testing a runtime properly is the fear of leaving something running, and at these prices the honest advice is to go and try it yourself rather than trust my numbers.
The manifests, the raw tick log and both charts are in a small repo if you want to reproduce it. Three commands is the whole of it:
doctl harness-runtime create -f agent.yaml
doctl harness-runtime pause <session-id>
doctl harness-runtime resume <session-id>
What to take from it
If you are evaluating any agent runtime, the pause claim is the one to test first and the one nobody tests, because the naive version of the test passes trivially. Put something in memory that is never written down, and see whether it is still counting on the other side.
And if it is, the operational consequences are larger than the feature description suggests. An agent that costs nothing while it is paused can wait for a human review instead of being torn down and rebuilt. An agent you can fork from a running state can be tried three ways from one expensive setup. Both of those change how you would structure a long-running job, and neither of them is the sort of thing you find out from a pricing page.


Top comments (4)
Love that you designed the test so a restored filesystem couldn't fake a pass — the write-only counter in a shell variable is the cleanest possible proof that it's live memory and not a volume replay. The
setsid nohup ... & disowndance is painfully familiar; the exec channel reaping its children is one of those things that silently kills "background" jobs on a lot of these sandboxes, and people blame the platform when it's really the session teardown.The part I'd be curious about: memory and processes surviving is the headline, but what happened to anything holding external state across the pause — open TCP connections, a socket to a DB, an in-flight HTTP request? A CRIU-style checkpoint can restore the process but the peer on the other end moved on four minutes ago. Did the fork behaving "stranger" have anything to do with duplicated file descriptors or PIDs? That's usually where these snapshot-resume systems get interesting (and where the sharp edges live for real agent workloads).
Pause and fork are invisible to the process. That is what makes them useful, and it is also what breaks anything the process holds that has a clock or an identity attached. Your counter shows it: nothing inside the VM noticed four and a half minutes go by.
For pause, the risk is leases. An agent that takes a 60-second lock or a short-lived token, gets paused for a human review, and then resumes will act as if it still holds both. Meanwhile the lock expired and another worker took it. It is the long GC pause problem, and the fix is the same one: a fencing token the resource checks, not a TTL the client trusts.
For fork, it is identity. HARBOUR-7742 now exists in three sandboxes. So does any idempotency key, request ID or nonce the agent generated before the fork, and each child will present it as its own. If two children go on to call the same payment or ticketing API, the provider sees one key arriving with two different payloads.
Two more rows in your table would show both: take a 60-second Redis lock before the pause and try a write with it after resume, and print a UUID generated before the fork from all three children.
The shell-variable counter is the right probe for a claim like this — one line, one number, no vendor dashboard in between. 47 to 48 across 4m28s says the process was parked, not kept warm, whatever the marketing page says about memory. What did forking do to the counter — duplicate it, or split the count in two?
"Resuming with its memory intact" is the whole problem, and I think the hardest part is one you get to skip in a sandbox: deciding what the resumed instance is allowed to assume about the world.
I built the same capability into a trading guard, from the opposite direction. It persisted its state so that a restart would not reset a daily limit - and the interesting bugs were never in the serialisation. They were in the assumptions:
The framing I have ended up using: the useful unit is not memory but a memory plus an expiry policy. Anything with a time component needs both halves written down at persist time, because the resumed process is a different process that happens to share a file.
Nice to see a write-up that treats it as a systems problem rather than a serialisation problem.