containerd 2.2 shipped a mount manager: a service that can format a file as ext4 or xfs, attach it as a loopback device, and hand the result to a runtime, all from a single Activate call instead of the usual truncate, mkfs, losetup, mount sequence. I wanted to know whether that call is actually faster than doing it by hand, so I wrote a small Go program against the manager's package directly and ran both paths three times each on the same machine.
The manual sequence came out at 23.5 to 25.9 milliseconds. The mount manager's Activate came out at 29.6 to 47.2 milliseconds. It was not faster. Along the way it also panicked once, leaked a raw BoltDB error once, and left a loop device attached with nothing able to find it again.
What the mount manager actually is
There's no ctr subcommand for any of this. The manager lives at github.com/containerd/containerd/v2/core/mount/manager and is meant to be embedded by a snapshotter or a runtime shim, not driven from a terminal. Its job is to let a mount type be built out of steps: a "transformer" can create a file, format it, and format directories, and a "handler" can attach it as a loopback device, before the result gets handed off as an ordinary system mount.
The container running this test was Docker Engine 29.3.1, whose bundled containerd reports itself as v2.2.2. I confirmed the mount manager package is at that exact version by pinning it in go.mod and building against it, not by trusting the daemon's version string.
A working activation for a 200MiB ext4 image looks like this, once I had the templating right:
mounts := []mount.Mount{
{
Type: "mkfs/loop",
Source: imgPath,
Options: []string{
"X-containerd.mkfs.size=200MiB",
"X-containerd.mkfs.fs=ext4",
},
},
{
Type: "format/ext4",
Source: "{{ mount 0 }}",
},
}
info, err := mgr.Activate(ctx, "demo1", mounts)
Activate returns two sets of mounts: info.Active, the ones it handled itself (the loopback attach), and info.System, the ones it expects the caller to mount with the ordinary mount(2) syscall (the ext4 filesystem on that loop device). The manager does not mount your rootfs for you. You still call Mount() on whatever comes back in info.System.
The speed comparison
Both paths format a 200MiB ext4 image, attach it as a loopback device, and mount it. I timed the manager's Activate call and, separately, the same steps run as four shell commands: truncate -s 200M, mkfs.ext4 -q -F, losetup -f --show, mount.
| Path | Run 1 | Run 2 | Run 3 |
|---|---|---|---|
Mount manager Activate
|
29.6ms | 39.0ms | 47.2ms |
| Manual (truncate + mkfs.ext4 + losetup + mount) | 24.4ms | 23.6ms | 25.9ms |
The manual path was faster on every run. That is not a criticism of the manager's design so much as an observation about what it's for: I checked the source of the mkfs transformer (core/mount/manager/mkfs.go) and it shells out to the real mkfs.ext4 and mkfs.xfs binaries with exec.CommandContext, exactly what the manual path calls directly. There's no custom fast-path formatter underneath. The manager's job is composability across snapshotters, not raw speed, and on this measurement it costs a small constant overhead (BoltDB writes, symlink creation, bookkeeping) rather than removing any.
Cleanup showed the same pattern: Deactivate plus umount took 32 to 45ms; umount plus losetup -d by hand took 11 to 13ms.
Disk usage was identical either way. du --apparent-size on both images read 200M; actual usage was 17M on both, since neither mkfs.ext4 nor truncate writes real data to most of the file.
What I got wrong on the way
My first version of the manual comparison used dd if=/dev/zero of=disk.img bs=1M count=200 to create the image before formatting it, because that's the version of this recipe I'd actually run in the past. It came out at 358ms to 1.6 seconds; against that, the mount manager looked ten to fifty times faster, which would have been the headline of this post.
dd was writing 200MB of real zero bytes to disk before mkfs.ext4 ever ran. truncate -s 200M creates the same size file as a sparse hole in under a millisecond, and mkfs.ext4 doesn't need the data pre-zeroed, it only writes its own metadata. The mount manager's mkfs transformer already does the equivalent of truncate, via os.OpenFile plus f.Truncate(size), then calls the real mkfs.ext4 binary on the result. Once I made the manual comparison do the same thing, the ten-times "win" disappeared and mildly reversed. The lesson: when a new API bundles three steps into one call, benchmark it against the shortest correct version of those three steps, not the version you happen to type from muscle memory.
Under concurrency
Ten goroutines calling Activate in parallel against the same manager instance, each formatting its own 50MiB image, completed in 70.1ms of wall time, with individual calls ranging from 30.5 to 69.8ms. All ten succeeded. If activations were serialized behind a lock, ten of them would have taken close to 300-400ms; they didn't, so the manager's internal locking (RLock during normal activation, held exclusively only during garbage collection) does allow real concurrent formatting.
What it refuses
An unsupported filesystem type is rejected before any file gets created:
unsupported filesystem "btrfs": invalid argument
A missing size option is also rejected cleanly:
mkfs requires mkfs.size option: invalid argument
A path outside the manager's configured root is rejected too, but with a misleading error class:
no root "/tmp/not-the-root/disk.img" configured for mkfs: not implemented
That comes back as errdefs.ErrNotImplemented, the same error category containerd uses for "this operation genuinely doesn't exist here." Code that checks errdefs.IsNotImplemented() to decide whether to fall back to a different mount path would treat "you forgot to allow this directory" the same as "this feature isn't built." I read the source (core/mount/manager/mkfs.go) to confirm this isn't a formatting quirk on my end; the transformer returns exactly that wrapped error whenever the source path doesn't match any configured root.
Activating a second time under the same name, without deactivating the first, doesn't get a clean "already exists" either:
bucket already exists
That's a raw bbolt error surfacing straight from the metadata store, with no containerd-level wrapping. It's accurate, but it tells you about the manager's storage engine rather than about your mistake.
The panic
The documentation's own examples always chain at least two mounts: something that produces a loopback device, and something that mounts a filesystem on it. I tried activating a single mkfs/loop mount on its own, with nothing consuming its output, expecting either a successful format-only activation or a clean validation error.
panic: runtime error: index out of range [1] with length 1
goroutine 1 [running]:
github.com/containerd/containerd/v2/core/mount/manager.(*mountManager).Activate(...)
.../core/mount/manager/manager.go:421 +0x2030
I read manager.go to find out why. The function tracks firstSystemMount, the index of the first mount it expects the caller to handle. When a mount is only ever a transform target (my single mkfs/loop mount, with no second mount to hand the loop device to), that index gets set to i+1, which in a one-mount list equals len(mounts). A later loop indexes into mountConv[firstSystemMount] to apply any pending format templating, and mountConv was allocated with len(mounts) elements. Index 1 into a slice of length 1 panics.
This matters beyond the crash itself. The loopback device gets attached to the backing file before the panic point, and because Activate never returns successfully, nothing gets written to the manager's BoltDB. There is no record of this activation to recover.
Crash recovery, and where it doesn't reach
The manager persists activation state in BoltDB specifically so a restarted process can find and clean up mounts from before a crash. I tested that separately from the panic: one process called Activate and then exited hard with os.Exit(0), skipping Deactivate entirely, to simulate a daemon that died mid-operation.
$ ./mmdemo -mode=crash-activate
activated in 30.7ms, err=<nil>
$ ./mmdemo -mode=crash-recover
List() after simulated crash: 1 activations, err=<nil>
recovered activation: crashy active=[{loop ... /tmp/mm-crash/targets/1/1}]
cleanup via Deactivate: <nil>
A second process, pointed at the same database and target directory, listed the orphaned activation and deactivated it cleanly, and the loop device it had been holding was released. That documented claim held up.
But it only works because the first Activate call returned successfully and got committed. Compare that against the one-mount panic above: I left that test running separately, and the loop device it opened is still attached to a deleted backing file with no BoltDB record anywhere pointing at it. losetup -a still shows it. No amount of restarting a mount manager pointed at any database will find it, because it was never written down. The crash-recovery mechanism protects against a daemon dying after an activation completes. It has no way to protect against the daemon crashing during one.
xfs, and a size limit that isn't the manager's
I ran the same mkfs/loop chain with X-containerd.mkfs.fs=xfs at 200MiB and got a full mkfs.xfs usage message back:
mkfs.xfs failed: Filesystem must be larger than 300MB.
That's mkfs.xfs itself refusing, not the manager. At 400MiB, formatting succeeded in 48 to 72ms. Mounting the result failed in this environment specifically:
mount source: ".../targets/1/1", target: ".../rootfs", fstype: xfs, flags: 0, data: "", err: no such device
/proc/filesystems on this container's kernel has no xfs entry and there's no modprobe to load one. That's a property of the sandbox this test ran in, not of containerd, and I'm noting it rather than counting it as a finding against the feature.
Run it yourself
This needs Go 1.24+, root (for loopback devices and mounts), and mkfs.ext4.
mkdir mounttest && cd mounttest
go mod init mounttest
go get github.com/containerd/containerd/v2@v2.2.2
Save as main.go:
package main
import (
"context"
"fmt"
"os"
"github.com/containerd/containerd/v2/core/mount"
"github.com/containerd/containerd/v2/core/mount/manager"
"github.com/containerd/containerd/v2/pkg/namespaces"
bolt "go.etcd.io/bbolt"
)
func main() {
base := "/tmp/mm-demo"
os.RemoveAll(base)
os.MkdirAll(base+"/targets", 0755)
db, _ := bolt.Open(base+"/meta.db", 0644, nil)
defer db.Close()
mgr, err := manager.NewManager(db, base+"/targets",
manager.WithMountHandler("loop", mount.LoopbackHandler()),
manager.WithAllowedRoot(base),
)
if err != nil {
panic(err)
}
ctx := namespaces.WithNamespace(context.Background(), "default")
mounts := []mount.Mount{
{Type: "mkfs/loop", Source: base + "/disk.img", Options: []string{
"X-containerd.mkfs.size=200MiB",
"X-containerd.mkfs.fs=ext4",
}},
{Type: "format/ext4", Source: "{{ mount 0 }}"},
}
info, err := mgr.Activate(ctx, "demo1", mounts)
fmt.Printf("info=%+v err=%v\n", info, err)
os.MkdirAll(base+"/rootfs", 0755)
for _, sm := range info.System {
if err := sm.Mount(base + "/rootfs"); err != nil {
fmt.Println("mount failed:", err)
}
}
// activating the same name again without deactivating first:
_, err = mgr.Activate(ctx, "demo1", mounts)
fmt.Println("second activate, same name:", err)
}
go build -o mmdemo .
sudo ./mmdemo
mount | grep mm-demo
sudo umount /tmp/mm-demo/rootfs
I verified every command in this section on a fresh checkout before publishing.
What to do with this
If you're building on the mount manager today, treat it as a composability primitive, not a performance one; it won't beat a shell one-liner on latency. Never construct a mount list that ends in a transform-only type like mkfs/loop without a mount that actually consumes its output, until this specific crash is fixed upstream. Don't branch on errdefs.IsNotImplemented() from this package without also checking the error text, because it currently covers both "unsupported" and "not configured." And if you're relying on its crash recovery for anything in production, test it against a process that dies mid-Activate, not just one that dies after — those are different guarantees, and only one of them is covered right now.
Top comments (11)
The firstSystemMount panic is the sharpest find here. The write ordering is the real problem: the loopback attach happens before the BoltDB commit, so the panic leaves a device attached with no record anywhere to recover it from. The crash-recovery path only covers 'died after a successful Activate' — and the window between 'device attached' and 'record committed' is exactly the gap a metadata-store design like this normally exists to close, which makes that ordering the thing I'd want fixed before the index check.
Two notes for anyone hitting this in the wild: the orphan is still findable — losetup -a lists devices whose backing file is unlinked, and cross-referencing against the manager's targets directory tells you which one was the manager's. Not a fix, but it makes cleanup scriptable instead of a manual audit. And since the failure is deterministic (a one-mount list always indexes len(mounts)), validating the chain client-side — every mkfs/loop output must be consumed by a following handler mount — would turn this panic into a rejected request.
The ErrNotImplemented for a path outside the configured root deserves its own bug report, honestly. Fallback logic keyed on errdefs.IsNotImplemented() would misroute 'you forgot to allow this directory' into 'feature not built,' and retry policy for those two situations is usually the opposite. Same category, opposite semantics — that's the kind of error class that ages worst.
Yeah, the ordering point is sharper than how I framed it. I called it "there's no record to recover from," you're right that the actual bug is attach-before-commit, and fixing just the index panic wouldn't close that gap, anything else that dies in that same window hits the same orphan.
The losetup -a plus targets-dir cross-reference is a good one, didn't think to write that up. I've still got the one orphaned device from testing sitting there, might actually go try it.
Client-side validation matches what I put in the last section, good to hear it lines up independently.
Agreed on ErrNotImplemented too, that's a real footgun. Didn't file anything upstream yet, might turn both of these into an issue if I get to it.
Nice. Try it. On the recovery attempt: before losetup -d on the orphan, snapshot /proc/self/mountinfo; if the shim still holds an fd on that device, a plain detach bounces off EBUSY and you end up debugging the wrong layer. lsof on the loop node shows exactly which process pins it, and stopping that shim releases it deterministically. For upstream: the single-mount mkfs chain is a minimal repro on its own, and the targets-dir mismatch is the evidence a maintainer needs; short enough to paste in an issue. Curious whether your orphan survives a clean shim restart; if it does, that is index-vs-device divergence on the journal side.
I'm not running this through real containerd, just a standalone test harness (mmdemo) calling the mount manager package directly. But same idea applies, I left that process running when I found the orphan, so it's entirely possible it's still sitting on an fd for that loop device. Honestly hadn't even thought to check, I just went straight for losetup -d in my head without considering it might bounce.
On the restart question, closest thing I've got to "clean shim restart" is killing mmdemo and starting a fresh one against the same DB and target dir. My genuine guess, and this is a guess, not a result yet, is that it won't get reaped, since the whole bug is that no BoltDB record exists for this activation in the first place. The documented recovery path only knows to look for things it has a record of, so a fresh process doing its normal startup reconciliation should have literally nothing to find. If that's right, that's exactly the index-vs-device divergence you're describing, the device is real, the index has no idea it exists. Going to go find out instead of just asserting it though.
Appreciate you pushing for the minimal repro framing for upstream too, you're right that scoping the issue around the single-mount chain keeps it small enough that nobody can wave it off as "well your setup was weird."
Fair on both — filing the issue is the right call, but scope it to the invariant, not the symptom: "no mount may be committed unless its backing loop device is attached and recorded." If the issue is phrased as "index panic," you will get a patch that guards the panic and leaves the window open.
Two cheap tests that will keep the fix honest:
Kill the process between attach and commit, then restart: assert the reconciler reaps the device (or leaves it with a clear log line), and never silently adopts a stale mount.
Run that same test twice in a row. Orphan handling that works once often leaks on the second pass, because the first reap leaves state that does not match either set.
If you file it, drop the link here — happy to read the repro before it gets bikeshedded.
Yeah, "no mount may be committed unless its backing loop device is attached and recorded" is the better framing, and you're right that "index panic" as the title basically invites a patch that just wraps the slice access in a bounds check and calls it done. That fixes the crash, not the thing that caused the crash to matter.
Good news is both tests are cheap because the recovery path already exists, mmdemo -mode=crash-recover is literally the tool from the post, I just haven't pointed it at this specific failure mode yet. It currently only proves the happy path, died after a successful Activate, record exists, second process finds it and cleans up. What you're asking for is the uglier case: kill it between attach and commit, where there's nothing in BoltDB at all, and see whether the same recovery logic even notices there's a device to deal with. My honest guess is it won't, since the whole mechanism is built around reading records that exist, but that's exactly the kind of thing I'd rather run than assume.
The run-it-twice test is the one I wouldn't have thought to add on my own, and it's the sneakier bug class, something that looks fixed because the first pass happened to clean up correctly, not because the logic is actually sound. Going to go build both into the harness before I file anything, and yeah, I'll drop the link here once it's up, would rather have you look at the repro before it goes anywhere near a maintainer's queue.
That orphan-cleanup path is the right next step, and the losetup -a cross-reference will tell you quickly whether it's recoverable or whether you're just going to rm the loop device and recreate the mount.
One thing I'd watch: if the process died between attach and commit, the device may still be referenced by the kernel but absent from the targets dir, so a naive sweep that only trusts the BoltDB records will miss it. A cheap guard is to reconcile the two sets on startup (loop devices vs committed mounts) and only reap the symmetric difference, with a short grace window so an in-flight attach isn't mistaken for an orphan.
Curious what you find with the existing orphan, a real repro beats the theoretical ordering any day.
Yeah, that matches what I'm seeing. The targets dir entry exists for this one (mkfs runs before the device gets attached, so the directory structure's already there), it's specifically the BoltDB commit that never happened, which is exactly the "kernel knows, index doesn't" gap you're describing. A sweep that only trusts BoltDB would walk right past this device and call the cleanup done.
The reconcile-the-two-sets idea is the right shape for a real fix too, loop devices on one side, committed mounts on the other, reap only what's in the device set and missing from the record set. The grace window is the detail I wouldn't have thought to add and definitely needed, without it you'd be racing your own in-flight attaches and reaping something that's two milliseconds from being committed just fine.
Going to go actually look at the orphan now instead of describing it secondhand, losetup -a plus the targets dir cross-reference like you said in the first comment, see whether it's cleanly recoverable or whether I'm just going to rm the loop device and eat the rebuild. Real repro over theoretical ordering, agreed, will report back with whatever I actually find instead of what I assume I'll find.
That distinction — targets dir present, BoltDB commit missing — is the exact shape the reconciler needs to handle, and this orphan is already your first real fixture for it. One shortcut for the kernel side before any rm: /sys/block/loop*/loop/backing_file shows whether the kernel still holds a live reference, independent of both the targets dir and BoltDB. Backing file present = the attachment is real and rm-ing the loop device is the whole fix; empty = you're in pure 'index forgot it' territory. Two files, one glance each. Looking forward to the repro — either answer is useful: a clean recovery path, or a device state nobody's filed yet.
That's a genuinely better first check than what I had planned, cuts straight to kernel ground truth instead of me inferring it from two other systems that might both be wrong in the same direction. Two files, one glance each, yeah, that's exactly the kind of shortcut I should've reached for before overthinking it with lsof and mountinfo.
Going to go read /sys/block/loop*/loop/backing_file for the orphaned device right now before I touch anything else. If it's non-empty, even pointing at the deleted path, that's your "attachment is real" case and rm plus a clean re-run is the whole story. If it's empty, that's the more annoying one, a device state with nothing left anywhere that explains it, which is probably the one actually worth filing on its own.
Either way beats guessing, appreciate you narrowing it down to something this cheap to check.
plot twist: uptime green without a signed tip is still a vibe.
1 cut: when the outage ticket opens, can a buyer GET a queryable hop, or only another status page?
receipts > seals. #marker0528