containerd 2.2 shipped a mount manager: a service that can format a file as ext4 or xfs, attach it as a loopback device, and hand the result to a r...
For further actions, you may consider blocking this person and/or reporting abuse
The firstSystemMount panic is the sharpest find here. The write ordering is the real problem: the loopback attach happens before the BoltDB commit, so the panic leaves a device attached with no record anywhere to recover it from. The crash-recovery path only covers 'died after a successful Activate' — and the window between 'device attached' and 'record committed' is exactly the gap a metadata-store design like this normally exists to close, which makes that ordering the thing I'd want fixed before the index check.
Two notes for anyone hitting this in the wild: the orphan is still findable — losetup -a lists devices whose backing file is unlinked, and cross-referencing against the manager's targets directory tells you which one was the manager's. Not a fix, but it makes cleanup scriptable instead of a manual audit. And since the failure is deterministic (a one-mount list always indexes len(mounts)), validating the chain client-side — every mkfs/loop output must be consumed by a following handler mount — would turn this panic into a rejected request.
The ErrNotImplemented for a path outside the configured root deserves its own bug report, honestly. Fallback logic keyed on errdefs.IsNotImplemented() would misroute 'you forgot to allow this directory' into 'feature not built,' and retry policy for those two situations is usually the opposite. Same category, opposite semantics — that's the kind of error class that ages worst.
Yeah, the ordering point is sharper than how I framed it. I called it "there's no record to recover from," you're right that the actual bug is attach-before-commit, and fixing just the index panic wouldn't close that gap, anything else that dies in that same window hits the same orphan.
The losetup -a plus targets-dir cross-reference is a good one, didn't think to write that up. I've still got the one orphaned device from testing sitting there, might actually go try it.
Client-side validation matches what I put in the last section, good to hear it lines up independently.
Agreed on ErrNotImplemented too, that's a real footgun. Didn't file anything upstream yet, might turn both of these into an issue if I get to it.
Nice. Try it. On the recovery attempt: before losetup -d on the orphan, snapshot /proc/self/mountinfo; if the shim still holds an fd on that device, a plain detach bounces off EBUSY and you end up debugging the wrong layer. lsof on the loop node shows exactly which process pins it, and stopping that shim releases it deterministically. For upstream: the single-mount mkfs chain is a minimal repro on its own, and the targets-dir mismatch is the evidence a maintainer needs; short enough to paste in an issue. Curious whether your orphan survives a clean shim restart; if it does, that is index-vs-device divergence on the journal side.
I'm not running this through real containerd, just a standalone test harness (mmdemo) calling the mount manager package directly. But same idea applies, I left that process running when I found the orphan, so it's entirely possible it's still sitting on an fd for that loop device. Honestly hadn't even thought to check, I just went straight for losetup -d in my head without considering it might bounce.
On the restart question, closest thing I've got to "clean shim restart" is killing mmdemo and starting a fresh one against the same DB and target dir. My genuine guess, and this is a guess, not a result yet, is that it won't get reaped, since the whole bug is that no BoltDB record exists for this activation in the first place. The documented recovery path only knows to look for things it has a record of, so a fresh process doing its normal startup reconciliation should have literally nothing to find. If that's right, that's exactly the index-vs-device divergence you're describing, the device is real, the index has no idea it exists. Going to go find out instead of just asserting it though.
Appreciate you pushing for the minimal repro framing for upstream too, you're right that scoping the issue around the single-mount chain keeps it small enough that nobody can wave it off as "well your setup was weird."
Fair on both — filing the issue is the right call, but scope it to the invariant, not the symptom: "no mount may be committed unless its backing loop device is attached and recorded." If the issue is phrased as "index panic," you will get a patch that guards the panic and leaves the window open.
Two cheap tests that will keep the fix honest:
Kill the process between attach and commit, then restart: assert the reconciler reaps the device (or leaves it with a clear log line), and never silently adopts a stale mount.
Run that same test twice in a row. Orphan handling that works once often leaks on the second pass, because the first reap leaves state that does not match either set.
If you file it, drop the link here — happy to read the repro before it gets bikeshedded.
Yeah, "no mount may be committed unless its backing loop device is attached and recorded" is the better framing, and you're right that "index panic" as the title basically invites a patch that just wraps the slice access in a bounds check and calls it done. That fixes the crash, not the thing that caused the crash to matter.
Good news is both tests are cheap because the recovery path already exists, mmdemo -mode=crash-recover is literally the tool from the post, I just haven't pointed it at this specific failure mode yet. It currently only proves the happy path, died after a successful Activate, record exists, second process finds it and cleans up. What you're asking for is the uglier case: kill it between attach and commit, where there's nothing in BoltDB at all, and see whether the same recovery logic even notices there's a device to deal with. My honest guess is it won't, since the whole mechanism is built around reading records that exist, but that's exactly the kind of thing I'd rather run than assume.
The run-it-twice test is the one I wouldn't have thought to add on my own, and it's the sneakier bug class, something that looks fixed because the first pass happened to clean up correctly, not because the logic is actually sound. Going to go build both into the harness before I file anything, and yeah, I'll drop the link here once it's up, would rather have you look at the repro before it goes anywhere near a maintainer's queue.
plot twist: uptime green without a signed tip is still a vibe.
1 cut: when the outage ticket opens, can a buyer GET a queryable hop, or only another status page?
receipts > seals. #marker0528