Bottom line first: when a deleted user avatar still loads, the object storage layer is almost never the culprit — the bytes are gone from the bucket, and what's still serving them is a CDN cache entry, the visitor's own browser cache, or a cached HTML document that carries the old asset URL. A purge call by itself won't settle it. The pattern that holds up is boring: content-addressed, versioned URLs for everything rendered into a page, plus a deletion path that ends in a verified 404 at the edge rather than a hopeful POST /purge.
The part nobody costs out is what that immutability does to the storage bill.
I ship image pipelines for a property management platform, and the primary consumer is not a human — it's a listing card. Every unit photo goes through background removal so appliances, furniture and signage sit on a uniform backdrop, and agent avatars ride the same pipeline. One upload fans out into six derived renditions. Multiply that by a portfolio of 40,000 photos and deletion stops being a one-row DELETE and becomes a fan-out problem with a cache bill attached.
Here's the flow in plain terms. An upload lands in object storage as an immutable original keyed by its SHA-256 digest; a worker built on libvips (ImageMagick is the other common choice, with a heavier memory profile per job) cuts the background and writes each rendition under a key derived from a hash of the source digest plus the transform parameters; a manifest row maps the entity to its renditions; the API emits rendition URLs; the CDN caches each one for a year with Cache-Control: public, max-age=31536000, immutable. Deletion has to unwind every one of those layers, in order, and the order matters more than the tooling.
How do you debug a deleted user avatar that still loads from the CDN cache?
Work down the stack, one layer per probe, instead of purging blindly and hoping.
curl -sI https://origin.internal/derived/9f2c8a1e.webp
curl -sI https://cdn.example.com/derived/9f2c8a1e.webp | grep -Ei 'age|cache'
curl -sI "https://cdn.example.com/derived/9f2c8a1e.webp?cachebust=1"
If the origin returns 200, the delete never reached storage and the CDN is innocent. If the origin 404s but the edge returns 200 with a non-zero Age, you're looking at a live cache entry — purge it, then check whether your provider runs a shield or second tier, because a single-PoP purge leaves the upper tier holding the object. If the cache-busting query 404s while the plain URL still resolves, that confirms it's an edge entry rather than a stale DNS or routing artifact. And if the asset only disappears once you tick "Disable cache" in devtools, the copy lives in the visitor's own browser, where you have no remote control whatsoever.
Two failure modes hide underneath that ladder. The first is that the document is cached, not the image: a JSON profile response or a server-rendered page still hands out the old URL, so the browser re-fetches an asset you believe is unreferenced. The second is stale-while-revalidate and stale-if-error, defined in RFC 5861 — those directives instruct the edge to keep serving the cached copy while it revalidates, and to keep serving it if the origin errors, which is exactly what a fresh 404 looks like to a conservative cache.
Also check your cache key. A Vary: Accept header, a resize query string and a signed-URL parameter each create separate entries, and a purge that names one URL clears one of them.
A deletion path that ends in a verified 404
Lead with the tombstone, finish with the probe. Everything in between is fan-out.
import hashlib
import os
import time
import boto3
import requests
CDN_BASE = os.environ["CDN_BASE"] # https://cdn.example.com
PURGE_ENDPOINT = os.environ["CDN_PURGE_ENDPOINT"] # provider-agnostic purge API
BUCKET = os.environ["ASSET_BUCKET"]
s3 = boto3.client("s3", endpoint_url=os.environ["S3_ENDPOINT"])
RENDITIONS = [("cutout", 96, "webp"), ("cutout", 192, "webp"), ("cutout", 384, "avif")]
def rendition_key(source_digest: str, kind: str, width: int, fmt: str) -> str:
# Deterministic keys: recomputable without the manifest, so an orphan sweeper
# can still find the blobs after the manifest row itself is deleted.
spec = f"{source_digest}|{kind}|{width}|{fmt}".encode()
return f"derived/{hashlib.sha256(spec).hexdigest()[:32]}.{fmt}"
def delete_avatar(db, user_id: str, source_digest: str) -> dict:
db.tombstone_avatar(user_id) # step 1: stop emitting the URL everywhere, first
keys = [rendition_key(source_digest, *spec) for spec in RENDITIONS]
keys.append(f"originals/{source_digest}")
s3.delete_objects(Bucket=BUCKET, Delete={"Objects": [{"Key": k} for k in keys]})
urls = [f"{CDN_BASE}/{k}" for k in keys]
requests.post(PURGE_ENDPOINT, json={"urls": urls}, timeout=10).raise_for_status()
return {"user_id": user_id, "purged": urls, "verified": verify_gone(urls)}
def verify_gone(urls: list[str], attempts: int = 6, delay: int = 5) -> bool:
pending = list(urls)
for _ in range(attempts):
pending = [
u for u in pending
if requests.head(u, headers={"Cache-Control": "no-cache"}, timeout=10).status_code != 404
]
if not pending:
return True
time.sleep(delay)
raise RuntimeError(f"still served at the edge after purge: {pending}")
The tombstone comes first because it's the only step that stops new references from being minted while the purge is in flight. Deleting the blob before the row leaves every cached document pointing at an object that no longer exists, and your error budget absorbs the difference. The retry loop in verify_gone exists because purges are asynchronous and not every edge honors a client-side no-cache on revalidation — as far as I can tell that behavior varies by provider, so a single probe proves nothing. Six attempts at 5 seconds is a starting point, not a law.
One more thing about that loop: it belongs in your test suite, not only in production. I keep a deletion eval that uploads a throwaway avatar, deletes it, and asserts a 404 through the real edge within 60 seconds, and it catches cache-key regressions long before a support ticket does.
What versioned URLs actually cost in storage and cache
Content addressing makes correctness cheap and storage expensive, which is the trade-off worth stating out loud. Every re-cut of a photo produces a new digest, so a background-removal model upgrade that touches a whole portfolio writes a second full generation of renditions while the first generation is still referenced by cached pages. You pay for both until a lifecycle rule reaps the old keys. On top of that, each new URL is a cold object at every PoP, so a re-render is also an origin egress event multiplied by your edge footprint — the cache-fill cost of immutability shows up on the bandwidth line, not the storage line, and it's the one people forget to model.
| URL strategy | Edge TTL that works | What deletion requires | Storage and fill cost |
|---|---|---|---|
| Content-addressed immutable path | up to a year | tombstone, purge, verify | two generations until lifecycle expiry; full cold fill per version |
Version query string (?v=3) |
up to a year, if the cache keys on the query | same, per variant | same, plus more cache keys to purge |
| Mutable path, short TTL | 30–60 seconds | wait out the TTL | one copy, higher revalidation traffic |
| Signed URL, short expiry | none, private by design | expiry does most of the work | one copy, no shared cache benefit |
Deterministic keys are what keep the storage side honest. If the manifest is the only record of which blobs exist, a failed delete halfway through leaves orphans nobody can enumerate, and you pay rent on them indefinitely; recomputing the key from the source digest plus transform parameters lets a nightly sweeper diff the bucket against live tombstones and reclaim whatever the online path missed.
The catch is that immutable URLs don't help for anything you can't rewrite. An avatar embedded in an email you already sent, a partner feed you don't control, an exported PDF lease packet — none of those documents can be updated, so a versioned URL there is a permanent reference to a specific generation. For those surfaces, stick with a mutable redirect endpoint on a 30–60 second TTL and accept the revalidation traffic. Immutability is also a poor fit for large video renditions, where keeping two generations of a multi-gigabyte encode dwarfs any purge-latency savings.
And the legal clock doesn't care about your TTL. GDPR Article 17 gives data subjects a right to erasure, so "the cache will expire within a year" is not an answer for a user photo; the purge and the verification are what make the deletion real, which is a good reason to log the verified 404 with a timestamp.
Running this without babysitting it
Treat deletion as a job with a result, not a fire-and-forget call. Every delete writes a record — tombstone time, keys targeted, purge request id, verification time — and anything that fails verification lands on a retry queue with the same idempotent payload, since re-deleting an already-deleted key is harmless and re-purging a cold URL costs a request. Alert on the age of the oldest unverified deletion rather than on the raw failure count; a single transient purge error is noise, while a deletion that stayed unverified for two hours means either the shield tier isn't in your purge scope or someone shipped a new cache key. Watch the orphan sweeper's reclaim volume as a cost signal — a sudden jump usually means the online delete path silently stopped fanning out over all six renditions. Keep the eval in CI against a staging distribution with the same cache configuration as production, because a deletion path validated against an origin with no CDN in front of it is validated against the wrong system.
None of this requires a specific vendor. It requires that deleting an asset be a state machine with a terminal state you can observe.
Sources
- MDN — Image file type and format guide: https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
- MDN — HTTP caching: https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Caching
- RFC 9111 — HTTP Caching: https://www.rfc-editor.org/rfc/rfc9111.html
- RFC 8246 — HTTP Immutable Responses: https://www.rfc-editor.org/rfc/rfc8246.html
- RFC 5861 — HTTP Cache-Control Extensions for Stale Content: https://www.rfc-editor.org/rfc/rfc5861.html
- Amazon S3 API — DeleteObjects: https://docs.aws.amazon.com/AmazonS3/latest/API/API_DeleteObjects.html
- libvips documentation: https://www.libvips.org/
- GDPR Article 17 — Right to erasure: https://gdpr-info.eu/art-17-gdpr/
Top comments (0)