An image import becomes expensive the moment a user has to guess what happened to 4,000 files. The useful design is a resumable batch with explicit checkpoints, a rate-limit-aware worker, and a cancellation record that can distinguish “stop now” from “this file was rejected.”
Short answer: treat a batch as a durable state machine, not a loop around an upload call. Give every image an idempotency key, persist progress after each state transition, and make cancellation a first-class command that can safely undo only work which is still reversible.
That choice matters in a small SaaS. I care about revenue per hour, and a support ticket about a half-imported library burns more of it than a slightly slower job. The system should let me ship weekly while outsourcing the undifferentiated plumbing to a queue and a database.
Why “batch” is a product boundary, not a loop
The word batch hides several different contracts. A user may mean “upload these files together,” “moderate them together,” or “show me one progress bar.” Those are not the same operation. A 429 response from an image-analysis service, a corrupt JPEG, and a browser tab closing each produce different recovery work.
I model the import as a parent record plus one child record per image. The parent owns a stable batchId; each child owns a digest of the source bytes and a sequence number. The digest prevents a retry from creating a second logical image when the transport repeats a request. The sequence number gives the UI deterministic ordering even when workers finish out of order.
Progress is a count of durable transitions, not a count of promises resolved in memory. If 37 of 100 children reached stored, the API can report 37 even after a deploy. A separate failed count keeps a transient retry from looking like a permanent rejection.
The state names should be boring: queued, fetching, stored, rejected, cancel_requested, and cancelled. Boring names make dashboards and SQL queries useful. They also make an audit trail possible when a customer asks why an image never appeared.
How should image batches handle rate limits, progress, and cancellation for wrong imports?
Start with a small state transition table. The worker is allowed to move a row forward only when the current version matches the version it read. That optimistic check stops two workers from both claiming the same image.
| State | Durable fact | Next action |
|---|---|---|
queued |
Source reference and digest exist | Claim work with a lease |
fetching |
A worker owns the lease | Fetch and validate bytes |
stored |
Object and metadata are committed | Increment completed progress |
rejected |
Validation or policy decision is final | Show a reason, do not retry |
cancel_requested |
The user asked to stop | Finish the safe boundary, then cancel |
cancelled |
No more work will start | Retain the audit record |
Rate limiting belongs at the worker boundary. A token bucket or leaky bucket can cap concurrency, but it cannot decide whether a response is retryable. Retry transport timeouts and 429 responses with bounded exponential backoff. Do not retry a malformed image forever; that is a data problem, not capacity pressure. Honor a Retry-After value when the upstream provides one, and add jitter so a fleet does not wake up on the same second.
Here is the smallest worker shape I use. It is intentionally provider-neutral: the fetchImage and storeImage functions sit behind interfaces, so changing a backend does not change the batch state machine.
type Status =
| "queued"
| "fetching"
| "stored"
| "rejected"
| "cancel_requested"
| "cancelled";
type Item = {
id: string;
batchId: string;
digest: string;
status: Status;
version: number;
};
async function processItem(item: Item): Promise<void> {
const claimed = await db.claim(item.id, item.version, "fetching");
if (!claimed) return; // Another worker won the lease.
try {
const bytes = await fetchImage(item.digest);
validateImage(bytes); // Check type and size before storage.
await storeImage(item.digest, bytes);
await db.transition(item.id, "fetching", "stored");
} catch (error) {
if (isRetryable(error)) {
await db.releaseForRetry(item.id, retryDelay(error));
return;
}
await db.transition(item.id, "fetching", "rejected");
}
}
The comment is the important part: a failed claim is normal under concurrency. It is not an error to log at warning level for every item. Good observability separates expected contention from a growing retry queue.
Cancellation needs a boundary. If the user imports the wrong directory and clicks cancel, mark the batch cancel_requested immediately. A worker checks that flag before fetching the next item and after a retry delay. It does not pretend an in-flight network request can always be interrupted. Once the current item reaches a safe state, the worker marks it cancelled or leaves a committed stored object alone, depending on the product's retention policy.
For a wrong import, “undo” should be explicit. Keep a manifest of object keys created by this batch, then offer a compensating delete job with its own audit entries. Do not issue a broad “delete everything from this user” command. The manifest is the difference between a reversible import and a frightening support script.
Ship the boring part.
The first time I test this flow, I use a deliberately wrong folder: a few valid PNGs, one file renamed to .jpg, a duplicate digest, and a cancellation click while the queue is waiting on backoff. That small fixture exposes the failure modes that a happy-path demo hides. The renamed file should become rejected with a reason tied to the detected media type; the duplicate should attach to the existing child record rather than create a second object; the waiting item should observe cancel_requested before its next attempt; and the already-stored files should remain discoverable through the manifest. If the UI reports “cancelled” while a worker can still start a new fetch, the state model is lying. If the cleanup job deletes by user prefix instead of by manifest key, a typo can remove an unrelated import. I would rather spend an hour making those transitions visible in a test database than spend a week explaining an ambiguous progress bar to customers.
What I would change before this runs at scale
The first version can use a relational table, a queue with visibility timeouts, and object storage. At higher volume, I would add leases with a heartbeat, a dead-letter stream for records that exceed a retry budget, and a reconciliation job that compares the manifest with stored objects. Those additions cost operations time, so I would wait until queue age or recovery time is a measured problem.
The progress endpoint should return a snapshot, not a stream that the browser must keep alive. A client can poll with If-None-Match, or subscribe to events when the UI truly needs live updates. Either way, compute progress from committed child states. A progress bar that reaches 100% before metadata is durable is worse than no progress bar.
Validation is another scale boundary. Read the file signature, not just the filename extension. The browser may call a file .jpg while the bytes are a different format, and image formats have different decoding and color-management behavior. MDN's image format guide is a useful map of those trade-offs; the storage contract should record the detected media type alongside the original name.
I am not sure every pipeline needs per-pixel normalization before storage. Your mileage may vary. If downstream moderation requires a canonical color space, normalize once and record that transform; if the product only needs thumbnails, delaying full-resolution decoding can keep the queue cheaper and easier to drain.
The trade-offs I keep visible
This design favors recoverability over minimum latency. That is usually right for customer imports, but it is not suitable when the job is disposable telemetry or a one-second preview. For those, an in-memory worker and a time limit may be the honest choice.
The database manifest also creates retention work. Keeping every source digest and cancellation event helps explain a dispute, but it increases privacy obligations. Set a retention period, encrypt sensitive metadata, and make deletion of the audit record a documented policy rather than an accidental cascade.
A queue is not a moderation policy. It can enforce rate limits and preserve order; it cannot tell you whether a generated image is acceptable. Keep policy decisions in a versioned service or ruleset, store the decision with the item, and make a policy change a new batch or an explicit reprocess. That separation lets the import machinery stay stable while the product's safety rules evolve.
The practical decision rule is simple: if a user can identify the wrong import after it starts, build checkpoints and a manifest before optimizing throughput. If the work has no user-visible side effect, keep the machinery lighter. Both choices are defensible. The mistake is promising cancellation without defining what has already become permanent.
Top comments (0)