DEV Community

ValerianBlack3895
ValerianBlack3895

Posted on

Why Image Batches Exist: Node.js Rate Limits, Progress, and Wrong-Import Cancellation

An image import becomes expensive the moment a user has to guess what happened to 4,000 files. The useful design is a resumable batch with explicit checkpoints, a rate-limit-aware worker, and a cancellation record that can distinguish “stop now” from “this file was rejected.”

Short answer: treat a batch as a durable state machine, not a loop around an upload call. Give every image an idempotency key, persist progress after each state transition, and make cancellation a first-class command that can safely undo only work which is still reversible.

That choice matters in a small SaaS. I care about revenue per hour, and a support ticket about a half-imported library burns more of it than a slightly slower job. The system should let me ship weekly while outsourcing the undifferentiated plumbing to a queue and a database.

Why “batch” is a product boundary, not a loop

The word batch hides several different contracts. A user may mean “upload these files together,” “moderate them together,” or “show me one progress bar.” Those are not the same operation. A 429 response from an image-analysis service, a corrupt JPEG, and a browser tab closing each produce different recovery work.

I model the import as a parent record plus one child record per image. The parent owns a stable batchId; each child owns a digest of the source bytes and a sequence number. The digest prevents a retry from creating a second logical image when the transport repeats a request. The sequence number gives the UI deterministic ordering even when workers finish out of order.

Progress is a count of durable transitions, not a count of promises resolved in memory. If 37 of 100 children reached stored, the API can report 37 even after a deploy. A separate failed count keeps a transient retry from looking like a permanent rejection.

The state names should be boring: queued, fetching, stored, rejected, cancel_requested, and cancelled. Boring names make dashboards and SQL queries useful. They also make an audit trail possible when a customer asks why an image never appeared.

How should image batches handle rate limits, progress, and cancellation for wrong imports?

Start with a small state transition table. The worker is allowed to move a row forward only when the current version matches the version it read. That optimistic check stops two workers from both claiming the same image.

State Durable fact Next action
queued Source reference and digest exist Claim work with a lease
fetching A worker owns the lease Fetch and validate bytes
stored Object and metadata are committed Increment completed progress
rejected Validation or policy decision is final Show a reason, do not retry
cancel_requested The user asked to stop Finish the safe boundary, then cancel
cancelled No more work will start Retain the audit record

Rate limiting belongs at the worker boundary. A token bucket or leaky bucket can cap concurrency, but it cannot decide whether a response is retryable. Retry transport timeouts and 429 responses with bounded exponential backoff. Do not retry a malformed image forever; that is a data problem, not capacity pressure. Honor a Retry-After value when the upstream provides one, and add jitter so a fleet does not wake up on the same second.

Here is the smallest worker shape I use. It is intentionally provider-neutral: the fetchImage and storeImage functions sit behind interfaces, so changing a backend does not change the batch state machine.

type Status =
  | "queued"
  | "fetching"
  | "stored"
  | "rejected"
  | "cancel_requested"
  | "cancelled";

type Item = {
  id: string;
  batchId: string;
  digest: string;
  status: Status;
  version: number;
};

async function processItem(item: Item): Promise<void> {
  const claimed = await db.claim(item.id, item.version, "fetching");
  if (!claimed) return; // Another worker won the lease.

  try {
    const bytes = await fetchImage(item.digest);
    validateImage(bytes); // Check type and size before storage.
    await storeImage(item.digest, bytes);
    await db.transition(item.id, "fetching", "stored");
  } catch (error) {
    if (isRetryable(error)) {
      await db.releaseForRetry(item.id, retryDelay(error));
      return;
    }
    await db.transition(item.id, "fetching", "rejected");
  }
}
Enter fullscreen mode Exit fullscreen mode

The comment is the important part: a failed claim is normal under concurrency. It is not an error to log at warning level for every item. Good observability separates expected contention from a growing retry queue.

Cancellation needs a boundary. If the user imports the wrong directory and clicks cancel, mark the batch cancel_requested immediately. A worker checks that flag before fetching the next item and after a retry delay. It does not pretend an in-flight network request can always be interrupted. Once the current item reaches a safe state, the worker marks it cancelled or leaves a committed stored object alone, depending on the product's retention policy.

For a wrong import, “undo” should be explicit. Keep a manifest of object keys created by this batch, then offer a compensating delete job with its own audit entries. Do not issue a broad “delete everything from this user” command. The manifest is the difference between a reversible import and a frightening support script.

Ship the boring part.

The first time I test this flow, I use a deliberately wrong folder: a few valid PNGs, one file renamed to .jpg, a duplicate digest, and a cancellation click while the queue is waiting on backoff. That small fixture exposes the failure modes that a happy-path demo hides. The renamed file should become rejected with a reason tied to the detected media type; the duplicate should attach to the existing child record rather than create a second object; the waiting item should observe cancel_requested before its next attempt; and the already-stored files should remain discoverable through the manifest. If the UI reports “cancelled” while a worker can still start a new fetch, the state model is lying. If the cleanup job deletes by user prefix instead of by manifest key, a typo can remove an unrelated import. I would rather spend an hour making those transitions visible in a test database than spend a week explaining an ambiguous progress bar to customers.

What I would change before this runs at scale

The first version can use a relational table, a queue with visibility timeouts, and object storage. At higher volume, I would add leases with a heartbeat, a dead-letter stream for records that exceed a retry budget, and a reconciliation job that compares the manifest with stored objects. Those additions cost operations time, so I would wait until queue age or recovery time is a measured problem.

The progress endpoint should return a snapshot, not a stream that the browser must keep alive. A client can poll with If-None-Match, or subscribe to events when the UI truly needs live updates. Either way, compute progress from committed child states. A progress bar that reaches 100% before metadata is durable is worse than no progress bar.

Validation is another scale boundary. Read the file signature, not just the filename extension. The browser may call a file .jpg while the bytes are a different format, and image formats have different decoding and color-management behavior. MDN's image format guide is a useful map of those trade-offs; the storage contract should record the detected media type alongside the original name.

I am not sure every pipeline needs per-pixel normalization before storage. Your mileage may vary. If downstream moderation requires a canonical color space, normalize once and record that transform; if the product only needs thumbnails, delaying full-resolution decoding can keep the queue cheaper and easier to drain.

The trade-offs I keep visible

This design favors recoverability over minimum latency. That is usually right for customer imports, but it is not suitable when the job is disposable telemetry or a one-second preview. For those, an in-memory worker and a time limit may be the honest choice.

The database manifest also creates retention work. Keeping every source digest and cancellation event helps explain a dispute, but it increases privacy obligations. Set a retention period, encrypt sensitive metadata, and make deletion of the audit record a documented policy rather than an accidental cascade.

A queue is not a moderation policy. It can enforce rate limits and preserve order; it cannot tell you whether a generated image is acceptable. Keep policy decisions in a versioned service or ruleset, store the decision with the item, and make a policy change a new batch or an explicit reprocess. That separation lets the import machinery stay stable while the product's safety rules evolve.

The practical decision rule is simple: if a user can identify the wrong import after it starts, build checkpoints and a manifest before optimizing throughput. If the work has no user-visible side effect, keep the machinery lighter. Both choices are defensible. The mistake is promising cancellation without defining what has already become permanent.

References

Top comments (0)