Pick the background removal API whose output you can afford to keep for as long as your policy says you have to keep it. For an ecommerce catalogue driven by a Node.js worker, the per-image fee is almost never the term that grows. The terms that grow are stored cutouts, stored source product images, and the moderation verdicts you are obliged to retain about both. Edge quality on a hero shot is a one-afternoon evaluation; retention is a multi-year liability, and it is the one you sign without noticing.
The ordering is counterintuitive until you write the arithmetic down.
The system costed out here is a customer support platform that renders short promo videos from a text prompt. A support agent types "show the walnut side table in a small apartment," and the pipeline pulls catalogue photographs, removes their backgrounds, composites the cutouts into generated scenes, and returns a twelve-second clip the agent can send in a reply. Moderation coverage is the axis every other decision hangs from, because a generated frame that reaches a customer is a published statement by the merchant, and a verdict you cannot produce later is the same as a verdict you never made.
Where the money goes in a cutout pipeline
Take a catalogue of 40,000 SKUs with three photographs each — 120,000 source images, which is an ordinary mid-market catalogue and not a stress test. A cutout has to carry transparency, and transparency narrows the format choice sharply: JPEG has no alpha channel at all, so the moment matting enters the pipeline you are in PNG, WebP or AVIF territory, and PNG is the only one of those three with no lossy mode. Stored as PNG at a 2000-pixel long edge, a cutout of a furniture photograph lands in the low single-digit megabytes. At 3.5 MB average, one full pass over the catalogue writes roughly 420 GB.
Then the matting model version changes and you write it again.
| Cost term | Scales with | Re-billed when |
|---|---|---|
| Removal calls | images × model versions | the catalogue is re-cut |
| Cutout storage | images × bytes × replicas | never released |
| Moderation calls | sampled frames per asset | policy version changes |
| Verdict retention | assets × retention window | never released |
| Metric series | unique label combinations | a new label value appears |
Three of those five terms are storage, and storage is the only one that compounds: last quarter's cutouts are still there while this quarter's are being written. Compute is a flow, bytes are a stock. A pipeline that bills $0.002 per removal and silently retains four copies of every output — origin, CDN cache, a backup bucket, and a "temporary" staging prefix that somebody created during a migration and nobody deleted — is a storage product wearing an API's pricing page.
How should an ecommerce catalogue pipeline budget background removal API calls in Node.js?
Budget in byte-years, not in calls. Multiply images by average output bytes by replica count by retention window, and compare that number against the removal fee for the same period before arguing about matting quality. If the storage term is more than about half the total, the correct optimization is not a cheaper API — it is storing less.
The change that moves the dominant term is to stop storing the composited cutout and store only the alpha matte. A matte is a single 8-bit channel at the same dimensions as the source, and because a matte of a product photograph is mostly saturated black and saturated white with a thin transition band at the object boundary, it compresses far better than the RGB it was cut from. The source photograph already exists in the catalogue; keeping a matte beside it, keyed by the source hash, means the composite becomes a request-time operation rather than a stored artifact. libvips and ImageMagick both composite a matte against a background quickly enough to sit behind a CDN, and hosted transformation services such as imgix and Cloudflare Images bill per stored original plus per transformation family, so this move shifts cost between two line items rather than deleting it outright. Measure both lines before committing.
Ask the API for the matte, not the composite.
curl -sS -X POST https://media.internal/cutouts \
-H "Authorization: Bearer $MEDIA_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sha256:9f2c4a1b7e03c41a" \
-d '{
"source_sha256": "9f2c4a1b7e03c41a",
"output": "matte",
"max_edge": 2000
}'
The Node.js side is thinner than the curl suggests — Node.js 22 LTS ships fetch, so there is no client library to pin, and the idempotency key is just the source hash, which makes a retried job free instead of double-billed. A well-behaved service answers 202 Accepted with a job identifier rather than blocking the connection for eight seconds, which is what RFC 9110 reserves that status for. The worker then polls or receives a webhook, and either way the response it stores is small:
{"status":"done","matte_sha256":"4d81ff02","model_version":"matte-3.2","bytes":186221}
Two hundred kilobytes instead of three and a half megabytes, keyed by content hash, with the model version recorded so that a re-cut is a diff rather than a full pass. The catch is real: request-time compositing moves latency onto the read path, and a catalogue page that renders forty thumbnails is now forty composites. If your traffic is read-heavy and your catalogue is small, the stored-composite design is cheaper and simpler, and you should stick with it.
Moderation coverage is a sampling decision, and the loss is asymmetric
Here is where the observability instinct — sample everything, keep a fraction — actively misleads you. A twelve-second clip at 24 fps is 288 frames. Moderating all 288 costs 288 calls per generated video; moderating every twelfth frame costs 24, a 92% reduction that looks like exactly the kind of win a sampling argument is supposed to produce.
It isn't, for one clip.
Sampling is sound when the loss function is symmetric and the thing you are estimating is a rate. Latency percentiles survive sampling because a missed sample costs you a little precision. Moderation does not behave that way: a single unmoderated frame that reaches a customer costs a takedown, a merchant relationship, and possibly a regulator's attention, and no amount of correctly-moderated neighbours compensates. The useful split is by path, not by percentage — moderate deterministically along the publication path, and sample freely along the evaluation path. Concretely, every asset that can reach a customer gets a verdict: the first frame, the last frame, every keyframe the encoder emits, and the composited still that becomes the thumbnail. ffmpeg emits keyframes deterministically given a fixed GOP setting, which is what makes a sampled pass reproducible rather than merely cheap. Prompt-side signals raise coverage to every frame for that job. Everything else — the drift dashboards, the per-model quality comparisons, the matting-regression counters — runs on a 1-in-100 sample and nobody complains.
What you store from that pass is the verdict, not the frame. This is the single largest retention decision in the pipeline, and it is worth spelling out as a record:
{"asset_sha256":"4d81ff02","policy_version":"2026.3","model_version":"mod-7.1","verdict":"pass","scores":{"violence":0.01,"adult":0.00},"ts":"2026-09-12T11:04:19Z"}
That is roughly 220 bytes. The frame it describes is 400 KB.
Four retention tiers, and the cardinality rule that goes with them
Four tiers have survived contact with an actual bill, and the boundaries are set by who asks the question and how long after the fact they ask it.
- Tier 0, seven days, everything. Full per-asset events for every stage, unsampled, including the matte bytes. This is the rollout-debugging tier; it exists so that a bad matting version is diagnosable while it is still deployed.
- Tier 1, ninety days, verdicts and hashes. Verdict records, source and matte hashes, model and policy versions. No pixels. This answers "what did we decide about this asset, and under which policy."
- Tier 2, thirteen months, aggregates. Daily counters by verdict class, model version and stage outcome. Thirteen rather than twelve, so that a year-over-year comparison has both endpoints.
- Tier 3, indefinite, the mapping. Content hash to verdict, policy version, and timestamp — around 100 bytes per published asset. At 120,000 assets that is 12 MB, which is free by any measure that matters.
The cardinality rule is the other half. SKU identifiers belong on log records and on the Tier 3 mapping; they do not belong in metric labels. Every unique label-value combination is a separate time series, and a sku_id label on a moderation counter turns one series into 40,000, multiplied again by verdict class and model version. Put the high-cardinality identifiers in the log body where they cost bytes, and keep metric labels to the bounded set — stage, model version, verdict class, outcome — where they cost series. OpenTelemetry's logs data model is explicit about this separation between attributes and the record body, and it is worth following even if you never ship a trace.
What I stopped keeping, and what that costs when something goes wrong
Two things went away, and both have a bill attached on the day they are missed.
The generated frames went first. Beyond the Tier 0 window there is a verdict and a content hash, and no pixels. When a merchant disputes a moderation decision four months later, the record proves that a verdict was rendered, under which policy version, by which model — but it cannot re-show the frame unless the generation is reproducible, which means the prompt, the seed, the matte hash and the model version all have to be in the Tier 1 record and the model version has to still be served. Retire a generation model without an archived weights snapshot and that reproducibility claim quietly becomes false. I'm not sure there is a clean answer here; the honest position is that reproducibility is a contract with your model registry, not a property of your log schema, and content provenance work such as C2PA's credentials is the direction that actually addresses it.
The per-SKU metric labels went second, and that one I miss more often. Nobody can graph a single SKU's matting failure rate anymore. The answer is still in the logs, but it is a query rather than a dashboard, and a query takes four minutes where a dashboard took four seconds.
That trade is worth making at 40,000 SKUs. It is not worth making at 400 — a small catalogue has no cardinality problem, storage measured in gigabytes is rounding error, and every tier boundary above is overhead you would be maintaining for its own sake. Sampling and tiering are tools for systems where the bill has stopped being obvious, and applying them early buys complexity with no return.
The decision rule I'd hand to somebody starting this week: compute byte-years before comparing matting quality, make the publication path's moderation coverage total and the evaluation path's coverage cheap, and keep the hash-to-verdict mapping forever because it is the only tier that is both tiny and irreplaceable.
Further reading
- MDN — Image file type and format guide: https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
- RFC 9110, HTTP Semantics (202 Accepted): https://www.rfc-editor.org/rfc/rfc9110.html
- Prometheus — metric and label naming best practices: https://prometheus.io/docs/practices/naming/
- OpenTelemetry — logs data model: https://opentelemetry.io/docs/specs/otel/logs/data-model/
- libvips documentation: https://www.libvips.org/
- C2PA specifications: https://c2pa.org/specifications/
Top comments (0)