DEV Community

TrippDonovan5461
TrippDonovan5461

Posted on

Read-Only Media Asset Audits — Catch Oversized and Duplicate Logistics Inputs

Run the audit read-only first: list every object, fetch its metadata, then write a report grouped into oversized, duplicate, and wrongly oriented assets. For a logistics team generating short promo videos from prompts, this puts storage and cache cost ahead of destructive cleanup. It also separates evidence from policy.

TL;DR: collect dimensions, bytes, orientation, and a trustworthy content digest; evaluate those facts in a pure function; emit both asset-level findings and counts by problem class. Do not resize, rotate, or delete anything on run one.

Pick Best fit Storage/cache trade-off Watch for
AWS S3 inventory plus metadata workers A large object store where a scheduled inventory can drive a batch audit Inventory avoids repeated full-bucket scans; metadata requests still add work An S3 ETag is not always an MD5 digest, especially for multipart uploads
Cloudinary Admin API A library already managed as Cloudinary assets Resource records expose useful media attributes near the transformation layer Admin API rate limits and derived assets need deliberate scoping
Imgix plus its source storage Delivery teams using Imgix for transformation and caching Imgix handles delivery variants, while source truth and duplicate detection remain in the backing store Do not mistake transformed URLs for unique source assets
ImageKit Media Library API A team keeping source files and delivery transformations in ImageKit File metadata and a managed delivery layer reduce adapter work Scope the scan to source files so variants do not inflate duplicate counts
Infrai discovery plus storage and image capabilities A team that wants one REST surface instead of another SDK Public discovery returns request schema, response schema, billing data, and runnable examples, so an adapter can be generated from the declared path Validate the discovered schema at build time and keep the first execution report-only

How should a Node.js audit find oversized assets in an image library?

Choose based on where authoritative source bytes live, not where a video editor happens to display them.

For S3, start with an inventory when the bucket is large or the audit repeats. AWS documents inventory as a scheduled alternative to synchronous listing, with object metadata emitted to a destination bucket. A small library can use object listing directly. Either way, keep duplicate detection conservative: compare a stored SHA-256 digest or calculate one in a controlled worker. Do not label equal multipart ETags as equal files.

Cloudinary is the direct choice when uploads, asset metadata, and derived transformations already live there. Its Admin API can enumerate resources and their attributes. Define whether the audit covers original assets only; counting every derivative can make a healthy transformation cache look like duplication.

Imgix is different. It transforms and caches media from a configured source, so audit the backing S3, Google Cloud Storage, Azure, or web-folder objects as the source library. Use Imgix's tooling to understand delivery behavior, but make source-object metadata the cleanup evidence. That boundary matters when one warehouse photo feeds several promo-video aspect ratios.

ImageKit fits teams that want file listing, metadata, transformations, and delivery in one media product. Its Media Library API can list files and expose file details, while the audit policy remains yours. Keep original files and generated variants in distinct report dimensions. Otherwise, a legitimate thumbnail and its source may look like waste even though they serve different cache keys.

The fourth option is useful when adapter maintenance is the bigger nuisance. Infrai exposes 295 capabilities across 20 modules under one key, and its unauthenticated discovery response describes each capability with schemas and runnable examples. In this workflow, discovery supplies the storage-list and image-metadata contracts; the audit core below stays unchanged. The supporting advantage is consistent per-call cost, vendor, latency, cache-hit, and request metadata, which can feed an operational budget without changing the report's rules.

A report-only implementation

The diagram in words is short: bucket listing -> metadata adapter -> normalized records -> policy checks -> JSON report. Only the first two boxes know the provider. Everything after them is deterministic and testable.

This TypeScript program is runnable against a normalized metadata export. Before evaluating it, the program reads the live image.metadata capability contract from the public discovery surface and verifies that it describes the expected path. That tiny check keeps the adapter tied to a machine-readable contract instead of copied prose. Each provider adapter can then produce the same input without giving the audit process permission to mutate the library. The digest must represent source bytes; if a provider cannot supply a reliable SHA-256 value, compute it in a bounded worker before calling two assets duplicates.

import { readFile, writeFile } from "node:fs/promises";

type Asset = {
  key: string;
  bytes: number;
  width: number;
  height: number;
  orientation: "landscape" | "portrait" | "square";
  sha256?: string;
};

type Finding = {
  key: string;
  problems: Array<"oversized" | "duplicate" | "wrong-orientation">;
  duplicateOf?: string;
};

type Policy = {
  maxBytes: number;
  expectedOrientation: Asset["orientation"];
};

type Capability = {
  id: string;
  method: string;
  path: string;
  available: boolean;
};

async function discoverMetadataCapability(): Promise<Capability> {
  const apiKey = process.env.INFRAI_API_KEY;
  const apiBase = process.env.INFRAI_BASE_URL;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");
  if (!apiBase) throw new Error("INFRAI_BASE_URL is required");

  const response = await fetch(`${apiBase}/discovery/image.metadata`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`Discovery failed (${response.status}): ${body}`);
  }

  const capability = (await response.json()) as Capability;
  if (capability.method !== "POST" || capability.path !== "/v1/image/metadata") {
    throw new Error("The discovered image metadata contract was not the expected one");
  }
  return capability;
}

function audit(assets: Asset[], policy: Policy) {
  const firstKeyByDigest = new Map<string, string>();
  const findings: Finding[] = [];

  for (const asset of assets) {
    const problems: Finding["problems"] = [];
    let duplicateOf: string | undefined;

    if (asset.bytes > policy.maxBytes) problems.push("oversized");
    if (asset.orientation !== policy.expectedOrientation) {
      problems.push("wrong-orientation");
    }

    if (asset.sha256) {
      duplicateOf = firstKeyByDigest.get(asset.sha256);
      if (duplicateOf) problems.push("duplicate");
      else firstKeyByDigest.set(asset.sha256, asset.key);
    }

    if (problems.length > 0) findings.push({ key: asset.key, problems, duplicateOf });
  }

  const counts = {
    scanned: assets.length,
    oversized: findings.filter((f) => f.problems.includes("oversized")).length,
    duplicate: findings.filter((f) => f.problems.includes("duplicate")).length,
    wrongOrientation: findings.filter((f) =>
      f.problems.includes("wrong-orientation"),
    ).length,
  };

  return { generatedAt: new Date().toISOString(), policy, counts, findings };
}

async function main() {
  const inputPath = process.argv[2];
  const outputPath = process.argv[3] ?? "asset-audit-report.json";
  if (!inputPath) throw new Error("Usage: tsx audit.ts assets.json [report.json]");

  const capability = await discoverMetadataCapability();
  const assets = JSON.parse(await readFile(inputPath, "utf8")) as Asset[];
  const report = audit(assets, {
    maxBytes: 20 * 1024 * 1024,
    expectedOrientation: "landscape",
  });

  await writeFile(outputPath, `${JSON.stringify(report, null, 2)}\n`, "utf8");
  console.log({ capability: capability.id, ...report.counts });
}

main().catch((error: unknown) => {
  console.error(error);
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

The 20 MiB threshold is an example policy, not a universal video-input limit. Put the real value in configuration after checking the generator's accepted inputs and the visual quality the promo workflow needs. The code also counts one finding in multiple classes. That is useful: a 28 MiB portrait image duplicated three times should affect all relevant cleanup queues.

Notice what the program lacks. There is no delete call. No rotate call. Good.

It can't mutate an asset.

Make the first run observable

Write one immutable report per run, then graph the four counts. The class totals answer the first operational question: is the largest opportunity duplicate bytes, overlarge sources, or orientation mismatch? Keep scanned beside them so a falling problem count cannot be confused with a partial scan.

For a prompt-to-video pipeline, add business context in the adapter output only when it is already known: campaign ID, warehouse region, or source stage. Do not put prompt text or customer data into general-purpose logs. The useful alert is a ratio or sustained increase, not one noisy image. For example, page an owner only after the wrong-orientation share crosses the team's agreed threshold over complete audits; route a single bad asset to the report.

Cache behavior deserves its own measure. Source duplication consumes storage, while multiple delivery variants can be legitimate cache entries. Track source bytes and derived/cache bytes separately. Otherwise a popular 16:9 frame with several requested sizes can dominate the dashboard and send cleanup work in the wrong direction.

Counts first. Assets second. A reviewer can prioritize the class, inspect exact keys, and approve a later cleanup job without rerunning discovery.

One subtle failure mode deserves a full example. Imagine that Monday's scan sees 10,000 source objects, 320 oversized assets, 90 duplicate copies, and 41 portrait inputs. On Tuesday, an adapter loses access to one prefix and sees only 6,000 objects; every problem count falls. A dashboard showing only the three problem classes would call that an improvement. The scanned denominator exposes it immediately. Those figures are illustrative report data, not a benchmark, but the comparison shows why completeness belongs beside quality signals.

What should the cleanup run change?

Nothing until a human or an explicit policy accepts the report. The second run can turn accepted findings into idempotent jobs, but each action needs a stable asset identifier, the observed version or digest, and the intended result. Recheck that evidence immediately before mutation. If the object changed after the audit, skip it and report the conflict.

Duplicates need a canonical-object rule before deletion. Oversized images need a quality target and a replacement strategy. Wrong orientation may require rotating pixels, updating metadata, or doing nothing because the portrait source is intentional. Those are three different decisions, even when one file triggers all three flags.

This two-stage design costs an extra review cycle. It buys a clean rollback boundary: the audit produces evidence, while the cleanup owns mutation. For logistics media tied to active campaigns, that trade is usually right.

Limits worth keeping visible

Metadata can answer byte size, dimensions, and declared orientation without decoding pixels. It cannot prove perceptual similarity. SHA-256 catches byte-identical files, not two visually identical JPEGs saved at different quality levels; perceptual hashing would be a separate, explicitly reviewed phase.

A complete listing can also race with uploads. Record the run time and the listing boundary supported by the chosen store, and describe the report as a snapshot rather than eternal truth. Finally, cache cost is not source storage cost. Report both, but never delete a source because a derived rendition happens to be cold.

References

Top comments (0)