Short answer
Short answer: for restaurant menu digitization, inspect metadata at upload, but defer expensive image cleanup until a dish is actually indexed or viewed. That split keeps searchable dish data predictable without making every upload wait on a full media pipeline.
It starts with one rule:
It failed once.
I keep a long-lived rule: metadata is a gate; pixels are a workload. A menu photo with the wrong orientation, an impossible color profile, or a 40 MB payload should be rejected or normalized immediately. Background removal, denoising, and several thumbnail sizes can wait. This is an architectural choice, not a preference for one image library.
The constraint that changes the design
A menu digitization system receives uneven material: phone photos, scanned PDFs, screenshots from delivery apps, and images exported by a designer. The OCR layer needs a stable input, while search needs stable dish names, prices, dietary labels, and provenance. If image cleanup runs synchronously for every upload, a slow source file turns a simple intake request into a queueing problem. If everything is deferred, bad metadata leaks into OCR and creates dishes that are hard to correct later.
The useful boundary is the first durable write. At that point, record the original object, its byte size, media type, pixel dimensions, orientation, and a content hash. Keep the original immutable. Store a normalized inspection record next to it. That record is what downstream workers trust.
This is where standards help. The browser and server should agree on media types, and the storage layer should not infer a file type from a filename alone. MDN's media format guide documents formats browsers commonly decode. Your OCR provider may accept fewer formats than a browser, so the acceptance list belongs in configuration and tests.
A tiny TypeScript inspection boundary is enough to make the decision explicit:
type MediaInspection = {
mime: string;
bytes: number;
width: number;
height: number;
orientation: number;
sha256: string;
};
type IntakeDecision =
| { kind: "reject"; reason: string }
| { kind: "accept"; inspection: MediaInspection; cleanup: "defer" };
const accepted = new Set(["image/jpeg", "image/png", "image/webp", "application/pdf"]);
export function decide(ins: MediaInspection): IntakeDecision {
if (!accepted.has(ins.mime)) return { kind: "reject", reason: "unsupported media type" };
if (ins.bytes > 25_000_000) return { kind: "reject", reason: "payload exceeds intake limit" };
if (ins.width < 800 || ins.height < 600) return { kind: "reject", reason: "insufficient pixels for OCR" };
return { kind: "accept", inspection: ins, cleanup: "defer" };
}
The numbers above are policy examples, not universal truths. Measure OCR accuracy on your own menu corpus before changing them. I am not sure a single minimum dimension works for every cuisine or camera, which is exactly why the threshold should be versioned and observable.
How should metadata inspection and image cleanup shape searchable dish data?
Treat the pipeline as two clocks. The intake clock protects data integrity in seconds. The enrichment clock can run for minutes and may be retried. Search should publish only a versioned dish record whose source image and OCR result are both traceable.
At intake, parse the actual header bytes, then compare the result with the declared MIME type. Read EXIF orientation and color information, but avoid copying GPS coordinates into a searchable record. Capture a hash before transformation so duplicate menu pages can be detected even when a later cleanup step changes pixels. Keep a failure reason such as metadata_invalid or ocr_input_unsupported; these are data states, not HTTP responses for readers to memorize.
During enrichment, create a deterministic derivative. Normalize orientation, convert to a format accepted by the OCR engine, cap the longest edge, and preserve enough contrast for small prices. A cleanup job should include the inspection version and source hash in its idempotency key. That prevents a retry from creating a second set of thumbnails when a worker is restarted.
The dish index needs more than extracted text. A useful row includes the restaurant identifier, menu version, source object key, OCR confidence, language, bounding boxes for name and price, and a human-review state. Search relevance suffers when a low-confidence price is treated as fact. Keep uncertain fields visible to review instead of silently guessing.
A worker contract can stay boring:
type CleanupJob = {
sourceKey: string;
sourceSha256: string;
inspectionVersion: number;
target: "ocr" | "thumbnail";
};
export async function processJob(job: CleanupJob): Promise<void> {
const source = await objectStore.read(job.sourceKey);
const normalized = await normalizeForTarget(source, job.target);
await objectStore.write(
`${job.sourceSha256}/${job.target}.webp`,
normalized,
{ sourceSha256: job.sourceSha256, inspectionVersion: job.inspectionVersion },
);
}
The storage calls are intentionally abstract.
No magic.
Ship it only when the key is stable. A queue, a cron runner, or a database-backed worker can implement them. What matters is that the contract carries provenance and that writing the same key is safe to retry.
Upload-time versus on-demand processing
Upload-time processing gives a clean user experience when every image needs the same derivative immediately. It also makes latency and capacity coupled: a burst of 10,000 menu pages consumes CPU before anyone searches for a dish. You need back-pressure, bounded concurrency, and a clear timeout policy.
On-demand processing shifts work toward actual demand. A popular dish gets a warm derivative; an archived seasonal menu may never pay the cleanup cost. The catch is first-request latency and more complicated cache invalidation. Pin the derivative to a source hash, and the invalidation rule becomes mechanical: a new source hash means a new key.
For a mixed restaurant workload, I use a hybrid: metadata inspection and a small OCR-ready derivative at upload, then high-resolution cleanup and display thumbnails on demand. That choice is unsuitable when a regulator or an offline kiosk requires every derivative before publication. In that case, stick with an upload-time batch and expose queue progress to operators.
Do not hide the choice behind a giant configuration file. Three explicit policies are easier to benchmark:
| Policy | Best fit | Cost to watch | Failure mode |
|---|---|---|---|
| Upload-time | Every asset is published immediately | Queue CPU during bursts | Intake latency grows |
| On-demand | Long-tail or archival menus | First-view latency | Missing derivative at read time |
| Hybrid | OCR now, rich images later | Two worker paths | Inconsistent version handling |
I benchmark p50 and p95 intake latency separately from derivative latency. A single average hides the exact incident you will debug at 2 a.m. It's a bad summary. Track queue age, rejected MIME types, bytes per source, OCR confidence, cache hit rate, and the percentage of dishes awaiting review.
What I would change at scale
Once volume grows, split the metadata ledger from the object store. The ledger records state transitions: received, inspected, queued, derived, indexed, and reviewed. Workers consume transitions, not ad-hoc database scans. Emit a correlation id with every transition so an operator can follow one dish from upload to search.
Add a corpus-based test set containing crooked phone shots, menus with transparent PNGs, CMYK scans, rotated pages, tiny prices, and duplicate uploads. Assert both acceptance decisions and extracted field confidence. Golden images are useful, but golden metadata is the stronger contract because it catches an accidental parser change before pixels make the diff noisy.
At this scale, deletion deserves as much design as creation. A restaurant may replace a menu, and a customer may request removal. Tombstone the source hash, remove derived objects asynchronously, and make search exclude tombstoned versions immediately. Keep audit records without retaining the image bytes longer than policy allows.
There is no universal winner between upload-time and on-demand processing. The right answer follows publication guarantees, traffic shape, and review capacity. Start with metadata as the hard gate, choose the cheapest derivative that protects OCR, and measure the rest as a background workload.
References
- https://developer.mozilla.org/en-US/docs/Web/Media/Guides/Formats
- https://www.w3.org/TR/exif/
- https://www.rfc-editor.org/rfc/rfc9110
Top comments (0)