Short answer: use an asynchronous PDF endpoint for final invoice generation during a document format migration, and keep a synchronous endpoint only for previews or other small, latency-sensitive renders. That split protects throughput under load without forcing every request through a queue. Fidelity remains a release criterion, not a reason to make customers wait on an open connection.
| Choice | Best fit | Fidelity control | Latency under load | Operational cost |
|---|---|---|---|---|
| Synchronous render | Preview and occasional one-off invoice | Immediate visual feedback | Caller waits for rendering capacity | Lower at first; retries and timeouts move into the request path |
| Asynchronous render | Final invoices, batches, and migration backfills | Stable template and renderer version per job | Queue absorbs bursts; completion is separate | Higher because jobs, storage, and observability need ownership |
| Client-side render | Non-authoritative preview | Depends on the user's browser and fonts | No server render queue | Low server burden, weak control for final documents |
The recommendation is deliberately uneven: choose asynchronous final rendering by default. Preserve synchronous rendering as a narrow fast path, with an explicit size limit and deadline. Don't make one endpoint serve incompatible jobs.
How should a US/EU SaaS balance PDF fidelity and latency under load?
Treat fidelity and latency as two separate service-level decisions. Fidelity asks whether the same order data, template revision, fonts, locale, page size, and renderer revision produce an acceptable invoice. Latency asks how long a caller waits and what happens when demand exceeds render capacity. Combining both into one average response-time target hides the failure mode that matters: a short burst can turn expensive render work into a wall of open requests, retries, and duplicate documents.
For final invoices, accept a render job, return a stable job identifier, and let the client inspect completion later. A webhook can reduce polling, but polling should still be possible because delivery and rendering are different concerns. Store the source payload or a tamper-evident reference to it, the template revision, locale, requested output, and an idempotency key with the job. The resulting PDF should be immutable once published; a correction becomes a new document revision rather than an in-place mutation.
For previews, a synchronous endpoint is useful because the human editing a template needs a quick visual loop. Put a deadline around that path. Also cap the input and reject work that belongs in the queue before rendering begins. Those controls are product decisions, so I'm not sure a universal timeout value exists; a load test with representative invoices, fonts, images, and concurrency is what resolves it.
Region belongs in the contract, not in scattered caller logic. Let the application select an approved processing region through configuration, then keep the job data, generated file, logs, and retry path aligned with that selection. Legal and security owners still need to decide the actual retention and transfer requirements. An endpoint shape can't make that decision for them.
Keep it boring.
Fidelity is a regression suite, not a screenshot
A format migration often looks complete when one clean invoice renders. The hard cases arrive later: a description wraps onto a second page, a tax label grows under localization, a logo has an unusual aspect ratio, or a line item contains a long unbroken identifier. The useful fidelity suite is therefore a versioned corpus of order inputs plus assertions about the resulting document. Include sparse and dense invoices, negative adjustments, multiple currencies, long addresses, missing optional fields, page breaks, and the locales the SaaS actually sells into. None of those cases needs a vendor-specific test harness.
Pixel comparison alone is too brittle for many document changes, while text extraction alone misses clipping and layout drift. Use layers: validate business data before rendering, verify that required text is present afterward, inspect page count and file type, and use image comparison for a small set of layout-critical fixtures. Human review remains appropriate when a template revision intentionally changes typography or pagination. The migration gate should record that approval alongside the template revision.
Fonts deserve explicit treatment because the renderer cannot preserve a typeface it cannot access. Package or otherwise make the approved fonts available in the render environment, confirm their licensing permits the intended use, and test the fallback behavior. External images create a similar dependency; fetching them during a render makes output depend on another system's latency and availability. Prefer validated, controlled assets for authoritative invoices.
The browser-facing side has one clean boundary. A completed PDF is binary data, and the Web Platform Blob interface represents immutable raw data that can be read as text or binary data or converted into a ReadableStream. That makes a Blob a suitable handoff for previewing or downloading a completed document in a web client. It does not decide how the server schedules the render.
A small TypeScript boundary keeps migration reversible
Application code should submit an invoice intent, not know which rendering engine processes it. The boundary below supports both modes and makes the choice visible at the call site. It also prevents a migration from leaking renderer-specific request fields into order code.
type Region = 'us' | 'eu';
type RenderMode = 'preview' | 'final';
type InvoiceRender = {
orderId: string;
templateRevision: string;
locale: string;
region: Region;
mode: RenderMode;
idempotencyKey: string;
};
type RenderReceipt =
| { kind: 'complete'; pdf: Blob }
| { kind: 'accepted'; jobId: string };
interface PdfRenderer {
render(input: InvoiceRender): Promise<RenderReceipt>;
result(jobId: string): Promise<Blob | undefined>;
}
async function generateInvoice(
renderer: PdfRenderer,
input: Omit<InvoiceRender, 'mode'>,
): Promise<RenderReceipt> {
return renderer.render({ ...input, mode: 'final' });
}
The interface is intentionally strict. A final render may be accepted before a PDF exists, while a preview may complete in the original call. Production code also needs cancellation rules, authentication, authorization, bounded retry behavior, and durable job state, but those concerns live behind the boundary. Callers shouldn't convert an uncertain response into a second invoice job; they should reuse the same idempotency key and inspect the original job.
During migration, implement the old and new renderers behind this interface. Run shadow renders from the fixture corpus, compare the outputs, and record the template and renderer revisions used for each artifact. Do not shadow every production request by habit: duplicate rendering raises cost and can duplicate external side effects if the boundary is poorly drawn. A controlled sample or an offline replay is easier to reason about. Promotion then becomes a configuration change after fidelity and load gates pass, while rollback keeps the application contract intact.
Load behavior needs admission control and evidence
Average latency is a weak capacity signal. Track queue age, render duration by template revision, completion rate, retry count, output size, and the number of active workers. Use percentiles for duration and queue age so a small slow tail is visible. Split client waiting time from render time; otherwise a network delay can be misdiagnosed as a document engine regression. Backpressure starts before the renderer: limit accepted payload size, validate required order fields, constrain concurrency, and reject excess preview work predictably. For asynchronous work, a bounded queue plus worker concurrency protects the rest of the SaaS from render bursts. Autoscaling may help, but only after the team understands startup time, font and asset loading, memory pressure, and the downstream limits of storage and notification systems. Your mileage may vary because invoice complexity changes the shape of the workload. Retries need a budget too. Retry only operations the application has classified as safe, add delay between attempts, and preserve the job identity. A dead-letter state is better than an infinite retry loop because an operator can inspect the input and decide whether a template or data correction is required. Record structured failure categories without putting customer invoice contents in routine logs.
Load-test both endpoint modes with the same representative corpus used for fidelity. Increase concurrency gradually, observe queue age and tail latency, and stop when the agreed service objective or resource ceiling is crossed. The result is a capacity curve for this workload, not a borrowed requests-per-second claim. Re-run it when templates gain large images, fonts change, or renderer configuration changes. This is the longer part of the work, and it pays rent: a one-person SaaS cannot spend release week manually untangling duplicate invoices because a pretty demo hid poor admission control.
Ship weekly, but measure first.
When should the synchronous runner-up win?
Stick with a synchronous endpoint when renders are small, infrequent, and immediately consumed by a person, and when the measured tail latency stays inside the product's interaction budget at expected concurrency. It is also reasonable during an early migration stage if the application already has strict request deadlines, idempotent retries, and enough spare render capacity. The operational surface is smaller because there is no separate job lifecycle to expose.
The catch is that synchronous simplicity transfers queueing to callers as load grows. It is not suitable for bulk invoice regeneration, scheduled billing peaks, or backfills where completion matters more than immediate response. In those cases, the asynchronous choice earns its extra moving parts by making admission, progress, retries, and capacity visible. Client-side rendering is the runner-up only for non-authoritative previews; keep final invoice production in a controlled environment when consistent templates, fonts, and records matter.
This decision rule is enough: synchronous for bounded interactive previews, asynchronous for authoritative output and bursty work. Outsource the undifferentiated renderer if that improves revenue per engineering hour, or operate one internally when control requirements justify the time. Either way, keep the application boundary portable and make fidelity tests plus load evidence the release gate.
Top comments (0)