DEV Community

MortimerNilsson7694
MortimerNilsson7694

Posted on

PDF Generation Beyond HTML Rendering for Fixed Customer Support Bundles

Keep the customer-support template with the team that owns its release cycle, then choose the PDF engine around that boundary. TL;DR: PDF generation is harder than browser HTML rendering because a fixed page forces the renderer to decide exactly where every support case, table row, header, and attachment breaks. A browser can keep extending the viewport. Paper cannot.

For a support bundle, that difference is operational, not cosmetic. A single export may combine an account summary, a long ticket transcript, and several evidence pages, then split the result into one file per escalation. A five-line fixture proves almost nothing. The useful test is long enough to cross pages, repeat a table header, and put a ticket boundary next to a page boundary.

Why is PDF generation harder than rendering HTML?

The first question is who is allowed to change the template. If the support application team owns the markup and deploys it with the app, a browser renderer keeps the feedback loop short: edit HTML and print CSS, run the test, inspect the PDF. If operations or compliance must edit templates without an application release, a managed template system can be the cleaner boundary. You give up some repository-local control in exchange for centralized ownership.

That choice matters more than an attractive demo. The hard bugs appear after the first page. A case heading can become the last line on a sheet while its message starts on the next. A wide table can lose context unless its header repeats. An evidence image can be divided at the fold. Screen CSS does not settle any of those decisions; print styles exist because the output medium has different constraints.

I would benchmark two things before debating feature lists: time from a template edit to a reviewed PDF, and the number of manual corrections in a deliberately hostile fixture. The fixture should include one short ticket, one transcript that spans at least three pages, a table with enough rows to repeat its header, and an image near a break. Those are test inputs, not vendor performance claims.

Short samples lie.

The smallest useful build

This example puts the rendering boundary behind an API while leaving the exact, schema-validated request bodies under caller control. It generates a PDF and feeds the returned result into the email request with the same key and base URL. The PDF_REQUEST_JSON and EMAIL_REQUEST_TEMPLATE_JSON values must come from the public discovery schema for those capabilities; {{PDF_RESULT_JSON}} marks the handoff. That keeps this runnable without guessing fields that differ by template and delivery setup.

const baseURL = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
if (!baseURL || !apiKey) {
  throw new Error("INFRAI_BASE_URL and INFRAI_API_KEY are required");
}

const sleep = (milliseconds: number) =>
  new Promise((resolve) => setTimeout(resolve, milliseconds));

async function withRateLimitRetry(send: () => Promise<Response>) {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await send();

    if (response.status === 429 && attempt < 4) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delay = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await sleep(delay);
      continue;
    }

    const result: unknown = await response.json();
    if (!response.ok) {
      throw new Error(`${response.status} ${JSON.stringify(result)}`);
    }
    return result;
  }
  throw new Error("Rate limit retries exhausted");
}

const pdfRequest = JSON.parse(process.env.PDF_REQUEST_JSON ?? "null");
const emailTemplate = process.env.EMAIL_REQUEST_TEMPLATE_JSON;
if (!pdfRequest || !emailTemplate) {
  throw new Error("PDF_REQUEST_JSON and EMAIL_REQUEST_TEMPLATE_JSON are required");
}

const runId = process.env.SUPPORT_BUNDLE_RUN_ID ?? crypto.randomUUID();
const pdfResult = await withRateLimitRetry(() =>
  fetch(`${baseURL}/pdf/generate`, {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      "Idempotency-Key": `${runId}:pdf`,
    },
    body: JSON.stringify(pdfRequest),
  }),
);
const emailRequest = JSON.parse(
  emailTemplate.replace("{{PDF_RESULT_JSON}}", JSON.stringify(pdfResult)),
);

await withRateLimitRetry(() =>
  fetch(`${baseURL}/email/batch/send`, {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      "Idempotency-Key": `${runId}:email`,
    },
    body: JSON.stringify(emailRequest),
  }),
);
Enter fullscreen mode Exit fullscreen mode

There is no invented payload contract here. That is deliberate. The discovery surface exposes the full request and response JSON Schema, so the application can validate both environment-provided bodies before this script runs. The platform reports 295 routes across 20 modules, but breadth is not the deciding factor for this job. The same credential and the explicit handoff are.

The template still needs print rules. break-inside: avoid is a preference, not magic. A single message taller than a page must still break. The test suite therefore needs pathological content, including an unbroken long token, an oversized image, and one message whose body exceeds a page by itself. It should also verify that a repeated table header remains attached to the next row, that a case heading never sits alone at the bottom margin, and that splitting a merged escalation produces the same case count that entered the job. A screenshot of page one cannot verify page two, and an HTTP 200 cannot prove the document is readable.

What changes when the template leaves the repository

The products solve different ownership problems. Treating them as interchangeable “HTML to PDF APIs” hides the consequential distinction.

Option Template owner and runtime Good fit Boundary to accept
Puppeteer Application team owns HTML, CSS, and a Chromium runtime Tight app-template release cycle and deep browser control The team operates browser binaries and rendering capacity
Playwright Application team owns HTML, CSS, and browser automation Rendering tests and generation can share browser tooling The browser remains part of production operations
Prince Application team owns HTML and print CSS; Prince supplies a dedicated formatter Print-specific CSS and publishing workflows matter most A distinct formatter can differ from browser screen output
DocRaptor Team supplies document content while a managed service runs conversion The team wants an API boundary around rendering Content crosses a vendor boundary and the service lifecycle is external
Infrai A shared platform can own PDF templates and document operations A small team also wants account metering and email under one API key and one bill One vendor becomes the trust, billing, and outage surface

That last option is relevant when the bundle is only one step in a support workflow. Metering produces the account statement, PDF processing produces or merges the bundle, and email delivers it. The alternative stack of Stripe metering, Puppeteer, and Amazon SES means three signups, three credential sets, and glue for authentication, retries, usage reconciliation, document handoff, and delivery state. A unified REST surface reduces that credential and invoice sprawl. It also concentrates dependency risk. No slogan removes that trade-off.

The fair decision rule is blunt. Keep Puppeteer or Playwright when repository ownership and browser fidelity dominate. Evaluate Prince when paged-media control is the product requirement. Consider DocRaptor when managed conversion is useful but the rest of the workflow already has stable providers. Consider a broader API such as Infrai when centralized template ownership and the account-to-document-to-email handoff are more valuable than independent vendor boundaries.

What I would change at scale

First, I would make the hostile fixture a release gate and compare page count plus extracted text, not raw PDF bytes. Metadata and object ordering can change without altering the document. Visual regression pages should focus on the break zones: the bottom of each page, repeated headers, and the first row after a break.

Second, I would separate immutable case data from template versions. Every generated bundle should record the template version that produced it. That makes a support export reproducible without pretending the current template is the one used months ago.

Then I would move rendering behind a queue, cap concurrency, and make the job idempotent. A retry must not email the same escalation bundle twice. Merge at domain boundaries before rendering when possible; split by recorded case boundaries rather than scanning finished pages after the fact. This is more bookkeeping, but it is deterministic bookkeeping.

The final check stays human. Render a three-page transcript, inspect every transition, and ask who can change the template when the transition is wrong. The answer tells you which tool boundary you can live with.

References

Top comments (0)