If the job is to watermark 30,000 documents before an external-sharing window closes, use the cheapest operation that still produces a correct page: draw the mark onto an already-typeset PDF, and stop re-running HTML layout once per recipient. Print layout is the expensive half of PDF generation. Rendering glyphs is not. A browser-based pipeline that regenerates from source markup pays the pagination cost on every single output file, which is why the batch gets slower in proportion to how many people you share with rather than how many documents you actually have.
That distinction is worth roughly two orders of magnitude in CPU time, and — less obviously — it changes what you are obliged to keep on disk for the next seven years.
Where the nightly bill actually comes from
Four terms, and usually only one of them matters. Layout CPU is the work of turning markup into a fixed sequence of pages. Stamping CPU is the work of drawing something on pages that already exist. Object storage holds whatever you produce. Retention is the multi-year obligation to keep what you produced, and it is the only term that grows without anyone deciding to grow it.
Put numbers into a concrete shape. Thirty thousand source documents, eighteen pages on average, each shared with about six external recipients a month, each recipient receiving a distinct watermark so a leak can be traced back to one account. That is 180,000 output files a month. If a headless browser needs 1.8 s of CPU to lay out and print an eighteen-page document, the regenerate-per-recipient path costs 324,000 CPU-seconds, or about 90 CPU-hours a month, before you account for the memory ceiling that forces you to cap concurrency well below core count. If drawing an overlay on an existing PDF costs 40 ms, the same 180,000 outputs cost roughly 2 CPU-hours.
Forty-five to one.
The ratio is what you should carry away, not my placeholder constants — measure your own p50 render time on your own page templates before you take this arithmetic anywhere near a capacity plan. Storage is the quiet term underneath it. At 1.4 MB per watermarked copy, 180,000 copies a month is 252 GB a month, a bit over 3 TB a year, and something like 21 TB by the time a seven-year records policy has run its course. Nobody reviews that line item. It arrives as a slope.
Why is PDF generation harder than rendering the same HTML in a browser?
A screen has one page and it is infinite. Print has a finite number of finite pages, and every one of them has to be decided before the file is written. That single constraint explains most of the difficulty, because pagination is a global problem: a font substitution on page 3 can push a heading onto page 4, which shifts every subsequent break, which changes the total page count, which changes the running footer that says "Page 7 of 214", which is itself content that participates in layout. Real engines resolve this with multiple passes and a fixed-point check, and when the fixed point does not converge they pick something and move on.
CSS gives you the vocabulary for this — @page, margin boxes, break-inside: avoid, orphans and widows are specified in the Paged Media and Fragmentation modules — but the vocabulary is not the implementation. Support differs by engine, and the differences are silent. A table with a repeating header row that splits cleanly in one engine loses its header in another, and nothing errors. You find out when a customer forwards a screenshot.
Fonts are the second trap. On screen, a fallback font is a cosmetic annoyance; in a PDF the substituted metrics change line breaking, which changes pagination, which changes the page numbers printed inside a document you may have to produce in a dispute two years later. The format wants fonts embedded and subsetted for exactly this reason, and subsetting is also why two runs over identical input do not produce identical bytes unless you pin things deliberately. ISO 32000-2 files carry a creation timestamp, a file identifier, and generated subset tags; the reproducible-builds community has documented this class of nondeterminism at length. If your compliance story includes the sentence "we can regenerate the exact artifact", byte-level determinism is not a nice-to-have, it is the entire claim.
Geometry is the third. CSS thinks in reference pixels, PDF user space is 1/72 inch by default, and a page in PDF 2.0 carries up to five boxes — media, crop, bleed, trim and art — that a print vendor will read and a browser will mostly ignore. Getting a 3 mm bleed right is not a rendering problem. It is a coordinate-system problem.
Stamping a typeset original instead of typesetting it again
The change that moves the dominant term is to separate the two operations that a browser pipeline fuses together. Lay the document out once, store the result as an immutable original addressed by its SHA-256, and treat every per-recipient watermark as an overlay drawn onto the existing page content streams. No text is re-measured. No line is re-broken. Cost becomes a function of page count and storage round-trips, and the batch turns from CPU-bound into I/O-bound, which is a much cheaper kind of bound to be.
| Pipeline shape | Where the cost lands | Typical implementations | Main limit |
|---|---|---|---|
| Re-run HTML layout per output | Layout CPU, memory ceiling | headless Chromium print path, Puppeteer, Playwright | pays the pagination cost on every copy |
| Dedicated paged-media engine per output | Layout CPU, better print fidelity | WeasyPrint, PrinceXML | same per-copy cost, stronger CSS paged media coverage |
| Overlay onto a typeset original | Storage round-trips | PDF object-model libraries such as pdf-lib or ReportLab | the mark cannot influence pagination |
Coming from ledger systems, the part I would not compromise on is that an issuance is an event and the artifact is a projection of it. Each watermarked copy gets a row, the row is keyed idempotently, and a retried batch converges instead of duplicating.
package issuance
import (
"crypto/sha256"
"encoding/hex"
"time"
)
// Record is the only per-recipient state we keep. The watermarked PDF is
// derivable from OriginalSHA256 plus these fields, so we stop storing the file.
type Record struct {
OriginalSHA256 string // content address of the typeset PDF
RecipientID string
PolicyVersion string // watermark template, opacity, placement rules
RendererDigest string // pinned image digest of the stamping toolchain
FontBundleSHA string // subsetted fonts embedded in the original
Pages int
IssuedAt time.Time
}
// Key is the idempotency key for one issuance. IssuedAt is deliberately not in
// the digest: a retry an hour later must land on the same key, not a new one.
func (r Record) Key() string {
h := sha256.New()
for _, part := range []string{r.OriginalSHA256, r.RecipientID, r.PolicyVersion, r.RendererDigest} {
h.Write([]byte(part))
h.Write([]byte{0x1f}) // separator, so ("ab","c") never collides with ("a","bc")
}
return hex.EncodeToString(h.Sum(nil))
}
The batch loop then has one job: reserve before you draw, and treat an already-reserved key as success rather than as an error, because a partially completed overnight run is the normal case and not the exceptional one.
// imports elided: context, errors, io, golang.org/x/sync/errgroup
// Stamper draws an overlay on every page of an already-paginated PDF. It never
// re-runs layout, so cost scales with page count rather than document complexity.
type Stamper interface {
Stamp(ctx context.Context, src io.ReaderAt, size int64, dst io.Writer, o Overlay) error
}
func StampBatch(ctx context.Context, jobs []Record, s Stamper, workers int) error {
g, ctx := errgroup.WithContext(ctx)
g.SetLimit(workers) // size this to storage round-trips, not to CPU cores
for _, job := range jobs {
job := job
g.Go(func() error {
switch err := ledger.Reserve(ctx, job.Key()); {
case err == nil:
return stampOne(ctx, s, job)
case errors.Is(err, ledger.ErrAlreadyIssued):
return nil // a retry of last night's partial batch
default:
return err // let errgroup cancel the rest, then back off and resume
}
})
}
return g.Wait()
}
The catch is real and it has two parts. An overlay cannot change pagination, so if the watermark has to reflow into the text, or the recipient gets a personalised cover page that renumbers everything behind it, you are back to full layout and you should stick with a paged-media engine for that tier of documents. And an overlay is trivially strippable — anyone with a PDF toolkit can delete the top content stream in a minute. Overlay watermarking is an attribution aid for accidental forwarding, not a control against a determined leaker. If that is your threat model, forensic marking inside the glyph positioning or in a rasterised layer is the honest answer, and it is slower per page by a wide margin.
What you stop keeping, and what that costs when a file leaks
Here is the deliberate deletion. The per-recipient watermarked files do not survive the sharing window; a lifecycle rule expires them after 30 days. What survives is the immutable original and one ledger row per issuance, and those rows are tiny — a few hundred bytes against 1.4 MB of PDF. The 21 TB projection collapses to the originals plus a table you can keep for a decade without noticing.
Then something leaks, and you pay for that choice.
Recovery is now a derivation rather than a lookup: minutes instead of milliseconds, and only correct if every input to the original stamping run is still reachable. That means the renderer digest has to be a real pinned digest you can still pull, the font bundle has to be content-addressed and retained alongside the originals, and the watermark policy has to be versioned rather than edited in place. Drop the font bundle during a cleanup, or upgrade the stamping library without bumping PolicyVersion, and you can still assert from the ledger which recipient received which copy — but you cannot reproduce the artifact byte for byte. An assertion is weaker evidence than a reproduction, and in an argument about a leaked contract that gap is the whole conversation.
There is also a hard legal floor under all of this, and it does not care about your storage bill. Records rules of the SEC Rule 17a-4 family require covered records to be preserved as they were, either on non-rewriteable media or under an audit-trail alternative, and regeneration on demand is not preservation. I am not the person who signs that off and neither are you, so get the carve-out in writing: tenants under a records obligation keep their artifacts, everyone else gets the derivation path.
So the decision rule I would write on the wall is narrow. If the mark participates in layout, re-typeset and accept the CPU. If it does not, typeset once, stamp many, keep the original and the ledger, and let the copies expire — unless a regulator has an opinion, in which case the regulator wins.
References and further reading
- ISO 32000-2, Portable Document Format — https://www.iso.org/standard/75839.html
- W3C CSS Paged Media Module Level 3 — https://www.w3.org/TR/css-page-3/
- W3C CSS Fragmentation Module Level 3 — https://www.w3.org/TR/css-break-3/
- MDN,
@page— https://developer.mozilla.org/en-US/docs/Web/CSS/@page - Chrome DevTools Protocol,
Page.printToPDF— https://chromedevtools.github.io/devtools-protocol/tot/Page/#method-printToPDF - Reproducible Builds, timestamps and build nondeterminism — https://reproducible-builds.org/docs/timestamps/
- 17 CFR 240.17a-4, records preservation — https://www.ecfr.gov/current/title-17/chapter-II/part-240/section-240.17a-4
Top comments (0)