DEV Community

AshtonBlake6879
AshtonBlake6879

Posted on

Filling AcroForm PDF Fields on Scanned Freight Docs (and Retaining Proof for Compliance)

If you just want a filled, searchable, tamper-evident copy of a scanned bill of lading, the least complex approach that survives an audit is four ordered steps: OCR the scan into an invisible text layer, fill the AcroForm fields programmatically, flatten the form, and only then sign. A Node.js service can drive all four, either through a remote API or a local library, and that choice matters far less than the ordering does.

Flatten before you sign. Never after.

I spend my week arguing about retention, so this is where I start: a document pipeline's bill is not made of API calls. It is made of page images you decided to keep, multiplied by a number of years a regulator picked for you.

Where the bytes actually go in a scanned-document pipeline

Take a freight forwarder running 40,000 scanned documents a month — bills of lading, delivery notes, customs declarations — averaging six pages at 300 dpi bitonal. Call the scan 300 KB. Every stage after that produces another artifact, and most teams keep all of them because nobody ever wrote down which ones are evidence and which ones are scratch.

Artifact per document Size on these assumptions Worth seven years of storage?
Source scan, 6 pages at 300 dpi 300 KB Yes — it is the evidentiary original
OCR sidecar, ALTO XML with per-word boxes 120 KB Sample only
Searchable PDF, image plus invisible text 320 KB No
Filled but unflattened PDF 330 KB No
Flattened, signed PDF/A 340 KB Yes
Audit events, 12 records of roughly 1.2 KB 14 KB Yes

That is about 1.42 MB per document, and 1.29 MB of it is four near-identical renditions of the same page images. The field values, the text layer and the whole audit chain together account for under 10% of the total. At 40,000 documents a month the archive grows by 56.8 GB monthly; over an 84-month retention window, with nothing ever deleted, that is roughly 4.8 TB of documents whose informational content would fit in a few hundred gigabytes.

Reducing that dominant term is not a compression problem. It is a decision about how many renditions of the same page you are willing to defend in front of an auditor, and the honest answer is two: the original you received, and the one you signed.

There is a second bill hiding behind the first, and it is the one I get called about. Each of those twelve audit events per document is usually emitted as a log line and a metric at the same time, and if the metric carries a document_id label, then 40,000 documents a month becomes 480,000 new time series a month in a store that charges by active series. Document identity belongs in the event record, where it is queried once a year during a dispute. The metric labels should stay at pipeline stage, form template and outcome — eight stages by thirty templates by four outcomes is 960 series, flat forever, and it answers every operational question you will actually ask at 3 a.m.

Should I fill and flatten AcroForm fields with an API or a local library in Node.js?

The deciding factor is not throughput or developer experience. It is where the signing key lives and which bytes cross a boundary.

An AcroForm is a document-level dictionary in the PDF catalog, described in ISO 32000-2, holding field dictionaries whose widget annotations carry the appearance streams a viewer draws. Filling a field means writing a value and generating the matching appearance stream. Skip that second half and set NeedAppearances, and your file renders correctly today in the viewer you tested and blank in the archival one two years from now, because you delegated rendering to software you don't control. Archival profiles expect the appearance to be in the file, not regenerated on open.

Flattening merges those appearance streams into the page content stream and removes the widget annotations, so the values become page content rather than editable state. No spec clause defines a single "flatten" operation — it's a composition of operations you can implement, buy, or get wrong.

curl -sS -X POST https://pdf-svc.internal/forms/fill \
  -H "Authorization: Bearer ${PDF_SVC_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "source_sha256": "9f2c1b...c41a",
    "flatten": true,
    "generate_appearances": true,
    "fields": {
      "bol_number": "MAEU-4471902",
      "gross_weight_kg": "18420",
      "consignee": "Nordfracht Logistik GmbH",
      "customs_ref": "DE-EORI-8841772"
    }
  }' \
  -o filled-flattened.pdf
Enter fullscreen mode Exit fullscreen mode

The remote option puts the customer's scanned document in someone else's memory space and adds a per-document line item that scales exactly with volume. The local option puts a pinned binary and its CVE surface in your container, and you own the upgrade treadmill. In Node.js, pdf-lib fills fields and flattens forms but does not produce a CMS signature, so signing comes from a separate component; Apache PDFBox covers both halves if a JVM in the path is acceptable to you. Either way, one thing to guard in CI: a flatten implemented by rasterizing pages erases the invisible text layer you just paid an OCR engine to produce, and the document silently stops being searchable.

My rule is narrow. If the signing key is in an HSM or a qualified service under a jurisdictional boundary, fill and flatten wherever you like, but keep the signature step on your side of that boundary. If your volume is under a few thousand documents a month and no key custody requirement applies, the operational cost of running your own rendering stack is probably not worth the unit-price saving, and a hosted call is the smaller system.

Why flattening after signing destroys the evidence

A PDF signature covers a ByteRange — the whole file except the hole where the signature container sits. Any byte written outside an incremental update changes the digest, and flattening rewrites the page content stream, which is about as outside as it gets. The signature that was supposed to prove the document's integrity is now the thing reporting that the document changed.

The order is fill, flatten, sign, timestamp.

This gets interesting when the shipper signs the blank form first, which happens more than the tidy diagram suggests. ISO 32000-2 defines a certification signature carrying a DocMDP transform with a permission value: 1 allows no changes at all, 2 allows form fill-in and signing, 3 additionally allows annotations. A form certified at 2 can still be filled by an incremental update that appends to the file without touching signed bytes — but flattening is off the table for the life of that document, because flattening is not a permitted change under any of the three values. You either keep the fields live forever, or you get the form unsigned and control the order yourself. I'd take the second, and I'd reject the first at intake rather than discover it in a batch job.

Add an RFC 3161 timestamp from a trusted authority. Certificates expire; a timestamp is what lets a verifier in 2033 conclude that the signature was valid when it was made rather than merely that it has since expired.

What US and EU retention rules force you to store

Both regimes ask for less than teams assume, and they ask for it more strictly.

Under ESIGN, 15 U.S.C. §7001(d), an electronic record satisfies a retention requirement when it accurately reflects the information and remains capable of accurate reproduction for later reference. Nothing there requires you to keep the intermediate renditions. It requires that the one you kept reproduces faithfully, which is the argument for a flattened archival profile rather than a live form whose displayed values depend on a viewer regenerating appearances.

In the EU, Regulation (EU) No 910/2014 sets the framework, and the PAdES baseline profiles in ETSI EN 319 142-1 define what a long-term PDF signature has to carry. The B-LT level embeds validation material so the signature can still be verified after the signing certificate expires, and B-LTA adds archival timestamps on top. That embedding is a storage decision disguised as a compliance checkbox: an OCSP response is a couple of kilobytes, while a full CRL from a large public CA can run into hundreds of kilobytes, and it is embedded per document because a Document Security Store cannot be shared across files. On a 40,000-document month, choosing CRLs over OCSP responses can quietly add more bytes than the page images do.

Retention length is the input I refuse to guess. Customs and carriage record requirements vary by jurisdiction and by document class; seven years is the number I used for the arithmetic above, and yours should come from your own counsel rather than from an engineering blog.

What I stop keeping, and what that costs on a bad day

I keep three things per document: the source scan with its SHA-256, the flattened and signed archival PDF, and a compact JSON manifest that maps each field to its value, its provenance, and the OCR engine version that produced it.

curl -sS -X PUT "https://archive.internal/objects/${DOC_ID}/manifest.json" \
  -H "Authorization: Bearer ${ARCHIVE_TOKEN}" \
  -H "Content-Type: application/json" \
  --data-raw '{
    "doc_id": "BOL-2026-0091883",
    "source_sha256": "9f2c1b...c41a",
    "signed_sha256": "37ab90...ee12",
    "ocr_engine": "tesseract 5.5.0",
    "fields": [
      {"name": "gross_weight_kg", "value": "18420", "source": "ocr", "confidence": 0.91},
      {"name": "consignee", "value": "Nordfracht Logistik GmbH", "source": "operator_correction"}
    ],
    "signed_at": "2026-09-12T09:14:22Z",
    "tsa": "rfc3161"
  }'
Enter fullscreen mode Exit fullscreen mode

The searchable-but-unsigned PDF and the filled-but-unflattened PDF are deleted as soon as the signed artifact verifies. The ALTO sidecar with per-word boxes and confidences is kept for a 2% random sample, plus every document that had a field below the confidence threshold or an operator correction — which is where disputes actually cluster, so the sample is not really 2% of the interesting population.

Here's the part I owe you, since it's the part that bites. When a consignee disputes a declared weight two years later, I can produce the signed document, the hash of the original scan, and a manifest saying the value came from OCR at 0.91 confidence with a named engine version. I cannot produce the alternative readings the engine considered, because I threw the sidecar away. If that document wasn't in the sample, reconstructing them means re-running OCR from the archived scan, and a different engine version may read the smudged digit differently — which is exactly why the version goes in the manifest, and why I'm not fully comfortable with this trade-off even though I keep making it.

This policy is wrong for two kinds of shop. Under roughly a thousand documents a month, storage isn't your constraint and deleting intermediates buys you nothing but risk. And if you need the complete processing record rather than the reproducible result — some regulators and insurers ask for exactly that — none of the above applies, so stick with keeping everything, put it in cold storage, and spend your engineering time on the lifecycle rules instead.

Everyone else is paying to store four pictures of the same page.

Further reading

Top comments (0)