DEV Community

DarkveilCorvyn26
DarkveilCorvyn26

Posted on

Receipt PDF Background Jobs: Signed Evidence for Attached Order Confirmations

The page says confirmation_without_receipt, and the on-call view is brutally small: an order identifier, a queue age, and a message stating that email delivery finished while its PDF attachment did not. The customer has a confirmation, support has a complaint, and a dashboard full of green averages cannot answer the first useful question: what page fired, and which durable fact proves the redacted receipt was the file actually sent?

TL;DR: accept the order in the web request, enqueue an immutable receipt job, render the PDF in a worker, redact personal data before the document crosses the sharing boundary, sign a digest of the final bytes plus the redaction policy version, and record that evidence before handing the email to a transport. Retry from durable state with idempotency keys. Never rebuild a supposedly identical attachment inside an email retry.

This is the operational shape a Node.js Express service needs even though the example below uses Go, because the safety boundary belongs in the protocol and persisted state rather than in one framework. The request handler should return after committing the order and an outbox record in one database transaction. A worker then owns render -> redact -> validate -> sign -> record -> send. Each transition must be recoverable after a process exit.

No signature, no send.

How should a Node.js background job attach a receipt PDF to an order confirmation?

Work backward from the bad page. confirmation_without_receipt is a late customer-impact signal. The earlier page should fire when the oldest ready job approaches the delivery objective while attempts continue to fail, or when an order remains in an impossible state such as email_sent without a recorded artifact digest. Queue depth alone is weak: a deep queue may drain quickly, while one poisoned job can wait forever in a shallow queue.

A useful event carries order_id, job_id, attempt, state, artifact_digest, policy_version, and timestamps for the last successful transition. It must not carry the patient's name, email address, medical record number, or unredacted document text. The alert points to an audit record; it does not become another copy of regulated data.

Instrument the state transitions, not just the worker process. Count completed and failed transitions by failure class, measure age from the durable enqueue time, and test the invariant joining order, artifact, and email records. Logs answer what one worker did. The invariant answers whether the system kept its promise.

This choice has a concrete trade-off: more persisted transitions mean more schema and recovery logic, but fewer transitions leave the on-call engineer guessing whether a retry can disclose the wrong bytes. For health data, that uncertainty is unacceptable.

Make the audit record bind the exact attachment

PDF is a file format standardized by ISO 32000-2:2020, but a valid PDF is not proof that disclosure was appropriate. Redaction must remove the underlying information rather than merely draw an opaque rectangle over visible text. For US health information, the HIPAA de-identification guidance describes two methods, Expert Determination and Safe Harbor; neither can be replaced by a home-grown list called sensitiveFields.

The evidence record should bind the final attachment bytes to the decision that permitted sharing. Hash the completed PDF after redaction and validation, then sign a canonical statement containing the digest, order identifier, policy version, signer key identifier, and creation time. NIST FIPS 186-5 specifies approved digital-signature techniques; key custody and rotation remain system responsibilities. A signature proves that signed data has not changed undetected and identifies the signing key. It does not prove that the redaction policy was correct.

This compact Go model makes the boundary explicit:

package receipt

import (
    "crypto/sha256"
    "encoding/hex"
    "time"
)

type Evidence struct {
    OrderID     string    `json:"order_id"`
    ArtifactSHA string    `json:"artifact_sha256"`
    Policy      string    `json:"policy_version"`
    KeyID       string    `json:"key_id"`
    CreatedAt   time.Time `json:"created_at"`
    Signature   []byte    `json:"signature"`
}

func Digest(finalPDF []byte) string {
    sum := sha256.Sum256(finalPDF)
    return hex.EncodeToString(sum[:])
}
Enter fullscreen mode Exit fullscreen mode

Sign a deterministic serialization, not whatever map ordering a runtime happens to emit. Store the signature beside the evidence statement and restrict access to the PDF separately. An audit trail full of patient data creates a second disclosure surface; identifiers in operational records should be opaque references wherever investigators do not need the underlying value.

Treat rendering and sending as separate failure domains

The background job needs durable checkpoints because PDF generation and email submission fail differently. Rendering can fail on malformed input, font handling, resource limits, or validation. Email submission can time out after the receiving service accepted the message, leaving the worker uncertain. A single failed state erases the distinction and encourages unsafe retries.

Retries lie.

Use explicit states such as queued, rendered, redacted, verified, signed, submitted, and delivery_failed. Persist the content digest when the artifact becomes immutable. On retry, compare it before reuse. If input or policy changes, create a new artifact version and new evidence record rather than overwriting history.

MIME defines how the PDF is represented as an attachment, while SMTP governs message transfer; neither provides exactly-once delivery. Give each logical confirmation a stable idempotency key, save the transport's accepted-message identifier when available, and reconcile uncertain submissions before sending again. The customer impact of a duplicate receipt is usually smaller than a privacy leak, but duplicates still trigger support work and can confuse a financial record. That trade-off should be written down, not hidden inside retry middleware.

The web route stays narrow: validate the request, commit the business record plus outbox entry atomically, and return an order response. The worker claims the outbox item with a lease, renews that lease during long rendering, and releases it only after a durable transition. Exponential backoff needs a cap and jitter; permanent policy or validation failures go to review rather than cycling forever.

Test the evidence path, then deploy it gradually

Unit tests should prove canonical serialization and digest stability. Integration tests should kill the worker after every transition, including the uncomfortable instant after email acceptance but before the database update. Feed the redactor fixtures containing text, metadata, annotations, embedded files, and image-only pages; then inspect extracted text and rendered pixels. No single check covers both the logical and visual layers of a PDF.

Run policy changes in shadow mode against representative, authorized fixtures before enforcing them. Compare decisions without emitting documents, review disagreements, then canary by a bounded cohort. Keep signing keys outside application configuration, record key identifiers rather than secrets, and rehearse verification with retired public keys before rotation day.

Retention deserves its own control. The order record, audit evidence, PDF, and mail-provider metadata do not need identical lifetimes. Define deletion and legal-hold behavior per class, then test that deletion removes every intended copy without destroying the minimal evidence required by policy.

The operational decision rule is plain: no email job may read an attachment unless a verified evidence record names the same digest. That makes a database invariant enforce the privacy boundary, and it gives an incident responder something stronger than a screenshot of a healthy dashboard.

Thresholds carry a bill. Page on every first failure and transient renderer or mail errors will train the on-call engineer to ignore the channel; wait until confirmations are already missing and the first actionable signal arrives from support. Start from the documented delivery objective, page on sustained job age plus failed progress, ticket low-rate permanent failures, and review the alert after every incident. A page must imply an action. Otherwise it is noise with a ringtone.

References

Further reading

Top comments (0)