Short answer: put a finite deadline around polling, model success and failure as explicit terminal states, and preserve an append-only event trail for every transition. For scanned freight documents, also bind the searchable output to the input digest and signature evidence. That is the least complex design that can distinguish a slow renderer from a lost job without waiting forever or weakening the audit trail.
Start with this decision table. It separates recovery choices by what the evidence says, rather than by how long an operator has been staring at a spinner.
| Observation | Likely boundary | Pick this action | Evidence to retain |
|---|---|---|---|
| State advances and the lease remains fresh | OCR or rendering is slow | Keep polling inside the caller's deadline | Attempt, state sequence, timestamps |
State is processing, but its lease expired |
Worker disappeared or lost ownership | Requeue with a new attempt identifier | Old lease, new lease, reason |
| Poll requests fail intermittently | Status path is unhealthy | Retry with bounded backoff and jitter | Error class and request correlation ID |
The same input repeatedly reaches failed
|
Deterministic input or processing error | Stop automatic retries | Input digest and structured failure code |
| Output exists, but signature evidence does not match | Integrity check failed | Quarantine the result | Digests, signer evidence, verification result |
The key rule is blunt: processing is not permission to poll forever. It is a temporary state whose owner, lease, and deadline must be observable.
How should you debug a PDF job stuck in progress forever?
An asynchronous pipeline has several clocks. The API accepted the job at one time. A queue made it visible later. A worker acquired it, renewed a lease, ran OCR, assembled searchable text with page images, and committed the result. The polling client sees only a projection of those events. If that projection says processing after the worker has vanished, the client cannot infer the truth from the label alone.
Picture the path in words: scanner to object store; object store to job record; job record to queue; queue to worker; worker to OCR and PDF assembly; result back to storage; status back to the caller. Every arrow is a failure boundary. Acknowledging a queue message before committing status can strand the record. Committing output before publishing the final state can hide a valid result. Letting two workers share an attempt can make an older worker overwrite newer evidence.
This gets sharper with signed freight paperwork. A bill of lading or delivery receipt may arrive as several scanned pages, and the useful result is both a viewable PDF and searchable text. The system must preserve which bytes were processed. Replacing the source beneath an existing job identifier produces a plausible document with the wrong lineage. A digest makes that mismatch visible; a display name does not.
Names drift. Bytes do not.
PDF conformance is a separate concern. ISO 32000-2 defines the PDF format, but a conforming file format does not define the lifecycle of an application job. Queue leases, polling deadlines, OCR progress, and retry policy belong to the surrounding system. Keep those contracts separate. It prevents a processing timeout from being reported as a malformed PDF merely because both paths end in failure.
Pick the recovery path from evidence
Keep polling when progress is recent and meaningful. A rising completed-page count, a renewed lease, or a newer stage timestamp can justify another observation. Do not reset the caller's total deadline each time progress moves. That turns many small extensions into an unlimited wait.
Requeue when ownership is stale, not merely because the task is old. A large scan may legitimately take longer than a small one. Lease expiry provides a stronger signal: the worker responsible for the current attempt stopped proving liveness. Create a new attempt, record why, and fence the old one from committing. The job identifier can remain stable while the attempt identifier changes.
Stop when evidence points to a deterministic failure. Repeating OCR against the same unreadable bytes without changing input or processing policy burns capacity and muddies the audit trail. A terminal failed state should carry a stable machine-readable code and a safe operator-facing summary. Raw document text and signature material do not belong in a log line.
Quarantine is different from ordinary failure. Use it when output was produced but the relationship among source bytes, rendered bytes, searchable text, and signature evidence cannot be established. The artifact may be readable. It is still unfit for downstream release.
Consider one diagnostic walk-through. A freight scan is accepted as job J, the queue assigns attempt A1, and a worker changes the state from queued to processing. Polling then keeps returning that state. The operator should resist pressing a generic retry button. First compare the current time with A1's lease, then inspect its last durable transition and look for an output committed under the same input digest. A fresh lease means the worker still owns the attempt, so the caller may continue only within its original polling budget. An expired lease with no output permits a fenced A2; the job record must reject any later commit from A1. An output whose digest is present but whose terminal transition is missing calls for reconciliation, not another render. And an output tied to a different input digest goes to quarantine. This sequence matters because four identical spinners lead to four different actions. The status word alone never supplied enough evidence.
That gives four real outcomes: completed, failed, cancelled, and quarantined. Timeout is a caller outcome, not necessarily a job outcome. A client may stop waiting while the server continues processing, provided the client receives a durable job identifier and can reconcile later.
Implement one bounded, auditable poller
The poller below uses generic interfaces. It has two budgets: a maximum number of observations and an absolute elapsed-time deadline. Either one can stop the wait. It also validates transitions, returns terminal failures as data, and treats an unknown state as a protocol error instead of silently looping.
type ActiveState = "queued" | "processing";
type TerminalState = "completed" | "failed" | "cancelled" | "quarantined";
type JobState = ActiveState | TerminalState;
type Snapshot = {
jobId: string;
attemptId: string;
state: JobState;
observedAt: string;
inputSha256: string;
outputSha256?: string;
failureCode?: string;
};
type PollOptions = {
deadlineMs: number;
maxPolls: number;
initialDelayMs: number;
maxDelayMs: number;
signal?: AbortSignal;
};
const allowed: Record<JobState, ReadonlySet<JobState>> = {
queued: new Set(["queued", "processing", "failed", "cancelled"]),
processing: new Set([
"processing",
"completed",
"failed",
"cancelled",
"quarantined",
]),
completed: new Set(["completed"]),
failed: new Set(["failed"]),
cancelled: new Set(["cancelled"]),
quarantined: new Set(["quarantined"]),
};
const terminal = new Set<JobState>([
"completed",
"failed",
"cancelled",
"quarantined",
]);
const sleep = (ms: number, signal?: AbortSignal) =>
new Promise<void>((resolve, reject) => {
const timer = setTimeout(resolve, ms);
signal?.addEventListener(
"abort",
() => {
clearTimeout(timer);
reject(signal.reason ?? new Error("Polling aborted"));
},
{ once: true },
);
});
export async function waitForRender(
read: (signal?: AbortSignal) => Promise<Snapshot>,
options: PollOptions,
): Promise<Snapshot> {
const started = Date.now();
let previous: Snapshot | undefined;
let delay = options.initialDelayMs;
for (let poll = 1; poll <= options.maxPolls; poll += 1) {
if (Date.now() - started >= options.deadlineMs) {
throw new Error("Polling deadline exceeded; reconcile by job ID");
}
const current = await read(options.signal);
if (previous && !allowed[previous.state].has(current.state)) {
throw new Error(`Invalid transition: ${previous.state} -> ${current.state}`);
}
if (terminal.has(current.state)) return current;
previous = current;
const jitter = Math.floor(Math.random() * Math.max(1, delay / 4));
await sleep(delay + jitter, options.signal);
delay = Math.min(delay * 2, options.maxDelayMs);
}
throw new Error("Polling budget exhausted; reconcile by job ID");
}
The client should log one structured observation per poll, but keep cardinality under control. state, failureCode, and processing stage are useful metric dimensions. A job ID, attempt ID, document digest, or shipment number belongs in traces or logs, not metric labels. Otherwise each document creates a new time series.
Use a correlation ID to join submission, queue delivery, worker attempt, and poll observations. The audit event itself needs more than a message string: record the job and attempt IDs, prior and next states, event time, actor or component, input digest, reason code, and a monotonically increasing version. A conditional update on that version prevents an expired worker from writing over the current attempt.
Sign the evidence envelope, not a mutable status page. The envelope can bind the source digest, output digest, OCR text digest, attempt, terminal state, and completion time. Verification then answers a precise question: do these artifacts and this decision belong together? Signature validity does not prove OCR accuracy, so keep content-quality checks separate.
Observe the state machine, not the spinner
A dashboard should make abandonment visible. Track jobs entering each state, terminal outcomes, lease expirations, retry attempts, and end-to-end latency. For latency, retain a distribution rather than only an average; a small set of very slow scans is exactly what an average conceals. Alert on symptoms that require action, such as an oldest active-job age beyond the operational objective combined with no recent lease renewal.
Logs explain individual transitions. Metrics show population behavior. Traces connect the handoffs. None can replace the durable job record, which remains the source for reconciliation after a client deadline.
Test the awkward boundaries deliberately. Kill a worker after it writes the output but before it commits completed. Deliver the same queue message twice. Expire a lease while OCR is running. Change the source object after submission and verify that the digest check rejects it. Abort the poller during backoff. Attempt a backward transition from processing to queued without creating a new attempt. These tests are more valuable than another happy-path snapshot because they exercise the places where “forever” begins.
A crisp before/after helps during an incident. Before: “the job is still processing.” After: “attempt 3 stopped renewing its lease at a recorded time; attempt 4 owns the job; the caller ended its wait after its budget; no terminal artifact has been released.” The second statement gives an operator a next action and preserves what happened.
Limits of this field guide
Bounded polling does not make rendering faster, and leases do not prove that a worker is healthy; they prove only that it renewed ownership. Progress counters can also lie if a stage reports work before its transaction commits. Treat them as diagnostic signals, not authority.
There is a real trade-off. This design adds durable events, lease fencing, and reconciliation work. It is not appropriate for a synchronous, disposable conversion where the caller can safely start over and no signed record or audit history must survive. In that case, a request timeout and one idempotent retry may be enough. It is also a poor fit when workers cannot make conditional writes; use a queue or database that can enforce single-attempt ownership before claiming that stale workers are fenced.
No state machine repairs bad OCR.
The release decision still depends on the organization's retention rules, signature policy, threat model, and acceptable OCR error. The durable pattern is narrower: explicit terminal states, fenced attempts, finite client budgets, content digests, and an audit trail that can be verified after the fact. That is enough to turn an endless spinner into a finite diagnosis.
Further reading
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
Top comments (0)