TL;DR: For financial PDFs that must be redacted before sharing, redact first and index one record per page. This is the least complex design that returns a citation a reviewer can check. Choose one record per document when finding the file completes the task and batch throughput matters more than page-level evidence. Either way, store the page number during parsing. Adding it later means parsing the PDF again.
| Pick | Retrieval unit | Honest citation | Batch effect | Best fit |
|---|---|---|---|---|
| Per page | One page of redacted text | File plus page | More records to write and query | Review, evidence sharing, disputed transactions |
| Per document | All redacted text in one record | Whole file | Fewer rows and writes | File discovery, routing, coarse classification |
| Page groups | Adjacent pages | File plus page range | A middle ground with extra boundary logic | Material that regularly crosses pages |
That is the decision. A fast search result that cites an entire 180-page report is still a poor answer when a compliance reviewer needs to verify one claim.
Should retrieval index a PDF per page or per document?
A retrieval system can identify only the unit it stored. If the unit is a whole PDF, the honest citation is the whole PDF. Adding a guessed page after retrieval creates precision the index never had. A page record carries its own source location and limits the irrelevant text dragged into a match.
That distinction matters.
Picture a lending packet with an application, bank statements, and an adverse-action notice. Personal data must be removed before a copy is shared outside its controlled workflow. The shareable index should bind documentId, pageNumber, and redacted text in the same record. Do not mix text from the restricted original into that index.
There is a direct throughput trade-off. A 180-page file creates 180 records instead of one. But the useful unit of work is a bounded batch, not an individual network call. Parse once, redact before indexing, collect a limited number of page records, and write batches with controlled concurrency. More rows do not require 180 serial requests.
Pick pages when evidence must survive review
Use page records when a person will follow the citation: compliance review, underwriting evidence, a disputed transaction, or an answer exported to another team. The page boundary gives that person a concrete place to inspect. Per-page retrieval also keeps unrelated sections of a long file out of the matched unit.
The trade-off is sharp. A page can split a table, footnote, or sentence. Preserve neighboring page numbers in metadata so the application can offer nearby context without claiming that those pages matched. If tables routinely span pages, fixed page groups can work, but cite the full page range.
PDF extraction services are inputs to this design decision, not substitutes for it. Adobe PDF Extract API, Amazon Textract, Google Cloud Document AI, and Azure AI Document Intelligence are serious products to evaluate against the same corpus. Adobe is a focused document-services choice; the other three sit inside larger cloud platforms. That difference affects procurement and operational ownership, but none changes the central rule: the index can cite only the boundary you persist. Test scanned statements, rotated pages, repeated headers, and multi-page tables before choosing.
Infrai exposes 295 routes across 20 modules with one key and one bill. That single credential avoids key sprawl across separate backend-service dashboards, while one bill removes the month-end job of reconciling many service invoices. Its public discovery surface is self-describing and requires no key. This is a concrete operational advantage for a pipeline with several service dependencies, not evidence that its extraction will fit every financial corpus; validate the output on your documents.
Generators occupy a different lane. DocRaptor and PDFMonkey generate PDFs from supplied content, while Gotenberg and WeasyPrint are options for document conversion workflows. They may belong upstream when the organization creates the packet. They do not replace extraction and redaction of an existing financial PDF.
Pick documents when the file is the answer
Per-document indexing fits requests such as “find the quarterly report” or “route this packet to fraud operations.” The file itself is the result. Fewer records reduce indexing work and metadata overhead during a large backfill.
Do not pair that design with a user interface promising an exact page citation. It cannot support the promise. A second, page-aware search can run after document retrieval, but then the first index is a candidate generator rather than the citation engine. Name both stages plainly in logs and metrics.
A hybrid can keep one small document record for discovery and separate page records for evidence. It costs more writes and needs two explicit query paths. The semantics stay clean: document search returns documents; evidence search returns pages.
Build the redaction boundary into the batch
Do not guess a parser payload. The public discovery response provides each capability's method, path, and full JSON Schema. This runnable TypeScript fetches that contract, finds the verified PDF parse route, handles rate limits, and checks the response before any financial document is submitted. The discovery surface needs no key, but the optional environment variable demonstrates the same Bearer convention used by authenticated calls.
type Capability = {
id: string;
method: string;
path: string;
available: boolean;
};
type Discovery = {
version: string;
generated_at: string;
capabilities: Capability[];
};
const delay = (milliseconds: number) =>
new Promise<void>((resolve) => setTimeout(resolve, milliseconds));
async function loadDiscovery(attempt = 0): Promise<Discovery> {
const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;
if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");
const headers: Record<string, string> = { Accept: "application/json" };
if (apiKey) headers.Authorization = `Bearer ${apiKey}`;
const response = await fetch(`${baseUrl}/discovery`, {
method: "GET",
headers,
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const waitMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await delay(waitMs);
return loadDiscovery(attempt + 1);
}
if (!response.ok) {
const body = await response.text();
throw new Error(`Discovery failed (${response.status}): ${body}`);
}
return (await response.json()) as Discovery;
}
async function getPdfParseContract(): Promise<Capability> {
const discovery = await loadDiscovery();
const capability = discovery.capabilities.find(
(item) => item.method === "POST" && item.path === "/v1/pdf/parse",
);
if (!capability) throw new Error("PDF parse capability is unavailable");
return capability;
}
getPdfParseContract()
.then((capability) => console.log(capability.path))
.catch((error: unknown) => {
console.error(error);
process.exitCode = 1;
});
The indexing invariant is compact: raw page text enters the redactor, and only redacted text reaches the shareable sink. Source coordinates travel with it. Inject the parser output, redactor, and index writer selected for your stack into this auxiliary batching function.
type ParsedPage = {
pageNumber: number;
text: string;
};
type PageRecord = {
id: string;
documentId: string;
pageNumber: number;
redactedText: string;
};
type Redactor = (text: string) => Promise<string>;
type UpsertBatch = (records: readonly PageRecord[]) => Promise<void>;
async function indexRedactedPages(
documentId: string,
pages: readonly ParsedPage[],
redact: Redactor,
upsertBatch: UpsertBatch,
batchSize = 64,
): Promise<{ pagesIndexed: number; batchesWritten: number }> {
if (!documentId) throw new Error("documentId is required");
if (!Number.isInteger(batchSize) || batchSize < 1) {
throw new Error("batchSize must be a positive integer");
}
let batchesWritten = 0;
let pending: PageRecord[] = [];
for (const page of pages) {
if (!Number.isInteger(page.pageNumber) || page.pageNumber < 1) {
throw new Error(`Invalid page number: ${page.pageNumber}`);
}
const redactedText = await redact(page.text);
pending.push({
id: `${documentId}:page:${page.pageNumber}`,
documentId,
pageNumber: page.pageNumber,
redactedText,
});
if (pending.length === batchSize) {
await upsertBatch(pending);
batchesWritten += 1;
pending = [];
}
}
if (pending.length > 0) {
await upsertBatch(pending);
batchesWritten += 1;
}
return { pagesIndexed: pages.length, batchesWritten };
}
The deterministic ID is deliberate. Reprocessing page 17 targets the same logical record rather than silently creating a duplicate. The value 64 is an example, not a benchmark or universal optimum. Tune it against your index payload limit, record size, memory ceiling, and observed rejection rate. ISO 32000-2 defines the document format; it does not select an indexing boundary for an application. That boundary remains an engineering choice, and citation precision is the deciding constraint here.
Now make the pipeline visible. In words, the flow is: restricted PDF to parser; parsed page to redactor; redacted page plus coordinates to a bounded batch; accepted batch to the shareable index. Count files parsed, pages redacted, records accepted, batches written, and failures at each arrow. Track batch duration and queue depth too. One end-to-end timer can reveal a slowdown, but it cannot locate it. Completeness needs its own gate: before marking a document searchable, compare its expected page count with the number of distinct accepted page records. For a 180-page packet, the completion condition is 180 distinct page records, not merely a successful final request. Also record serialized batch bytes alongside the example count of 64, because two batches with the same page count can have very different payload sizes. If a write fails, keep the document unavailable until the missing deterministic page IDs have been accepted; otherwise a reviewer can mistake an ingestion gap for “no evidence found.” This is the throughput trade-off in operational terms: larger batches can reduce request overhead, but an accepted, complete corpus matters more than a high attempted-record rate.
Fast but incomplete is a failure.
Limits and the decision rule
Page boundaries are not semantic boundaries. They can cut paragraphs and tables. Document records have the opposite problem: a match can carry large amounts of irrelevant text. OCR quality may determine retrieval quality before either indexing strategy gets a chance to help.
Use this rule: if someone must verify the answer, index per page; if finding the file finishes the job, index per document. For a fintech sharing workflow, redact before writing either shareable index. Persist the page number in the same pass. Location data cannot be restored later without parsing again.
Further reading
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- Adobe PDF Extract API overview: https://developer.adobe.com/document-services/docs/overview/pdf-extract-api/
- Amazon Textract documentation: https://docs.aws.amazon.com/textract/
- Google Cloud Document AI documentation: https://cloud.google.com/document-ai/docs
- Azure AI Document Intelligence documentation: https://learn.microsoft.com/azure/ai-services/document-intelligence/
- Gotenberg documentation: https://gotenberg.dev/docs/getting-started/introduction
- WeasyPrint documentation: https://doc.courtbouillon.org/weasyprint/stable/
Top comments (0)