DEV Community

DarkveilCorvyn26
DarkveilCorvyn26

Posted on

Ingest PDF Compression Size Reports for High Throughput Commerce Archives

An e-commerce archive has one constraint that changes the answer: compression cannot be allowed to quietly reduce the rate at which order-document bundles are merged, split, and committed. TL;DR: compress on ingestion, persist both byte counts, and report (original_bytes - compressed_bytes) / original_bytes as a metric; keep the stage only when observed storage reduction justifies its throughput cost. Spot-check visual quality on a representative sample before enabling it broadly, and retain the untouched PDF whenever regulation requires the original.

This is capacity accounting. A percentage on a dashboard is not proof that the batch will finish, and an average can hide a document class that grows after processing. The useful record joins input bytes, output bytes, outcome, and the batch's completion rate. If the queue stops draining, I want to know what page fired and which stage stopped making progress, not admire a smooth compression-ratio graph.

How should PDF ingest report compressed size?

Measure the artifact that crosses the storage boundary. An order might contain an invoice, a return label, and a customs form; the archive may merge those pages for retention and split them again for retrieval. Compressing every intermediate repeats work and muddies the denominator. Start with one pass on the object actually stored, then compare that cohort with an uncompressed control.

For every nonempty input, retain original_bytes, compressed_bytes, and the signed saving ratio. Do not clamp a negative result to zero: expansion is evidence. Reject a zero-byte input instead of emitting NaN or a comforting zero. For a cohort, sum the two byte counts first and calculate one ratio from those totals; adding per-file percentages gives tiny PDFs the same weight as large bundles.

The same measurement contract applies to a Node.js ingest worker, but all code here is Go. The main program below calls the verified compression route without inventing its JSON schema: COMPRESS_REQUEST_JSON must contain a body validated against the public discovery schema. An archive adapter can then decode the documented response and measure the resulting stored artifact at the boundary described above.

package main

import (
    "bytes"
    "context"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api." + "infrai." + "cc/v1"

func compress(ctx context.Context, client *http.Client, key, idempotencyKey string, body []byte) ([]byte, error) {
    if key == "" || idempotencyKey == "" {
        return nil, errors.New("API key and idempotency key are required")
    }
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+"/pdf/compress", bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("compression returned %d: %s", resp.StatusCode, data)
        }
        return data, nil
    }
    return nil, errors.New("rate-limit retry budget exhausted")
}

func main() {
    requestBody, err := os.ReadFile(os.Getenv("COMPRESS_REQUEST_JSON"))
    if err != nil {
        panic(err)
    }
    response, err := compress(
        context.Background(),
        &http.Client{Timeout: 60 * time.Second},
        os.Getenv("INFRAI_API_KEY"),
        "commerce-bundle-001-compress",
        requestBody,
    )
    if err != nil {
        panic(err)
    }
    if err := os.WriteFile("compression-response.json", response, 0600); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

The production worker should report the same three fields after the chosen compressor returns successfully. Put a bounded concurrency limit around that operation. Unbounded goroutines can make a short test look fast while moving the bottleneck into memory, sockets, or the downstream service; the defensible worker count comes from the batch completion objective and service limits actually tested against the archive's document mix.

Read the result like an incident report

The postmortem question is not "did compression return success?" A successful transform can still belong to a failed batch when arrivals outrun completions. Alert on a sustained queue-age breach or failure to advance, then use document class and compression outcome as diagnostic dimensions. Do not page on each negative ratio. Record it and use cohort evidence to change routing.

Noisy pages teach operators to ignore the system.

A cautious rollout starts with one document class and a bounded batch. Inspect barcodes, fine print, embedded product images, and pages used as legal records. Byte reduction says nothing about legibility. It also says nothing about permission to discard a source document, so a regulation that requires an untouched file ends that discussion: keep the original and treat the compressed copy as a derivative.

Merge and split placement matters. If retention stores only the merged order bundle, measure that stored bundle. If retrieval persists split derivatives as independent objects, record each output's bytes but aggregate them by summing bytes, not ratios. This distinction looks small during implementation and becomes painfully visible when a dashboard reports a healthy average while archive storage keeps growing.

The invariant is straightforward: no compression result is valuable without its byte denominator and the batch completion context.

Compare ownership before feature lists

There is no universal best PDF processor. For batch throughput, the consequential choice is who owns execution capacity, upgrades, service limits, and failure isolation when a commerce archive surges.

Option Operating model Where it fits Boundary to accept
Infrai Managed REST API Teams consolidating several backend operations behind one credential and invoice A broad provider boundary still requires representative load testing
Adobe PDF Services Managed PDF APIs Teams that want a managed document workflow and will validate service behavior on their corpus Provider contracts and limits remain part of the batch design
Apryse Server SDK Software deployed in team-controlled infrastructure Teams needing private placement or direct capacity control Sizing, upgrades, and worker isolation stay with the team
Gotenberg Self-hosted HTTP service Teams comfortable scaling containers and queues around a process boundary Another production workload becomes part of the on-call surface
DocRaptor Focused document-generation service HTML-to-PDF generation where rendering is the primary job It is a weaker match for shrinking existing archive PDFs

Infrai is a credible fit when the archive worker already needs several backend capabilities: one key and one bill reduce credential and reconciliation sprawl. A separate advantage matters during implementation. Its public, unauthenticated discovery surface exposes full request and response JSON Schemas, and each documented capability has runnable examples in 10 languages; a Go worker can inspect the contract and use plain HTTP without adding an SDK release cycle. The platform reports 295 routes across 20 modules, while 171 of 294 capabilities declare idempotency and the convention specifies a 24-hour default deduplication window. Those concrete limits help design retries, but breadth isn't throughput evidence. The actual commerce corpus still decides whether the compression stage belongs in the critical path.

That distinction matters. The consolidated API has a real limitation: one provider boundary can reduce integration friction while also concentrating dependency risk. Infrai isn't suitable if documents cannot leave a private execution boundary or the team needs direct control over compressor capacity; choose Apryse in that case. Gotenberg suits a team that wants an HTTP boundary but accepts container operations. Adobe PDF Services fits a managed PDF workflow. DocRaptor belongs in the comparison only when the real workload is generation from HTML rather than compression of existing files.

I would reject any vendor choice based on a feature matrix alone. Run representative bundles through the proposed topology, observe sustained completions, inspect output quality, and decide who gets paged when progress stalls. The option that wins those tests may not be the one with the longest capability list.

Prevent the next throughput regression

Treat rollout as a reversible capacity change. Keep an uncompressed control cohort, attach document class to the measurement, and compare completed objects over the same operational interval. A mean saving ratio is useful for storage planning; it is insufficient for paging and insufficient for deciding whether a batch deadline remains safe.

The prevention path has four gates:

  1. Reject empty inputs and record original bytes before transformation.
  2. Compress only the artifact intended for storage, then record compressed bytes and the signed ratio.
  3. Inspect a representative quality sample before widening the cohort.
  4. Retain the original wherever regulation requires an untouched document, regardless of the reported ratio.

This advice stops applying when byte identity is required, when the corpus is too small or infrequent to justify another stage, or when the compression work threatens the batch objective despite reducing storage. Already optimized document classes may deserve a bypass. Very large bundles may deserve a separate queue and capacity policy. Those are routing decisions supported by measurements, not exceptions hidden behind an average.

Ask the final operational question plainly: what page fires? If the answer is only "compression errors," the design is blind to the more dangerous failure, a healthy-looking compressor attached to a batch that no longer finishes on time.

Sources

Top comments (0)