DEV Community

KillianBerg5391
KillianBerg5391

Posted on

FastAPI Storefront Cropping: OCR-Safe Hero Images Across Responsive Slots

The least complex reliable design is to run OCR and text moderation on the original property photo, reject or redact disallowed text there, and only then derive every banner crop from that approved source. Use one focal region plus explicit slot dimensions to produce variants. Never let a desktop or mobile crop become a second moderation boundary.

Short answer: moderate the full-resolution source once, generate slot-specific candidates in Python, and accept a crop only when an offline eval shows that it preserves the property focal region without clipping approved text.

That ordering matters for property-management storefronts. A leasing photo can contain a phone number on a yard sign, a lockbox code near a door, a resident name on a parcel, or a street number. A narrow mobile banner might hide that text and pass OCR even though the wider desktop image exposes it. Cropping first creates a coverage gap; source-first moderation closes it.

This is a small pipeline, but it has two distinct decisions: whether the source is publishable, and how an approved source should be framed. Keep them separate.

Source first.

How should storefront hero images be cropped for desktop and mobile slots?

Treat each requested slot as a view onto the same approved image, not as an independent asset. The request should carry an image identifier, a focal box selected by a person or detector, and the exact pixel dimensions of the destination slots. OCR runs against the uncropped source. A policy function then classifies the extracted text. If the source fails policy, no crop is generated; if it passes, the cropper searches valid rectangles at each target aspect ratio.

The focal box is better than a single center point for property photography. A point can keep the front door visible while cutting off half the building. A box expresses the region that needs to survive. OCR boxes serve a different purpose: approved text should not be sliced through the middle, because half a word looks broken and can change what downstream OCR reads. The candidate score can reward focal-region coverage and penalize intersected text boxes.

One rule stays firm: a crop score cannot overrule moderation. If the original contains text prohibited by the storefront policy, reject or redact the source before generating banners. A beautiful crop isn't evidence that the underlying asset is safe.

The data flow is therefore plain: decode and validate the source, run OCR over the whole image, apply the text policy, find or receive a focal region, enumerate slot-shaped rectangles, score them, resize the winners, and record the source hash plus policy version beside every output. Desktop and mobile share the moderation result but get different crop coordinates.

Build the source-first crop path in FastAPI

The following example keeps OCR behind a callable interface. That makes the crop math runnable in a notebook with fixture boxes, while production can inject whichever OCR implementation the team has evaluated. It also prevents a model response schema from leaking through the rest of the application.

from __future__ import annotations

from dataclasses import dataclass
from io import BytesIO
from typing import Callable

from fastapi import FastAPI, File, HTTPException, UploadFile
from PIL import Image, ImageOps


@dataclass(frozen=True)
class Box:
    left: int
    top: int
    right: int
    bottom: int
    text: str = ""

    @property
    def area(self) -> int:
        return max(0, self.right - self.left) * max(0, self.bottom - self.top)


SLOTS = {
    "desktop": (1600, 600),
    "mobile": (750, 900),
}

DISALLOWED_LABELS = {"access_code", "resident_name", "personal_phone"}


def intersection_area(a: Box, b: Box) -> int:
    width = max(0, min(a.right, b.right) - max(a.left, b.left))
    height = max(0, min(a.bottom, b.bottom) - max(a.top, b.top))
    return width * height


def candidate_boxes(width: int, height: int, out_width: int, out_height: int) -> list[Box]:
    target_ratio = out_width / out_height
    if width / height >= target_ratio:
        crop_height = height
        crop_width = round(height * target_ratio)
        max_offset = width - crop_width
        return [
            Box(round(max_offset * step / 20), 0,
                round(max_offset * step / 20) + crop_width, crop_height)
            for step in range(21)
        ]

    crop_width = width
    crop_height = round(width / target_ratio)
    max_offset = height - crop_height
    return [
        Box(0, round(max_offset * step / 20),
            crop_width, round(max_offset * step / 20) + crop_height)
        for step in range(21)
    ]


def choose_crop(
    image_size: tuple[int, int],
    slot: tuple[int, int],
    focal: Box,
    approved_text: list[Box],
) -> Box:
    def score(candidate: Box) -> float:
        focal_coverage = intersection_area(candidate, focal) / max(1, focal.area)
        clipped_text = sum(
            1
            for box in approved_text
            if 0 < intersection_area(candidate, box) < box.area
        )
        center_x = (candidate.left + candidate.right) / 2
        focal_x = (focal.left + focal.right) / 2
        center_distance = abs(center_x - focal_x) / max(1, image_size[0])
        return 10 * focal_coverage - 3 * clipped_text - center_distance

    candidates = candidate_boxes(*image_size, *slot)
    return max(candidates, key=score)


def render_variants(
    source: Image.Image,
    focal: Box,
    approved_text: list[Box],
) -> dict[str, Image.Image]:
    outputs = {}
    for name, size in SLOTS.items():
        crop = choose_crop(source.size, size, focal, approved_text)
        region = source.crop((crop.left, crop.top, crop.right, crop.bottom))
        outputs[name] = ImageOps.fit(region, size, method=Image.Resampling.LANCZOS)
    return outputs


OCR = Callable[[Image.Image], list[tuple[Box, str]]]


def fixture_ocr(_: Image.Image) -> list[tuple[Box, str]]:
    return [
        (Box(80, 720, 410, 805, "Leasing Office"), "business_sign"),
    ]


app = FastAPI()


@app.post("/hero-crops")
async def create_hero_crops(file: UploadFile = File(...)) -> dict[str, object]:
    raw = await file.read()
    try:
        source = ImageOps.exif_transpose(Image.open(BytesIO(raw))).convert("RGB")
    except Exception as exc:
        raise HTTPException(status_code=422, detail="Invalid image") from exc

    findings = fixture_ocr(source)
    blocked = [label for _, label in findings if label in DISALLOWED_LABELS]
    if blocked:
        raise HTTPException(status_code=422, detail={"blocked_labels": blocked})

    focal = Box(
        source.width // 4,
        source.height // 5,
        source.width * 3 // 4,
        source.height * 4 // 5,
    )
    variants = render_variants(source, focal, [box for box, _ in findings])
    return {
        "source_size": source.size,
        "variants": {name: image.size for name, image in variants.items()},
    }
Enter fullscreen mode Exit fullscreen mode

The sample returns dimensions rather than storing files, so the interesting boundary remains visible. In a real service, the write happens only after both variants have been produced, and the database record should link them to the same immutable source digest and policy version. That prevents a later crop request from silently using different moderation rules.

The fixture_ocr function is deliberately boring. Replace it with an adapter whose output has been normalized to pixel coordinates and internal labels. Don't let provider-specific confidence fields flow directly into policy; calibration can differ by model and document type, while the application needs stable meanings such as personal_phone and access_code.

Score moderation coverage before visual quality

“Smart” often gets reduced to saliency: find the visually interesting area and center it. That is insufficient here. The primary metric is moderation coverage, because a visually plausible banner can still disclose sensitive text. Run the text policy over the entire decoded source, including orientation correction, before any geometry changes. Then measure crop quality only among sources that passed.

A useful eval set should be shaped like the incoming catalog, not like a general image benchmark. Include wide exterior photos, tall phone captures, dim hallway shots, oblique signs, parcel labels, reflections, small lockboxes, and images with no text. For each source, store the expected moderation labels, an acceptable focal region, and the required desktop and mobile slots. Keep the labels under version control. This is the notebook-to-prod bridge: the exact fixtures used to tune weights become regression tests in the service.

Pretty comes later.

Start with four measurements:

Measurement What it catches Release rule
Source text-policy recall Sensitive text missed before cropping Must not regress against the approved baseline
Focal-box coverage Building or room subject cut away Set a minimum per slot
Partial text intersections Words sliced at crop edges Prefer zero; review exceptions
Decode and orientation failures Inputs that never reach OCR correctly Reject explicitly and count them

The first row deserves more weight than a prettier composition. A team can debate whether 92% or 95% of a facade belongs in frame; it cannot use that debate to excuse an exposed access code. Keep moderation thresholds and visual weights in separate configuration, and require a new eval report when either changes.

I'm not sure a single crop score can represent every merchandising preference, because a luxury apartment landing page and a maintenance-services page may value different regions. The uncertainty is easy to resolve: segment the eval results by page template and property-photo class. If one set of weights fails a segment, use a named profile rather than stacking more hidden exceptions into the scoring function.

There is a second catch. OCR cost is paid on the source image, while crop search is cheap local geometry; rerunning OCR for every slot wastes work and can yield inconsistent moderation decisions. Cache the normalized OCR result by content digest and policy version, but don't cache a final approval forever. A policy revision should invalidate the decision even when the pixels have not changed.

Ship responsive assets without hiding the trade-offs

Generate actual slot-sized files rather than asking the browser to download one huge source and imitate cropping with CSS. The page can still use object-position as a presentation hint, but it should not become the only record of the editorial decision. Persist crop coordinates. They make previews reproducible and give reviewers a concrete diff when a focal region changes.

Format negotiation belongs after crop selection. Encode the same approved pixels into the formats your delivery path and target browsers support, then use responsive markup to supply appropriate candidates. The media format guide in the references is a practical starting point for checking browser support and format characteristics. Measure encoded size and visual quality on the catalog itself; hero photography with foliage, brick, signage, and interior gradients does not behave like a synthetic benchmark.

This approach is not suitable when the composition must be art-directed independently for each breakpoint. A promotional banner with embedded typography, legal copy, or a deliberately off-center subject needs separately authored desktop and mobile assets, each moderated as its own source. Stick with manual art direction for those cases. Also keep human review for low-confidence OCR or policy labels whose consequence is disclosure; crop automation should route ambiguity, not conceal it.

Operationally, log the source digest, decoder result, OCR model or engine version, policy version, focal box, chosen crop coordinates, output dimensions, and encoded byte count. Do not log the extracted sensitive text in ordinary application logs. Alert on changes in rejection rate, empty OCR rate, partial-text intersections, and per-slot focal coverage. A sudden shift after a decoder or OCR update is more actionable when the event includes the version that produced it. The production checklist is short enough to remain prose: validate MIME type by decoding, apply orientation before OCR, moderate the complete source, and fail the whole request before writing outputs if policy blocks it. Generate every required slot, verify exact dimensions, write variants under immutable keys, and commit their metadata together. Run the labeled eval suite before changing OCR, policy, crop weights, or image encoding. Then inspect a small stratified sample, because aggregate metrics can hide a failure concentrated in tall mobile captures. That sample should deliberately include the least common input classes rather than another random draw dominated by ordinary landscape exteriors; otherwise, a clean aggregate can mask the one portrait-oriented doorway image where the mobile focal region and a small parcel label compete for the same narrow frame.

Do that, and desktop versus mobile stops being a moderation argument. It becomes what it should have been all along: a deterministic framing decision over an already approved image.

References

Further reading

Top comments (0)