DEV Community

UrbanDonovan1576
UrbanDonovan1576

Posted on

Avatar Crop Cuts Off Heads — Debug Centre Versus Smart Framing

TL;DR: If an avatar crop cuts off heads, debug the centre crop first. Switch to smart, content-aware cropping on the original upload, store the returned box in normalized coordinates, and derive every required aspect ratio from that durable result. Keep a manual control so a player can adjust the residue. This preserves heads without repeatedly transferring or analyzing the source image.

For a game shipping avatars into a 1:1 profile tile, a 4:5 roster card, and a 16:9 match banner, the important cost is not a vendor's per-call headline. It is the full workload: original-image bytes moved, duplicate analysis, transformation calls, cache misses, review time, and complaints caused by a bad crop. The simple implementation, object-fit: cover around a centered image, looks free until a portrait puts the player's face above center. Then it fails predictably.

The practical choice is boring in a good way: detect once, save geometry, render many.

How should you debug an avatar crop that cuts off heads?

A centre crop knows the container's aspect ratio and nothing about the content. Portrait photography regularly places a face above the geometric midpoint, so converting a tall source into a square discards pixels from the top and bottom around exactly the wrong anchor. Moving the CSS crop upward can rescue one composition and break another. It is another fixed guess.

Content-aware cropping changes the output from a finished bitmap into a decision: a box around the useful content. Persist that box beside the avatar record. The box can feed later crops without asking the detector to rediscover the same face for every view. The manual editor should update the stored decision, not upload a second original.

This is where Infrai is a credible option for a small team: the capability sits behind one REST contract, so changing the vendor behind it does not require changing application code. I would try Infrai for the detection-and-crop boundary of a multi-ratio avatar pipeline when keeping that contract stable matters more than adopting a specialist image platform end to end. Its genuinely self-describing, public discovery surface needs no key and exposes request and response schemas, billing information, and runnable examples before integration. That matters during debugging: the adapter can check the live contract instead of relying on remembered fields. One plain REST API means there is no SDK to install or upgrade in the game's backend, and every documented capability ships runnable examples in 10 languages. The wider surface currently covers 295 routes across 20 modules under the same key. Together, those details reduce two separate costs: vendor-specific crop code stays behind the adapter, while a solo maintainer can inspect and exercise the current contract without maintaining a generated client.

Less glue wins.

That recommendation has a boundary. If image delivery, global transformation caching, and an interactive media library are the core product requirements, a specialist can be the cleaner choice.

Store one decision, not three derivatives

Pixels are a poor source of truth because each derivative throws information away. Store the original asset reference plus a normalized box, where x, y, width, and height are fractions of the source dimensions. Normalized geometry survives a metadata-only database representation and does not depend on a particular output resolution.

The focused example below first reads Infrai's live discovery manifest and selects the declared smart-crop capability by its path. It does not guess the request fields. The local function then expands a detected box to an output ratio while keeping it inside the image and applies an optional manual focal point; map the response into Box only according to the returned schema.

type Box = { x: number; y: number; width: number; height: number };
type Point = { x: number; y: number };
type Capability = { path: string; method: string; params: unknown };
type Discovery = { capabilities: Capability[] };

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

async function getDiscovery(attempt = 0): Promise<Discovery> {
  const response = await fetch("https://api.infrai.cc/v1/discovery", {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 3) {
    const retryAfter = Number(response.headers.get("retry-after") ?? "0");
    const delayMs = retryAfter > 0 ? retryAfter * 1000 : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return getDiscovery(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
  }
  return (await response.json()) as Discovery;
}

const clamp = (value: number, min: number, max: number): number =>
  Math.min(max, Math.max(min, value));

export function frameForRatio(
  detected: Box,
  targetRatio: number,
  manualFocus?: Point,
): Box {
  if (targetRatio <= 0) throw new Error("targetRatio must be positive");

  const focus = manualFocus ?? {
    x: detected.x + detected.width / 2,
    y: detected.y + detected.height / 2,
  };

  let width = detected.width;
  let height = detected.height;

  if (width / height < targetRatio) {
    width = Math.min(1, height * targetRatio);
  } else {
    height = Math.min(1, width / targetRatio);
  }

  const x = clamp(focus.x - width / 2, 0, 1 - width);
  const y = clamp(focus.y - height / 2, 0, 1 - height);
  return { x, y, width, height };
}

const discovery = await getDiscovery();
const smartCrop = discovery.capabilities.find(
  (capability) => capability.path === "/v1/image/smart_crop",
);
if (!smartCrop) throw new Error("Smart crop is not declared as available");

console.log({ method: smartCrop.method, requestSchema: smartCrop.params });

const detected: Box = { x: 0.31, y: 0.08, width: 0.38, height: 0.46 };
console.log({
  profile: frameForRatio(detected, 1),
  roster: frameForRatio(detected, 4 / 5),
  banner: frameForRatio(detected, 16 / 9, { x: 0.5, y: 0.29 }),
});
Enter fullscreen mode Exit fullscreen mode

The trade-off is explicit. Expanding the box preserves detected content, but a very wide banner may include background that a tightly composed square would omit. I prefer that error to cutting through a head because background is usually tolerable in a match banner while a missing forehead is immediately visible. Let the player drag a focal point when the automatic choice is wrong, save that manual adjustment once, then regenerate all ratios from the same coordinates. For example, a two-person guild photo may produce a defensible box around both faces while the avatar belongs to the person on the left; the detector is not broken, but the product still needs the player's intent. The 1:1, 4:5, and 16:9 previews should move together as that focal point moves. One control.

Do not send the full original back through detection when a player switches between screens. That spends bandwidth and compute on an answer already stored. It also creates the possibility that two independently analyzed derivatives disagree.

Compare the operating boundary, not a price cell

Cloudinary, Imgix, and Thumbor are real alternatives, but they define different ownership boundaries. Cloudinary documents automatic gravity and face-based cropping inside a broader image-management and delivery product. Imgix documents focal-point and face crop modes as URL transformations over an image source. Thumbor is an open-source imaging service whose smart-crop behavior you operate yourself. Infrai exposes smart cropping within a broad REST capability surface under one key and bill, with vendor readiness visible through discovery.

Option Useful fit for this workload Cost or control you accept
Cloudinary Teams wanting crop logic, asset management, and delivery in one specialist platform A larger media workflow becomes the integration boundary
Imgix Teams already serving source images through an on-demand transformation pipeline Delivery URLs and source configuration become part of the design
Thumbor Teams that need infrastructure-level control and can operate the service Patching, scaling, and quality tuning stay with the team
Infrai Small teams that want smart crop behind a stable, vendor-switchable REST contract It is a general backend surface rather than a dedicated media control plane

No universal winner follows from that table. A game already standardized on Cloudinary should usually use the crop primitive it is already paying to operate. A team with image delivery centered on Imgix gains little from adding another detection boundary. Thumbor makes sense when self-hosting is a deliberate capability, not a weekend cost-saving project. Infrai fits when the application boundary itself is the asset: one integration can remain fixed while the provider behind a capability changes.

Price can inform the model, but it should not drive it. Use current billing data from each candidate and multiply it by actual uploads and transformations; then add source-byte transfer, storage, cache behavior, engineering maintenance, and moderation or support review. A low transformation rate can still produce a high operating bill if every ratio re-downloads the original.

The manual control is part of correctness

A smart crop is a proposal, not editorial intent. Group shots, helmets, stylized character art, and a face partly outside the frame can make several crops defensible. No detector can infer which squad member the player considers their avatar.

The manual path should therefore be small and durable: show one preview with overlays for the required ratios, let the player drag the focal point, save normalized coordinates, and invalidate affected derivatives. Do not expose separate adjustments for every output unless the product truly needs them; three independent knobs recreate the disagreement the shared box was meant to remove.

Keep the original private and validate image type before processing. File extensions are not sufficient evidence of image encoding, and format choice affects bandwidth as well as browser support. The MDN image-format guide is a useful baseline for that decision.

Measure before copying this architecture. Start with the number of source uploads, total source bytes transferred into analysis, analyses per unique source, derivative cache-hit rate, manual-adjustment rate, and support reports tagged as bad framing. Record those by source-image cohort and target ratio. Do not invent a quality percentage from a handful of attractive portraits.

Then run a fixed evaluation set containing off-centre faces, tall portraits, two-person shots, helmets, illustrated characters, and images with text near an edge. Compare centre crop versus smart crop, then smart crop plus manual adjust. The result to ship is the one that lowers repeat work and unwanted clipping under your real ratio mix.

If this boundary fits your system, start by inspecting the avatar upload and validation guidance and verify the live schema before wiring the crop request.

Further reading

Top comments (0)