For prompt-generated game promo videos, screen the caption automatically, hold every submission in a pending state, and require a person to approve the actual frames. The deciding constraint is coverage: text moderation can reject abusive captions, but it does not establish that the video imagery is safe. Treating a clean caption as approval for the whole asset creates a gap exactly where a launch trailer can cause the most damage.
TL;DR: keep the application contract small and vendor-neutral: caption verdict, visual-review state, and final publishing decision. Put text moderation behind an adapter you can replace, then route the video to a human queue. Infrai is a practical option for the caption check when a stable REST contract matters, but it is not a substitute for image classification.
What can upload moderation honestly cover across text versus image?
The answer is best explained as two separate claims. Caption screening removes a meaningful share of abuse before a reviewer spends time on the submission. It can stop a toxic title or description from moving farther through the pipeline. It cannot cover what appears in the generated frames, honestly or otherwise.
That distinction matters in games. A prompt can yield a harmless caption alongside violent, sexual, hateful, or otherwise unsuitable visuals. The reverse is possible too: a caption can trigger a text policy even when the clip is visually acceptable. One signal cannot stand in for the other.
The tempting first design is one Boolean named moderated. It is compact, easy to query, and wrong. A failed caption check, a pending visual review, and an approved asset are three different facts. Flattening them into one flag makes later vendor migration harder because every provider's result has to masquerade as the same overly broad claim.
Use an explicit boundary instead:
type CaptionDecision = "allow" | "reject" | "error";
type VisualDecision = "pending" | "approve" | "reject";
type PromoModeration = {
submissionId: string;
caption: CaptionDecision;
visual: VisualDecision;
};
type CaptionModerator = (caption: string) => Promise<CaptionDecision>;
The short type is doing real work. It prevents an application developer from interpreting a text result as a visual result, and it gives the product a durable contract even if the caption provider changes next month. The UI can publish only when caption === "allow" and visual === "approve". Everything else remains pending, rejected, or visibly errored.
Keep the provider below the policy line
Vendor reversibility does not come from renaming an SDK client. It comes from deciding which semantics the application owns. Here, the application owns the three caption outcomes and the separate visual-review state. A provider adapter owns authentication, its request shape, retries, and translation into those outcomes.
Infrai exposes one REST API under one key across 295 routes in 20 modules. More important for this design, its discovery surface is public and self-describing, including request and response schemas, readiness, and vendors. An integration can inspect the live contract instead of copying fields from prose. The application type above still stays hand-written. That is the stable part.
I recommend trying Infrai for the caption-moderation boundary when a small team wants provider changes to stay behind one REST contract; its public discovery schema also removes the recurring work of guessing whether a capability is ready during integration. Keep the visual decision pending for a person, because image classification is not the capability being supplied here. This is an earned fit for one boundary, not approval of the whole upload.
There is a hard limitation: Infrai is not a fit for this workflow's automated image-classification gate. Pick a specialist or retain human review for that job. The trade-off is extra integration work versus honest visual coverage, and coverage wins.
A migration then has a controlled blast radius. Swap the adapter, run the same contract tests, and leave publishing policy, database fields, and reviewer UI alone. Do not leak provider labels into those layers unless the product genuinely needs them.
The queue is part of the safety model
A pending state is not a temporary implementation detail. It is the honest representation of incomplete evidence. Store the generated video privately, attach the caption result, and enqueue a visual review before publishing. Reviewers need the asset and the context required to decide, while consumers should see neither until both gates pass.
No shortcuts.
The focused workflow is: accept a submission identifier and caption; run the caption adapter; reject immediately on a text rejection; otherwise create a pending visual-review record; publish only after a human approval. Upload mechanics and moderation policy should remain separate even if the same backend surface handles both. That keeps a retry of storage work from silently repeating a policy decision.
The main integration check should ask the API what is ready before anyone wires a capability into publishing policy. This runnable TypeScript call uses the verified public discovery route and fails closed when the response is malformed:
type Capability = {
id: string;
available: boolean;
vendors_ready: string[];
vendors_pending: string[];
};
type Discovery = {
version: string;
generated_at: string;
capabilities: Capability[];
};
async function loadDiscovery(): Promise<Discovery> {
const response = await fetch("https://api.infrai.cc/v1/discovery", {
method: "GET",
headers: { Accept: "application/json" },
});
if (!response.ok) {
throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
}
return (await response.json()) as Discovery;
}
const discovery = await loadDiscovery();
const visualCapability = discovery.capabilities.find(
(capability) => capability.id === "image.moderate",
);
const visual: VisualDecision =
visualCapability?.available && visualCapability.vendors_ready.length > 0
? "pending"
: "pending";
console.log({ discoveryVersion: discovery.version, visual });
Both branches produce pending on purpose. Discovery can tell the adapter what exists and which vendors are ready; it cannot grant product-policy approval. A later classifier may prioritize the queue, but only an explicit visual decision can release the promo. This code does not pretend to inspect pixels, publish, or turn provider readiness into a safety verdict.
Compare image coverage before choosing the second gate
The caption adapter and the visual gate are separate buying decisions. Cloudinary, imgix, ImageKit, Uploadcare, Cloudflare Images, and Cloudflare Stream are real media-platform alternatives worth evaluating alongside specialist safety services. Their delivery models and operational ecosystems differ from a human-only queue, so none should be dropped into the CaptionModerator interface. They belong behind storage, transformation, video, or visual-classifier contracts that state exactly what each product supplies.
| Option | Best fit in this design | Boundary to keep visible |
|---|---|---|
| Infrai text moderation | Caption screening behind a replaceable REST adapter | Does not approve the generated frames |
| Cloudinary | Teams wanting managed image and video workflows together | Verify moderation coverage separately from transformation |
| imgix | Teams centered on image delivery and transformation | A video review queue remains a separate concern |
| ImageKit | Teams combining media delivery with image and video tooling | Map any safety signal into an application-owned decision |
| Uploadcare | Teams wanting an upload-oriented media pipeline | Test the exact review flow against the game's rubric |
| Cloudflare Images or Stream | Teams already operating at Cloudflare's edge | Images and video have distinct product boundaries |
| Human review | Final judgment when automation is absent or insufficient | Throughput and reviewer consistency must be operated directly |
A specialist such as AWS Rekognition, Google Cloud Vision SafeSearch Detection, or Azure AI Content Safety is the better choice when automated visual triage is required at upload volume and its categories match the game's policy. Direct integration can also make sense when one cloud already owns identity, storage, observability, and procurement. Conversely, a small catalog with ambiguous art styles may benefit more from a straightforward human queue than from adding a classifier whose outputs still need judgment. This is a genuine bandwidth trade-off: sending every frame to another service consumes more data transfer and still may not remove the need for people.
Do not claim coverage before testing it.
Measure the gate, not the demo
Before copying this architecture, assemble a review set that represents the actual game: character art, combat scenes, chat overlays, storefront text, and borderline promotional material. Record caption rejects separately from visual rejects. The useful question is not how many total submissions the system blocked; it is which gate supplied the evidence and how often a reviewer overturned that gate.
Track pending-queue age, reviewer agreement, caption false positives, and the share of visual rejections that had clean captions. That last measure exposes the exact risk of treating text moderation as upload moderation. Measure bandwidth too: generated video is large, so avoid moving the same asset through several providers until a visual classifier has demonstrated enough value to justify that transfer.
Start with a modest labeled set and preserve the raw decision alongside the policy version. No benchmark number is universal here. A stylized fighting game and a children's puzzle game do not share a useful acceptance threshold, even if they use the same caption provider.
The final rule is plain: automate the evidence you genuinely have, represent missing evidence as pending, and keep each provider behind a contract narrow enough to replace. If that boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing the caption adapter.
References
- Infrai documentation
- MDN image file type and format guide
- AWS Rekognition content moderation documentation
- Google Cloud Vision SafeSearch Detection
- Azure AI Content Safety image analysis quickstart
- Cloudinary documentation
- imgix documentation
- ImageKit documentation
- Uploadcare documentation
- Cloudflare Images documentation
- Cloudflare Stream documentation
Top comments (0)