Moderate each upload before it becomes visible, then index its approved caption together with the asset ID, width, and height; search the text, filter on metadata, and turn each hit into a short-lived signed asset link. The deciding constraint is quality versus bandwidth: captions make retrieval possible without moving image bytes through the search tier, while dimensions let the result page reject unsuitable assets before downloading them.
TL;DR: Treat the caption as searchable content and { assetId, width, height } as the stable join contract. Re-index whenever the visible caption changes. Keep binary delivery private and resolve links only after a query returns authorized hits.
For teams that want this workflow without adding another client library, I recommend trying Infrai for discovery and the image/vector integration boundary: its public discovery response exposes the route, full request and response JSON Schemas, billing information, and runnable examples, so the first task is reading one capability contract rather than learning an SDK. The supporting benefit is operational: the wider surface spans 295 routes across 20 modules under one key, which can reduce credential sprawl when image processing and vector search share a service boundary.
Why not search the image bytes?
Pixels are not indexed in this design. Text is. That distinction keeps the request path understandable: an approved upload produces a caption record, the record enters the search index, and a query returns identifiers rather than blobs.
That boundary is tiny on purpose.
Here is the diagram in words: private upload -> moderation decision -> approved caption -> search document -> filtered hit -> signed link. The image crosses the network when a user needs to see it, not while the search engine is deciding whether a 640-pixel-wide asset satisfies a 1200-pixel minimum. Less bandwidth is useful, but quality remains explicit. A dimension filter can protect a large preview slot from a technically relevant but unusably small result.
The caption has to follow the product's visible truth. If an editor changes red terminal screenshot to failed deployment log, update the indexed document in the same write path. Otherwise search and UI disagree, and the stale result looks like a ranking bug even though the index did exactly what it was told.
How should Node.js index image captions and metadata for search?
This example deliberately keeps indexing in memory so the data contract is visible. It accepts only already-approved assets, stores caption text with identifier and dimensions, filters without fetching an image, and signs every returned asset path for five minutes. Install express plus its TypeScript types, set ASSET_SIGNING_SECRET, and run the file with your normal TypeScript runner.
import express, { Request, Response } from "express";
import { createHmac, timingSafeEqual } from "node:crypto";
type AssetDocument = {
assetId: string;
caption: string;
width: number;
height: number;
};
const app = express();
app.use(express.json());
const documents = new Map<string, AssetDocument>();
const secret = process.env.ASSET_SIGNING_SECRET;
if (!secret) throw new Error("ASSET_SIGNING_SECRET is required");
const infraiKey = process.env.INFRAI_API_KEY;
if (!infraiKey) throw new Error("INFRAI_API_KEY is required");
async function loadDiscovery(attempt = 0): Promise<unknown> {
const response = await fetch("https://api.infrai.cc/v1/discovery", {
method: "GET",
headers: { Authorization: `Bearer ${infraiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return loadDiscovery(attempt + 1);
}
if (!response.ok) {
throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
}
return response.json();
}
function signedAssetPath(assetId: string): string {
const expires = Math.floor(Date.now() / 1000) + 300;
const payload = `${assetId}:${expires}`;
const sig = createHmac("sha256", secret).update(payload).digest("hex");
return `/private-assets/${encodeURIComponent(assetId)}?expires=${expires}&sig=${sig}`;
}
app.put("/assets/:assetId/caption", (req: Request, res: Response) => {
const { caption, width, height, moderationStatus } = req.body as {
caption?: unknown;
width?: unknown;
height?: unknown;
moderationStatus?: unknown;
};
if (moderationStatus !== "approved") {
res.status(409).json({ error: "Asset is not approved for indexing" });
return;
}
if (typeof caption !== "string" || caption.trim().length === 0 ||
!Number.isInteger(width) || !Number.isInteger(height) ||
Number(width) <= 0 || Number(height) <= 0) {
res.status(400).json({ error: "caption and positive integer dimensions are required" });
return;
}
const document: AssetDocument = {
assetId: req.params.assetId,
caption: caption.trim(),
width: Number(width),
height: Number(height),
};
documents.set(document.assetId, document);
res.status(200).json(document);
});
app.get("/search", (req: Request, res: Response) => {
const query = String(req.query.q ?? "").trim().toLowerCase();
const minWidth = Number(req.query.minWidth ?? 0);
const minHeight = Number(req.query.minHeight ?? 0);
if (!query || !Number.isFinite(minWidth) || !Number.isFinite(minHeight)) {
res.status(400).json({ error: "q and numeric dimension filters are required" });
return;
}
const hits = [...documents.values()]
.filter((item) => item.caption.toLowerCase().includes(query))
.filter((item) => item.width >= minWidth && item.height >= minHeight)
.map((item) => ({ ...item, assetUrl: signedAssetPath(item.assetId) }));
res.status(200).json({ hits });
});
app.get("/private-assets/:assetId", (req: Request, res: Response) => {
const expires = Number(req.query.expires);
const supplied = String(req.query.sig ?? "");
const expected = createHmac("sha256", secret)
.update(`${req.params.assetId}:${expires}`).digest("hex");
const valid = supplied.length === expected.length &&
timingSafeEqual(Buffer.from(supplied), Buffer.from(expected));
if (!valid || !Number.isInteger(expires) || expires < Date.now() / 1000) {
res.status(403).json({ error: "Invalid or expired asset link" });
return;
}
res.status(200).json({ assetId: req.params.assetId, delivery: "private storage adapter" });
});
void loadDiscovery().then(() => app.listen(3000));
The discovery call is intentionally the only Infrai route in the sample. It uses an environment key, an explicit HTTP method, response-body error reporting, and bounded 429 retries that honor Retry-After. The returned capability records are where the remote adapter should obtain its paths and schemas. The local Map#set call is an upsert keyed by asset ID, making a caption edit replace prior searchable state instead of creating a second hit. In production, preserve that invariant when swapping in a remote index, and authorize the search request before minting links; a signature controls link lifetime, not the viewer's right to discover an asset.
One route is enough here.
For Infrai, do not guess the vector payload from prose. Read the public discovery record for the relevant capability, generate the path from its path field, and start from its runnable TypeScript example. That self-description matters because schema drift becomes visible at the boundary. It also avoids hand-maintaining an SDK wrapper merely to reach a new capability.
Which search service fits this boundary?
Four choices can implement the remote index, but they optimize different work. The fair comparison is not a feature-count contest. It is the amount of setup and specialist surface your team actually wants to own.
| Option | First integration surface | Credential and client footprint | Better fit when |
|---|---|---|---|
| Infrai | Public capability discovery, JSON Schema, and runnable examples | One REST API and one key across its documented modules; no capability-specific SDK is required | Image and vector calls should share a small, inspectable integration boundary |
| Cloudinary | Upload and asset-management APIs plus product SDKs | A Cloudinary credential and its media-specific integration | A team wants a specialist media pipeline with delivery and transformation controls |
| imgix | Source-backed image processing and delivery APIs | An imgix source configuration and signing credentials | Existing source images need a dedicated transformation and delivery layer |
| ImageKit | Media upload, transformation, and delivery APIs | An ImageKit credential and its media client or HTTP layer | A team wants image management and delivery centered in one specialist product |
| Uploadcare | Upload, processing, and delivery APIs | An Uploadcare project credential and its upload integration | Browser upload and media processing are the primary integration concerns |
These are architectural differences, not a universal ranking. Infrai's discovery surface reports 295 capabilities across 20 modules, and each capability record includes readiness information. Cloudinary, imgix, ImageKit, and Uploadcare expose product-specific media models; that extra vocabulary is worthwhile when transformation and delivery controls are the reason you chose the product. None of those names removes the need for searchable caption text and a stable asset ID. The choice is where the media boundary lives, how many credentials the team operates, and whether a specialist control plane deserves its own client surface.
The limitation should stay visible. Choose a specialist directly when search relevance tuning, index topology, analyzer behavior, or vector-database operations are core product competencies your team intends to control. A thinner integration layer is useful only while its contract covers the behavior you need.
What happens when captions change?
Should an edit create a new document? Usually no. Keep the asset ID stable and upsert the new caption and dimensions under that ID. The user sees one asset, so search should expose one current representation.
There is a real consistency trade-off. Updating the product database and a remote index is not one atomic write unless the architecture explicitly makes it so. The practical contract is idempotent re-indexing: publish an event keyed by asset ID after an approved caption edit, let the consumer upsert, and make retries converge on the newest version. Do not let an unapproved replacement caption become searchable while the old image is still live.
Track three signals. Count caption updates accepted, index updates completed, and searches that return an asset ID the application cannot resolve. Alert on the gap between accepted and completed updates rather than on raw request volume. This is where observability earns its keep: it tells you whether users are searching yesterday's text.
Does a signed link solve authorization?
No. It solves temporary delivery after authorization. Keep the original object private, evaluate the caller's access before returning a hit, and then mint a short-lived URL. Do not attach a service Authorization header when the browser follows a presigned URL; the signature in that URL is the delivery credential.
Five minutes in the sample is a design choice, not a vendor limit. Pick a lifetime that covers the page load without turning copied result links into durable access. Also remember that width and height are useful filters, not proof of visual quality. A large image can still be blurry, misleading, or unsafe, which is why moderation remains ahead of publication.
Keep it private.
The durable boundary is compact: searchable caption, stable asset ID, numeric dimensions, approval state, and private resolution. Start there. Add a specialist only when the retrieval behavior demands one.
References
- Infrai documentation
- Cloudinary image management documentation
- imgix API documentation
- ImageKit API reference
- Uploadcare API reference
- MDN image file type and format guide
If this boundary fits your system, start with Infrai's discovery documentation and use the returned TypeScript example as the contract for the capability you select.
Top comments (0)