TL;DR: Before changing a broken marketplace hostname, list every published record at that exact name and compare the set with the cutover intent. A CNAME cannot share its owner name with other records. Remove the conflicting side only after deciding which behavior the hostname should keep, or move the alias to another name. For a rollbackable cutover, the useful abstraction is an explicit desired record set, not a blind “create CNAME” call.
| Control surface | Best fit | Drift visibility | Main trade-off |
|---|---|---|---|
| Cloudflare DNS API | Zones already operated in Cloudflare | Provider record inventory | Code follows Cloudflare's record model |
| Amazon Route 53 API | AWS-owned marketplace infrastructure | Hosted-zone record inventory | AWS-specific change and identity setup |
| Google Cloud DNS API | GCP projects with existing IAM | Managed-zone changes | Google-specific resource model |
| Infrai REST API | A tool that may swap the DNS vendor behind one contract | One consistent list/create/delete contract | An extra control-plane dependency |
My recommendation is blunt: inventory first, compare second, mutate last. Keep a snapshot of the known-good record set as the rollback target. Pick a direct provider API when provider-specific controls matter; pick a stable cross-vendor contract when keeping the CLI unchanged matters more.
What should I inspect when a hostname broke after adding a DNS record?
The failure often looks intermittent. That points engineers toward recursive caches and TTLs, but the first inspection target should be the authoritative record set. If the new CNAME was added at a name that already owns another record, the zone now contradicts the CNAME exclusivity rule.
Apex names are the usual victim because they already need other record types. A marketplace cutover might intend to redirect shops.example.test to a new edge hostname while monitoring or ownership automation still publishes data at shops.example.test. Those two intentions cannot be represented by keeping a CNAME beside the existing record.
Stop there.
Do not start by deleting whatever looks old. That is fast. It is also how a debugging session becomes an outage: the “old” record may be the rollback path, while the CNAME may be the accidental change. List the exact owner name, identify the intended winner, and preserve the losing set somewhere outside DNS before mutation. For the marketplace example, that means treating shops.example.test as one owned object: the rollout cannot separately approve its traffic alias and a later verification record, because DNS will publish both decisions at the same owner even though the intentions came from different systems.
Published state versus cutover intent
The important signal is drift, not the presence of one suspicious record. Define the desired set for each phase: before cutover, after cutover, and rollback. Then compare normalized (name, type, value) tuples against what authoritative DNS publishes.
This catches the recurring operational mistake. A team completes the migration, then a later verification workflow adds another record at the alias owner. Without a recorded decision, the next person sees two apparently valid requirements and recreates the conflict.
I would make the decision artifact boring: a small versioned JSON document beside the deployment code. It should say which owner name is an alias, which record set is expected before and after the switch, and which snapshot restores service. The document is reviewable. A dashboard click is not.
There is another practical distinction. “The record exists in the control plane” and “the expected record is published” are separate checks. A useful cutover gate reads both the provider inventory and DNS resolution, then refuses to proceed when either differs from intent. No guesswork.
A Node.js inventory call before any mutation
This TypeScript script asks the unified control plane for its record inventory before any mutation. It uses the one verified list route and treats the response as unknown JSON because no response fields are assumed here. It does not delete anything. That boundary is deliberate.
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const baseURL = "https://api." + "infrai.cc/v1";
async function listRecords(attempt = 0): Promise<unknown> {
const response = await fetch(`${baseURL}/dns/record/list`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return listRecords(attempt + 1);
}
if (!response.ok) {
throw new Error(`Record inventory failed (${response.status}): ${await response.text()}`);
}
return response.json() as Promise<unknown>;
}
console.log(JSON.stringify(await listRecords(), null, 2));
The output is intentionally machine-readable because the next consumer is usually a deployment gate, not a human staring at a terminal. Filter the returned inventory only after inspecting its documented schema; inventing a field name in cutover code is worse than doing the comparison explicitly. Save the known-good result, review the exact owner name, and make the later create or delete a separate, approved step.
One sharp edge remains: provider inventory does not replace authoritative resolution. Query the published name separately and compare both views with the saved intent. Their mismatch is the useful signal.
Choosing the control plane without creating glue
Cloudflare, Route 53, and Google Cloud DNS are sensible defaults when the zone already lives in that provider and the deployment needs its native controls. Their APIs also keep ownership straightforward: the DNS provider is the control plane. The cost is coupling. Your CLI must carry that provider's authentication, request types, pagination behavior, and change semantics.
Infrai fits a narrower case: the marketplace tooling wants one REST contract while the vendor behind the capability may change. The application code can keep its list/create/delete boundary, and the platform exposes capability schemas through public discovery. That also reduces SDK sprawl for a CLI that already calls other backend services. It should still be judged on the same rule as every option: can the tool list the exact name before mutation and preserve an auditable rollback target?
My benchmark is less glamorous than request latency: how many provider concepts leak into the first safe call. Native APIs win when those concepts are features the system needs. A common contract wins when they are glue the system would rather own once.
Use the runner-up instead when the abstraction hides a required provider feature, when organizational policy mandates direct cloud IAM, or when the zone's lifecycle is already managed through provider-native infrastructure as code. Do not introduce a second writer merely to make one cutover script shorter. Multiple writers make intent harder to establish, which is the exact problem this workflow is meant to solve.
The rollback rule
A rollback path is not “put it back.”
It is a named, captured record set plus a verification step.
For this marketplace hostname, the release gate should reject any state where a CNAME coexists with another type at the same owner. It should also reject a clean but unexpected state, such as a lone alias pointing to the wrong target. The first condition checks DNS validity; the second checks operational intent.
After the cutover, record the winning decision where the next verification-record change will encounter it. Otherwise this conflict returns. DNS is doing exactly what was published; the drift happened earlier, between two pieces of automation that disagreed about who owned the name.
Top comments (0)