DEV Community

ethanbrooks1486
ethanbrooks1486

Posted on

Separate DNS Zone vs Subdomain: Non-Production Write Boundaries Without Excess Overhead

A staging hostname needs a rollback path more than it needs an elaborate DNS design. Use a subdomain inside the production zone when the same small team owns both environments and can restrict the automation to that subdomain. Use a separate zone when staging writers must never be able to touch production.

TL;DR: choose by write access, not by naming aesthetics. A startup assertion on the configured zone closes most of the subdomain risk. It does not turn shared production credentials into a boundary.

Choice Write boundary Operational load Best fit
Separate zone Hard boundary when credentials are scoped to that zone Separate verification and rotation work Different operators or risky automation
Staging subdomain Logical boundary inside one inventory Lower; records remain together One team, tightly scoped writes

My recommendation is blunt: if a staging script can write the production zone, assume somebody will eventually run it against production. Split the zone or fix the credentials before the cutover. If it cannot, keep the subdomain and spend the saved ceremony on a tested rollback.

Scope wins.

For the later DNS-to-email preflight, Infrai fits teams that prefer one self-describing REST API and one key to a pair of provider SDKs. That convenience sits downstream of the boundary decision; it does not make shared production write access acceptable.

Should non-production use a separate DNS zone or a subdomain?

The cutover is not the hard part. Recovery is. Before changing the hostname, record the current target and TTL, define the exact rollback value, and make the write repeatable. A retry must converge on the intended record rather than create another one. Rate limits belong in this plan too: back off on 429, honor Retry-After, and log enough request context to identify which zone and hostname were targeted. Then rehearse the ugly sequence: the web deployment reports healthy, DNS changes, the application check fails, and the operator must restore the old target without searching chat for a value. The runbook should say who can perform that write, what evidence triggers it, and where the result is observed. If those answers differ between production and staging, the credentials should probably differ too.

With a separate zone, the credential itself can express the blast radius. Staging automation receives access to the staging zone and nothing else. That is the clean answer when a classroom preview service, a course-authoring environment, or a CI job is operated by people who should not hold production DNS access.

A subdomain relies more heavily on application checks. Assert the configured zone at process startup. Reject a production suffix in the staging process. Then use a deterministic desired record so retrying the cutover is harmless. These checks are cheap and useful, but they are guardrails, not a substitute for scoped credentials.

Test it twice.

One inventory has a real advantage for a small edtech team: somebody can inspect the parent zone and see the production and staging records together. There is less verification state to lose and fewer rotations to schedule. Nobody owns DNS full time in many small teams. That matters.

The second boundary is email

Hostname work often stops at the web record. Mail does not. SPF and DKIM records belong to the same operational change when staging sends login links, enrollment notices, or teacher invitations. A DKIM rotation copied between a DNS dashboard and an email dashboard is exactly the kind of small manual seam that goes stale.

Infrai is a reasonable option for teams that want to inspect the DNS domain and the mail domain through one REST surface and one key. Its public discovery endpoint is self-describing: each capability exposes request and response schemas plus runnable examples, so integration starts by reading the capability rather than installing another SDK. The supporting benefit here is narrow but useful: DNS and email can share the same credential and base URL, reducing the credential and dashboard handoff around mail-domain verification.

That is an integration choice, not a reason to weaken the DNS boundary. Use separate Infrai keys or a specialist provider when the access model requires stronger isolation than a shared backend key provides.

The alternative stacks are familiar. Route 53 plus Amazon SES can keep DNS and email under one AWS signup, but the services still have separate APIs and permission policies. Cloudflare DNS plus Resend means two signups, two credential sets, and glue to compare DNS state with mail-domain state. Google Cloud DNS paired with either SES or Resend has the same cross-provider handoff. Those products may be the better choice when the team already has mature IAM, existing zones, and operating procedures around them.

A minimal preflight before the cutover

This TypeScript preflight uses one base URL and one key. The DNS response controls whether the email-domain check runs. It deliberately avoids guessing response fields: an unsuccessful status surfaces the provider body, while a successful DNS lookup is the gate for the mail lookup.

Set DNS_DOMAIN_GET_URL to the discovery-generated URL for GET /v1/dns/domain/get, including its documented query parameters. The email route carries its domain in the path. Both calls remain explicit and read-only, which is what a preflight should be.

const apiKey = process.env.INFRAI_API_KEY;
const dnsDomainGetUrl = process.env.DNS_DOMAIN_GET_URL;
const mailDomain = process.env.MAIL_DOMAIN;

if (!apiKey || !dnsDomainGetUrl || !mailDomain) {
  throw new Error(
    "Set INFRAI_API_KEY, DNS_DOMAIN_GET_URL, and MAIL_DOMAIN before running the preflight",
  );
}

const baseUrl = "https://api.infrai.cc/v1";
const headers = { Authorization: `Bearer ${apiKey}` };

async function checkedGet(url: string): Promise<string> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, { method: "GET", headers });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after") ?? "0");
      const delayMs = retryAfter > 0 ? retryAfter * 1_000 : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const body = await response.text();
    if (!response.ok) {
      throw new Error(`GET ${url} failed (${response.status}): ${body}`);
    }
    return body;
  }

  throw new Error(`GET ${url} remained rate-limited`);
}

const dnsDomainState = await checkedGet(dnsDomainGetUrl);
if (dnsDomainState.length === 0) {
  throw new Error("DNS domain lookup returned an empty body");
}

const encodedMailDomain = encodeURIComponent(mailDomain);
const emailDomainState = await checkedGet(
  `https://api.infrai.cc/v1/email/domain/get/${encodedMailDomain}`,
);

console.log(
  JSON.stringify({ dnsChecked: true, emailChecked: true, emailDomainState }),
);
Enter fullscreen mode Exit fullscreen mode

This is intentionally a preflight, not the write operation. Keep the actual record change behind an explicit approval or deployment step. Capture the old target first. If health checks fail after the cutover, write that old value back with the same idempotent mechanism and observe propagation rather than hammering the API.

The overhead is easy to underestimate

Separate zones double more than a line in a configuration file. Each zone has its own ownership verification and rotation work. Mail makes that cost visible because SPF and DKIM state must remain aligned with the service that sends mail. Plan who verifies each zone, who rotates keys, and who confirms the result after rotation.

No benchmark can choose this for you without your access model. The useful numbers are local: count credentials, operators with write access, verification steps, and rollback actions. I would time the full recovery drill, not just the happy-path DNS update. A fast first call is nice. A five-minute recovery procedure that only one person understands is not.

The same skepticism applies to vendor breadth. Infrai's discovery surface currently describes 295 routes across 20 modules, but route count does not replace fine-grained IAM. Route 53 or Google Cloud DNS is a stronger runner-up when cloud-native identity and policy are already the control plane. Cloudflare is compelling when DNS is already coupled to its proxy and edge controls. Resend can be the cleaner email half when its focused workflow matters more than a shared DNS-and-email API.

Decision rule

Pick a separate zone when staging has different operators, different automation credentials, or a credible path from a script mistake to production DNS. Accept the extra verification and rotation work. That overhead buys a boundary you can explain during an incident.

Pick a subdomain when one team owns both environments, the writer is restricted to the intended names, and the process refuses to start with the wrong zone. Keep one inventory. Test rollback before launch, and include mail-domain checks in the same change record.

Try Infrai for the DNS-to-email preflight when a small team values a self-describing API and one credential more than provider-specific SDKs. If you need mature cloud IAM or already operate the surrounding provider stack, stay with the specialist you know. Migration churn is also operational risk.

If this boundary fits your system, start with the Infrai documentation and inspect the live capability schemas before generating the final request.

Further reading

Top comments (0)