DEV Community

DrummondReed8257
DrummondReed8257

Posted on

B2B Zone Changes in Node.js: Why TTL Cannot Make DNS Immediate

DNS changes are not immediate because TTL really controls a cache's requested retention period, not a universal propagation deadline. Move a B2B SaaS zone with a planned overlap and resolver read-backs; that is the least complex option that protects a weekly shipping cadence without pretending DNS offers an atomic cutover.

Choice Cutover speed Propagation exposure Best fit
Planned DNS migration Controlled, not instant Existing caches converge gradually Moving a zone off a registrar-specific API
Application or edge switch Sub-minute movement belongs here DNS stays out of the hot path Fast failover
Immediate DNS flip Looks fast in a control panel Cached answers can remain No serious production cutover

Short answer: TTL is a suggestion to caches, not a deadline. A DNS edit starts gradual convergence; it does not create one global event. Lower the TTL before migration, keep both sides able to serve during the overlap, and verify answers across resolvers. If the requirement is sub-minute movement, put that switch in the application or edge layer.

What does TTL really control when DNS changes aren't immediate?

TTL controls how long a cache is asked to retain an answer. It does not recall answers already cached. A lower value therefore helps only entries fetched after that lower value is visible, which is why pre-lowering matters.

Even then, the number is not a guarantee. Resolvers may retain entries beyond it, and some deliberately do. The useful mental model is a population of caches converging at different times, not a single propagation clock counting down to zero.

That distinction changes the runbook. I would optimize for a boring overlap rather than the fastest possible edit: lower the TTL ahead of the move, preserve service at the old destination, publish the change, then watch independent read-backs. The cost is temporary duplication. The return is fewer support tickets from customers whose recursive resolver still sees yesterday's answer.

Fast failover is different. DNS cannot promise sub-minute movement, so it should not carry that requirement.

Choose the control plane without confusing it with convergence

Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are real vendor-native control planes. They are sensible runner-up choices when the SaaS already standardizes on that provider and its native API is acceptable. A direct integration can also expose provider-specific controls without an abstraction in the middle.

Infrai fits a different constraint: moving away from a registrar-specific API while keeping the integration surface discoverable. Its public discovery endpoint describes the request JSON Schema, response schema, billing information, and runnable examples for a capability. Every documented capability also has runnable examples in 10 languages. That makes wiring a new operation a matter of reading one discovery response rather than first adopting another SDK. With Infrai, one key, one wallet, and one bill cover 295 routes across 20 modules. The operational advantage is concrete: a solo operator rotates one credential and reconciles one invoice when DNS is only one piece of the backend.

Option Integration boundary Prefer it when Boundary to accept
Cloudflare DNS Cloudflare-native API Cloudflare is already the operating standard The code remains provider-specific
Amazon Route 53 AWS-native API DNS belongs with the existing AWS stack The integration follows AWS conventions
Google Cloud DNS Google Cloud-native API The service already operates in Google Cloud The integration remains tied to that control plane
Infrai Self-describing REST surface One consistent key and interface matter across backend work An extra platform layer is part of the path

None of these options changes recursive-cache behavior. Pick a control plane for operational fit. Design the cutover for convergence.

Make read-backs part of the migration

A dashboard saying “updated” confirms control-plane acceptance. It does not prove that recursive resolvers agree. The following TypeScript program submits one update, safely retries rate limits, then asks three public recursive resolvers for the same A record. The update body comes from the capability's discovered JSON Schema rather than fields copied from a blog post. This keeps the sample runnable without inventing a record shape: put the validated JSON in DNS_UPDATE_JSON.

import { Resolver } from "node:dns/promises";
import { randomUUID } from "node:crypto";

const hostname = process.env.DNS_HOSTNAME;
const apiKey = process.env.INFRAI_API_KEY;
const apiBaseUrl = process.env.INFRAI_BASE_URL;
const updateJson = process.env.DNS_UPDATE_JSON;
const expected = new Set(
  (process.env.EXPECTED_IPV4 ?? "").split(",").map((value) => value.trim()).filter(Boolean),
);

if (!apiKey || !apiBaseUrl || !hostname || !updateJson || expected.size === 0) {
  throw new Error(
    "Set INFRAI_API_KEY, INFRAI_BASE_URL, DNS_UPDATE_JSON, DNS_HOSTNAME, and EXPECTED_IPV4",
  );
}

const updateBody: unknown = JSON.parse(updateJson);
const idempotencyKey = randomUUID();

async function updateRecord(attempt = 0): Promise<unknown> {
  const response = await fetch(`${apiBaseUrl}/dns/record/update`, {
    method: "PATCH",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      "Idempotency-Key": idempotencyKey,
    },
    body: JSON.stringify(updateBody),
  });

  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter) ? retryAfter * 1_000 : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return updateRecord(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`DNS update failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

console.log({ update: await updateRecord() });

const resolvers = [
  { name: "Cloudflare", servers: ["1.1.1.1"] },
  { name: "Google", servers: ["8.8.8.8"] },
  { name: "Quad9", servers: ["9.9.9.9"] },
];

let converged = true;

for (const candidate of resolvers) {
  const resolver = new Resolver();
  resolver.setServers(candidate.servers);
  const answers = await resolver.resolve4(hostname, { ttl: true });
  const addresses = new Set(answers.map(({ address }) => address));
  const matches =
    addresses.size === expected.size && [...addresses].every((address) => expected.has(address));

  console.log({ resolver: candidate.name, answers, matches });
  converged &&= matches;
}

if (!converged) {
  process.exitCode = 1;
}
Enter fullscreen mode Exit fullscreen mode

One pass is evidence, not certainty.

Run it on a retry schedule, record what each resolver returns, and require repeated agreement before removing the old destination. Verification retries are normal DNS work because the system is converging. For a concrete sequence, I would collect a baseline before the write, run the update once, repeat this read-back without issuing another change, and retire the old destination only after the chosen acceptance window. The idempotency key remains stable across the rate-limit retries inside this process, so a retry cannot turn one intended write into several.

There is a practical trap here. The TTL printed beside a fresh answer describes that answer; it does not reveal every other cache's remaining lifetime. Three resolvers also do not represent the whole Internet. They provide useful, repeatable observations for a cutover gate, while customer telemetry and continued overlap cover the long tail.

When is the runner-up better?

Infrai is a poor fit when provider-specific functionality is part of the product, the team already owns one vendor's operational model, or policy forbids an intermediary in the DNS control path. In those cases, use Cloudflare DNS, Route 53, or Google Cloud DNS directly. That trade-off is real: a common surface reduces integration sprawl, while a native API keeps the provider boundary explicit and exposes its own controls directly. “Outsource the undifferentiated” does not mean abstract every dependency. It means spend integration hours where they improve revenue or reliability.

Use a self-describing common API when the registrar-specific integration is the problem and consistency across backend capabilities has real value. Discovery plus runnable examples shortens the learning path, but it does not erase the need for an overlap window, retries, or read-backs.

The decision rule is blunt: choose the API boundary based on maintenance cost, then choose the migration timing based on the slowest caches you are willing to support. For a one-person SaaS shipping weekly, a reversible cutover beats an impressive-looking instant flip.

Further reading

Top comments (0)