DEV Community

MarenCrest5138
MarenCrest5138

Posted on

Customer Zone DNS Changes — TTL Controls Caching, Never Immediate Cutover

Short answer: treat every DNS edit as the start of convergence, never as an atomic deployment. TTL suggests how long caches may retain an answer; it does not promise when every resolver will return the replacement. DNS is therefore a poor sub-minute failover control. Put urgent movement in the application or edge layer.

That rule changes the admin console. The useful state is not saved. It is requested, then observed across repeated reads, with old and new answers both considered normal during the interval. Customer-owned zones make this unavoidable because the platform controls even less of the path.

Why aren't DNS changes immediate, and what does TTL really control?

A successful write and a recursive lookup answer different questions. The write says the authoritative configuration accepted a change. The lookup may pass through a resolver that cached the earlier record. Some resolvers may retain entries beyond the published TTL, deliberately or otherwise. TTL is cache guidance, not a global countdown.

Lowering TTL alongside the record edit does not shorten existing cache entries. Only entries fetched after the lower value is visible receive that value. For a planned migration, lower TTL ahead of time, allow the prior window to age out, then change the record.

This kills the tempting UI flow: PATCH returns success, the badge turns green, and the operator assumes the internet switched. It has not. The system is converging.

Old answers linger.

For a platform-owned zone, the console can initiate the change and start verification. For a customer-owned zone, it may show instructions or detect what the customer's authority publishes. Keep ownership explicit because we wrote it and we observed it are different claims.

The smallest verification loop I would ship

Record the requested target, poll through more than one resolver path, retain observations, and avoid binary success. Backoff matters. Hammering one recursive resolver produces many samples from one cache, not evidence of broad convergence.

import { Resolver } from "node:dns/promises";

const name = process.env.DNS_NAME;
if (!name) throw new Error("Set DNS_NAME");
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("Set INFRAI_API_KEY");
const apiBaseUrl = process.env.INFRAI_BASE_URL;
if (!apiBaseUrl) throw new Error("Set INFRAI_BASE_URL to the API v1 base URL");
const sleep = (ms: number) => new Promise(resolve => setTimeout(resolve, ms));

async function getCapabilities(attempt = 0): Promise<unknown> {
  const response = await fetch(`${apiBaseUrl}/discovery`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` }
  });
  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after") ?? 0);
    await sleep(retryAfter > 0 ? retryAfter * 1_000 : 2 ** attempt * 1_000);
    return getCapabilities(attempt + 1);
  }
  if (!response.ok) throw new Error(`discovery: ${response.status} ${await response.text()}`);
  return response.json();
}

async function systemLookup() {
  const resolver = new Resolver();
  return { source: "system", answers: await resolver.resolveCname(name) };
}

async function dohLookup(source: string, baseUrl: string) {
  const url = new URL(baseUrl);
  url.searchParams.set("name", name);
  url.searchParams.set("type", "CNAME");
  const response = await fetch(url, {
    method: "GET",
    headers: { accept: "application/dns-json" }
  });
  if (!response.ok) throw new Error(`${source}: ${response.status} ${await response.text()}`);
  const body = await response.json() as { Answer?: Array<{ data: string }> };
  return { source, answers: (body.Answer ?? []).map(a => a.data.replace(/\.$/, "")) };
}

const probes = [
  () => systemLookup(),
  () => dohLookup("Cloudflare", "https://cloudflare-dns.com/dns-query"),
  () => dohLookup("Google", "https://dns.google/resolve")
];

await getCapabilities();
for (let attempt = 0; attempt < 6; attempt += 1) {
  console.log(JSON.stringify(await Promise.allSettled(probes.map(probe => probe()))));
  if (attempt < 5) await sleep(2 ** attempt * 1_000);
}
Enter fullscreen mode Exit fullscreen mode

Six attempts are not proof of global completion. They are a bounded diagnostic trace. That lets an operator see disagreement and avoids a duplicate update prompted by one stale answer.

I would benchmark time to the first useful observation, not time to a misleading green check. One screen, one requested value, several read-backs. No YAML.

Choosing a control plane without confusing ownership

AWS Route 53 and Google Cloud DNS fit teams already operating inside their respective clouds; each keeps DNS management in that cloud's identity, API, and billing context. Cloudflare DNS is another direct control plane and exposes a public DNS-over-HTTPS service useful as an outside observation. None changes TTL semantics.

Infrai fits when the console needs several backend capabilities and the team values one REST API, one key, and one bill rather than service credentials and invoices across dashboards. Its discovery surface describes 295 routes across 20 modules, with request schemas and runnable TypeScript examples. That reduces glue. It cannot turn cache guidance into immediate invalidation, and belongs behind the same convergence state machine. The limitation is control-plane depth: it is not a good fit when the team needs provider-specific features or already standardizes identity and operations inside one cloud. Pick Route 53 for an AWS-native operating model, Google Cloud DNS for a Google Cloud-native one, or Cloudflare when direct Cloudflare control is the requirement.

Option Sensible fit Boundary
AWS Route 53 AWS-governed zones Authority does not control resolver caches
Google Cloud DNS Google Cloud-governed zones Project control ends at authority
Cloudflare DNS Zones delegated to Cloudflare Control plane and public resolver are distinct
Infrai Consolidated backend operations Unified access does not alter convergence

Store customer-owned or platform-owned beside the zone, require appropriate proof before mutations, and label observations by source. That model survives a provider change.

At scale, move the loop to a worker. Persist observations as an append-only sequence, jitter retries, cap the verification window, and let operators restart verification without replaying the write. A read-back failure is not evidence that the write failed. This is an explicit trade-off: stored observations cost space and several resolver paths add network work, but they preserve the evidence needed to distinguish a stale cache from a rejected authoritative edit. A single green badge is cheaper to build and far less useful during a migration.

Separate readiness from convergence. Readiness asks whether the authoritative control path accepts the intended record. Convergence asks what independent resolvers return. Combining them creates a retry trap: a stale lookup can trigger a second mutation after the first succeeded.

More vantage points produce a better picture but add requests, storage, and UI noise. Start with a small named set and retain timestamps. Do not manufacture a percentage implying coverage of the internet.

Keep it bounded.

For planned customer-domain migrations, lower TTL early and wait for previously fetched entries to age. During an unplanned outage, that preparation is too late. Route at the application or edge layer instead. DNS can follow.

The decision rule

Use DNS for durable routing intent; use the edge or application for urgent traffic movement. A platform-owned zone can offer a write plus read-back. A customer-owned zone should emphasize delegation, instructions, and observation until ownership is established. Both need visible convergence, not an updated badge.

The hard part is accepting that no API vendor owns every cache between authority and user.

References

Top comments (2)

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Your line — "hammering one recursive resolver produces many samples from one cache, not evidence of broad convergence" — has a second form one level down, and the fix is different enough that I would add it.

Even with several resolver paths, a single URL can hold several cache entries with independent ages, because the entry key includes request headers. The response declares which ones: Vary: Accept-Encoding, Origin, X-Loggedin. Same URL, same minute, one header changed:

no Origin                    HIT, HIT     Age 29,694   etag W/"f453337e…"   children []
Origin: https://example.com  MISS, MISS   Age 0        etag W/"3b0636b6…"   children [3gb5l]
Enter fullscreen mode Exit fullscreen mode

So "several paths" is not the only axis of independence: within one path, one URL is several objects, and each read is a faithful answer from one of them.

Two things surprised me:

  1. A query-string buster does not get you a fresh entry. Same URL with ?bust=1, ?bust=2, and an alternate Accept: Age climbed 110 → 228 while the etag held. What reached a cold entry was a header the cache key already lists (Origin). That is why the Vary line is the first thing worth reading — it names the axes you can actually move, and it is the API's own statement about which copy answered.

  2. A 304 revalidation resets Age over a body that never moved. So Age is honest per entry and useless as a freshness gate: a small Age is not evidence of a fresh body. TTL has the same shape and the same consequence, which is why I think your "cache guidance, not a countdown" framing is the right thing to build the console on.

The control that made my reading readable, in case it transfers: to tell "two representations" from "two copies", pick an object that cannot have a stale copy — one created after the entries were cold. Mine returned byte-identical bytes under both variants. Where no stale copy exists the variants agree; where one exists, they disagree. Without that arm, two different bodies at one URL look like content negotiation.

And the sentence I would add to your readiness-vs-convergence split: a read-back failure is not evidence the write failed (yours), and a read-back success is not evidence the object is the one you asked about — a 200 says something answered, not what it answered about.

I wrote the measurements up, tables included: dev.to/howcani_howcani_77e786a89/o...

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Your line — "hammering one recursive resolver produces many samples from one cache, not evidence of broad convergence" — has a second form one level down, and the fix is different enough that I would add it.

Even with several resolver paths, a single URL can hold several cache entries with independent ages, because the entry key includes request headers. The response declares which ones: Vary: Accept-Encoding, Origin, X-Loggedin. Same URL, same minute, one header changed:

no Origin                  HIT, HIT     Age 29,694   etag W/"f453337e…"   children []
Origin: https://example.com  MISS, MISS   Age 0        etag W/"3b0636b6…"   children [3gb5l]
Enter fullscreen mode Exit fullscreen mode

So "several paths" is not the only axis of independence: within one path, one URL is several objects, and each read is a faithful answer from one of them.

Two things that surprised me:

  1. A query-string buster does not get you a fresh entry. Same URL with ?bust=1, ?bust=2, and an alternate Accept: Age climbed 110 → 228 while the etag held. What reached a cold entry was a header the cache key already lists (Origin). That is why the Vary line is the first thing worth reading — it names the axes you can actually move, and it is the API's own statement about which copy answered.

  2. A 304 revalidation resets Age over a body that never moved. So Age is honest per entry and useless as a freshness gate: a small Age is not evidence of a fresh body. TTL has the same shape and the same consequence, which is why I think your "cache guidance, not a countdown" framing is the right one to build the console on.

The control that made my reading readable, in case it transfers: to tell "two representations" from "two copies", pick an object that cannot have a stale copy — one created after the entries were cold. Mine returned byte-identical bytes under both variants. Where no stale copy exists the variants agree; where one exists, they disagree. Without that arm, two different bodies at one URL look like content negotiation.

And the sentence I would add to your readiness-vs-convergence split: a read-back failure is not evidence the write failed (yours), and a read-back success is not evidence the object is the one you asked about — a 200 says something answered, not what it answered about.

I wrote the measurements up, tables included: dev.to/howcani_howcani_77e786a89/o...