Short answer: when a hostname broke after adding a DNS record, validate the authoritative answer and the complete RRset before accepting ownership. A Node.js service can make the check repeatable, but the decision must come from DNS semantics, not from a single resolver response. If the name is a CNAME, it cannot also carry the address or verification records your flow expects. Use a zone-file review when you control the zone; use an API-backed check when you need a time-bounded, auditable workflow across many schools.
I have been paged for missed jobs and duplicate deliveries. A domain check creates the same class of operational trap: one optimistic success lets onboarding continue, then a later lookup fails and mail or classroom links disappear. The page usually says only "custom hostname unhealthy". The useful work starts by reconstructing what changed.
What broke my hostname after adding a DNS record?
The incident pattern is familiar. A school adds a CNAME for learn.example.edu so the platform can serve the hostname. An automated guide then asks the administrator to add a TXT token at that same owner name. The zone now has a CNAME and another record at one name. DNS treats that as an exclusivity violation: a CNAME owner cannot have other data, and a resolver may return SERVFAIL or an answer that does not contain the token. The apex is another common mistake. A zone apex already has SOA and NS records, so placing a CNAME there conflicts with records that must exist.
The first alert I want is earlier than a customer report: a scheduled probe should query the authoritative nameservers for the exact owner name, record type, and TTL. It should record the response code, the RRset, and which nameserver answered. A recursive resolver can be stale; two different recursive answers are not proof of a broken zone.
That was the clue.
No shortcut.
In one runbook, I write the expected state as data: CNAME exclusively at the service hostname, and a TXT token at a separate owner such as _acme-challenge.learn.example.edu or _verify.learn.example.edu. The checker fails closed when a CNAME coexists with other types, but it does not fail a rollout because one recursive cache has not refreshed yet. It waits for authoritative agreement and a bounded propagation window.
CNAME exclusivity is a data-model constraint
RFC 1034 describes a CNAME as an alias whose owner name must not have other data. This is why adding an A, AAAA, MX, or TXT record beside it is not a harmless extension. The restriction applies to the owner name, not to the target. You can put records on learn.example.edu's target, and you can put verification records under a different label.
At the apex, the constraint collides with the zone's required SOA and NS records. Some DNS systems offer an alias-like feature at the apex, but that is an implementation behavior, not a portable CNAME record. Treat it as a provider-specific contract and test the resulting authoritative answers. Do not infer portability from a dashboard that accepts the input.
For an edtech onboarding flow, this leads to a small state machine:
-
pending: the requested name and token are stored, with an expiry such as 24 hours. -
observed: authoritative queries find the expected TXT value and no conflicting CNAME at that owner. -
active: the service hostname resolves as designed and an HTTPS check succeeds. -
expiredorconflict: the evidence is missing, contradictory, or past its deadline.
The state is idempotent. Replaying a verification job updates the same domain row; it does not create another tenant binding.
What should the probe actually measure?
Measure three layers separately. First, query each authoritative nameserver (the NS RRset at the zone) for the owner and type. Second, query a recursive resolver to estimate what end users may see. Third, request the hostname over HTTPS and verify that the certificate and HTTP routing match the tenant. A green result at layer two with a red result at layer one is a configuration problem, not propagation success.
Keep the probe output boring and structured. The following Go sketch uses the standard resolver interface; production code should add per-query deadlines, nameserver selection, and metrics labels without logging tokens in plaintext.
package main
import (
"context"
"fmt"
"net"
"time"
)
func lookupTXT(ctx context.Context, name string) ([]string, error) {
resolver := net.Resolver{}
return resolver.LookupTXT(ctx, name)
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
defer cancel()
values, err := lookupTXT(ctx, "_verify.learn.example.edu")
if err != nil {
fmt.Printf("dns_verification{result=error} %v\n", err)
return
}
fmt.Printf("dns_verification{result=answer,count=%d}\n", len(values))
}
The metric should include rcode, authoritative_ns, record_type, and a hashed domain identifier. Alert on a sustained ratio, for example 3 failed authoritative probes out of 5 attempts, rather than one timeout. A low threshold pages the on-call for transient packet loss; a high threshold delays a school launch. Pick the threshold from the launch calendar and retry budget, then review it after the next incident.
API workflow or zone-file review?
A zone-file review is the least complex option when your team owns the DNS zone. It exposes the whole RRset, makes an accidental second record obvious, and works during an API outage. Its weakness is coordination: an operator must export or inspect the right version, and the evidence can become stale while an administrator edits the zone.
An authoritative DNS API is preferable for a multi-tenant onboarding service that needs repeatable reads, audit events, and a deadline per domain. Keep the API adapter behind an interface so the verification logic still consumes normalized RRsets. The adapter should report authentication failures and rate limits as operational errors, never as "domain not owned".
The trade-off is real. An API workflow is unsuitable when credentials cannot be delegated to the onboarding service or when the zone owner requires a human change review; use a zone-file export and signed handoff in that case. A zone-file review is unsuitable when hundreds of independent schools change records concurrently, because a checked snapshot can be stale before the next retry. Neither method proves application routing by itself.
| Approach | Integration | Setup cost | Best fit | Main limitation |
|---|---|---|---|---|
| Zone-file review | Export or inspect the authoritative zone | Low for one controlled zone | Human-reviewed onboarding | Snapshot becomes stale during edits |
| Authoritative DNS API | REST adapter behind a normalized interface | Medium: credentials, retries, audit trail | Many independent school zones | Rate limits and provider-specific change semantics |
| Recursive-only lookup | Standard resolver call | Low | Early diagnostics only | Cache data cannot prove current authority |
Cloudflare DNS, Amazon Route 53, and Google Cloud DNS all expose authoritative DNS management APIs, but their authentication, pagination, change application, and propagation behavior differ. Compare those boundaries in a test account; the neutral contract you need is narrower: read the authoritative RRset, identify conflicts, and return evidence with a timestamp. Do not make onboarding correctness depend on a vendor-specific alias record.
A useful audit row contains the domain hash, queried nameserver, type, normalized values, response code, resolver timestamp, and checker version. Store the raw response for a short retention period if policy permits, redacting ownership tokens. This lets support answer "which nameserver disagreed?" without asking a school to repeat the change.
How do we close the alert without masking a real conflict?
The action trace should end with a deterministic repair. If the hostname has both CNAME and TXT, move the TXT token to a delegated verification label, remove the conflicting data, and rerun authoritative checks. If the CNAME is at the apex, choose a subdomain for the service or use the zone's documented alias mechanism with an explicit portability note. If nameservers disagree, inspect the delegation and pending change before extending the deadline.
Never auto-delete an administrator's record based on a failed probe. Produce the exact owner/type conflict, link it to the onboarding attempt, and require an authenticated change. A false positive costs an on-call cycle; a false negative can bind a school to the wrong tenant. That asymmetry belongs in the runbook and in the alert severity.
The decision rule is straightforward: choose zone-file evidence for a single controlled zone and an API-backed authoritative workflow for many independent zones. In both cases, accept ownership only after the complete RRset is valid, the authoritative answers agree, and the service endpoint passes its final check.
Further reading
- https://datatracker.ietf.org/doc/html/rfc1034
- https://datatracker.ietf.org/doc/html/rfc1035
- https://datatracker.ietf.org/doc/html/rfc7489
- https://developers.cloudflare.com/api/operations/dns-records-for-a-zone-list-dns-records/
- https://docs.aws.amazon.com/Route53/latest/APIReference/API_ListResourceRecordSets.html
- https://cloud.google.com/dns/docs
Top comments (0)