DNS Best Practices: Resilience, Security, and Hygiene for a Critical Service
The Domain Name System is the quiet foundation on which nearly every online service depends.
- Content owner
- Insyto Content Team
- Editorial reviewer
- Ritesh Mhatre
- Next review
- To be scheduled
- Technical reviewer
- Navish Ansari
- Last reviewed
- Review pending
- Technical level
- Intermediate · CISOs, security teams, IT directors
Executive Summary
The Domain Name System is the quiet foundation on which nearly every online service depends. Every website visit, email delivery, cloud API call, and software update begins with a DNS lookup that translates a human-friendly name into the numeric address a computer can reach. When DNS works — which is almost always — no one thinks about it. When it fails, everything fails at once: the website is unreachable, email stops flowing, applications time out, and to users it looks as though the whole business has vanished from the internet. Precisely because it is so foundational and so reliable, DNS is chronically under-managed, treated as a set-and-forget utility rather than the critical, targeted infrastructure it actually is.
That neglect is risky on two fronts. Operationally, DNS is a single point of failure for an entire domain: one lapsed registration, one misconfigured record, or one name-server outage can take a company offline in minutes, and the cached nature of DNS can make such failures both sudden and slow to reverse. On security, DNS is a favorite target because it is trusted by default — attackers who can forge, redirect, or hijack DNS answers can silently reroute users and email to systems they control, often without tripping any other alarm. Getting DNS right is therefore one of the highest-leverage, lowest-cost reliability and security investments an organization can make.
This vendor-neutral guide covers what it takes to run DNS well. It explains how resolution actually works and why caching and TTLs matter, the records that make up a zone and the mistakes each invites, the principal DNS threats and the established defenses that counter them — including DNSSEC and encrypted DNS — how to architect redundancy so no single failure takes you offline, and the operational hygiene that keeps DNS boringly reliable. Grounded in long-standing internet standards and public-sector deployment guidance, the goal is a DNS practice that is resilient, secure, and well-documented, so the service stays invisible for the right reasons.
How DNS Resolution Works
To manage DNS well, it helps to understand what happens in the fraction of a second between typing a name and reaching a server. Resolution is a walk down a distributed hierarchy, short-circuited by caching.
How a DNS lookup actually works
When an application needs the address for a name, it asks a recursive resolver — typically run by an ISP, an internal server, or a public provider. The resolver first checks its cache; if it has a recent answer, it returns it instantly. If not, it does the legwork: it asks a root server, which points it to the servers for the relevant top-level domain (such as .com); it asks the TLD servers, which point it to the domain’s authoritative name servers; and it asks those authoritative servers, which hold the real answer. The resolver returns the address to the application and caches it for the duration of the record’s time-to-live (TTL), so subsequent lookups are answered immediately without repeating the walk. This caching is why DNS is fast and resilient — the full hierarchy is only traversed when the cache is cold — but it is also why changes take time to propagate and why TTL is one of the most important dials in DNS management.
The Records You Manage
A DNS zone is assembled from a small set of record types, and most misconfigurations trace back to misunderstanding what one of them does. Knowing the building blocks prevents the majority of self-inflicted outages.
The DNS records you actually manage
The A and AAAA records are the workhorses, mapping a name to an IPv4 or IPv6 address respectively. A CNAME aliases one name to another — useful, but it must never sit on the zone apex, and a CNAME pointing at a resource that no longer exists (a “dangling” record) is a classic route to subdomain takeover. MX records direct email to the correct mail servers in priority order and should always be paired with the TXT-based SPF, DKIM, and DMARC records that protect against spoofing; a wrong MX record silently loses mail. TXT records carry that free-form verification and anti-spoofing data and are critical to email deliverability. NS and SOA records are the backbone of every zone — NS delegates the zone to its authoritative servers, while SOA holds the serial number and refresh and TTL parameters. And PTR records provide reverse lookups from IP to name, which mail servers and logging systems rely on; a missing PTR can cause a mail server’s messages to be rejected.
| Record | Purpose | Common pitfall |
|---|---|---|
| A / AAAA | Name to IPv4 / IPv6 address | Pointing at a decommissioned or wrong IP |
| CNAME | Alias one name to another | Dangling CNAME → subdomain takeover; apex misuse |
| MX | Route email to mail servers | Wrong priority or host → silent mail loss |
| TXT | SPF, DKIM, DMARC, verification | Missing/incorrect records → spoofing, poor deliverability |
| NS / SOA | Delegation and zone parameters | Inconsistent NS sets, stale serials |
| PTR | Reverse lookup (IP to name) | Missing PTR → mail rejected |
Cutting across all of them is TTL, which governs how long resolvers cache an answer. Long TTLs make DNS fast and resilient but slow to change; short TTLs make it agile but increase query load. The practical discipline is to lower TTLs a day or two before a planned change and raise them again once the change is stable.
DNS Threats and Defenses
DNS is attractive to attackers precisely because it is trusted implicitly — a forged answer is acted on without question. Fortunately, each major threat has an established, specific countermeasure.
DNS threats and the defenses that counter them
Cache poisoning and spoofing inject forged answers into a resolver so that users are redirected to attacker-controlled sites; the defense is DNSSEC, which cryptographically signs records so resolvers can verify a chain of trust and reject forgeries. Distributed denial-of-service attacks flood name servers to take a domain dark; the defense is redundancy combined with anycast and rate limiting, spreading the load across many geographically diverse servers that absorb and blunt the flood. Domain and registrar hijacking — where a stolen registrar account is used to change name servers or transfer the domain — is countered with registry and registrar locks, multi-factor authentication, and monitoring that alerts on any change. And subdomain takeover and eavesdropping, which exploit dangling records and plaintext queries, are addressed by record hygiene and by encrypting DNS queries in transit with DNS over HTTPS (DoH) or DNS over TLS (DoT).
| Threat | What it does | Primary defense |
|---|---|---|
| Cache poisoning / spoofing | Forged answers redirect users | DNSSEC (signed records, chain of trust) |
| DDoS against DNS | Floods servers, domain goes dark | Redundancy, anycast, rate limiting |
| Registrar / domain hijacking | Alters NS or transfers the domain | Registrar lock, MFA, change monitoring |
| Subdomain takeover | Exploits dangling records | Record hygiene, remove stale entries |
| Eavesdropping | Reads plaintext queries | Encrypted DNS (DoH / DoT) |
DNSSEC deserves particular emphasis because it addresses the one problem the original DNS design did not: it provides a way to verify that an answer genuinely came from the authoritative source and was not altered in transit, closing the door on the spoofing attacks that DNS is otherwise defenseless against.
Designing for Resilience
Because DNS is a single point of failure for the entire domain, resilience is not optional — it must be engineered so that no single server, site, or administrative lapse can take the business offline.
Building DNS that does not go down
The foundation is multiple authoritative name servers — at least two, ideally on separate networks or even separate providers, with a primary and secondary kept in sync through zone transfers so the loss of one does not lose the zone. Anycast and geographic diversity serve the same IP address from many locations worldwide, which lowers latency, absorbs DDoS traffic, and lets the service survive the loss of an entire site. Split-horizon (or split-brain) views present an internal view for private names and an external view for public ones, so internal records are never leaked to the internet. And registrar protection guards the administrative layer: a registry or registrar lock prevents unauthorized transfers, MFA secures the account, and auto-renewal plus monitoring of expiry and name-server changes prevents the depressingly common outage caused by a simply forgotten domain renewal.
| Resilience measure | What it protects against |
|---|---|
| Multiple authoritative NS | Loss of a single name server or network |
| Anycast + geo-diversity | Site outages, latency, DDoS |
| Secondary provider | Failure of an entire DNS provider |
| Split-horizon views | Leaking internal records publicly |
| Registrar lock + MFA + auto-renew | Hijacking and forgotten-renewal outages |
Running DNS Well
Beyond architecture, DNS reliability comes from operational discipline. Four pillars keep the service invisible and dependable.
Running DNS well — four operating pillars
Hygiene and change control matter most because the majority of DNS outages are self-inflicted edits: document every zone and record, remove stale and dangling entries, make changes through review rather than ad hoc, and version or back up zone files. Security applies the defenses above as standing practice — sign zones with DNSSEC, lock the registrar and enforce MFA, use DNS filtering to block known-malicious domains, and encrypt queries with DoH or DoT. Resilience is the architecture just described: multiple authoritative servers, anycast and geographic diversity, a secondary provider for failover, and automatic domain renewal. And monitoring and TTL management close the loop: watch resolution success and response time, alert on any name-server or record change, log queries to spot anomalies, and tune TTLs — lowering them ahead of planned changes. Together these turn DNS from an invisible risk into a measured, managed service.
DNS Best Practices Checklist
- Run at least two authoritative name servers on separate networks or providers.
- Use anycast and geographic diversity for authoritative DNS where possible.
- Enable DNSSEC to sign zones and protect against spoofing and cache poisoning.
- Lock the domain at the registrar and enforce MFA on the registrar account.
- Enable auto-renewal and monitor domain and certificate expiry well in advance.
- Document every zone and record; remove stale and dangling entries regularly.
- Make DNS changes through a review process; version or back up zone files.
- Pair MX records with SPF, DKIM, and DMARC to protect email.
- Use split-horizon views so internal records are never exposed publicly.
- Manage TTLs deliberately — lower before planned changes, raise once stable.
- Monitor resolution success, response time, and unauthorized NS/record changes.
- Consider DNS filtering to block malicious domains and DoH/DoT to encrypt queries.
Best Practices
Treat DNS as critical infrastructure. It is not a utility to configure once and forget; it is a single point of failure for the whole business. Give it the redundancy, monitoring, and change control you would give any tier-one system.
Never rely on a single name server or provider. Two or more authoritative servers on independent networks, ideally spanning providers, is the baseline. Anycast extends that resilience globally and blunts DDoS at the same time.
Sign your zones with DNSSEC. The original DNS has no way to prove an answer is authentic. DNSSEC closes that gap and is the definitive defense against spoofing and cache poisoning.
Lock down the registrar. Many of the most damaging DNS incidents are not clever attacks but hijacked registrar accounts or lapsed renewals. Registrar lock, MFA, auto-renewal, and change monitoring prevent both.
Keep records clean. Dangling records invite subdomain takeover, and stale entries cause confusing failures. Document everything, review regularly, and remove what is no longer needed.
Manage TTLs with intent. Lower TTLs before a planned change so it propagates quickly, then raise them afterward for performance and resilience. Do not leave everything on a single default.
Common Mistakes
Treating DNS as set-and-forget. The service that “just works” for years is exactly the one that fails catastrophically when a record, renewal, or server is neglected. DNS needs ongoing management.
Running only one name server. A single authoritative server means a single outage takes the domain offline. Redundancy across networks and providers is non-negotiable.
Skipping DNSSEC. Without signed zones, there is no way for resolvers to detect a forged answer, leaving users exposed to silent redirection.
Forgetting domain renewal. A lapsed registration can take an entire business offline and is entirely avoidable with auto-renew and expiry monitoring.
Leaving dangling records. CNAMEs and other records pointing at decommissioned resources are a well-known path to subdomain takeover. Remove them promptly.
Making undocumented, ad hoc changes. Unreviewed edits are the leading cause of DNS outages. Change control and zone backups turn a mistake into a quick rollback rather than an incident.
Frequently Asked Questions
Why is DNS so critical? Almost every online action — web, email, cloud, updates — starts with a DNS lookup. If DNS fails, all of those fail simultaneously, and to users the business appears to have disappeared from the internet.
What is DNSSEC and do we need it? DNSSEC cryptographically signs DNS records so resolvers can verify an answer is authentic and unaltered. It is the primary defense against spoofing and cache poisoning, and it is strongly recommended for any domain that matters.
How many name servers should we have? At least two authoritative name servers, ideally on separate networks or providers. Anycast and geographic diversity add further resilience and help absorb denial-of-service attacks.
What is TTL and how should we set it? TTL controls how long resolvers cache a record. Longer TTLs improve speed and resilience but slow changes; shorter TTLs do the reverse. Lower TTLs a day or two before a planned change, then raise them once it is stable.
What is split-horizon DNS? Split-horizon (or split-brain) DNS serves different answers to internal and external clients, so private, internal records are never exposed to the public internet while public records remain available to everyone.
How do we protect our domain from hijacking? Enable registrar lock, require multi-factor authentication on the registrar account, turn on auto-renewal, and monitor for any changes to name servers or records. Most hijackings exploit weak registrar security rather than DNS itself.
Conclusion
DNS earns its reputation as invisible infrastructure only when it is managed with the seriousness its importance warrants. It is simultaneously a single point of failure for the entire business and a prime target for attackers who exploit the trust placed in it — which makes resilience, security, and hygiene not optional refinements but core requirements. Understand how resolution and caching work, keep your records correct and clean, sign your zones with DNSSEC, and architect redundancy so no single server, provider, or lapsed renewal can take you offline.
The investment is modest and the return is outsized. Run multiple authoritative servers, lock down the registrar, automate renewals, document and review changes, manage TTLs deliberately, and monitor resolution and unauthorized changes. Do that, and DNS goes back to being what it should be: a reliable, secure foundation that no one has to think about — because it simply, quietly, always works.
References
- NIST SP 800-81r2 — Secure Domain Name System (DNS) Deployment Guide
- RFC 1034 — Domain Names: Concepts and Facilities
- RFC 1035 — Domain Names: Implementation and Specification
- RFC 9364 — DNS Security Extensions (DNSSEC) BCP
- RFC 4033 — DNS Security Introduction and Requirements
- RFC 7858 — DNS over TLS (DoT)
- RFC 8484 — DNS Queries over HTTPS (DoH)