Troubleshooting

NXDOMAIN or SERVFAIL? Investigate before changing DNS

A read-only DNS incident workflow for negative answers, inconsistent nameservers, DNSSEC errors and service failures.

Prepared with AI assistance. Read our editorial and corrections policy.

When a domain stops working, the first useful output is a precise failure description. “DNS is broken” hides important differences between a missing name, an unreachable nameserver, a validation failure and an application refusing a correct address. Gather evidence before editing the zone.

Define the scope of the incident

Record the exact hostname and record type, the time, the resolver used and the observed error. Ask whether all users are affected, whether IPv4 and IPv6 differ and whether only one service fails. A successful homepage does not prove that the MX host or DKIM selector works. A failed browser connection does not prove that DNS resolution failed.

Use read-only queries from more than one network when possible. Keep the full output, including status, flags, answer and authority sections. The short answer format is convenient once you understand the incident, but it can omit the detail needed to distinguish a negative response from a validation problem.

dig example.com A
dig example.com AAAA
dig example.com MX
dig +trace example.com
dig @ns1.example.net example.com SOA +norecurse

Understand negative responses

NXDOMAIN states that the queried name does not exist in the returned DNS view. Check spelling and the exact label before adding records. In contrast, a NOERROR response without the requested type may describe an existing name with no data of that type. Both kinds of negative evidence can be cached according to the zone's negative-caching parameters.

If you have just created a hostname, compare the authoritative response with the recursive resolver. When authority has the new record and a resolver still returns a cached negative response, repeatedly deleting and recreating the record adds risk without addressing the cache. Record the negative TTL and retest later.

Trace SERVFAIL to a component

SERVFAIL is a broad resolution failure. Check whether the delegated nameservers respond, whether they serve the zone authoritatively and whether the DNSSEC chain validates. Query every delegated server; one missing zone or broken route can create intermittent failures.

If a validating recursive query fails but a diagnostic query with checking disabled returns data, DNSSEC warrants investigation. This comparison is a clue, not proof that the returned data is safe. Do not tell users to disable validation as a permanent fix. Check DS records in the parent, DNSKEY records at the child and the provider's signing status.

Compare authority with caches

Different recursive resolvers can legitimately return different cached versions within a TTL window. Different authoritative servers publishing contradictory data require attention at the authoritative service. Compare SOA serials and important records. Equal serials are useful evidence, but they do not replace comparison of actual answers.

If only the local machine fails, inspect its configured resolver, VPN, local hosts file and network policy. A corporate split-DNS setup may intentionally return a different private answer. Do not overwrite public records to solve a local resolver configuration problem.

Once DNS works, test the service

A correct A record can still lead to an HTTP error if the server does not recognize the Host header. A correct AAAA record may lead to a blocked IPv6 port. HTTPS additionally needs a valid certificate for the requested hostname. For mail, an MX target must lead to a reachable mail service; a website listening on that address is not enough.

When testing web routing directly, preserve the original hostname for TLS and HTTP rather than replacing it with the IP address in the browser. Otherwise you may create an unrelated certificate or virtual-host error. Use a diagnostic tool that supports an explicit hostname-to-address override.

Create a useful escalation record

  • Exact hostname, query type and time with timezone.
  • Parent delegation and all authoritative results.
  • Recursive results, including TTLs and status codes.
  • Recent change history and expected value.
  • Service checks and the smallest reproducible failure.

Share redacted output with the relevant provider. Exclude private keys, passwords, tokens and message bodies. Once the failing component is identified, use a targeted change and the change checklist to verify recovery.

Sources and further reading

Examples use reserved documentation names and addresses. Confirm current provider requirements before changing production systems.