# DNS_PROBE_FINISHED_NXDOMAIN: Find and Fix DNS Failures

> Fix DNS_PROBE_FINISHED_NXDOMAIN by comparing authoritative and cached answers, then checking VPN DNS, hosts files, Docker, CoreDNS and DNSSEC.

Source: https://howhttpworks.com/debug/dns-probe-finished-nxdomain
Last reviewed: 2026-10-05

Error messages this page covers:
- `DNS_PROBE_FINISHED_NXDOMAIN`
- `ERR_NAME_NOT_RESOLVED`
- `curl: (6) Could not resolve host: no-such-host.hhw.invalid`
- `getaddrinfo ENOTFOUND no-such-host.hhw.invalid`
- `socket.gaierror: [Errno -2] Name or service not known`
- `socket.gaierror: [Errno 8] nodename nor servname provided, or not known`
- `api.example.com could not be resolved (3: Host not found)`

> **TL;DR:** DNS couldn't turn the hostname into an address, so the browser never reached a server. Run `dig api.example.com. A` on the failing machine, then the same query against `@1.1.1.1` and `@8.8.8.8`, and read the **status** line plus the answer and authority sections (`+short` hides them). Public DNS works but your app doesn't? Look at VPN DNS, `/etc/hosts` and Chrome Secure DNS. The authoritative server has the record but a resolver still says NXDOMAIN? That's a cached negative answer: check the SOA TTL before you touch the record again.

## What it means

The client failed to resolve the hostname, so it never got as far as connecting to the HTTP server. Chromium defines [`ERR_NAME_NOT_RESOLVED`](https://github.com/chromium/chromium/blob/main/net/base/net_error_list.h) as a hostname resolution failure. `DNS_PROBE_FINISHED_NXDOMAIN` is something else: a [browser probe diagnosis](https://github.com/chromium/chromium/blob/main/components/error_page/common/net_error_info.h) that runs after the failure and concludes that DNS works but this name seems to be missing. It's Chrome's interpretation, not a capture of the original DNS reply.

Here's how the same failure looks for `no-such-host.hhw.invalid` on macOS with **curl 8.7.1**, **Node 26.10.0** and **Python 3.14.8**, in that order:

```text
curl: (6) Could not resolve host: no-such-host.hhw.invalid
getaddrinfo ENOTFOUND no-such-host.hhw.invalid
socket.gaierror: [Errno 8] nodename nor servname provided, or not known
```

On Linux with glibc, Python prints the line below instead. Unlike the captures above, it's an illustrative reconstruction, built from Python's [gaierror definition](https://docs.python.org/3/library/socket.html#socket.gaierror) and glibc's [numeric value](https://github.com/bminor/glibc/blob/master/resolv/netdb.h) and [message](https://github.com/bminor/glibc/blob/master/sysdeps/posix/gai_strerror-strs.h):

```text
socket.gaierror: [Errno -2] Name or service not known
```

Node's [`ENOTFOUND`](https://nodejs.org/api/errors.html#common-system-errors) covers both `EAI_NONAME` and `EAI_NODATA`. [`dns.lookup`](https://nodejs.org/api/dns.html#dnslookuphostname-options-callback) goes through the OS resolver, and it can return this code for failures other than a missing hostname. A common one is passing a URL: hostname lookup APIs want `api.example.com`, not `https://api.example.com/path`.

If nginx resolves a dynamic upstream, the failure shows up in its error log instead. The fragment below is reconstructed from nginx's [upstream formatting](https://github.com/nginx/nginx/blob/master/src/http/ngx_http_upstream.c) and [resolver error strings](https://github.com/nginx/nginx/blob/master/src/core/ngx_resolver.c):

```text
api.example.com could not be resolved (3: Host not found)
```

Here nginx's own upstream lookup is failing, and the browser usually sees a [502 Bad Gateway](https://howhttpworks.com/debug/nginx-502-bad-gateway). Run your DNS checks from nginx's host or container, not your laptop.

## NXDOMAIN, SERVFAIL and no answer

Swap `api.example.com` for your failing hostname everywhere below. The trailing dot makes the name absolute, so no search suffix gets appended.

```bash
dig api.example.com. A
dig api.example.com. AAAA
nslookup api.example.com.
```

- **`status: NXDOMAIN`:** the name doesn't exist in that resolver's view of DNS. If the answer includes a CNAME, check its target too, because a CNAME can point at a name that doesn't exist.
- **`status: SERVFAIL`:** the resolver gave up. A down authoritative server, a broken delegation or a DNSSEC validation failure can all cause this. It tells you resolution failed, not that the name is missing.
- **`status: NOERROR` with no address:** this is NODATA. The name exists, but not with the record type you asked for, so an AAAA query can come back empty while A works. Check the authority section to tell NODATA apart from a referral.
- **No response before the tool times out:** that's neither NXDOMAIN nor NODATA. Check that you can reach the configured resolver and that nothing is dropping DNS traffic.

These categories come from [RFC 1035's response codes](https://www.rfc-editor.org/rfc/rfc1035#section-4.1.1) and [RFC 2308's negative-response definitions](https://www.rfc-editor.org/rfc/rfc2308#section-2). `dig +short` strips out exactly the fields you need to tell them apart.

## Fix it, in diagnostic order

### 1. A typo, missing record or broken CNAME target

Start by checking the exact failing name against the zone you actually manage:

```bash
dig @1.1.1.1 api.example.com. A
dig @8.8.8.8 api.example.com. A
dig +trace api.example.com. A
# Replace ns1.example.net with a server authoritative for the zone
dig @ns1.example.net api.example.com. A +norecurse
```

[`dig +trace`](https://bind9.readthedocs.io/en/latest/manpages.html) walks the delegation chain itself instead of asking a recursive resolver for the final answer. That means it needs to reach every authoritative server on the path, so on a network that blocks direct DNS queries it fails even when the zone is fine. It also doesn't validate DNSSEC.

If the authoritative server says NXDOMAIN, fix the spelling or create the record in the delegated zone. If there's a CNAME, query its target. And remember that each name is its own record: creating `www.example.com` gives you nothing for `example.com` or `api.example.com`.

A BIND zone-file record looks like this. The address is from the documentation range, so use your service's real one:

```text
api.example.com. 300 IN A 203.0.113.10
```

Publish the change through your DNS provider or authoritative server, then query each authoritative nameserver. If they give different answers, sort that out first. The recursive cache isn't your problem yet.

### 2. The record is new, but a resolver cached its earlier absence

```bash
dig @1.1.1.1 api.example.com. A +noall +comments +answer +authority
dig @ns1.example.net api.example.com. A +norecurse
dig @ns1.example.net example.com. SOA +norecurse
```

Authoritative server returns the A record, recursive resolver returns NXDOMAIN with an SOA in the authority section? You're looking at a cached negative answer. Check the TTL on that SOA. [RFC 2308 Sections 3 and 5](https://www.rfc-editor.org/rfc/rfc2308#section-3) set the initial negative TTL to **the smaller of the SOA record's TTL and its MINIMUM field** (the last number in the SOA data). The TTL you see counts down while the answer sits in cache. It has nothing to do with the TTL on the A record you just created.

Take this zone-file SOA as an example. The record TTL is 900 seconds and MINIMUM is 300, so resolvers cache "doesn't exist" for 300 seconds:

```text
example.com. 900 IN SOA ns1.example.net. hostmaster.example.com. (
  2026100501 3600 600 604800 300
)
```

Either wait out the remaining TTL or flush a resolver you control. On Linux with systemd-resolved:

```bash
sudo resolvectl flush-caches
resolvectl query api.example.com
```

That flushes [systemd-resolved's local cache](https://github.com/systemd/systemd/blob/main/man/resolvectl.xml) only. A public resolver upstream keeps its copy. Lowering the SOA values now won't shorten an answer that's already cached either, so set the negative TTL you want before you create records in future. There's no universal "DNS takes 48 hours" rule; the SOA tells you the real number.

### 3. VPN DNS, a split DNS view or a hosts override

Look at the resolver configuration on the machine where the app actually fails:

```bash
# Linux with systemd-resolved
resolvectl status
resolvectl query api.example.com
# macOS: includes supplemental resolver configurations
scutil --dns
# OS lookup, including hosts-file handling
node -e 'require("node:dns").lookup("api.example.com", {all:true}, console.log)'
grep -n 'api\.example\.com' /etc/hosts
```

On macOS, [`scutil --dns`](https://github.com/apple-oss-distributions/configd/blob/main/scutil.tproj/scutil.8) shows the full DNS setup, including the supplemental resolvers a VPN adds. Then query the VPN's resolver directly by its real address and compare it with public DNS:

```bash
dig @10.0.0.53 api.example.com. A
dig @1.1.1.1 api.example.com. A
```

If only the VPN resolver knows the record, it's a private name. Reconnect the VPN and get its DNS routing working again; switching the machine to public DNS throws that private view away. If only the public answer works, find out which resolver the VPN picked for that domain.

An [`/etc/hosts`](https://man7.org/linux/man-pages/man5/hosts.5.html) entry can make the OS lookup disagree with `dig`, because `dig` asks DNS and never reads the hosts file. Remove a stale override or fix its address. A temporary override looks like this:

```text
203.0.113.10 api.example.com
```

Point it at an address that actually answers. If the entry sends the name to a dead machine, DNS is fine but you'll hit a [connection timeout](https://howhttpworks.com/debug/err-connection-timed-out) next.

### 4. Chrome uses a different resolver path

OS lookup and curl both work, but Chrome fails? Check **Settings → Privacy and security → Security → Use secure DNS**. A custom DoH provider can answer differently from your VPN or local resolver, and Google [documents that a custom provider does not fall back to unencrypted DNS](https://support.google.com/chrome/answer/10468685?hl=en&co=GENIE.Platform%3DDesktop).

Switch temporarily to the system or VPN resolver and test again. If Chrome cached the failure, open `chrome://net-internals/#dns` and click **Clear host cache** (see Chromium's [DNS view code](https://github.com/chromium/chromium/blob/main/chrome/browser/resources/net_internals/dns_view.js)), then retry. This only clears Chrome's cache. It won't create a missing record or flush the upstream resolver's negative cache.

### 5. The application runs inside Docker or Kubernetes

Run the lookup from inside the failing container, with whatever tools that image has:

```bash
docker exec my-app cat /etc/resolv.conf
docker exec my-app cat /etc/hosts
docker exec my-app nslookup api.example.com.
kubectl -n web exec POD_NAME -- cat /etc/resolv.conf
kubectl -n web exec POD_NAME -- nslookup api.example.com.
kubectl -n web exec POD_NAME -- nslookup kubernetes.default
```

Docker [uses different DNS paths depending on the network](https://docs.docker.com/engine/network/#dns-services). Containers on the default bridge get a copy of the host's resolver config. Containers on custom networks use embedded DNS at `127.0.0.11`, which forwards external lookups to the host's configured DNS servers. Either way, entries in the host's `/etc/hosts` don't carry over. If the app needs a private resolver, give the container one it can reach:

```bash
docker run --dns 10.0.0.53 my-app-image
```

`--dns 127.0.0.1` won't reach the host's resolver. Inside the container, that's the container's own loopback.

Kubernetes usually sets `options ndots:5`, as its [DNS debugging guide](https://kubernetes.io/docs/tasks/administer-cluster/dns-debugging-resolution/) shows. Under the [`resolv.conf` search rules](https://man7.org/linux/man-pages/man5/resolv.conf.5.html), any name with fewer than five dots gets tried with each search suffix before the absolute lookup. `api.example.com.` skips that expansion in DNS tools. So an NXDOMAIN for one of the suffixed candidates is expected noise; check whether the final absolute query failed.

If search expansion is hurting external lookups, add this to the workload's Pod template. It lowers the threshold and keeps cluster DNS:

```yaml
spec:
  dnsPolicy: ClusterFirst
  dnsConfig:
    options:
      - name: ndots
        value: "1"
```

This [Pod DNS configuration](https://kubernetes.io/docs/concepts/services-networking/dns-pod-service/#pod-dns-config) changes lookup order, so check that short service names still resolve before you roll it out. It won't help with a record that's genuinely missing.

If `kubernetes.default` fails too, the problem is cluster DNS. Check the DNS service and CoreDNS before touching application settings:

```bash
kubectl -n kube-system get pods -l k8s-app=kube-dns
kubectl -n kube-system get service kube-dns
kubectl -n kube-system get endpointslice -l kubernetes.io/service-name=kube-dns
kubectl -n kube-system logs -l k8s-app=kube-dns --tail=100
kubectl -n kube-system get configmap coredns -o yaml
```

Work through the [CoreDNS debugging procedure](https://kubernetes.io/docs/tasks/administer-cluster/dns-debugging-resolution/): confirm the endpoints exist, read the forwarding config and, if you need to see queried names and response codes, temporarily add the `log` plugin inside the existing Corefile server block. If cluster names resolve but public names don't, look at CoreDNS's upstream resolver and the network path to it.

### 6. SERVFAIL comes from DNSSEC validation

Query the **same validating resolver** twice, once normally and once with checking disabled:

```bash
dig @1.1.1.1 api.example.com. A +dnssec
dig @1.1.1.1 api.example.com. A +dnssec +cdflag
dig example.com. DS +dnssec
dig example.com. DNSKEY +dnssec
```

SERVFAIL normally but data with checking disabled points to a validation failure. [RFC 4035](https://www.rfc-editor.org/rfc/rfc4035#section-5.5) has a validating resolver return server failure when validation fails and CD is unset, and Cloudflare [uses this same CD comparison in its troubleshooting procedure](https://developers.cloudflare.com/dns/dnssec/troubleshooting/).

The usual culprits are a stale DS at the parent after a nameserver migration, a DS that no longer matches the current DNSKEY, and expired signatures. Fix the signing setup with your DNS provider and registrar. Note that `+dnssec` only requests the DNSSEC records; plain `dig` doesn't validate the chain for you. Turning validation off for good hides the broken chain instead of fixing it.

## Related

- [ERR_CONNECTION_TIMED_OUT](https://howhttpworks.com/debug/err-connection-timed-out) — the hostname resolved, but connection establishment did not finish.
- [ERR_CONNECTION_REFUSED](https://howhttpworks.com/debug/err-connection-refused) — the address and port actively rejected the attempt.
- [nginx 502 Bad Gateway](https://howhttpworks.com/debug/nginx-502-bad-gateway) — a proxy may be the component failing to resolve its upstream.
