Skip to main content
All case studies
Warning DNS cache poisoning October 2, 2026

The cache that would not forget

After a DNS migration, one subdomain went dark while everything else resolved. Two record fixes landed in zone files the query path never reads. The real fault: one resolver cached a public-root does-not-exist for a private name, and held it for twenty-four hours.

Time to root cause
25 min
20:41 to 21:06, through two false summits
Duration
1 h 11 m
20:41 to 21:52, symptom to green
Commands
314 / 24 h
48 messages in session
Data lost
None
one A record and one cache clear

The symptom

The operator had just moved DNS from AdGuard to a pair of Technitium containers. Most services resolved fine, but one subdomain refused to load: the Gotify notification server behind a Caddy reverse proxy. The Caddy entry was still there, and the name had worked before the change.

The first volley found a split answer. Asked directly, resolver one returned the reverse proxy's internal address. Resolver two returned a public IP that belonged to nobody in the lab. Classic migration leftover: one box had the record, one did not. Or so it looked.

The investigation

The obvious fix was adding the missing record on resolver two. The operator did it, and the name still failed. The smell moved downstream: the public name now resolved on both boxes, but Caddy's upstream, an internal-only name, did not resolve through the proxy's primary resolver. One early TLS probe got thrown out as junk, because a curl to a bare IP sends no SNI, so the failed handshake said nothing about the server. With a correct probe the picture was clean: TLS healthy, certificates fresh, and the proxy answering HTTP/2 502. The upstream was unresolvable, full stop.

The operator then reframed the hunt. This is split-horizon DNS, and the internal zone is DHCP-driven: records should register themselves, no static addressing. The question became where the dynamic-update pipe breaks. The zone files answered with a surprise: neither resolver's zone contained the record at all. The boxes were not answering from their zones. They were forwarding.

The router settled it. Queried directly, it returned authoritative answers with TTL 0 for every DHCP name, including the one that would not resolve. Lease registration had worked the whole time. Both manual fixes were real, they just landed in files the live query path never reads.

Root cause

Negative cache poisoning by fallback. The internal suffix sits on .internal, a TLD reserved for private use that does not exist on the public internet. During the migration window, a query on resolver one missed the forwarder path, recursed out to the public root, and came back NXDOMAIN carrying the root's own SOA: a definitive does-not-exist, cached for 86,400 seconds. A definitive negative kills failover. Clients that asked resolver one gave up on the first answer.

That asymmetry is what made the incident look random. Resolver two carries the same flaw but had never cached a negative for this particular name, so it forwarded to the router and served the correct address. One box jailed the name for a day, the other answered it fine. Clients configured with both resolvers got a coin flip.

The reboot test proved how stubborn the jail was. With the operator's approval, resolver one was rebooted, and the poison survived: Technitium writes its cache to disk on a clean shutdown and reloads it at boot. Worse, the router itself mints fresh root-grade negatives on every miss, so the factory kept running either way.

The fix

Five seconds, one button: Clear Cache on the resolver's dashboard. The router already held the right answer, so the name resolved the moment the jailed entry died. No proxy restart, nothing else touched.

The durable fixes came next, in order of value. On the router, a UniFi UDM, declare the private domain local with a one-line dnsmasq addition delivered by an onboot script, because UniFi regenerates its DNS config and hides domain-local behavior from the UI. Misses then die at the router instead of minting twenty-four-hour poison. The conditional forwarder zones that pin the private name to the router turned out to need nothing: they had been configured correctly on migration day.

The agent could not push the button itself. The Technitium API wants the admin token, and secrets stay out of the model's hands. After the fix, the operator attached a read-only token to the API connectors, which gave future sessions live eyes into the zone tables, and no hands.

The result

End to end, green on the first probe after the clear: client to Caddy to the internal name to the app, which answered its version JSON through the full chain. Root cause was isolated roughly twenty-five minutes in, through two false summits, with a total session of one hour and eleven minutes.

Two honest notes went on the record. The agent's early worry about a stale shadow zone was wrong: the API showed the forwarder zones had existed since the migration. The whole incident was one cached negative, nothing structural. And the second resolver got a cache clear too, since the same fresh-probe test had left a harmless but pointless root negative sitting in it.

Incident timeline

  • Symptom reported

    20:41 · one subdomain dead after the AdGuard to Technitium move; Caddy entry intact

  • Split answer found

    20:42 · resolver one returns the proxy address, resolver two a public IP

  • First fix lands

    20:45 · missing A record added on resolver two; the public name resolves everywhere

  • Upstream unresolvable

    20:50 · correct-SNI probe: TLS healthy, HTTP/2 502, upstream name dead

  • Operator reframes

    20:55 · split-horizon DHCP should self-register; find the broken pipe

  • Zone files hold nothing

    21:03 · no zone contains the record; the boxes answer by forwarding

  • Root cause named

    21:06 · router authoritative with TTL 0; resolver one holds a root-grade cached negative

  • Reboot behind a gate

    21:10 · pct reboot 105 approved; the persisted cache file reloads the poison

  • Poison factory confirmed

    21:14 · the router itself mints fresh root NXDOMAIN on every miss

  • Chain green

    21:52 · cache cleared; the app answers its version JSON end to end

Command log

Signature commands

dig +short gotify.example.ca @10.10.1.3
exit 0 The split answer: resolver two hands back a public IP nobody in the lab owns
dig @10.10.1.1 gotify.example.internal A +noall +comments +answer
exit 0 The router answers authoritatively, TTL 0. DHCP registration never broke
dig @10.10.1.2 zzz-fresh-probe-7319.example.internal A +noall +comments +authority
exit 0 A never-queried name draws a root NXDOMAIN. The fallback is live, not history
curl -sS --resolve gotify.example.ca:443:10.10.1.50 https://gotify.example.ca/version
exit 0 Proof of fix: the app answers its version JSON through the full chain
Command log
Time Agent Command Exit
20:41:40 SCAIRE dig +short gotify.example.ca @10.10.1.2 resolver one: the reverse proxy address exit 0
20:41:41 SCAIRE dig +short gotify.example.ca @10.10.1.3 resolver two: a public IP nobody in the lab owns exit 0
20:42:29 SCAIRE dig example.ca AXFR @10.10.1.3 transfer refused: the two boxes drift independently exit 0
20:47:40 SCAIRE dig @10.10.1.2 gotify.example.internal NOERROR with zero records, root SOA in authority exit 0
21:01:53 SCAIRE pct exec 105 -- grep -ac gotify /etc/dns/zones/example.internal.zone zero matches: the record never lived in any zone exit 1
21:03:02 SCAIRE dig @10.10.1.1 gotify.example.internal A +noall +comments +answer the router: authoritative answer, TTL 0. DHCP registration works exit 0
21:03:03 SCAIRE dig @10.10.1.2 gotify.example.internal SOA +noall +answer +authority resolver one backs its NODATA with the root servers SOA exit 0
21:05:02 SCAIRE dig @10.10.1.2 zzz-fresh-probe-7319.example.internal A +noall +comments +authority never-queried name: NXDOMAIN plus root SOA. The fallback is live exit 0
21:10:09 SCAIRE pct reboot 105 approval gate: risk stated, operator signed off exit 0
21:10:32 SCAIRE curl --resolve gotify.example.ca:443:10.10.1.50 -sS -o /dev/null -w "HTTP %{http_code} / %{http_version}" https://gotify.example.ca/ after reboot: still 502. The poison survived exit 0
21:12:24 SCAIRE dig @10.10.1.2 gotify.example.internal +noall +answer +comments still NODATA, root SOA, about 15 hours of cache left exit 0
21:13:43 SCAIRE dig @10.10.1.1 freshprobe-58231.example.internal +noall +authority +comments the router holds a fresh root negative: it is the poison factory exit 0
21:49:49 SCAIRE Technitium DNS1: GET /zones/list read-only token: forwarder zones existed since migration day exit 0
21:52:03 SCAIRE curl -sS --resolve gotify.example.ca:443:10.10.1.50 https://gotify.example.ca/version end to end: the app answers its version JSON exit 0

From the session

Operator
I changed the DNS servers from AdGuard to Technitium and most things are working, but gotify.example.ca is not resolving. There is an entry in Caddy for it, and it worked before the change.
SCAIRE
Look at the split answer. Resolver one returns the proxy's address, resolver two returns a public IP that belongs to nobody here. Classic migration leftover: add the missing record on resolver two, then flush the client caches.
Operator
I updated the other server, but that was not the issue.
SCAIRE
Your fix landed, and a new smell appeared. The public name resolves on both boxes now, but the proxy's upstream, an internal-only name, does not resolve through resolver one. With a correct SNI probe the proxy answers 502: TLS is fine, the upstream is dead.
Operator
It is split-horizon DNS. The internal name is DHCP and should update the DNS server automatically, like every other internal address. That is what we need to troubleshoot.
SCAIRE
The router settled it. It answers every DHCP name authoritatively, including the one you thought was broken. Your auto-registration works, and neither zone file contains the record at all. Resolver one is holding a cached negative: it asked the public internet about a private name, the root said the TLD does not exist, and it cached that insult for twenty-four hours.
Operator
My router is a UniFi UDM. Does that mean rebooting the router?
SCAIRE
No reboot needed. Clear the cache on the resolver, that is the green light. The durable fix is config: a one-line local-domain declaration on the router so misses die there. I cannot push the button myself, the API wants the admin token, and secrets stay out of my hands.
SCAIRE
Case closed. The app answers its version JSON through the whole chain. One cached negative, nothing structural, and the migration was set up correctly all along.

Lessons

.internal is a reserved private TLD. A resolver that falls back to public recursion on a forwarder miss will cache a root-grade NXDOMAIN for twenty-four hours. Keep private zones local on the router and pinned with conditional forwarders.
A reboot is not a cache flush when the resolver persists its cache file across clean shutdowns. Know what a restart actually clears before you trust it.
Follow the query path, not the config. Two correct record fixes landed in zone files the live path never reads, while the truth sat in a TTL 0 lease registration on the router.