The cache that would not forget
After a DNS migration, one subdomain went dark while everything else resolved. Two record fixes landed in zone files the query path never reads. The real fault: one resolver cached a public-root does-not-exist for a private name, and held it for twenty-four hours.
The symptom
The operator had just moved DNS from AdGuard to a pair of Technitium containers. Most services resolved fine, but one subdomain refused to load: the Gotify notification server behind a Caddy reverse proxy. The Caddy entry was still there, and the name had worked before the change.
The first volley found a split answer. Asked directly, resolver one returned the reverse proxy's internal address. Resolver two returned a public IP that belonged to nobody in the lab. Classic migration leftover: one box had the record, one did not. Or so it looked.
The investigation
The obvious fix was adding the missing record on resolver two. The operator did it, and the name still failed. The smell moved downstream: the public name now resolved on both boxes, but Caddy's upstream, an internal-only name, did not resolve through the proxy's primary resolver. One early TLS probe got thrown out as junk, because a curl to a bare IP sends no SNI, so the failed handshake said nothing about the server. With a correct probe the picture was clean: TLS healthy, certificates fresh, and the proxy answering HTTP/2 502. The upstream was unresolvable, full stop.
The operator then reframed the hunt. This is split-horizon DNS, and the internal zone is DHCP-driven: records should register themselves, no static addressing. The question became where the dynamic-update pipe breaks. The zone files answered with a surprise: neither resolver's zone contained the record at all. The boxes were not answering from their zones. They were forwarding.
The router settled it. Queried directly, it returned authoritative answers with TTL 0 for every DHCP name, including the one that would not resolve. Lease registration had worked the whole time. Both manual fixes were real, they just landed in files the live query path never reads.
Root cause
Negative cache poisoning by fallback. The internal suffix sits on .internal, a TLD reserved for private use that does not exist on the public internet. During the migration window, a query on resolver one missed the forwarder path, recursed out to the public root, and came back NXDOMAIN carrying the root's own SOA: a definitive does-not-exist, cached for 86,400 seconds. A definitive negative kills failover. Clients that asked resolver one gave up on the first answer.
That asymmetry is what made the incident look random. Resolver two carries the same flaw but had never cached a negative for this particular name, so it forwarded to the router and served the correct address. One box jailed the name for a day, the other answered it fine. Clients configured with both resolvers got a coin flip.
The reboot test proved how stubborn the jail was. With the operator's approval, resolver one was rebooted, and the poison survived: Technitium writes its cache to disk on a clean shutdown and reloads it at boot. Worse, the router itself mints fresh root-grade negatives on every miss, so the factory kept running either way.
The fix
Five seconds, one button: Clear Cache on the resolver's dashboard. The router already held the right answer, so the name resolved the moment the jailed entry died. No proxy restart, nothing else touched.
The durable fixes came next, in order of value. On the router, a UniFi UDM, declare the private domain local with a one-line dnsmasq addition delivered by an onboot script, because UniFi regenerates its DNS config and hides domain-local behavior from the UI. Misses then die at the router instead of minting twenty-four-hour poison. The conditional forwarder zones that pin the private name to the router turned out to need nothing: they had been configured correctly on migration day.
The agent could not push the button itself. The Technitium API wants the admin token, and secrets stay out of the model's hands. After the fix, the operator attached a read-only token to the API connectors, which gave future sessions live eyes into the zone tables, and no hands.
The result
End to end, green on the first probe after the clear: client to Caddy to the internal name to the app, which answered its version JSON through the full chain. Root cause was isolated roughly twenty-five minutes in, through two false summits, with a total session of one hour and eleven minutes.
Two honest notes went on the record. The agent's early worry about a stale shadow zone was wrong: the API showed the forwarder zones had existed since the migration. The whole incident was one cached negative, nothing structural. And the second resolver got a cache clear too, since the same fresh-probe test had left a harmless but pointless root negative sitting in it.
Incident timeline
Command log
Signature commands
dig +short gotify.example.ca @10.10.1.3
dig @10.10.1.1 gotify.example.internal A +noall +comments +answer
dig @10.10.1.2 zzz-fresh-probe-7319.example.internal A +noall +comments +authority
curl -sS --resolve gotify.example.ca:443:10.10.1.50 https://gotify.example.ca/version
| Time | Agent | Command | Exit |
|---|---|---|---|
| 20:41:40 | SCAIRE |
dig +short gotify.example.ca @10.10.1.2
resolver one: the reverse proxy address
|
exit 0 |
| 20:41:41 | SCAIRE |
dig +short gotify.example.ca @10.10.1.3
resolver two: a public IP nobody in the lab owns
|
exit 0 |
| 20:42:29 | SCAIRE |
dig example.ca AXFR @10.10.1.3
transfer refused: the two boxes drift independently
|
exit 0 |
| 20:47:40 | SCAIRE |
dig @10.10.1.2 gotify.example.internal
NOERROR with zero records, root SOA in authority
|
exit 0 |
| 21:01:53 | SCAIRE |
pct exec 105 -- grep -ac gotify /etc/dns/zones/example.internal.zone
zero matches: the record never lived in any zone
|
exit 1 |
| 21:03:02 | SCAIRE |
dig @10.10.1.1 gotify.example.internal A +noall +comments +answer
the router: authoritative answer, TTL 0. DHCP registration works
|
exit 0 |
| 21:03:03 | SCAIRE |
dig @10.10.1.2 gotify.example.internal SOA +noall +answer +authority
resolver one backs its NODATA with the root servers SOA
|
exit 0 |
| 21:05:02 | SCAIRE |
dig @10.10.1.2 zzz-fresh-probe-7319.example.internal A +noall +comments +authority
never-queried name: NXDOMAIN plus root SOA. The fallback is live
|
exit 0 |
| 21:10:09 | SCAIRE |
pct reboot 105
approval gate: risk stated, operator signed off
|
exit 0 |
| 21:10:32 | SCAIRE |
curl --resolve gotify.example.ca:443:10.10.1.50 -sS -o /dev/null -w "HTTP %{http_code} / %{http_version}" https://gotify.example.ca/
after reboot: still 502. The poison survived
|
exit 0 |
| 21:12:24 | SCAIRE |
dig @10.10.1.2 gotify.example.internal +noall +answer +comments
still NODATA, root SOA, about 15 hours of cache left
|
exit 0 |
| 21:13:43 | SCAIRE |
dig @10.10.1.1 freshprobe-58231.example.internal +noall +authority +comments
the router holds a fresh root negative: it is the poison factory
|
exit 0 |
| 21:49:49 | SCAIRE |
Technitium DNS1: GET /zones/list
read-only token: forwarder zones existed since migration day
|
exit 0 |
| 21:52:03 | SCAIRE |
curl -sS --resolve gotify.example.ca:443:10.10.1.50 https://gotify.example.ca/version
end to end: the app answers its version JSON
|
exit 0 |