Link Decay on Reference Pages: Eight Failure Modes and the Checks That Catch Them
Eight documented failure modes in reference-page links, and the verification each one gets past.
A reference page is only as good as the places it sends you. A municipal energy guide, a university cost handbook, a campus security syllabus or a curated resource list exists to point elsewhere, and its value equals the value of its destinations. When a destination disappears, the page keeps its layout, its headings and its confident anchor text, and delivers nothing to the reader who follows the link.
Direct observation of reference pages and their outbound targets turned up eight distinct failure modes. Several of them get past the check most maintainers run, which is a request against the linked domain instead of the linked route. Each finding below is stated with the identifiers observed at the time of testing.
A federal agency removed an entire section
energy.gov/energysaver/* returns 404. The condition was confirmed on five routes,
among them /energysaver/air-sealing-your-home and /energysaver/energy-saver.
The home page of energy.gov responds 200.
The shape of that result is the problem. A verifier that resolves and requests only the
registered domain records energy.gov as healthy, reports zero issues, and leaves
every link into the Energy Saver section broken. A second property makes the miss worse: the
error pages keep the generic <title> "Department of Energy". Verifiers
that compare titles, or that accept a plausible institutional title as evidence of a live
document, classify these 404 responses as live pages. Two independent heuristics fail on the
same set of URLs, which is why removals like this stay on reference pages for a long time.
A national laboratory left DNS
nrel.gov, www.nrel.gov, pvwatts.nrel.gov, and
rredc.nrel.gov return NOERROR with no A record. The result was reproduced against
three independent public resolvers: 8.8.8.8, 1.1.1.1, and 9.9.9.9.
NOERROR without an A record differs from NXDOMAIN in a way that matters in practice. The
domain exists in the .gov zone; no name under it carries an address. Tools that
test for NXDOMAIN alone, or that treat any non-error DNS response as success, report the
domain as present. A control query confirms the failure is specific to the target and isn't
an artifact of the test environment: afdc.energy.gov, another subdomain run by
the same department, resolves normally.
The consequence for reference pages is broad. PVWatts was the standard solar calculation reference cited by municipal and university pages, so every citation of that host now ends in a resolution failure instead of an HTTP status a checker can read. What that result does not establish is whether the calculator itself is gone. The next section is about that distinction, because it is the one that most often sends a maintainer to the wrong fix.
A dead host is not a dead resource
The PVWatts case carries a second lesson that the DNS result hides, and it's the one most
likely to make a maintainer leave the page worse than they found it. The laboratory was
renamed. What was the National Renewable Energy Laboratory now operates as the National
Laboratory of the Rockies at nlr.gov, and PVWatts is live at
pvwatts.nlr.gov, serving the same calculator. The old zone was withdrawn from
DNS. The resource stayed on the web.
No DNS query can tell those two apart. A resolver reports that a host has no address, which is true and complete as far as it goes, and says nothing about whether the thing that lived there still exists somewhere else. A maintainer who reads "dead" and deletes the citation, or swaps in a different tool, has acted on a correct observation and reached the wrong remedy. The check that closes this gap needs no tooling: search for the resource by name before you conclude it's gone.
This failure mode is more common than outright disappearance, because institutions rebrand, merge and consolidate far more often than they shut down. It's also the one automated link checkers can never resolve by themselves, since the successor host shares no string with the one that failed.
Anchor text outlives its destination
The energy efficiency page of the City of Bainbridge Island, Washington, carries a link with
the anchor text "Caulking and weatherstripping" pointing to
energy.gov/energysaver/air-sealing-your-home, which returns 404.
The anchor text is still an accurate description of a subject that no longer has a destination. A person reviewing the page sees nothing wrong, because the visible words are right. The defect stays invisible until someone follows the link. That is why editorial review alone doesn't catch link decay: the signal a reader uses to judge relevance survives the loss of the resource.
An expired domain acquired a new owner
The construction cost guide published by the University of Missouri lists an entry labeled
"Get-A-Quote" pointing to get-a-quote.net. The domain changed hands. It now serves
an online casino in Russian, with the page <title> "Пинко Казино". Confirmed
by direct access on August 17, 2026.
This case returns HTTP 200. Every status-code check passes. Every DNS check passes. In practice the link is worse than a 404, because an institutional page now sends readers to unrelated commercial content under an anchor that implies a vetted construction resource. Catching it means comparing the retrieved page against what the anchor text or the surrounding context leads you to expect, and no status-based tool does that.
A publishing platform disappeared as a block
statistics.about.com, math.about.com, psychology.about.com,
sociology.about.com, and canadaonline.about.com do not resolve.
Educational pages that cited these hosts lost their references in a single event instead of
one at a time.
Platform-level shutdowns produce correlated failures. A page citing four About.com subdomains lost four links at once, so the decay rate observed on a page is not a smooth function of age. A maintenance schedule built on the assumption of gradual attrition underestimates the damage one platform retirement does.
Crawler identity changes the response
app.dimensions.ai returns 404 to an identified crawler and 202 to a browser
user-agent. In one batch of 26 URLs, 6 were false positives caused by this behavior.
That's roughly a quarter of the batch. Any dead-link survey that skips rechecking failures with a browser user-agent overstates the problem, and by enough to invalidate a report. It follows that decay figures produced by unattended crawlers should be treated as upper bounds until each failure has survived a second request under a different identity.
Template links multiply the count
owasp.org/index.php/Cross-Site_Request_Forgery_(CSRF) returns 404 because OWASP
migrated away from MediaWiki. The same URL appeared in six distinct guides on a single campus,
because it sat in a shared page element.
A naive count records six problems. One URL needs one decision and one edit. Counting by occurrence instead of by unique URL inflates the reported severity by a factor equal to the template's reach, and it points the repair work at pages when the fix belongs in the shared component.
Failure modes and detection
| Failure mode | Observed instance | Why a naive check misses it | Detection method |
|---|---|---|---|
| Section removed, domain alive | energy.gov/energysaver/* 404, root 200 | Check targets the domain, not the route | Request every distinct path; never infer path health from root health |
| Generic title on error page | 404 pages titled "Department of Energy" | Title comparison accepts an institutional string | Treat HTTP status as authoritative; ignore title as a liveness signal |
| NOERROR with no A record | nrel.gov, pvwatts.nrel.gov, and two others | Test looks only for NXDOMAIN | Query A records across at least three resolvers; treat empty answer as failure |
| Anchor survives destination | "Caulking and weatherstripping" on the Bainbridge Island page | Visible text still reads correctly | Resolve every link; editorial review is insufficient |
| Domain changed owner | get-a-quote.net serving "Пинко Казино" | Returns 200 and passes all status checks | Compare retrieved title and content against anchor text and page context |
| Platform-wide shutdown | Five about.com subdomains unresolvable | Failures assumed independent and gradual | Group failures by registrable domain to spot correlated loss |
| Template repetition | One OWASP URL in six campus guides | Occurrences counted instead of URLs | Deduplicate by normalized URL before counting; fix the shared component |
| Crawler-specific rejection | app.dimensions.ai: 404 to crawler, 202 to browser | Single request under one identity | Recheck every failure with a browser user-agent before reporting |
Method
Each URL was requested directly over HTTP and the returned status code recorded. Path-level requests were issued individually; no path status was inferred from any other path or from the domain root. Response titles were inspected separately from status codes so the two signals could be compared instead of conflated.
DNS behavior was tested by querying A records against 8.8.8.8, 1.1.1.1, and 9.9.9.9, and by
distinguishing NXDOMAIN from NOERROR with an empty answer section. A control host under the
same parent organization, afdc.energy.gov, was queried to confirm that a negative
result reflected the target and not the test path.
Every failure was rechecked with a browser user-agent before classification. The 26-URL batch that produced 6 false positives showed that this step is mandatory. URLs were normalized and deduplicated before any count was produced, which reduced the six OWASP occurrences to a single unique defect.
Limitations
The observations record a single point in time. There is no longitudinal series here, so no decay rate per year can be derived from this data, and none is claimed. The total number of reference pages examined and the total number of outbound links inspected are not reported here.
The sample was not drawn at random. Pages were selected because they are reference-heavy, which biases the set toward domains with many outbound institutional citations. Nothing here supports generalizing the proportions observed in the 26-URL batch to any larger population.
Removal dates are unknown. The dates on which the Energy Saver section was withdrawn, the NREL
zone lost its A records, and get-a-quote.net changed hands are not established
here. Whether the NREL condition is permanent or a passing operational state can't be
determined from three resolver queries at one moment.
The content served at get-a-quote.net may change again. The finding records what
that host returned on August 17, 2026 and makes no claim about its state before or after.
Conclusion
If you maintain a reference page, check by route, not by domain. A 200 at the root says nothing about a 404 five path segments down. Recheck every failure with a browser user-agent before you act on it; one batch produced a 6-in-26 false positive rate without that step. Separate NXDOMAIN from NOERROR with no A record, because the second is a live zone with no reachable host, and it passes tests written for the first. Deduplicate by URL before you count, since one bad link in a shared template shows up as six. And run at least one content-level comparison, because an expired domain under new ownership returns 200 while serving a casino.
Frequently asked questions
Why does requesting the home page of a linked domain fail to detect broken links?
Because removal happens at the section level. energy.gov responds 200 while energy.gov/energysaver/* returns 404 on all five routes tested. A domain-level check reports the site as healthy and leaves every path-level citation broken.
What distinguishes NXDOMAIN from a NOERROR response with no A record?
NXDOMAIN means the name does not exist. NOERROR with an empty answer means the name exists in the zone and carries no address. nrel.gov, www.nrel.gov, pvwatts.nrel.gov, and rredc.nrel.gov produce the second condition across 8.8.8.8, 1.1.1.1, and 9.9.9.9. A checker that tests only for NXDOMAIN reports these hosts as present. A missing host still doesn't prove a missing resource: the laboratory was renamed and PVWatts is live at pvwatts.nlr.gov. DNS answers where a name points. It never says whether the thing that lived there still exists.
Why does a URL return 404 to a checker and load normally in a browser?
Some hosts serve different responses by user-agent. app.dimensions.ai returns 404 to an identified crawler and 202 to a browser. In a batch of 26 URLs, 6 failures had this cause. Recheck with a browser user-agent before reporting any failure.
Does accurate anchor text indicate a working link?
No. The Bainbridge Island energy page carries the anchor "Caulking and weatherstripping" pointing to a URL that returns 404. The text describes the subject correctly and the destination does not exist. Only resolving the target tells you whether the link is alive.
How should the same broken URL appearing on many pages be counted?
By unique URL. The OWASP CSRF address appeared in six campus guides through a shared template. Counting occurrences records six defects where one edit to the shared component fixes all of them.