Home / Knowledge / Technical SEO and Indexing / The No Orphan Pages Rule

Technical SEO and Indexing

The No Orphan Pages Rule

An orphan page is one that no other page on your site links to. Search engines find pages mainly by following links, so an orphan is hard to discover, hard to recrawl, and starved of the internal authority links pass. The rule is simple: every indexable page must have at least one descriptive internal link pointing to it.

An orphan page is a page that exists but that nothing links to. You can reach it if you know the URL, and it might sit in your sitemap, but no other page on the site points at it. On a small site orphans are an accident. On a large directory they are an epidemic waiting to happen, because every automated import, every retired category, and every template change can quietly cut pages loose. At Kings Hospitality Group the no orphan rule is one of our firmest, and this is why.

How search engines actually find pages

Crawlers discover the web by following links. They land on a page, extract its links, queue those, and repeat. Sitemaps supplement this, but the link graph is the primary mechanism and the one that carries meaning. A page with no inbound internal links is a dead end approached from nowhere. The crawler has no path to it, no recurring reason to revisit it, and no context for what it is about.

That last point matters as much as discovery. Internal links do not just help a page get found, they tell the engine what the page is and how important the rest of the site thinks it is. An orphan receives none of that. Even if it gets indexed via the sitemap, it sits there isolated, passing and receiving no internal authority, competing with one hand tied. Links are the circulatory system of a site, and an orphan is a limb with no blood supply.

Why directories breed orphans

Directory sites are especially prone to orphans because so much of the page creation is programmatic. A bulk listing import creates ten thousand pages, but the linking that should connect them is often an afterthought. A category gets deprecated and its links removed, stranding every listing that was only reachable through it. A seasonal landing page is unpublished from navigation but the page itself lingers. Each of these is a routine operation, and each can orphan pages in bulk if no one is checking.

This is exactly why we pair the no orphan rule with a deliberately flat architecture. When the path from the homepage to any page is short and the linking is systematic rather than incidental, orphans have fewer places to hide. The structure does much of the enforcement for you.

Finding orphans before they hurt you

You cannot fix what you cannot see, and orphans are by definition invisible to a normal crawl, because a crawl follows links and an orphan has none. The way to find them is to crawl the site and then compare three lists: the URLs your crawler reached by following links, the URLs in your sitemap, and the full list of URLs from your database or server logs. Anything that appears in the sitemap or logs but never in the crawl's link graph is an orphan.

  • Crawl from the homepage following internal links only, and record every URL reached.
  • Export the full known URL set from your content system or sitemap.
  • Diff the two. URLs in the known set but missing from the crawl are your orphans.
  • Check server logs for pages receiving crawler hits that have no internal links, another reliable orphan tell.

We run this comparison before any site launches and on a schedule afterward. It is the same instinct as submitting section split sitemaps and watching coverage per section: make the invisible measurable, then act on it.

Fixing and preventing orphans

Fixing an existing orphan is straightforward once found: give it at least one genuine, contextual link from a related page. Not a token link buried in a footer, but a descriptive in body or navigational link from a page that actually relates to it. A stranded listing should be reachable from its category, its location, and ideally a related listing or two. A stranded guide should be linked from its cluster pillar and its siblings.

Prevention is the better game, and it is structural. The most reliable defence is to make linking a property of your templates rather than a manual task. Every listing template automatically links to its category, its location, and related listings. Every article template links up to its pillar and across to siblings. When linking is generated by the same system that generates the page, an import cannot create orphans because the links are born with the page.

The difference between reachable and well linked

A page reachable only through a thousand link mega footer is technically not an orphan, but it is barely connected. The strength of an internal link depends on its context and relevance. A link from a closely related page in body copy is worth far more than a generic site wide footer link. So the rule we actually hold to is not merely no orphans, it is that every page has at least one relevant, contextual inbound link. That higher bar is what turns a connected site into a strong one, and it is the same standard we apply to the knowledge hub, where every article carries descriptive links up to its pillar and across to its siblings.

Orphans and crawl budget

On large sites there is a second cost. Crawlers allocate a finite amount of attention to each site. When that attention is spent rediscovering and recrawling poorly connected pages, or when orphans force the engine to lean entirely on the sitemap, the budget is used inefficiently. A clean link graph helps crawlers spend their time on the pages that matter, which is part of why we treat internal linking as infrastructure rather than decoration. The broader case for that mindset runs through our building thesis.

The rule, stated plainly

No indexable page ships without at least one descriptive, contextual internal link pointing to it. We verify it before launch by crawling the site against its own known URL set, and we keep verifying it as the site grows. It is not glamorous work. It is the kind of quiet discipline that does not announce itself but quietly compounds, because a site where every page is genuinely connected discovers faster, recrawls more reliably, and distributes its hard won authority where it belongs.

Orphans are entropy. On a directory that publishes constantly, entropy is the default, and the no orphan rule is simply the habit of pushing back against it on a schedule. Build the linking into your templates, audit the gap between what you crawl and what you know, and fix the strays before the search engines have to wonder why so many of your pages sit alone.

Kings Hospitality Group framework

Kings Hospitality Group runs every site against a no orphan rule: no indexable page ships without at least one contextual internal link pointing to it, verified by crawling the site against its own sitemap before launch.

Common questions

How do I find orphan pages?

Crawl your site, then compare the crawled URLs against your sitemap and your full URL list. Pages that appear in the sitemap or server logs but never in the crawl's internal link graph are orphans.

Is a sitemap entry enough?

No. A sitemap helps discovery but passes no internal authority and is a weak signal on its own. A page needs real contextual links from related pages, not just a sitemap listing.

Can navigation links solve it?

They help, but a page reachable only through a giant footer or mega menu is still weakly connected. The strongest signal is a relevant in body link from a related page.

Subscribe to The Portfolio Brief

Get our field notes on building directory and hospitality brands that last. A few considered letters a year.

MA
Morten Andersen
Founder, Kings Hospitality Group
More from this author