Site structure and sitemap checks

Triggers and fixes for orphan and weakly linked pages, dead ends, and sitemap entries that redirect, error, or are non-canonical or noindex.

Updated 8 October 2026View as Markdown

These checks look at the whole site at once: which pages link to which, and what your XML sitemap lists. None of them is recheckable, because fixing one page changes counts on others. Run a full audit after the fix.

  • Sitemap: SEOFix reads /sitemap.xml at the root of the audited site, following a sitemap index to its child sitemaps (gzip supported), up to 50 sitemap files and 100 MiB of XML. Sitemaps declared only in robots.txt under another path are not read. If any child sitemap fails or a limit is hit, the sitemap counts as incomplete, and "not in the sitemap" checks are skipped.
  • Links: internal <a href> links from every crawled page. Pages disallowed by robots.txt are not crawled, so links on them are not counted.
  • Incomplete crawls: when the audit did not see the whole site (it hit the page cap, pages were firewall-blocked, sections were skipped, or a preview hit its time limit), incoming-link counts can't be trusted. ORPHAN_PAGE, ONE_INCOMING_LINK and REDIRECT_NO_INCOMING_LINKS are then not reported.

Indexable below means: no noindex in the robots meta tag, and no canonical pointing to another URL.

  • Severity: notice. Recheckable: no (needs a full audit).
  • Trigger: the URL is listed in the sitemap, answered 200, is not the start page, and no crawled page links to it. Any internal link counts, including nofollow ones. Not reported when the crawl was incomplete.
  • Why it matters: pages only reachable through the sitemap are hard for crawlers to discover and get no link equity, so they rarely rank.
  • How to fix: link to the page from at least one relevant page: navigation, a category or hub page, breadcrumbs or related-content links. If the page is obsolete, remove it from the sitemap and redirect or retire it.
  • Severity: notice. Recheckable: no (needs a full audit).
  • Trigger: an indexable 200 HTML page is linked from exactly one other crawled page with a followed link (no rel="nofollow"). Links from a page to itself are not counted. Not reported when the crawl was incomplete.
  • Details: linked_from, the one linking page.
  • Why it matters: a page with a single internal link gets little link equity and is easy for crawlers to miss.
  • How to fix: link to it from more relevant pages: hub or category pages, related links, breadcrumbs, or "more like this" blocks in the template.
  • Severity: warning. Recheckable: no (needs a full audit).
  • Trigger: an indexable 200 HTML page has no internal links to any other URL (nofollow links count as links).
  • Why it matters: a dead-end page passes no link equity on and gives visitors and crawlers nowhere to go.
  • How to fix: add links from the page to related pages on the site: breadcrumbs, navigation, related content. If the page renders its navigation with JavaScript only, render the links in the HTML.
  • Severity: notice. Recheckable: no (needs a full audit).
  • Trigger: a crawled URL answered 3xx, is not the start URL, is not in the sitemap, is not the target of another redirect, and no crawled page links to it. Not reported when the crawl was incomplete.
  • Details: redirect_to.
  • Why it matters: a redirect nothing links to is usually a leftover URL.
  • How to fix: nothing on your site needs it. Keep the redirect if external sites or old bookmarks may still use the URL; otherwise you can drop it. Make sure no sitemap or feed lists it.

XML sitemap

The sitemap checks run only when the site has a sitemap that SEOFix read.

INDEXABLE_NOT_IN_SITEMAP — Indexable page not in sitemap

  • Severity: notice. Recheckable: no (needs a full audit).
  • Trigger: an indexable 200 HTML page is not listed in the sitemap. Only reported when the whole sitemap was read.
  • Why it matters: the sitemap tells search engines which pages matter; indexable pages missing from it are discovered later or not at all.
  • How to fix: add the page to the sitemap. Generated sitemaps often miss a content type or paginated pages; fix the generator. If the page should not rank, add noindex or a canonical to the preferred URL instead.

SITEMAP_URL_NOT_200 — Sitemap URL is not a 200 page

  • Severity: warning. Recheckable: no (needs a full audit).
  • Trigger: a URL listed in the sitemap answered 3xx, 4xx or 5xx (firewall-blocked URLs excluded).
  • Details: status_code, redirect_to for redirects.
  • Why it matters: redirects and errors in the sitemap waste crawl budget and make search engines trust the sitemap less.
  • How to fix: list only final 200 URLs. Replace each redirect with its target (redirect_to) and remove broken URLs. Check that the generator builds URLs exactly as the server serves them (scheme, www, trailing slash).
<!-- before: redirects to the trailing-slash version -->
<url><loc>https://example.com/jobs</loc></url>
<!-- after -->
<url><loc>https://example.com/jobs/</loc></url>

SITEMAP_URL_NOT_CANONICAL — Non-canonical page in sitemap

  • Severity: warning. Recheckable: no (needs a full audit).
  • Trigger: a URL listed in the sitemap answered 200 but its canonical points to a different URL (and it is not noindex; that case is NOINDEX_IN_SITEMAP).
  • Details: canonical_url.
  • Why it matters: listing a URL that says another URL is the real one sends contradictory signals.
  • How to fix: list the canonical URL instead, or fix the canonical if the listed URL is the one that should rank.

NOINDEX_IN_SITEMAP — Noindex page listed in sitemap

  • Severity: warning. Recheckable: no (needs a full audit).
  • Trigger: a URL listed in the sitemap has noindex in its robots meta tag.
  • Details: meta_robots.
  • Why it matters: the sitemap asks search engines to index a page that asks not to be indexed.
  • How to fix: remove noindex pages from the sitemap, or drop the noindex if the page should rank.

More in Issue reference

Still stuck? Email [email protected] with your site and what you expected to see.