How the crawler works

How SEOFixBot finds and fetches your pages, how fast it goes and when it slows down, what it fetches, and what counts as a page and a credit.

Updated 8 October 2026View as Markdown

SEOFix starts at your site's address, reads /sitemap.xml, and follows the internal links on every HTML page, at 2 requests per second with at most 2 requests in flight. It honours robots.txt, slows down when your site slows down or returns errors, and stops at the audit's page limit. Every URL it fetches counts as one page and costs one credit.

Identity

Requests come from the user agent SEOFixBot/1.0 (+https://seofix.ai/bot). Each page request carries a Referer header with the page the link was found on, so you can trace the crawl in your logs. For robots.txt rules and firewalls, see /help/seofixbot and /help/firewall-allowlisting.

Speed and politeness

Rule Value
Default speed 2 requests per second per host
Requests in flight 2 (the speed is a hard ceiling regardless)
Fastest without a verified site 3 requests per second
Fastest with a verified site 10 requests per second
Slowest after backing off 1 request every 10 seconds
Request timeout 15 seconds
Largest response read 5 MB (the rest is cut off)

Latency backoff. SEOFix learns your site's normal response time from its first 20 successful (2xx) responses. If responses then get more than twice as slow as that and at least 300 ms slower, it halves its request rate, at most once every 10 responses. It speeds back up gradually once responses are back under 1.5 times the normal time. If your site stays slow for 30 responses at the slowest rates, SEOFix takes that as the new normal and returns to the configured speed.

Error backoff. A network error, a 429 Too Many Requests or any 5xx answer at least halves the request rate at once. A page that failed this way is retried once.

The report's stats show what happened: max_rps (the configured ceiling), effective_rps (pages per second actually crawled), slowdowns (how often SEOFix slowed down for your site), and speed_clamped with speed_clamp_reason when a speed above 3 was lowered because the site is not verified.

robots.txt

  • SEOFix reads /robots.txt for each host and scheme it crawls and applies the rules for SEOFixBot, or the * group when there is none.
  • A disallowed URL is not fetched. It is listed as a ROBOTS_BLOCKED notice and does not count as a page.
  • If robots.txt is missing or does not answer 200, everything is allowed.
  • Crawl-delay is not read; SEOFix uses its own pacing above.

How URLs are found

  1. Start URL. For a site, its address (for example https://example.com/). If the start URL redirects (for example apex to www, or http to https), SEOFix follows up to 5 redirects on the same site to find where it lands.
  2. Sitemap. SEOFix fetches /sitemap.xml under the start URL and the sitemap index files it lists: up to 50 files, 100,000 URLs, 50 MB per file and 100 MB in total (measured after decompressing gzip). Sitemap locations listed only in robots.txt are not read. Sitemap URLs enter the crawl at depth 1.
  3. Links. Every <a href> on a crawled HTML page that points to the same site is queued, including rel="nofollow" links. Links to /cdn-cgi/ paths are ignored.
  4. Redirects. A 3xx answer is recorded as a page. If its target is on the same site, the target is queued as a new URL.

URLs are crawled shallowest first, up to 25 links deep.

Same site means the start host and its www or apex twin: for example.com, both example.com and www.example.com. Other subdomains such as blog.example.com are treated as external.

URL normalisation. Before a URL is queued, SEOFix resolves it against the page, lowercases the scheme and host, drops a default port (:80, :443) and the #fragment, and turns an empty path into /. The query string and the path's case and trailing slash are kept, so /page and /page/ are two URLs. javascript:, mailto:, tel: and data: links are skipped.

What is fetched

  • Pages: every queued URL is requested with GET. Only 200 responses with a text/html content type are parsed for titles, headings, links and the other on-page checks. Other files linked with <a href>, such as PDFs, are fetched and recorded with their status, but not parsed.
  • Not fetched: images, stylesheets, scripts and fonts. The crawler reads their addresses from the HTML (for example for the mixed-content and image alt checks) but does not download them.
  • After the crawl, on every audit: /llms.txt, and the robots.txt rules for AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot).
  • After the crawl, on full audits only (not previews):
    • external links: up to 500 distinct external URLs, one at a time per host, at least 0.5 seconds apart on the same host;
    • the optional samples described in /help/javascript-rendering.

When the crawl stops

  • The page limit is reached (the report is then partial).
  • No URLs are left to crawl.
  • The first 50 pages were all blocked by a firewall. A URL section (such as /jobs) whose first 50 pages were all blocked is skipped for the rest of the crawl.
  • Someone cancels the audit.

Page limits

Audit Page limit
Preview (no account) 50
Free audit 500
Free plan 500
Starter plan 10,000 to 500,000 per audit, by page tier
Growth plan 250,000 to 500,000 per audit, by page tier
Any audit above 500 pages Needs a verified site

A site's own Page limit (see /help/crawl-settings) applies within the plan's limit.

What counts as a page and a credit

A page is one URL the crawler fetched: any status code (200, 3xx redirect, 404, 500), any content type, firewall-blocked or not, and URLs that failed with a network error. These do not count:

  • URLs disallowed by robots.txt;
  • URLs skipped because their section was blocked by a firewall;
  • robots.txt, sitemap and llms.txt requests;
  • external link checks and the optional samples.

Credits: 1 credit per page. A page that answers 304 Not Modified to an incremental request costs 0.2 credit, rounded up once per audit (see /help/comparing-audits). When an audit starts, credits equal to its page limit are reserved; when it ends, whatever it did not use goes back to your balance. Previews and the free audit use no credits. See /help/plans-and-credits.

Start an audit as an agent

curl -X POST https://api.seofix.ai/v1/sites/42/crawls \
  -H "Authorization: Bearer $SEOFIX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"max_pages": 2000}'
{ "crawl_id": 1843, "estimated_credits": 2000, "max_rps": 2 }

MCP: start_site_audit for a registered site (results are compared with its previous audit), or start_audit with a url. Then poll get_audit_status (GET /v1/crawls/{id}).

  • /help/seofixbot
  • /help/crawl-settings
  • /help/javascript-rendering
  • /help/audit-statuses
  • /help/plans-and-credits

More in Audits & crawling

Still stuck? Email [email protected] with your site and what you expected to see.