AI crawlers
Automated bots run by AI companies that fetch web pages, either to help answer a user's question in real time or to gather text for training future models. GPTBot, ClaudeBot, PerplexityBot, and Google-Extended are examples.
AI crawlers are automated programs, run by companies that build AI systems, that fetch web pages so those systems can use the content. They work much like the crawlers search engines have always used, but they feed AI products rather than a search index. Named examples include GPTBot from OpenAI, ClaudeBot from Anthropic, PerplexityBot from Perplexity, and Google-Extended, a control Google offers for its AI features. Each identifies itself by a user-agent name when it requests a page.
Retrieval versus training
The most important distinction is why a crawler is fetching the page, because two different purposes hide behind the same activity.
Retrieval means a system fetches a page at the moment a user asks something, reads it, and uses it to compose an answer there and then. The page is being consulted as a live source. Training means text is collected in bulk and used to help build or update a model, so the content contributes to what the model knows in general rather than being quoted in one specific answer.
A single company may run separate crawlers for each purpose, and a site can decide to allow one and not the other. That is why the two ideas are worth keeping apart rather than treating all AI crawling as one thing.
How a site controls them
The main control surface is robots.txt, the long-standing file that tells well-behaved crawlers which parts of a site they may request. Because most named AI crawlers publish their user-agent names and state that they honour robots.txt, a site can allow or disallow each one by name. This depends on the crawler choosing to comply; the file expresses a preference rather than enforcing it.
A newer and still emerging convention is llms.txt, a file that points AI systems to the content a site considers most useful. It is a map rather than a gate, and it is not yet a settled standard.
Where it connects
Whether a crawler can make good use of what it fetches depends partly on how the content is stored; see structured content. For the broader practice of optimising for AI-generated answers, see GEO / AEO.
We're booking content platform
engagements for 2026.
Twenty-five minutes to walk through the work and decide if we're the right team for it. Scoping and a fixed price come after.