Skip to main content

AI crawler

A bot that fetches web pages for an AI system — to train a model, to build a retrieval index, or live while answering a question.

Updated

An AI crawler is a bot that fetches web pages on behalf of an AI system rather than for a search results page.

Nine of them decide most of what today's assistants can read: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, CCBot and meta-externalagent.

Do they all do the same job?

No — they fall into three jobs, and the distinction matters because blocking one does not block the others:

  • Training crawlers collect pages that may end up in a model's weights: GPTBot, ClaudeBot, CCBot and meta-externalagent.
  • Retrieval crawlers build the index an assistant searches while answering: OAI-SearchBot and PerplexityBot. Google-Extended is the control surface for Google's AI surfaces rather than a separate fetcher.
  • User-triggered fetchers pull a single page live because somebody in a chat asked about it: ChatGPT-User and Perplexity-User.

How is access controlled?

In robots.txt, and a group that names a crawler beats the wildcard group — which is how real crawlers resolve it. So a site can refuse scrapers in general and still admit the engines it cares about.

Most sites that block them did not decide to. A security plugin, a copied robots.txt or a well-meaning ops ticket did it, and nobody has read the file since. Excluding an engine on purpose is a legitimate business decision; the failure worth finding is the one nobody in the company knows about.

One question

Find out what ChatGPT says about you.

Same seven checks, same ten seconds, still no signup.

Free forever. The score is never gated.