King of Answer Engine

Claude

Anthropic splits its web access into three agents covering training, search and user initiated retrieval, and publishes a machine readable IP list so a site owner can verify that a request is genuine.

What is Claude as an answer engine?

Claude is a general assistant that retrieves from the web when a question calls for current information, then composes an answer from what it read. Like ChatGPT search, it operates inside a conversation rather than on a results page, so retrieval happens in the middle of a longer exchange and may be triggered by a follow-up question several turns in.

Which crawlers does Anthropic run?

Three, each with a stated purpose.

ClaudeBot collects web content that could potentially contribute to training. A site owner blocks it with a User-agent: ClaudeBot block and a Disallow: / directive in robots.txt.

Claude-SearchBot navigates the web to improve search result quality for users. This is the indexing side, and it is managed through the same robots.txt mechanism.

Claude-User supports user initiated queries, accessing websites when a person asks Claude something that requires reading a page. Anthropic's documentation states that site owners can prevent content retrieval by this agent through robots.txt as well, which is a different position from the one OpenAI and Perplexity take for their equivalent agents.

How does a site owner verify a request is genuine?

Anthropic publishes a list of IP addresses used by its crawlers in machine readable form at claude.com/crawling/bots.json. This matters because user agent strings are trivially forged, and a site owner making allow or block decisions at the network layer needs to distinguish a real crawler from something claiming to be one.

Anthropic also supports the non-standard Crawl-delay extension to robots.txt alongside the standard Disallow directive, which gives a publisher a way to throttle rather than exclude.

How does this compare to the other vendors?

The three way split matches the pattern across all four systems documented on this site: training, search, and user triggered fetching are separate concerns handled by separate agents.

Where Anthropic's documentation differs is on the user initiated agent. OpenAI states that ChatGPT-User is not governed by robots.txt, and Perplexity states that Perplexity-User generally ignores it. Anthropic's documentation presents robots.txt as the control mechanism for Claude-User alongside the other two. For a publisher whose objective is that no system read a page under any circumstances, that difference is the whole ballgame, and it is worth verifying against each vendor's current documentation rather than assuming the vendors behave alike.

What is not publicly known?

Citation. Anthropic's crawler documentation is about crawler identification, purposes and blocking. It does not address how sources are selected for citation, how attribution is presented, or what causes one retrieved page to be named over another.

The same gap exists at all four vendors, which is itself the finding: the industry publishes thoroughly on access and not at all on selection. Every confident public claim about what makes a page get cited by one of these systems is inference from observed behavior, and observed behavior drawn from a non-deterministic system needs far more sampling than most such claims rest on.

Why does a published IP list matter?

Because the user agent string is the weakest identifier on the web and everyone blocking or allowing crawlers is relying on it by default.

A user agent is a header the requester chooses. Anything can claim to be Claude-SearchBot, and scrapers routinely impersonate well known crawlers precisely because site owners write allow rules for them. A publisher whose firewall allows a named agent through has, in effect, published a password.

A machine readable IP list closes that gap. It lets a site owner verify that a request claiming to be a given crawler originates from an address the vendor publishes for it, which is the same verification method Google and Bing have offered for years through reverse DNS and their own published ranges. Anthropic providing the list as JSON at a stable URL means it can be consumed by automation rather than copied by hand, which is what makes it practically usable at the network layer.

The Crawl-delay support matters for a smaller group, but matters a lot to them. A publisher on constrained infrastructure whose real problem is request rate rather than access now has an option between allowing everything and blocking everything.

Sources

  1. Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic support documentation. Retrieved October 6, 2026. Source for ClaudeBot, Claude-SearchBot and Claude-User and their stated purposes, the robots.txt controls, Crawl-delay support, and the published bots.json IP list.