Perplexity
Perplexity was built around cited answers rather than adding them later, and its crawler documentation reflects that: the indexing agent is stated to exist for surfacing and linking sites, not for collecting training data.
What is Perplexity?
Perplexity is a standalone answer engine. A user goes to it with a question and receives a composed answer with inline citations, rather than encountering an answer inside a product built for something else. Of the four systems documented here, it is the one whose entire interface is organized around the answer and its sources.
That origin shows in the product's treatment of attribution, which is structural rather than supplementary. It also shows in the crawler documentation, which frames indexing in terms of surfacing and linking rather than ingestion.
Which crawlers does Perplexity run?
Two, with a sharp division.
PerplexityBot is described as designed to surface and link websites in search results on Perplexity, and the documentation states explicitly that it is not used to crawl content for AI foundation models. This is the agent that governs whether a site can appear in Perplexity's results, and Perplexity's own recommendation to site owners who want to be included is to allow it and to permit requests from the published IP ranges.
Perplexity-User handles user initiated requests. When someone asks a question, the system may visit a web page to help produce an accurate answer and include a link to that page in its response. The documentation states that this agent generally ignores robots.txt rules, on the reasoning that the fetch was triggered directly by a user.
What does the training statement actually commit to?
It is worth reading precisely. The documentation says PerplexityBot is not used to crawl content for AI foundation models. That is a statement about what that specific crawler is for.
What it does not do is describe a separate training crawler, as OpenAI and Anthropic both do. The absence is notable and this page will not interpret it, because interpreting a gap in documentation is exactly the kind of inference that gets presented as fact elsewhere. What can be said accurately: Perplexity's published crawler documentation describes two agents and attributes neither to foundation model training.
What control does a site owner have?
Meaningful control over PerplexityBot through robots.txt, and limited control over Perplexity-User, which the documentation states generally ignores those rules.
This produces the same asymmetry OpenAI's documentation produces, and it is the single most important thing for a publisher to internalize about this generation of systems. A publisher who blocks the indexing crawler has removed themselves from the index. They have not necessarily stopped their page being fetched and read when a user asks a question that leads to it.
Perplexity also publishes IP ranges, and recommends permitting them alongside the robots.txt allowance. A publisher using firewall or WAF rules rather than robots.txt alone needs that list, because a bot allowed in robots.txt and blocked at the edge is still blocked.
What is not publicly known?
Selection, again. The documentation describes access and says that a link to a visited page is included in the response, but nothing about how a page is chosen for citation out of the candidates retrieved, how many citations a given answer carries, or what distinguishes a cited source from a retrieved but unused one.
The frequency and depth of recrawling are also undocumented, which matters for publishers whose content changes often and who want to know how stale the indexed version of a page might be when it is cited.
What does ignoring robots.txt mean in practice?
It means the file is a statement about indexing rather than a boundary around the content.
The reasoning Perplexity gives is that a fetch by Perplexity-User happens because a person asked a question whose answer requires that page, which the company treats as closer to a person following a link than to automated collection. Whether a publisher accepts that framing is a separate matter from whether it describes the behavior, and it does describe the behavior.
The consequence for a publisher is that robots.txt answers one question and not another. It answers whether your pages are gathered into an index that can surface them. It does not answer whether a page can be read at the moment someone asks about it. A publisher whose requirement is genuinely that content never be read by such a system is looking at authentication, paywalls or network level controls, not at a robots file.
For most publishers that distinction is academic, because most want to be read and cited. It becomes concrete for anyone with licensing arrangements, embargoed material, or content whose value depends on being behind a decision rather than behind a convention.