AI crawlers, sorted by what blocking them costs
Treating these as one group is the mistake. Some build the index an assistant answers from; some read a single page because a person asked; some only collect text for training. They are separate names precisely so you can decide separately, and almost nobody does.
They decide whether you can be cited
These build the indexes assistants answer from. Blocking one of these removes you from that assistant's answers. This is where the expensive mistakes happen.
OAI-SearchBotOpenAIClaude-SearchBotAnthropicPerplexityBotPerplexityGooglebotGooglebingbotMicrosoft
They read one page because a person asked
Live fetches, with a human waiting. They collect nothing and train on nothing. They are also the easiest to block by accident, because they arrive from data centres.
ChatGPT-UserOpenAIClaude-UserAnthropicPerplexity-UserPerplexity
They only collect text for training
Blocking these costs you no visibility at all. Refusing to be trained on is a normal choice and it is entirely separate from whether you appear in answers.
GPTBotOpenAIClaudeBotAnthropicGoogle-ExtendedGoogleApplebot-ExtendedApplemeta-externalagentMetaCCBotCommon CrawlBytespiderByteDance