comiza

AgentwatchCrawlers › Claude-SearchBot

Claude-SearchBot

Claude-SearchBot is how Anthropic builds the index that Claude answers from. When it cannot reach your pages, they are not in that index, and an assistant cannot cite a page it has never read. This is the crawler people block by accident while meaning to block something else.

Run byAnthropic
robots.txt tokenClaude-SearchBot
Purposefeeds answers
Cost of blockingYou disappear from its answers.

What blocking Claude-SearchBot actually costs

Your pages stop appearing in Claude. Not immediately and not visibly: what happens is that the index goes stale and then thin, and the traffic that would have come never arrives. There is no notification, no warning in any console, and nothing in your analytics that says why.

If you want to refuse training but stay quotable, this is precisely the crawler you must keep allowed.

Anthropic runs three of these, and they do different things

This is the part that costs people money. One Disallow aimed at the wrong name here is the difference between refusing to be trained on and vanishing from the answers.

robots.txt tokenWhat it is forWhat blocking it costs
Claude-SearchBot ← this pagefeeds answersYou disappear from its answers.
Claude-Userlive fetchIt cannot look at your page when somebody asks about you.
ClaudeBottraining onlyNothing. Your visibility is unaffected.

The rules, exactly

To allow it

User-agent: Claude-SearchBot Allow: /

An explicit Allow is only needed when a broader rule would otherwise catch it. If your robots.txt does not disallow anything, this crawler is already allowed and you need no rule at all.

To block it

User-agent: Claude-SearchBot Disallow: /

Put it in its own group. A named group replaces the * group entirely for that crawler and inherits nothing from it, which surprises almost everyone.

Three things about robots.txt that catch people out

A crawler obeys exactly one group
It picks the group whose User-agent value is the longest one that matches its name, and ignores every other group, including *. If you write a rule under * and a separate group for Claude-SearchBot, the rules under * do not apply to it at all.
Matching is by prefix, not by exact name
A group headed User-agent: Google matches Googlebot, and a group headed with a partial name matches more than you intended. Write the full token.
A server error on robots.txt blocks everything
If /robots.txt returns a 5xx, the documented behaviour is that crawlers stop crawling the whole site until it recovers. A missing file returning 404 is safe; a broken one is not.

Why your robots.txt may say yes while Claude-SearchBot still gets nothing

robots.txt is a request. A firewall is not. Bot protection at your CDN answers before your site does, and it has never read your robots.txt.

Every one of these crawlers arrives from a data centre, which is exactly what bot rules are tuned to refuse. The result is a site with a perfectly permissive robots.txt sitting behind a wall, and no checker that only reads robots.txt can see it, because it never makes the request.

The free scan on this site asks for your page as Claude-SearchBot, from a real server, and compares the answer against what an ordinary browser gets. If they differ, you have found the wall.

Check your site Free, no account, no email.

Making sure it is really them

A User-Agent is a claim, not proof. Anyone can send any name, so a rule that trusts the name alone can be walked straight through.

Anthropic documents its crawlers in the support article linked below. Treat the name alone as a claim.

Official documentation: https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler

Every crawler in the catalogue