Agentwatch › Crawlers › ChatGPT-User
ChatGPT-User
ChatGPT-User fetches your page live, at the moment a person asks ChatGPT something about it. There is a human waiting for the answer while it happens. It is not collecting anything and it is not training on anything; it is reading one page, once, because somebody asked.
| Run by | OpenAI |
| robots.txt token | ChatGPT-User |
| Purpose | live fetch |
| Cost of blocking | It cannot look at your page when somebody asks about you. |
What blocking ChatGPT-User actually costs
ChatGPT cannot open your page when a person asks about it. The user gets a worse answer about you, assembled from whatever else is available — which frequently means somebody else's description of your business rather than yours.
This one is the easiest to block by accident, because live fetches come from ordinary data-centre addresses and look to bot protection like exactly the traffic it was installed to stop.
OpenAI runs three of these, and they do different things
This is the part that costs people money. One Disallow aimed
at the wrong name here is the difference between refusing to be trained on
and vanishing from the answers.
| robots.txt token | What it is for | What blocking it costs |
|---|---|---|
OAI-SearchBot | feeds answers | You disappear from its answers. |
ChatGPT-User ← this page | live fetch | It cannot look at your page when somebody asks about you. |
GPTBot | training only | Nothing. Your visibility is unaffected. |
The rules, exactly
To allow it
User-agent: ChatGPT-User
Allow: /
An explicit Allow is only needed when a broader rule would
otherwise catch it. If your robots.txt does not disallow anything, this
crawler is already allowed and you need no rule at all.
To block it
User-agent: ChatGPT-User
Disallow: /
Put it in its own group. A named group replaces the
* group entirely for that crawler and inherits nothing
from it, which surprises almost everyone.
Three things about robots.txt that catch people out
- A crawler obeys exactly one group
-
It picks the group whose
User-agentvalue is the longest one that matches its name, and ignores every other group, including*. If you write a rule under*and a separate group forChatGPT-User, the rules under*do not apply to it at all. - Matching is by prefix, not by exact name
-
A group headed
User-agent: Googlematches Googlebot, and a group headed with a partial name matches more than you intended. Write the full token. - A server error on robots.txt blocks everything
-
If
/robots.txtreturns a 5xx, the documented behaviour is that crawlers stop crawling the whole site until it recovers. A missing file returning 404 is safe; a broken one is not.
Why your robots.txt may say yes while ChatGPT-User still gets nothing
robots.txt is a request. A firewall is not. Bot protection at your CDN answers before your site does, and it has never read your robots.txt.
Every one of these crawlers arrives from a data centre, which is exactly what bot rules are tuned to refuse. The result is a site with a perfectly permissive robots.txt sitting behind a wall, and no checker that only reads robots.txt can see it, because it never makes the request.
The free scan on this site asks for your page as ChatGPT-User, from a real server, and compares the answer against what an ordinary browser gets. If they differ, you have found the wall.
Making sure it is really them
A User-Agent is a claim, not proof. Anyone can send any name, so a rule that trusts the name alone can be walked straight through.
Address ranges at https://openai.com/chatgpt-user.json.
Official documentation: https://platform.openai.com/docs/bots
Every crawler in the catalogue
OAI-SearchBotfeeds answersGPTBottraining onlyClaude-SearchBotfeeds answersClaude-Userlive fetchClaudeBottraining onlyPerplexityBotfeeds answersPerplexity-Userlive fetchGoogle-Extendedtraining onlyGooglebotfeeds answersbingbotfeeds answersApplebot-Extendedtraining onlymeta-externalagenttraining onlyCCBottraining onlyBytespidertraining only