OpenAI
GPTBot
Collects web content that may be used to train OpenAI foundation models.
AI trainingTHE SIGNAL
Successful retrieval is one step toward discoverability. It does not prove that an answer cited your page or that a visitor converted.
Check whether useful product and documentation URLs return readable content. Compare failures with your intended search policy.
Look for the OAI-SearchBot identifier in a request’s User-Agent. This is a name match, not identity verification. Version strings may change.
Compare the request IP with the current operator-published ranges linked in the source. A matching name alone does not verify identity.
Disallow OAI-SearchBot to exclude content from ChatGPT search answers. Navigational links may still appear.
Optional full-site opt-out. Merge with existing rules only if intended. robots.txt does not secure private content.
User-agent: OAI-SearchBot
Disallow: /403 means access was denied; 404 means the resource was not found. Compare the public path, request time, edge security event and origin response to find the cause.
Separate intended restrictions and secret-file probes from pages that should work. A User-Agent name alone does not justify allowing a request.
Identity and purpose are based on these sources. Analytics interpretation and suggested checks are Apostl guidance.
See their requests. Find the pages that matter.
See your agent traffic