AI training

GPTBot

Collects web content that may be used to train OpenAI foundation models.

See your agent traffic
PurposeAI training
HTTP identifierGPTBot
IdentityVerify beyond the name

THE SIGNAL

What this visit tells you

A crawl indicates content collection, not a ChatGPT recommendation or a prospective customer. Keep this series separate from user-requested visits.

Crawl volume
Content paths
Response status

Your next move

Choose your training policy independently of search visibility. Review public URLs and response codes before changing access.

The technical details

Markdown
How to identify it

Look for the GPTBot identifier in a request’s User-Agent. This is a name match, not identity verification. Version strings may change.

How to check identity

Compare the request IP with the current operator-published ranges linked in the source. A matching name alone does not verify identity.

Access and robots.txt

Use the GPTBot robots.txt group to control training collection. This setting is independent of OAI-SearchBot search access.

Optional full-site opt-out. Merge with existing rules only if intended. robots.txt does not secure private content.

User-agent: GPTBot
Disallow: /
Seeing 403 or 404 responses?

403 means access was denied; 404 means the resource was not found. Compare the public path, request time, edge security event and origin response to find the cause.

Separate intended restrictions and secret-file probes from pages that should work. A User-Agent name alone does not justify allowing a request.

SourcesReviewed 2026-09-14

Identity and purpose are based on these sources. Analytics interpretation and suggested checks are Apostl guidance.

APOSTL Pulse

See GPTBot in context

See their requests. Find the pages that matter.

See your agent traffic