AI training

meta-externalagent

Collects content for uses including Meta AI training.

See your agent traffic
PurposeAI training
HTTP identifiermeta-externalagent
IdentityVerify beyond the name

THE SIGNAL

What this visit tells you

Collection indicates access to content, without proving a specific downstream use.

Collected paths
Crawl volume
Bytes served

Your next move

Set a content-use policy for this crawler independently of Meta link previews.

The technical details

Markdown
How to identify it

Look for the meta-externalagent identifier in a request’s User-Agent. This is a name match, not identity verification. Version strings may change.

How to check identity

A User-Agent match identifies a claimed client. Check the source IP and any operator-published verification method before granting access.

Access and robots.txt

Use a meta-externalagent robots.txt group to control collection. Meta may cache the rules for up to 24 hours.

Optional full-site opt-out. Merge with existing rules only if intended. robots.txt does not secure private content.

User-agent: meta-externalagent
Disallow: /
Seeing 403 or 404 responses?

403 means access was denied; 404 means the resource was not found. Compare the public path, request time, edge security event and origin response to find the cause.

Separate intended restrictions and secret-file probes from pages that should work. A User-Agent name alone does not justify allowing a request.

SourcesReviewed 2026-09-14

Identity and purpose are based on these sources. Analytics interpretation and suggested checks are Apostl guidance.

APOSTL Pulse

See meta-externalagent in context

See their requests. Find the pages that matter.

See your agent traffic