# Diffbot

Operator: Diffbot
Purpose: AI search
Reviewed: 2026-09-14
HTTP identifier: Diffbot
robots.txt token: Diffbot

## What is Diffbot?

Builds the search index behind Diffbot’s knowledge services.

## How to identify it

Look for the Diffbot identifier in a request’s User-Agent. This is a name match, not identity verification. Version strings may change.

## How to check identity

A User-Agent match identifies a claimed client. Check the source IP and any operator-published verification method before granting access.

## Access and robots.txt

Diffbot documents robots.txt and Crawl-delay compliance by default. A configured override can apply when a customer has an agreement with the site owner.

Optional full-site opt-out; merge with existing rules only if intended. Not a security control.

```text
User-agent: Diffbot
Disallow: /
```

## What it means in your analytics

Proactive crawling is separate from a user-triggered extraction request.

## Useful measures

Fetched pages; Failed retrievals; Repeat fetches.

## What to check next

Review the public pages you want indexed and confirm the applicable access rules.

## Sources

- [docs.diffbot.com](https://docs.diffbot.com/docs/does-crawl-respect-robotstxt)

Operator facts are source-backed; interpretation and actions are Apostl guidance.

## Related profiles

- [Diffbot-User](https://apostl.dev/bots/diffbot-user)
- [OAI-SearchBot](https://apostl.dev/bots/oai-searchbot)
- [Claude-SearchBot](https://apostl.dev/bots/claude-searchbot)

[Start measuring agent traffic](https://apostl.dev/start/agent-analytics)
