Skip to content
Seoptist

Glossary

GPTBot

In short: GPTBot is OpenAI's web crawler that collects content for training its models; it is distinct from OAI-SearchBot and ChatGPT-User.

GPTBot is the user agent OpenAI uses to crawl public web pages for training future models. It respects robots.txt, identifies itself with a string containing GPTBot, and publishes the IP ranges it crawls from so site owners can verify requests.

It is one of three OpenAI agents and they do different jobs:

  • GPTBot gathers training data.
  • OAI-SearchBot indexes pages so ChatGPT can surface and cite them in search results.
  • ChatGPT-User fetches a specific page on demand when a user asks about it.

Blocking GPTBot does not stop ChatGPT citing you in search-enabled answers, and allowing it does not guarantee you will be mentioned. What it does affect is whether future models learn about your business from your own pages rather than from third-party descriptions of you.

Many sites block GPTBot without meaning to, through a wildcard rule added years ago or a security plugin's "block AI bots" default. A CDN firewall rule can also return 403 regardless of what robots.txt says.

Seoptist's audit checks your robots.txt rules for GPTBot and the other major AI user agents, tests the live response, and puts any block in your checklist with the lines to change.

Find out what AI assistants say about you today.

Run the free check in two minutes, or start a trial and get your full SEO and GEO checklist this week.

No card needed for the free check. Prices exclude VAT.