Skip to content
Intermediate Time: 12 min
Step 1 of 8

The AI crawler landscape

Before changing robots.txt, identify which systems request your pages and what each request is for. That classification lets you set access policy by purpose instead of assuming every crawler from the same vendor or product serves the same function in practice.

Core reading

  1. AI Crawlers
    Wiki
  2. ChatGPT-User
    Wiki

Self-check

  • Which major AI crawlers appear in your logs, and how can you identify their User-Agents?
  • How can an engine cite you without ever crawling you?
  • Which crawlers train models, and which fetch your page at answer time?