Intermediate Time: 12 min
Step 1 of 8
The AI crawler landscape
Before changing robots.txt, identify which systems request your pages and what each request is for. That classification lets you set access policy by purpose instead of assuming every crawler from the same vendor or product serves the same function in practice.
Core reading
- AI Crawlers Wiki
- ChatGPT-User Wiki
Self-check
- Which major AI crawlers appear in your logs, and how can you identify their User-Agents?
- How can an engine cite you without ever crawling you?
- Which crawlers train models, and which fetch your page at answer time?