AI crawlers do not read your site the way Googlebot does — and most robots.txt files still have no idea GPTBot, ClaudeBot or PerplexityBot exist. This is a hands-on technical guide to configuring both files correctly, with Nevatrix's own live setup as a real example.
robots.txt tells crawlers which parts of your site they may access. llms.txt is a newer, separate file that tells AI systems which pages are worth reading and how to understand your business at a glance. In 2026, both matter for GEO — if AI crawlers cannot access your site, or cannot quickly understand what it is about, you will not be cited in ChatGPT, Perplexity or Google AI Overviews, no matter how good your content is.
Allow AI crawlers explicitly in robots.txt (GPTBot, ClaudeBot, PerplexityBot, Applebot-Extended, Google-Extended and others), and publish a structured llms.txt file at yourdomain.com/llms.txt summarising your business, services, pricing and key pages — most default robots.txt files silently block several of these bots.
What Is llms.txt and Why Does It Exist?
llms.txt is a proposed plain-text standard, placed at yourdomain.com/llms.txt, that gives AI systems a clean, structured summary of your site — company facts, services, pricing, FAQs and key page links — instead of forcing them to parse and interpret your full HTML. Think of it as a business card for AI crawlers: robots.txt controls access, llms.txt improves comprehension.
What Is robots.txt and How AI Crawlers Use It Differently Than Google
robots.txt is the decades-old standard that tells any crawler which paths it may or may not access, using User-agent and Disallow/Allow rules. Googlebot has followed robots.txt for 25+ years — but AI crawlers are newer, each has a different name, and if you have never explicitly addressed them, many default robots.txt configurations either accidentally block them or simply have no rule for them at all. Most AI crawlers default to allowing access when no specific rule exists, unless blocked by a blanket "Disallow: /" rule, but explicit Allow rules increase crawl frequency and demonstrate intent.
The Major AI Crawlers You Need to Know in 2026
| Crawler | Operated By | Purpose |
|---|---|---|
| GPTBot | OpenAI | Trains ChatGPT's models on public web content |
| ChatGPT-User | OpenAI | Fetches live pages when a user asks ChatGPT to browse |
| OAI-SearchBot | OpenAI | Powers ChatGPT Search result indexing |
| ClaudeBot / anthropic-ai | Anthropic | Crawls and trains Claude's models |
| Claude-Web | Anthropic | Fetches live pages for Claude's browsing feature |
| PerplexityBot / Perplexity-User | Perplexity AI | Indexes and fetches pages cited in Perplexity answers |
| Google-Extended | Controls use of your content for Gemini models and grounding — not Google Search or AI Overviews, which use Googlebot | |
| Applebot-Extended | Apple | Controls use of your content for Apple Intelligence features |
| Amazonbot | Amazon | Crawls for Alexa and Amazon AI features |
| CCBot | Common Crawl | Public dataset used to train many third-party AI models |
Why this table matters
- Blocking Googlebot means losing Google Search traffic, including AI Overviews. Blocking Google-Extended is different: it only stops your content being used for Gemini models and grounding, and does not affect Google Search rankings or AI Overviews
- Each AI company runs at least one "training" crawler and one "live fetch" crawler — you can allow one and block the other if you want citations without training use
- A default, unedited robots.txt typically has no explicit rule for most of these — meaning your access policy is accidental, not intentional
How to Configure robots.txt to Allow AI Crawlers
Add an explicit User-agent block for each crawler you want to allow. A minimal AI-friendly addition looks like this:
Minimal robots.txt block to allow AI crawlers
- User-agent: GPTBot → Allow: /
- User-agent: ChatGPT-User → Allow: /
- User-agent: ClaudeBot → Allow: /
- User-agent: PerplexityBot → Allow: /
- User-agent: Google-Extended → Allow: /
- User-agent: Applebot-Extended → Allow: /
If you want citations and visibility but do not want your content used for model training, allow the "live fetch" bots (ChatGPT-User, Claude-Web, Perplexity-User) while disallowing the "training" bots (GPTBot, anthropic-ai, CCBot) — this is a legitimate, increasingly common configuration, though it means you may be cited less often since some tools rely on training data rather than live fetches.
How to Write an llms.txt File
Place a plain Markdown file at yourdomain.com/llms.txt. Structure it with a one-line company summary, key company facts, your services with pricing ranges, your team/author credentials, and a curated list of your most important pages — organised so an AI system can extract facts in seconds rather than crawling your entire site.
Real Example: Nevatrix's Own robots.txt and llms.txt
We do not just recommend this setup — it is what runs on nevatrix.com. Our robots.txt explicitly allows every major AI crawler (GPTBot, ChatGPT-User, OAI-SearchBot, Google-Extended, anthropic-ai, Claude-Web, ClaudeBot, PerplexityBot, Perplexity-User, Applebot, Applebot-Extended, Amazonbot, CCBot and more), while explicitly blocking known scraper bots that offer no citation value (DotBot, BLEXBot, PetalBot, Bytespider). Our llms.txt lists company facts, every service with real pricing, author credentials with LinkedIn links, and a curated map of our blog content by topic cluster — structured exactly the way this guide recommends.
Common Mistakes That Block AI Visibility
- Using a wildcard "Disallow: /" for User-agent: * without adding explicit Allow rules for each AI crawler you want to permit — this silently blocks all of them
- Never updating robots.txt since it was first generated, missing every AI crawler that has launched since (most robots.txt files predate GPTBot)
- Publishing llms.txt with marketing fluff instead of extractable facts — AI systems cite specifics (prices, credentials, data), not adjectives
- Forgetting to keep llms.txt updated as services, pricing or team members change — a stale llms.txt can actively mislead AI citations
- Blocking AI crawlers at the CDN/firewall level (e.g. via a security plugin) even after correctly configuring robots.txt — check both layers