Skip to content

Playbooks/August 20, 2026

AI Crawler AccessThe Silent Reason Your Brand Is Missing From AI Answers

Robin Pautigny

Robin Pautigny

Co-founder, Refine

AI Crawler Access: The Silent Reason Your Brand Is Missing From AI Answers

Summary

AI assistants can only cite what they can fetch. Each major provider now operates several distinct crawlers - one that gathers training data, one that builds a search index, and one that fetches a page live when a user asks - and they must be handled separately. Blocking a training crawler is a legitimate business decision. Blocking a retrieval crawler removes your pages from the pool a model can quote, which is a self-inflicted visibility problem. This guide explains the difference, lists the user agents worth knowing, and gives you a 30-minute audit.

The short answer

If your brand is absent from AI answers, check crawler access before you write another word of content. Every large provider now splits its bots by purpose: OpenAI separates GPTBot (training) from OAI-SearchBot (search indexing) and ChatGPT-User (live, user-triggered fetch); Anthropic and Perplexity follow the same pattern. A single blanket rule in robots.txt, a bot-management default at your CDN, or a WAF rule that 403s unfamiliar user agents can block all of them at once. Opting out of training while staying eligible for citation is a valid, deliberate configuration - but it only works if you write the rules per agent.

The Two Kinds of AI Crawler (And Why the Difference Decides Your Visibility)

The most useful mental model in 2026 is that AI crawlers come in two families, and they do completely different jobs for your brand.

Training crawlers collect pages that may end up in the corpus used to train or refine a foundation model. The payoff is slow, diffuse and impossible to attribute: months later, a model may have absorbed the fact that your product exists and what it does. Whether you want that is a genuine business question, and reasonable companies answer it differently - publishers with licensing revenue at stake often say no, and that is a defensible position.

Retrieval crawlers are a different animal. They fetch pages so an assistant can ground an answer it is generating right now. When someone asks ChatGPT or Perplexity for the best tools in your category, the assistant runs a search, pulls a handful of URLs, reads them, and synthesises an answer with citations. If your pages cannot be fetched at that moment, you are not in the answer. There is no ranking penalty, no gradual decline, no diagnostic in your analytics. You are simply not part of the material the model had to work with.

That asymmetry is why the conversation about "blocking AI bots" is so often mis-framed. The question is not whether to let AI companies in. It is which of their agents to let in, and for what purpose.

The User Agents Worth Knowing in 2026

The list below covers the agents that account for most AI traffic to a typical B2B or ecommerce site. Names and behaviour change, so treat this as a starting point and verify against each provider’s published crawler documentation before you finalise a policy.

  • OpenAI - GPTBot collects data that may be used for model training. OAI-SearchBot indexes pages for ChatGPT search. ChatGPT-User fetches a specific page because a user or a ChatGPT action asked for it.
  • Anthropic - ClaudeBot is the training crawler. Claude-SearchBot supports search indexing. Claude-User fetches a page in response to a direct user request inside Claude.
  • Perplexity - PerplexityBot indexes pages so they can be surfaced and cited in answers. Perplexity-User retrieves a page a user has explicitly clicked through to or asked about.
  • Google - Googlebot remains the crawler behind Search, AI Overviews and AI Mode. Google-Extended is a separate control that governs use of your content for Gemini model training and grounding, and does not affect Search indexing.
  • Apple - Applebot powers Siri and Spotlight. Applebot-Extended is the opt-out signal for generative model training specifically.
  • Others worth allowlisting or reviewing - Meta-ExternalAgent, Amazonbot, MistralAI-User, Bingbot (which feeds several downstream assistants), and CCBot, the Common Crawl bot whose archives are used far more widely than most teams realise.

The pattern to internalise: within a single provider, blocking one agent does not block the others, and allowing one does not allow the others. Every rule you write needs an explicit user-agent target. This is the single most common configuration mistake we see, and it is usually invisible until someone goes looking.

Why Well-Optimised Sites Still Get Blocked

Most blocked sites did not choose to be blocked. The block came from somewhere in the stack nobody on the marketing team owns.

  • CDN and bot-management defaults. Several providers now ship AI-bot mitigation switched on by default or behind a single toggle. A security-minded engineer flips it, nothing visibly breaks, and your citation rate degrades over the following weeks.
  • WAF rules that reject unfamiliar user agents. A rule written to stop scrapers will happily return 403 to OAI-SearchBot, which looks nothing like a browser.
  • Inherited robots.txt. A "User-agent: GPTBot / Disallow: /" line copied from a 2024 blog post, never revisited, now blocking one provider while its siblings walk straight past.
  • Aggressive rate limiting. Retrieval bots often request several pages in a burst. If your limits are tuned for human sessions, the crawler gets throttled into failure and quietly gives up.
  • Client-side rendering. If the substance of your page only appears after JavaScript executes, a crawler that fetches raw HTML sees an empty shell. It was allowed in; it just found nothing to quote.
  • Interstitials and soft walls. Cookie banners that block content, geo-redirects, and "read the rest" gates all reduce what a retrieval agent can actually extract.

None of these show up as errors in Google Search Console, which is why they persist. Search indexing and AI retrieval are now separate pipelines with separate failure modes, and only one of them has a mature diagnostic console.

How to Audit Your AI Crawler Access in 30 Minutes

You do not need a project for this. Work through the following in order and you will know where you stand.

  • Read your own robots.txt properly. Fetch it, list every user-agent block, and write down what each one allows. Pay particular attention to wildcard rules that catch more agents than intended.
  • Simulate each crawler. Request a key page - your homepage, a product page, a comparison page - while sending each user agent string, and record the HTTP status. Anything other than 200 is a finding. A 403 or 429 means your edge is refusing them regardless of what robots.txt says.
  • Check what a bot actually sees. Compare the raw HTML response against the rendered page. If your value proposition, pricing or feature list only exists in the rendered version, that content is not available for citation.
  • Grep your server and CDN logs for the user agents above over the last 30 days. Zero hits from a provider is a strong signal you are blocked, not that nobody is asking about you.
  • Confirm your key pages are reachable without JavaScript, without a cookie wall, and without a login.
  • Write the policy down. Decide explicitly, per provider, what you allow for training and what you allow for retrieval - then put it in a document that the next engineer to touch the WAF will find.

Budget half an hour for the first pass and re-run it quarterly, plus any time you change CDN, hosting or security vendor. Access is not a one-off setting; it drifts.

Access is the input, citation is the outcome

Fixing crawler access tells you the door is open. It does not tell you whether anyone walked through it. The measurable outcome is whether your brand is actually named and cited when real prompts are run - which is what Refine tracks continuously across ChatGPT, Gemini, Perplexity, Claude, Copilot and Mistral. Run the access audit first, then watch mention and citation rates over the following four to eight weeks. If access was the bottleneck, that is where you will see it move.

Four Crawler Policies That Actually Make Sense

There is no universally correct answer, but there are four coherent postures. Pick one deliberately rather than ending up somewhere by accident.

  • Fully open. Allow training and retrieval agents everywhere. Best for most B2B SaaS, service businesses and challenger brands whose main problem is being unknown rather than being copied.
  • Retrieval-only. Block the training crawlers, allow the search and user-fetch agents. This is the sweet spot for companies with a genuine IP concern who still want to be citable. It requires per-agent rules and regular review, because a new agent name defaults to whatever your wildcard says.
  • Selective. Open your documentation, comparison pages, glossary and public blog; restrict gated research, customer data and anything behind a paywall. Pairs well with the fact that models cite explanatory content far more often than sales pages.
  • Closed. Block everything. Only rational when content licensing is a real revenue line and you have the leverage to negotiate. Understand what you are trading away: in your category, the answer still gets generated - just with someone else in it.

What Crawler Access Does Not Fix

It would be neat if crawler access were the whole story. It is not. Access is a precondition, not a strategy - it puts you in the pool of eligible sources and nothing more.

Once a model can read you, it still has to prefer you. That comes from the things GEO practitioners have been converging on for two years: content that answers a question directly in its first paragraph, clean extractable structure, specific claims with numbers and dates attached, and - most importantly - corroboration elsewhere. Models weight third-party agreement heavily. A claim that appears on your site, a review platform, a community thread and an industry roundup is treated very differently from the same claim appearing only in your own marketing copy.

So the sequence matters. Confirm the retrieval agents can fetch your pages. Confirm those pages contain extractable, factual, current answers. Then build the third-party corroboration that makes a model comfortable naming you. Skipping straight to step three is why so many well-funded content programmes produce nothing measurable in AI answers - and checking step one takes an afternoon.

Short on time? Have an assistant summarise this page for you.