Artificial intelligence

How to Check Whether AI Search Engines Can Access Your Website Content

AI Search Engines

To check whether AI search engines can access your website, look at three layers in turn: your robots.txt file, which tells each named crawler what it may fetch; your server or CDN, which may block automated visitors regardless of what robots.txt says; and the page itself, which must deliver its main content in HTML that a crawler can read without running heavy JavaScript. A block at any one of these layers can keep your content out of AI-generated answers.

Many site owners discover they are blocking AI crawlers they never meant to block, or allowing ones they assumed were shut out. An AI search readiness checker reads your robots.txt against the main AI user agents, fetches your page as a crawler would and reports what each one can and cannot see.

Training crawlers, search crawlers and user-triggered fetchers

The most important thing to understand is that AI companies now run several distinct bots, and blocking one does not block the others. Broadly, they fall into three groups.

  • Training crawlers collect content that may be used to train AI models. Examples include OpenAI’s GPTBot and Anthropic’s ClaudeBot.
  • Search crawlers index pages so an AI search product can find and cite them in answers. Examples include OpenAI’s OAI-SearchBot, Anthropic’s Claude-SearchBot and PerplexityBot.
  • User-triggered fetchers visit a page when a person asks the assistant something that requires it, such as ChatGPT-User, Claude-User and Perplexity-User. Because a person initiated the request, some providers state that these fetchers may not follow robots.txt in the same way as automated crawlers.

Where Google fits in

Google works differently. It does not use a separate crawler for AI. Instead, it offers a robots.txt control token, Google-Extended, that governs whether content Google crawls may be used for training its Gemini models and for certain grounding uses. Google states that Google-Extended does not affect inclusion or ranking in Google Search. AI features within Google Search itself rely on the regular Googlebot, so blocking Googlebot to avoid AI features would also remove you from ordinary search results.

The practical upshot: you can choose to opt out of model training while still allowing AI search engines to find and cite you. Many publishers and businesses make exactly that choice.

Crawler names and their purposes change as providers launch new products. Check each provider’s official crawler documentation periodically rather than treating any list, including this one, as permanent.

Worked example: a sample AI access check, line by line

A US B2B software company asked why its product guides never appeared as sources in AI search answers, even though they ranked well in Google. Here is a sample readiness check on one guide.

Line Crawler or check Purpose Result
1 Googlebot Google Search Allowed
2 Google-Extended Gemini training and grounding control Blocked (wildcard rule)
3 GPTBot OpenAI model training Blocked (robots.txt)
4 OAI-SearchBot ChatGPT search Blocked (wildcard rule)
5 ClaudeBot Anthropic model training Blocked (wildcard rule)
6 PerplexityBot Perplexity search 403 Forbidden (firewall)
7 Main content in initial HTML Readability No: loaded by JavaScript
8 Robots meta tag Indexing directives index, follow

The robots.txt file behind lines 1 to 5 looked like this:

User-agent: Googlebot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: *
Disallow: /

Line 1. Googlebot has its own group allowing everything, which is why the guides rank in Google.

Line 2. Google-Extended is not named, so it falls under the catch-all User-agent: * group. Crawlers that follow the robots.txt standard obey the most specific group that matches them, and fall back to the wildcard group only when nothing names them. Here that means opting out of Gemini training and grounding, which does not affect Google Search. The company may be happy with that, but it should be a deliberate choice rather than a side effect.

Line 3. GPTBot is deliberately blocked. That was the intended policy: the company did not want its content used for model training.

Line 4 is the unintended consequence. OAI-SearchBot is not named, so it falls back to the catch-all User-agent: * group, which disallows everything. The company blocked ChatGPT’s search crawler by accident, which explains why its guides were never cited there.

Line 5. ClaudeBot is also caught by the wildcard. That matches the company’s no-training policy, but Claude-SearchBot would be blocked too.

Line 6. PerplexityBot receives a 403 error from the firewall’s bot protection before robots.txt is even considered. Server-level blocks are invisible in robots.txt, which is why a crawler-perspective fetch matters.

Line 7. Even where access is allowed, the guide’s text loads through JavaScript after the page shell arrives. Some AI crawlers do not run JavaScript, so they would see an almost empty page. Server-side rendering or static HTML for the main content fixes this.

Line 8. No noindex directive, so nothing at page level is excluding the content.

A revised robots.txt that keeps training crawlers out while letting search crawlers in might read:

User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /

User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: *
Allow: /
Disallow: /account/

Check your own site: enter your website address in the free AI crawler access checker to see which AI search and training bots your robots.txt allows, whether your server lets them through and whether your main content is readable without JavaScript.

Common reasons AI search engines cannot see a site

  • A catch-all disallow. As in the example, User-agent: * with Disallow: / blocks every crawler without its own group, including new ones you have never heard of.
  • Bot protection on the CDN or firewall. Some services offer one-click options to block AI bots, or challenge all automated traffic. Review those settings alongside robots.txt.
  • Content behind logins, pop-ups or scripts. If text appears only after a click, a login or a JavaScript render, crawlers may never see it.
  • robots.txt in the wrong place. It must sit at the root of the host, such as https://www.example.com/robots.txt. A file in a subfolder is ignored, and each subdomain needs its own.
  • Case and path errors. Paths in robots.txt rules are case-sensitive, so Disallow: /Guides/ does not block /guides/.

Access is the starting point, not the whole story. AI search systems still favour pages that load quickly, use clear headings and answer questions directly. A website SEO checker will catch indexing directives, missing titles and weak heading structure, and a website speed test will show whether slow responses could be limiting crawling.

Frequently asked questions

How do I know if my website is blocking AI crawlers?

Open yoursite.com/robots.txt and look for groups naming AI user agents, such as GPTBot, ClaudeBot or PerplexityBot, and for any catch-all User-agent: * group with Disallow: /. Then check your CDN or firewall settings for bot blocking. A readiness checker tests both layers from a crawler’s perspective at once.

Will blocking GPTBot remove my site from ChatGPT search?

Not on its own. OpenAI documents GPTBot as its training crawler and OAI-SearchBot as the crawler for ChatGPT search, and each is controlled separately in robots.txt. You can block GPTBot to opt out of training while allowing OAI-SearchBot, provided a wildcard rule is not blocking it.

Does blocking Google-Extended affect my Google rankings?

Google states that Google-Extended does not affect a site’s inclusion in Google Search and is not used as a ranking signal. It controls whether content may be used for Gemini model training and certain grounding uses. Your ordinary search visibility depends on Googlebot, which is controlled separately.

Is robots.txt enough to stop AI companies using my content?

robots.txt is a widely respected convention, and major AI providers state that their automated crawlers follow it. It is not an enforcement mechanism, though, and user-triggered fetchers may be treated differently. For content that must stay private, use authentication rather than relying on robots.txt alone.

Being visible to AI search is now a deliberate choice rather than a default, and a single wildcard line can make the choice for you without your noticing. Decide which crawlers you want to welcome, write rules that say exactly that, and confirm the result with the TechBullion AI visibility checker whenever your robots.txt or firewall settings change.

Comments

TechBullion

FinTech News and Information

Copyright © 2026 TechBullion. All Rights Reserved.

To Top

Pin It on Pinterest

Share This