All posts

SEO4 min read

5 reasons ChatGPT can't see your website — and how to fix them

ChatGPT, Claude and Perplexity can't cite a page they can't see. Five common blockers for AI bots, and how to check and remove each one in an afternoon.

A small white robot figure behind frosted glass in front of a laptop showing code in a dark room

More and more people ask ChatGPT, Claude or Perplexity instead of Google. These tools answer from whatever their robots (crawlers) managed to read. If a robot can't reach your site, or can't see any text on it, your business effectively doesn't exist for that tool.

The good news: most of these blockers aren't intentional, and you can fix them in an afternoon. These are the five that come up most often.

1. Robots.txt shuts out the wrong bots

The robots.txt file tells robots where they may go. The problems are usually one of three:

  • A "block AI" list copied from the internet. It may also shut out the bots that cite you in answers, not only those that collect training data.
  • Blanket bans such as Disallow: / left over from the development version of the site.
  • Misunderstood groups. A robot follows only the group that names it most specifically. Once there is a User-agent: GPTBot block, the rules under User-agent: * no longer apply to it.

OpenAI runs separate bots for separate purposes. OAI-SearchBot powers search in ChatGPT, while GPTBot collects content for model training. The settings are independent, so you can stay visible in answers and still decline training:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: *
Allow: /
Sitemap: https://your-domain.com/sitemap.xml

OpenAI says its systems take about 24 hours to pick up a robots.txt change.

2. Your CDN or firewall blocks AI bots for you

Today this is the most common cause site owners don't know about. Cloudflare, which carries around a fifth of the web, has since July 2025 asked new domains whether to allow AI crawlers, with blocking as the default. Since 15 September 2026, new domains block training bots and AI agents by default on pages that show ads. Search bots stay allowed.

One important detail: if you block training, you also block bots that do several things at once (Cloudflare names Googlebot, Applebot and Bingbot). Your site then disappears from places where you expected it to show up.

Fix: in Cloudflare, open Security → Settings and review the AI bot settings. The same goes for other CDNs and hosting plans with a "block bots" switch.

3. The content only exists after JavaScript runs

Many modern sites send an empty HTML shell, and JavaScript draws the text in the browser. People don't notice, but Vercel's analysis found that the ChatGPT and Claude crawlers download JavaScript files but don't execute them. They see only what is in the raw HTML. Google's Gemini and Applebot do render pages, so Google can see you while ChatGPT can't.

Check: open View page source in your browser (Ctrl+U) and search for a sentence from the top of the page. If it isn't there, AI can't see it either.

Fix: render text, headings, prices and contact details on the server (SSR) or ahead of time (static generation). Next.js, Nuxt and Astro do this on their own when set up that way. Animations and 3D can stay in JavaScript, but the content must not depend on them.

4. "Are you human?" checks on public pages

A CAPTCHA, a "checking your browser" screen, an aggressive security plugin or a rate limit keeps attackers out, but a robot can't get past it either. The same goes for content behind a login and pages that only open after a cookie banner is clicked.

Fix: keep challenges for forms, login and checkout, not for articles and service pages. If your firewall needs an allowlist, OpenAI publishes the IP ranges of its bots. For other bots, check the provider's documentation.

5. The page itself says "no": noindex, errors and slowness

Some blockers live in the page itself:

  • noindex left over from development, WordPress's "Discourage search engines" option switched on, or an X-Robots-Tag: noindex header added by the server.
  • Wrong status codes: a page that returns 403 or 500 to robots, redirects to a login, or goes through a chain of several redirects.
  • Slowness: robots don't wait long. If the server takes several seconds to respond, a robot may give up before it gets any content.

Fix: search the page source for noindex, review the Pages report in Google Search Console, and make sure the server's first response arrives in under a second.

How to check all of this in ten minutes

Pose as an AI bot and see what you get:

curl -s -o /dev/null -w "%{http_code} %{time_total}s\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
  https://your-domain.com/

A 200 in under a second is a good sign. A 403, a 503 or a challenge page means the firewall is the problem. Repeat the test with ClaudeBot and PerplexityBot.

This test only checks rules based on the bot's name. CDNs also recognise real bots by IP address, so look at Cloudflare's security events or your server logs too: are the bots arriving, and what do they get?

For comparison, anim.hr returns a 200 to GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot in under 0.2 seconds, with the full article text in the HTML.

AI can't recommend you if it can't read you. Before writing new content, make sure your existing content actually reaches it.

Read next