Free tool

Robots.txt and AI crawlability checker.

Check whether Google and the AI crawlers can reach your site, read it, and understand what it says. Robots directives, heading structure, semantic HTML and structured data, in one report.

The basics

What should a robots.txt file contain?

At minimum: a line saying who the rules apply to, a line saying what they cannot visit, and a pointer to your sitemap. Most sites need nothing more complicated than that.

The mistakes we see are not exotic. A Disallow: / left in place after launch, which blocks everything. Rules written for a staging site that travelled to production with it. A file that returns an HTML error page instead of plain text, so crawlers either ignore it or misread it entirely.

One rule worth committing to memory: robots.txt controls crawling, not indexing. A blocked page can still show up in search results if other sites link to it, just without a description, which looks worse than not appearing at all. If you genuinely want something out of the index, that is a noindex tag, and a noindex tag only works if the crawler is allowed in to read it.

That last point catches people out constantly. Blocking a page in robots.txt and adding noindex means the crawler never sees the noindex.

Is your robots.txt blocking GPTBot?

GPTBot is OpenAI's crawler. It is what reads your site so that ChatGPT can talk about you. There are others: Google-Extended, ClaudeBot, PerplexityBot, CCBot.

Blocking them is a legitimate choice. Plenty of publishers do it deliberately, because they do not want their work training a model that competes with them, and that is a reasonable position to take.

The problem is doing it by accident, which is what usually happens. These rules arrive in a theme default, a plugin, or a snippet somebody copied from an article written in 2023 when the advice was different.

If you want to be recommended by AI, you have to let the crawlers in. You cannot be cited by something that has never read you. Whether that trade is worth it is your call, but it should be a call rather than an accident. Our AI search visibility service is the work behind getting cited once they can reach you.

The check

What we analyse

Four things that decide whether an AI crawler can read a page at all. A site can be perfectly readable to a person and still be a wall to GPTBot.

01 Heading hierarchy H1-H6 structure, heading depth, and whether AI can parse your content logically.
02 Semantic HTML structure HTML5 semantic elements, content-to-code ratio, and structural clarity.
03 Structured data Schema markup, JSON-LD, and how well AI can extract structured information.
04 Robots.txt and crawl directives Your robots file, meta tags, and the signals that help or hinder Google and the AI crawlers.
How it works

Three steps, one report

No account, no card. We fetch the page the way an AI crawler would and report what it could and could not read.

  1. Enter a URL

    Paste the URL of the website you want to check for AI crawlability.

  2. We analyse AI crawlability

    We check heading hierarchy, semantic HTML, structured data, meta tags, and crawl signals.

  3. Get your report

    See a detailed breakdown of how AI-ready your site is with actionable recommendations.

Structure

What is semantic HTML, and why does it matter to an AI?

Semantic HTML means using elements that describe what content is rather than how it looks. An <article> rather than a <div>. A real <h2> rather than bold text at 24 pixels.

To a person the two are identical. To a machine they are not remotely the same thing. One is a document with a structure it can follow. The other is a wall of boxes it has to guess at.

This matters more for AI than it ever did for traditional search, because a model extracting an answer needs to work out which part of your page is the answer. Clean headings and semantic elements are how it does that. It is also why a beautifully designed site built entirely from styled divs can be close to unreadable to the thing you want quoting you. It is one of the things we fix on a technical SEO audit.

Our research

What we found when we checked 89 Suffolk companies

In June 2026 we ran crawl and robots checks across every findable website belonging to Suffolk's 100 largest companies.

6.7% were blocking at least one major AI crawler. 5.6% were blocking GPTBot specifically. These are substantial businesses with real marketing budgets, and we would be surprised if any of them made that choice deliberately.

11.2% had no robots.txt at all, which is not fatal but tells you nobody has looked at the technical side in a while. One served an HTML page in place of the file, meaning any crawler reading it would either ignore it or misparse the instructions.

None of this is catastrophic on its own. It is the kind of thing nobody checks because nobody owns it. Read the full study.

Key takeaways

  • Robots.txt controls crawling, not indexing. Blocking a page does not remove it from search.
  • Never block a page and add noindex. The crawler cannot read the noindex.
  • Blocking GPTBot is a valid choice. Doing it by accident is not.
  • Semantic HTML is how a model works out which part of your page is the answer.
  • 6.7% of Suffolk’s biggest companies are blocking an AI crawler right now.

Frequently asked questions

What is AI crawlability?

AI crawlability refers to how easily AI platforms like ChatGPT, Gemini, and Perplexity can access, read, and understand your website content. It encompasses heading structure, semantic HTML, structured data, and crawl directives that determine whether LLM crawlers can efficiently extract and index your content.

Why does AI crawlability matter?

LLM crawlers have limited crawl budgets and different priorities than traditional search engine bots. If your site is difficult for AI crawlers to parse, they may skip your content entirely or misunderstand its context. Good AI crawlability ensures your content is accurately represented in AI-generated answers.

How is AI crawlability different from SEO crawlability?

Traditional SEO crawlability focuses on Googlebot and similar search engine crawlers. AI crawlability targets LLM-specific crawlers like GPTBot, Google-Extended, and PerplexityBot, which have different parsing behaviours, different content preferences, and different signals they use to evaluate and extract content.

How can I improve my AI crawlability?

Focus on clean heading hierarchy, semantic HTML elements (article, section, nav, aside), comprehensive structured data via JSON-LD, and clear crawl directives. Remove unnecessary JavaScript rendering barriers, ensure content is accessible without client-side rendering, and use schema markup to provide explicit context for AI systems.

Next step

Need help improving AI crawlability?

Our GEO service includes crawlability work as a matter of course: structure, schema, and the crawl directives that decide whether an AI engine reads you at all.

Speak to an expert