Skip to content
ScriptGarden
Now1 live tool · 10 coming soon

AI crawler control

Decide which AI bots may train on your content and which may cite it.

The short answer

AI crawler control means deciding, bot by bot, whether AI companies may use your pages to train models, to answer questions live, or to build a search index. You set it in robots.txt, and well-behaved bots follow it.

Why it matters

Training and search are now separate bots. For example you can block a training crawler such as GPTBot yet stay eligible for AI search citations through a search crawler such as OAI-SearchBot.

robots.txt is a request, not a lock. One 2025 sample found roughly 13% of AI bot requests ignored it, so sensitive content needs real access control.

Cloudflare proposed Content Signals in 2025: a line such as Content-Signal: search=yes, ai-train=no that states your preferences for search, live AI answers and training. It is a proposal, not a universal standard, but it is being adopted.

Tools in this category

Available now

Coming soon

  • AI Crawler Access Checker

    Soon

    Shows which AI bots your robots.txt and server allow or block.

  • Content Signals Generator

    Soon

    Writes the Content-Signal line for search, AI input and training.

  • AI Bot Blocker Generator

    Soon

    Builds robots.txt rules for GPTBot, ClaudeBot and similar bots.

  • Robots.txt Tester

    Soon

    Tests whether a URL is allowed for a chosen crawler.

  • AI Bot User-Agent Lookup

    Soon

    Identifies AI crawlers from the user-agent text in your logs.

  • Server Log AI Bot Analyzer

    Soon

    Counts AI bot visits from a pasted access log.

  • AI Training Opt-Out Checker

    Soon

    Checks the signals your site sends about training use.

  • AI Bot Rules Generator

    Soon

    Creates Apache, Nginx and Cloudflare rules to block chosen bots.

  • Robots.txt Change Monitor

    Soon

    Alerts you when a site's robots.txt changes.

  • TDM Opt-Out Header Generator

    Soon

    Makes the headers that reserve text and data mining rights.

These are planned ideas. Names and order may change.

What to do today

  1. 1

    Decide what you want: be cited by AI search, stay out of model training, or both.

  2. 2

    Build the rules with the Robots.txt Generator, which can add the common AI crawlers.

  3. 3

    Allow search-type bots if you want citations, and block training bots if you do not want your pages used for training.

  4. 4

    Check that your server or firewall is not blocking bots that robots.txt allows.

  5. 5

    Read what is robots.txt for the basics.

Frequently asked questions

Will blocking AI bots hurt my Google rankings?

Blocking AI training bots does not affect normal Google search. Google-Extended, for example, controls use for some Google AI products and is separate from Googlebot. Blocking Googlebot itself would remove you from search.

Which bots should I block?

It depends on your goal. Publishers who want to stay out of training often block GPTBot, ClaudeBot and Google-Extended while allowing search bots. If you want maximum AI visibility, allow them all.

Do all AI bots obey robots.txt?

The major companies say theirs do, but some requests from other bots do not. Treat robots.txt as a polite request.

What are Content Signals?

A proposed robots.txt line that says whether content may be used for search, for live AI answers, and for training. It is new and not yet followed by every bot.

Sources and further reading

Some figures in these sources come from vendor and industry studies and may change. Read the originals before relying on them.

Other SEO categories

See all 10 categories →