100% In-Browser · Zero-Upload
Crawlers: 16 AI Agents
Standard: RFC 9309
// AI CRAWLER POLICY & ROBOTS.TXT

Decide who may crawl. In plain English.

Pick which AI crawlers may read your site, see exactly what each user-agent token does, and download a clean robots.txt. Everything is generated in your browser.

mousy://robots-ai/presets POLICY READY

Presets

A preset only flips the checkboxes on the right. Fine-tune any crawler afterwards and the preset is marked as custom.

Options

Written inside every allowed group. Google ignores Crawl-delay; Bing, Yandex and some others honour it.


Only used for the comment on the first line. It does not change any crawler rule and crawlers ignore it entirely.

Good to know

Google-Extended is not Googlebot. Blocking it does not affect Google Search, indexing, snippets or ranking — it only opts your content out of Gemini training and grounding. Only a Googlebot rule would remove you from Search, and this tool never writes one.
Blocking is not a delete button. A robots.txt rule stops future fetches. It cannot remove text or files an AI company already downloaded. Opt-outs apply from the next crawl onwards.

AI crawlers

—

Filtering only hides rows from view. Every crawler in the list is still written to the file.

robots.txt output


        
—

Checks

What a robots.txt file can and cannot do

robots.txt is a plain-text file that lives at the root of a host — https://example.com/robots.txt — and tells crawlers which parts of the site they may request. It is part of a voluntary protocol (RFC 9309). Well-behaved crawlers fetch it, read it and follow it. Malicious scrapers ignore it entirely, and a few legitimate agents deliberately do not apply it to user-initiated fetches, because the request came from a person rather than from an automated crawl. Treat robots.txt as a clear statement of intent, not as a firewall.

Two distinctions explain most of the confusion around AI crawlers. The first is crawling versus indexing: robots.txt controls fetching, while indexing and ranking are decided by search engines and by page-level tags such as noindex. The second is training versus search: one company often runs several crawlers under the same brand with completely different jobs, and a single token rarely governs all of them.

Three families of AI user agents

Training crawlers (GPTBot, ClaudeBot, CCBot, meta-externalagent and others) collect pages to build datasets. Blocking them opts your content out of future training runs. Search and answer crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, YouBot) build the index behind an assistant's search feature so it can quote and link you. Blocking those removes you from the answers, which usually costs visibility rather than protecting anything. User-initiated fetchers (ChatGPT-User, Claude-User, Perplexity-User) open one specific page because a person asked about it. They are not crawlers and are not used for training.

Frequently Asked Questions

Does blocking Google-Extended hurt my Google rankings?↓

No. Google-Extended only governs the use of your content for Gemini training and grounding. Google Search crawling, indexing and ranking are controlled by Googlebot and are unaffected by this token.

If I block GPTBot, will my site disappear from ChatGPT?↓

Not necessarily. Training collection and search are separate. Blocking GPTBot opts you out of future training data while OAI-SearchBot can still index and cite you in ChatGPT's search results. If you block both, you leave the training corpus and the search index.

Do AI companies actually obey robots.txt?↓

The documented crawlers from OpenAI, Anthropic, Google, Apple, Perplexity, Amazon and Meta state that they honour it, and an opt-out group is the recognised way to decline. Two caveats are real: some agents have a poor compliance record, and user-initiated fetchers generally ignore robots.txt by design. Verify with your logs and add rate limiting or a WAF if volume becomes a problem.

Will blocking AI crawlers remove content that was already collected?↓

No. robots.txt is not retroactive. It stops future fetches from the moment the crawler re-reads the file; it cannot remove pages from a dataset that has already been built or from a model that has already been trained.

Should I also add a User-agent: * group?↓

Only deliberately. A wildcard group applies to every crawler that has no group of its own, and User-agent: * with Disallow: / removes your whole site from Google, Bing and everyone else. This generator never writes a wildcard group, so the file it produces cannot accidentally de-index you.

Related tools

llms.txt generator · Meta tag inspector · All MousyTools

MousyDev

Hecho por MousyDev

Creo herramientas web que respetan tu privacidad: tus datos se procesan en tu propio navegador y nunca pasan por mis servidores.