AI search · GEO & AEO

Robots.txt Generator

Choose which crawlers may read your site and this generator writes a valid robots.txt for it. Blocking an AI crawler takes two lines — User-agent: GPTBot, then Disallow: / — and the generator groups every crawler you block that way, keeps Google and Bing on the rules for all crawlers, and adds your sitemap at the end.

Start from a preset — the recommended one blocks model-training crawlers but keeps search and AI answer crawlers allowed — then test any URL against the result or against the live robots.txt of any site.

FreeNo signupRuns in your browserUpdated October 2026

Robots.txt Generator

The file is built in your browser. "Fetch from a site" has our server read the public robots.txt and keep nothing.

Start from a preset

Pick the policy closest to yours, then adjust single crawlers below.

AI and search crawlers

Each blocked crawler gets Disallow: / for the whole site. Allowed ones need no rule of their own: they follow the paths you set for all crawlers.

Fetch the pages assistants cite or open at a user's request. Block them and you are not cited.

0 of 5 blocked
  • OAI-SearchBotOpenAI

    Builds the index ChatGPT search cites from. Blocking it removes you from ChatGPT answers.

  • ChatGPT-UserOpenAI

    Fetches a page when a user explicitly asks ChatGPT to open your link. Blocking breaks that on-demand fetch.

  • Claude-UserAnthropic

    Fetches a page when a Claude user asks for your link directly.

  • PerplexityBotPerplexity

    Indexes pages Perplexity cites in answers. Blocking removes you from its citations.

  • AmazonbotAmazon

    Feeds Alexa answers and Amazon's AI features.

Model training

Collect text to train AI models. Blocking them does not affect your Google or Bing rankings.

7 of 7 blocked
  • GPTBotOpenAI

    Collects pages to train OpenAI models. Blocking keeps you out of training data but does not affect ChatGPT search citations.

  • ClaudeBotAnthropic

    Anthropic's crawler for model training data.

  • Google-ExtendedGoogle

    Controls use of your content for Gemini training and grounding. It does NOT affect normal Google Search ranking.

  • Applebot-ExtendedApple

    Controls use of your content for Apple Intelligence training. Regular Applebot indexing is separate.

  • meta-externalagentMeta

    Meta's crawler for AI products and model training.

  • BytespiderByteDance

    ByteDance crawler. Widely blocked for ignoring crawl-rate expectations; blocking costs little outside TikTok's ecosystem.

  • CCBotCommon Crawl

    Common Crawl's archive feeds many open datasets and models indirectly.

Classic search index

Google's and Bing's own crawlers. AI Overviews and Copilot are built on these indexes, so blocking them removes you from search and from those answers.

0 of 2 blocked
  • GooglebotGoogle

    Classic Google crawler. AI Overviews are built on this index, so blocking it removes you from both Search and AI Overviews.

  • BingbotMicrosoft

    Bing's crawler. Microsoft Copilot answers lean on the Bing index.

Paths for all crawlers

These rules go under User-agent: * and apply to every crawler you allowed. One path per line.

* matches any characters and $ marks the end of the URL, as in /*.pdf$

Re-opens part of a disallowed folder. The longer, more specific rule wins.

Sitemap

The full URL of your XML sitemap, so crawlers find every page.

File options

robots.txt
# Generated with seonix.ai

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Bytespider
User-agent: CCBot
Disallow: /

User-agent: *
Disallow:

Upload it to your site root so it opens at yourdomain.com/robots.txt

No conflicts: Googlebot and Bingbot can crawl the site, and AI search crawlers stay allowed.
Add your sitemap URL so crawlers find every page, including ones nothing links to yet.

Uploaded it? Check the live file with the AI Crawler Checker

Robots.txt tester

Testing: generated file

See whether a crawler may fetch a URL under these rules, and which line decides it.

Blocked

GPTBot may not crawl /

Deciding rule
Disallow: /
Group used
User-agent: GPTBot (its own group)
How it works

How to use it

  1. 01

    Pick a preset

    Allow all crawlers, block AI training only (the recommended middle ground) or block every AI crawler. Then switch single crawlers between Allow and Block.

  2. 02

    Add paths and your sitemap

    List the folders no crawler should visit, such as /admin/ or /cart/, any exceptions to re-open inside them, and the full URL of your XML sitemap.

  3. 03

    Test before you publish

    In the tester, pick a crawler and a path to see whether it is allowed and which line decides it. You can load any site's live robots.txt there too.

  4. 04

    Upload it to your site root

    Download robots.txt and place it at the root of your domain so it opens at yourdomain.com/robots.txt. Crawlers do not look for it anywhere else.

Honest notes

What robots.txt can and cannot do

robots.txt is a request, not a lock. Reputable crawlers follow it; a bot that ignores it can only be stopped at your server, CDN or firewall.

Blocking a URL does not remove it from search results: Google can still list a blocked page it finds through links, just without a description. To keep a page out of the index, let it be crawled and add a noindex meta tag or an X-Robots-Tag header instead.

The file is public — anyone can open yourdomain.com/robots.txt — so listing a private folder there only advertises it. Protect private areas with a login, not with Disallow.

There is no crawl-delay option because Google ignores that directive. A robots.txt also covers only the host it sits on, so every subdomain needs its own file.

Everything runs in your browser. The one exception is the optional Fetch from a site in the tester: our server reads that site's public robots.txt, returns it and keeps nothing.

Seonix does this on autopilot

Seonix catches robots.txt mistakes before they cost you traffic

The free Seonix site audit flags AI crawlers blocked in robots.txt, a robots.txt that blocks your sitemap, a missing robots.txt and meta tags that restrict AI use of your pages. Then Seonix finds what your buyers search for and publishes articles that answer it on your site, in 50+ languages and on a schedule.

For your stack

SEO on autopilot for the site you have

Seonix researches, writes and publishes articles — and keeps the technical side in check — wherever your site lives.

Frequently asked questions

How do I block GPTBot in robots.txt?

Add a group with User-agent: GPTBot followed by Disallow: /. GPTBot only collects training data for OpenAI models, so blocking it does not remove you from ChatGPT search, which uses OAI-SearchBot. To block several AI crawlers at once, put each one on its own User-agent line above a single Disallow: /.

Should I block AI crawlers?

Block the training crawlers if you do not want your text used to train models, and keep the AI search crawlers allowed if you want to be cited in AI answers — that is what the recommended preset does. Blocking every AI crawler keeps you out of ChatGPT search and Perplexity answers, while Google's AI Overviews still draw on the regular Googlebot index.

Does robots.txt stop AI training?

Yes, for crawlers that honour it, and only from then on. Blocking GPTBot, ClaudeBot, Google-Extended or CCBot tells those companies not to use your pages for training from that point. It does not remove text collected earlier, and a crawler that ignores robots.txt has to be blocked at your server or CDN.

Where does the robots.txt file go?

At the root of the host: https://yourdomain.com/robots.txt. Crawlers only look there — a robots.txt in a subfolder is ignored — and every subdomain, such as blog.yourdomain.com, needs its own. The name is robots.txt in lowercase, served as plain text.

What is the difference between robots.txt and noindex?

robots.txt controls crawling, noindex controls indexing. A page blocked in robots.txt can still appear in results if other pages link to it, because the crawler never fetches it and never sees a noindex. To keep a page out of search, allow crawling and add <meta name="robots" content="noindex"> or an X-Robots-Tag: noindex header.
Free tools

More free tools

All free tools