๐Ÿค– robots.txt Generator

Path rules
robots.txt
User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://example.com/sitemap.xml

Configure crawl allow/disallow rules, a crawl delay, and a sitemap URL to generate a robots.txt file for search engines. In addition to a general allow/disallow setting for the whole site, you can add fine-grained rules for specific paths.

How to use

  1. Choose whether to allow crawling of the entire site.
  2. Add as many allow/disallow rules for specific paths as you need.
  3. Enter your sitemap URL and an optional crawl delay โ€” the robots.txt content is generated automatically.

How the calculation works

robots.txt tells search engine crawlers which pages of a site they may crawl. It was standardised as RFC 9309 in 2022 and lives at the site root (https://example.com/robots.txt). This tool builds: โ€ข User-agent: * (applies to all crawlers) โ€ข Allow: / (when "allow all" is on) โ€ข Allow / Disallow lines for each path you add โ€ข Crawl-delay: seconds between requests โ€ข Sitemap: your sitemap URL Under RFC 9309, when several rules match a URL, the rule with the longest matching path wins. So "Allow: /" together with "Disallow: /admin/" blocks only /admin/ and below.

Worked example

Default output User-agent: * Allow: / Disallow: /admin/ Sitemap: https://example.com/sitemap.xml Here /admin/ and /admin/users are not crawled; everything else is.

Things to be aware of

  • robots.txt is a request, not access control. Malicious crawlers ignore it, and anyone can read the file. Protect private pages with a login or password.
  • A disallowed page can still appear in search results as a bare URL if other sites link to it. To keep a page out of results, use noindex on the page (and do not disallow it, or the crawler cannot see the noindex).
  • Google ignores Crawl-delay; only some search engines such as Bing use it.

FAQ

Where should robots.txt be placed?

It needs to be placed at your site's root directory (e.g. https://example.com/robots.txt).

Does Crawl-delay affect all search engines?

No, some search engines, including Google, ignore the Crawl-delay directive. Its effect varies by crawler.

Can I set different rules per crawler (User-agent)?

The current version only supports generating rules that apply to all crawlers (User-agent: *).