๐ค robots.txt Generator
User-agent: * Allow: / Disallow: /admin/ Sitemap: https://example.com/sitemap.xml
Configure crawl allow/disallow rules, a crawl delay, and a sitemap URL to generate a robots.txt file for search engines. In addition to a general allow/disallow setting for the whole site, you can add fine-grained rules for specific paths.
How to use
- Choose whether to allow crawling of the entire site.
- Add as many allow/disallow rules for specific paths as you need.
- Enter your sitemap URL and an optional crawl delay โ the robots.txt content is generated automatically.
How the calculation works
robots.txt tells search engine crawlers which pages of a site they may crawl. It was standardised as RFC 9309 in 2022 and lives at the site root (https://example.com/robots.txt). This tool builds: โข User-agent: * (applies to all crawlers) โข Allow: / (when "allow all" is on) โข Allow / Disallow lines for each path you add โข Crawl-delay: seconds between requests โข Sitemap: your sitemap URL Under RFC 9309, when several rules match a URL, the rule with the longest matching path wins. So "Allow: /" together with "Disallow: /admin/" blocks only /admin/ and below.
Worked example
Default output User-agent: * Allow: / Disallow: /admin/ Sitemap: https://example.com/sitemap.xml Here /admin/ and /admin/users are not crawled; everything else is.
Things to be aware of
- robots.txt is a request, not access control. Malicious crawlers ignore it, and anyone can read the file. Protect private pages with a login or password.
- A disallowed page can still appear in search results as a bare URL if other sites link to it. To keep a page out of results, use noindex on the page (and do not disallow it, or the crawler cannot see the noindex).
- Google ignores Crawl-delay; only some search engines such as Bing use it.
FAQ
Where should robots.txt be placed?
It needs to be placed at your site's root directory (e.g. https://example.com/robots.txt).
Does Crawl-delay affect all search engines?
No, some search engines, including Google, ignore the Crawl-delay directive. Its effect varies by crawler.
Can I set different rules per crawler (User-agent)?
The current version only supports generating rules that apply to all crawlers (User-agent: *).