Robots.txt Generator
Compose groups of User-agent rules with Allow/Disallow, optional Crawl-delay, and global Sitemap/Host.
Preview
Fill the form and click Generate to see your robots.txt here.
What is robots.txt and why it matters
- Controls crawling: tells crawlers which URLs they may fetch. It doesn’t guarantee indexing control.
- Protects crawl budget: prevents bots from wasting requests on duplicate or non-public areas.
- Reduces load: disallowing heavy or unhelpful paths can reduce server strain.
Official references:
Important
robots.txtmust live at the site root:https://example.com/robots.txt(per origin).- Blocked URLs can still appear in search if linked elsewhere (URL-only listings). Use
noindex(meta/X-Robots-Tag) or 404/410 to remove from index.noindexin robots.txt is not supported. - Wildcards:
*matches any sequence;$anchors the end. Tie-breaking typically uses the longest matching path; on ties, Allow wins. Crawl-delayandHostare non-standard; some crawlers ignore them.
FAQs
It controls crawling, not indexing. Disallowed URLs may still be indexed if discovered by links, but without content. To remove from index, use
noindex or return 404/410.
Group rules under each
User-agent. Crawlers usually pick the group with the longest matching user-agent token, then apply the longest matching path rule; ties tend to favor Allow.
Avoid blocking assets required for rendering. Search engines render pages; blocking critical CSS/JS can harm rendering-based understanding and features.
- Gate access (best practice): protect staging with HTTP auth/SSO/VPN or IP allow-listing. Auth walls return
401/403, preventing fetching and indexing. - Add
noindexas a second layer:- HTML:
<meta name="robots" content="noindex, nofollow"> - HTTP:
X-Robots-Tag: noindex, nofollow
- HTML:
- Optional robots.txt block:
Helps compliant crawlers, but robots.txt alone cannot stop indexing if URLs leak via links or sitemaps.User-agent: * Disallow: /
Quick server examples:
# Nginx (noindex + optional basic auth)
add_header X-Robots-Tag "noindex, nofollow" always;
# auth wall
auth_basic "Restricted";
auth_basic_user_file /etc/nginx/.htpasswd;
# Apache (noindex + optional basic auth)
Header always set X-Robots-Tag "noindex, nofollow"
AuthType Basic
AuthName "Restricted"
AuthUserFile /var/www/.htpasswd
Require valid-user
Summary: use auth first, then noindex, and optionally Disallow: / in robots.txt for defense-in-depth.