Robots.txt Generator

Compose groups of User-agent rules with Allow/Disallow, optional Crawl-delay, and global Sitemap/Host.

Group 1
Comma or newline separated. Choose from presets below or type custom tokens.
Non-standard; some crawlers respect it.
Tip Use * as wildcard; $ anchors the end of the path/query string.
Use absolute URLs.

Preview

Fill the form and click Generate to see your robots.txt here.

What is robots.txt and why it matters

  • Controls crawling: tells crawlers which URLs they may fetch. It doesn’t guarantee indexing control.
  • Protects crawl budget: prevents bots from wasting requests on duplicate or non-public areas.
  • Reduces load: disallowing heavy or unhelpful paths can reduce server strain.
Important
  • robots.txt must live at the site root: https://example.com/robots.txt (per origin).
  • Blocked URLs can still appear in search if linked elsewhere (URL-only listings). Use noindex (meta/X-Robots-Tag) or 404/410 to remove from index. noindex in robots.txt is not supported.
  • Wildcards: * matches any sequence; $ anchors the end. Tie-breaking typically uses the longest matching path; on ties, Allow wins.
  • Crawl-delay and Host are non-standard; some crawlers ignore them.

FAQs

It controls crawling, not indexing. Disallowed URLs may still be indexed if discovered by links, but without content. To remove from index, use noindex or return 404/410.

Group rules under each User-agent. Crawlers usually pick the group with the longest matching user-agent token, then apply the longest matching path rule; ties tend to favor Allow.

Avoid blocking assets required for rendering. Search engines render pages; blocking critical CSS/JS can harm rendering-based understanding and features.

  1. Gate access (best practice): protect staging with HTTP auth/SSO/VPN or IP allow-listing. Auth walls return 401/403, preventing fetching and indexing.
  2. Add noindex as a second layer:
    • HTML: <meta name="robots" content="noindex, nofollow">
    • HTTP: X-Robots-Tag: noindex, nofollow
    This defends against accidental link exposure.
  3. Optional robots.txt block:
    User-agent: *
    Disallow: /
    Helps compliant crawlers, but robots.txt alone cannot stop indexing if URLs leak via links or sitemaps.
Quick server examples:
# Nginx (noindex + optional basic auth)
add_header X-Robots-Tag "noindex, nofollow" always;
# auth wall
auth_basic "Restricted";
auth_basic_user_file /etc/nginx/.htpasswd;
# Apache (noindex + optional basic auth)
Header always set X-Robots-Tag "noindex, nofollow"
AuthType Basic
AuthName "Restricted"
AuthUserFile /var/www/.htpasswd
Require valid-user

Summary: use auth first, then noindex, and optionally Disallow: / in robots.txt for defense-in-depth.

Why this output matters

Robots.txt controls crawler access at the site level. A correct output helps SEO and engineering teams manage crawl paths and sitemap discovery while avoiding accidental blocks of pages or assets that search engines need to access.

Related tools

All tools