Robots.txt Generator

Build a robots.txt block by block: per-crawler Allow and Disallow rules, crawl delay, sitemap lines and comments, with live preview.

The file is assembled locally in your browser. Nothing ever leaves your device.

Sitemaps

One Sitemap: line per XML sitemap — full absolute URLs are the safe choice.

robots.txt preview
—

Starting from an existing file? Paste its lines into the blocks manually — the preview replaces the old file once you are done.

How It Works

The generator walks your ruleset blocks top to bottom and writes one User-agent: group per block: the agent line first, then every Allow:, Disallow: and optional Crawl-delay: line, followed by your Sitemap: entries and any comment lines you added. The preview updates as you type, and copy or download hands you a finished file ready for your site root at /robots.txt.

Directive order and most-specific match
A crawler finds the group whose User-agent: it matches; if several match, Google uses the most specific one (a User-agent: Googlebot-news group beats a plain * group). Inside a group, the longest path that matches the URL wins, so Allow: /private/public-part can carve an exception out of Disallow: /private. Order does not create priority — specificity does.
Disallow: / versus Disallow: (empty)
Disallow: / blocks the entire site for that agent. An empty value, Disallow:, means the opposite: crawl everything. A group with no Disallow line at all behaves like the empty one, which is why the default block in this tool ships an empty Disallow row rather than none.
Wildcards, end anchors and why this is not security
Google and Bing extend the syntax with * matching any characters and $ anchoring the end, so Disallow: /*.pdf$ skips every PDF path. Remember the limit: robots.txt is a courtesy protocol honored by search engines, not enforced by anything — and the file itself is public, so a Disallow line advertises exactly what you wanted hidden. Sensitive paths need authentication, not directives.

Frequently Asked Questions

Will a blocked URL disappear from Google?

Not necessarily. Disallow stops Google from fetching a page, but if other sites link to that URL Google can still list it in search results — usually with the title, a "URL may be blocked by robots.txt" notice and no description. To truly remove a page you need it reachable, with a noindex meta tag or an authenticated 410/401 response, then robots.txt can keep it out afterwards.

Is robots.txt better than noindex for hiding pages?

Usually the opposite is true. Google has stated that robots.txt blocking is the weakest de-indexing method: a crawl-blocked page can stay in the index, and Google cannot see a noindex tag on a page it is forbidden to fetch. The reliable order is: allow the crawler in, serve a noindex meta tag (or an HTTP noindex header), and only add the robots.txt Disallow once the page has left the index.

Is the Sitemap: line required?

No — search consoles and RSS-style feeds can discover sitemaps another way — but it is the cheapest self-description a site can offer: one free line telling every compliant crawler where your XML sitemaps live. Include at least the canonical URL of your main sitemap.

Can robots.txt protect private pages or admin panels?

No, and treating it that way is a classic mistake. robots.txt sits at a public URL, so publishing a Disallow entry is the same as printing a guest list of your hidden paths — Google and Bing both index the file itself, and scrapers read it daily. Anything secret belongs behind authentication, not behind a directive that well-behaved crawlers agree to follow.

Is anything uploaded to a server?

No. The file is assembled from your form inputs with plain JavaScript in the browser, and the copy and download actions work entirely client-side — you can build it with networking turned off. Nothing is pasted anywhere unless you upload the result to your own site.