This free robots.txt generator and validator builds a correct robots.txt from guided rules, checks an existing file line by line, and shows exactly which URLs a crawler may fetch. It follows the same rules as RFC 9309 and Google’s own parser, and it runs entirely in your browser.
A robots.txt file is only a few lines long, yet one wrong character can hide your whole site from search or leave an admin area open to every bot. Use the robots.txt generator and validator before you upload a new file, after a migration, or whenever Search Console reports a page as “Blocked by robots.txt”.
What This robots.txt Generator and Validator Does
The robots.txt generator and validator has two modes that share one parser, so a file you build is checked with the same logic you would use to debug a live one.
- Guided generator: start from a preset (allow everything, block everything, WordPress default), then add user-agent groups with
AllowandDisallowrules and an optionalCrawl-delay. - AI crawler block: tick GPTBot, ClaudeBot, Google-Extended, CCBot and other AI crawlers to give them one shared
Disallow: /group without touching Googlebot or Bingbot. - Sitemap declaration: list your sitemap URLs and the generator writes the
Sitemap:lines, which must be full URLs. - Paste-in validator: paste or drop a robots.txt and get errors, warnings and notes for each line, with the reason and the fix.
- URL tester: pick a crawler (Googlebot, Bingbot, Applebot, GPTBot or a custom token), enter paths or full URLs, and see Allowed or Blocked plus the group and rule that decided it.
- Group overview: see the groups the way a crawler reads them, including user agents that were merged by accident.
How to Use It
- Open the Generator tab of the robots.txt generator and validator and pick a preset, or edit the first group. Use
*for every crawler without its own group. - Add
Disallowrules for paths crawlers should skip andAllowrules for exceptions inside them, then tick any AI crawlers you want to opt out. - Paste your sitemap URLs, read the Checks panel, and copy or download the finished
robots.txt. - Click Test URLs to open the file in the validator, or paste the current file from
https://your-domain/robots.txt. - Enter the URLs you care about and confirm each crawler gets the verdict you expect before you upload the file to your web root.
How the robots.txt Generator and Validator Reads Your Rules
Most robots.txt surprises come from a few rules that the specification defines precisely and that humans rarely remember. The robots.txt generator and validator applies them for you and explains them in the report.
- Groups: a group is one or more
User-agentlines followed by rules. A blank line does not end a group, so a user agent with no rules of its own silently shares the rules of the next group. The validator flags this, and the generator writes an emptyDisallow:for rule-less groups to keep them separate. - One group per crawler: a crawler uses the group that names it and ignores the
*group. If it appears in several groups, their rules are combined. - Longest match wins: the rule with the longest matching path decides, and
Allowwins a tie. That is why the WordPress default works. - Wildcards:
*matches any run of characters and$anchors the end of the URL, so/*.pdf$blocks PDF files. Paths are case-sensitive, and the query string counts. - Encoding: non-ASCII characters are compared in their UTF-8 percent-encoded form, so
/café/and/caf%C3%A9/are the same rule.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Here /wp-admin/admin-ajax.php is allowed because the Allow rule is longer than the Disallow rule, while the rest of /wp-admin/ stays blocked. The details are in RFC 9309 and in Google’s robots.txt documentation.
What robots.txt Cannot Do
Being honest about limits saves you from false confidence, so the robots.txt generator and validator says it plainly when a line cannot do what it looks like it does.
- It is not access control. Reputable crawlers honour it, others may not, and anyone can read the file. Protect private areas with authentication.
- It does not remove pages from search. A blocked URL can still be indexed from links. Use a
noindexmeta tag and keep that page crawlable.Noindex:inside robots.txt has not worked in Google since 2019. - Crawl-delay is not universal. Googlebot ignores it; Bingbot and some other crawlers honour it.
- It covers one host. Every subdomain, protocol and port needs its own file.
The robots.txt generator and validator also does not fetch your live file for you: browsers block most cross-site requests, so you paste it instead.
Practical Use Cases
- 🚧 Staging sites: confirm a staging copy blocks everything, and that
Disallow: /did not follow the site to production. - 📝 Self-hosted WordPress: after running WordPress and MariaDB in Docker on OpenMediaVault, start from the WordPress preset and add your sitemap.
- 🤖 AI training opt-out: block GPTBot, ClaudeBot, Google-Extended and CCBot while search crawlers keep full access, then prove it in the URL tester.
- 🛡️ Noisy bots: robots.txt only asks. For crawlers that ignore it, pair it with CrowdSec and Nginx Proxy Manager at your reverse proxy.
- 🔍 Search Console debugging: paste the file into the robots.txt generator and validator, test the reported URL as Googlebot, and see the exact line that blocks it.
Privacy, Ads, and Data Policy
- ✅ 100% free: no registration, no limits on how many files you build or check.
- ✅ No data storage: your robots.txt, sitemap URLs and test URLs never leave the browser.
- ✅ No ads in results: no watermarks, no tracking pixels, nothing added to your file.
- ✅ Client-side only: no external API calls; parsing and matching run in plain JavaScript on your device.
Open Source and Self-Hosting
The robots.txt generator and validator lives in the vahac-tools repository on GitHub, next to every other tool on this site. It has no dependencies: one folder with an HTML, a CSS and a JavaScript file that you can host on any static web server. The repository also has unit tests for the parser and the matcher.
Forks and pull requests are welcome, whether that is a new crawler preset or a smarter check. Need to decode an encoded path first? Try the URL Encoder and Decoder. Looking for more utilities like this robots.txt generator and validator? Browse the full list on the Tools page.
Built by VahaC — 100% client-side, no data sent anywhere.
