Robots.txt Generator
- Start from a template: allow all, block all, WordPress, Magento or custom.
- Adjust the rules per user agent: which paths to allow and which to block.
- Enter your sitemap URL — that line is how search engines find your list of pages.
- Copy the result into a file named robots.txt at the root of your domain.
It controls **crawling**, not indexing. Blocking a path stops the bot from fetching the page, but does not stop the URL appearing in results if other pages link to it — the search engine shows the address with no description, because it could not read the content.
To keep a page out of results, the right instrument is a noindex directive on the page itself. And there is a trap: if the page is blocked in robots.txt, the bot never fetches it and therefore never sees the noindex.
The file is also not a security mechanism. It is public, anyone can read it, and listing a sensitive path there points straight at it.
- Blocking CSS and JavaScript folders: the search engine needs to render the page as a user sees it, and without those files the assessment goes wrong.
- Blocking the whole site in production — usually a leftover from a staging environment promoted without review.
- Relying on crawl-delay: major search engines ignore the directive. Crawl rate adjusts to what your server can handle.
- Forgetting the sitemap line, the cheapest way to expose your full list of URLs.
The file must sit exactly at the domain root. In a subdirectory it is ignored. Subdomains have their own independent files.
Rules are grouped by user agent. When a specific block exists for a bot, that bot ignores the generic asterisk block entirely rather than combining the two.
Among overlapping rules, the most specific one wins — the longest path — not the one that appears first.
Frequently asked questions
Not necessarily. robots.txt prevents crawling, not indexing. The URL can still appear, without a description, if other pages link to it. To remove it, use noindex on the page and leave crawling allowed.
No. The file is public, and listing a path there reveals that it exists. Use authentication to protect content.
At the domain root, reachable as your site address followed by slash robots.txt. In a subdirectory it is ignored, and every subdomain needs its own.
Major search engines ignore the directive. Crawl rate is adjusted automatically based on how your server responds.
It is not mandatory, but it is the cheapest way to make sure a crawler finds your complete list of URLs without relying only on internal links.