What does a robots.txt tester check?
Use this free robots.txt checker to test whether a selected crawler is allowed or blocked from an exact URL path before you publish the file. It evaluates a supplied rules file for one crawler and one path, makes matching allow and disallow rules visible, shows the winning rule, and lets you build draft rules for review.
A robots.txt file is a crawl-control document for a specific host, protocol, and port. It belongs at the top level—for example, https://example.com/robots.txt—and its path rules are relative to that origin.
How to check a robots.txt file
- Paste the exact robots.txt content you plan to serve.
- Choose the crawler you want to evaluate, such as Googlebot.
- Enter the complete path you need to test, including its leading slash.
- Review the result, then verify the deployed file at
/robots.txtand check important URLs separately.
User-agent: *
Disallow: /private/
Allow: /private/public-guide/
Sitemap: https://example.com/sitemap.xmlGooglebot robots.txt checklist
- Test the exact host, scheme, and top-level path that will be deployed; a subdomain has its own robots.txt scope.
- Test the actual crawler token you care about, then review the matching user-agent group and the most specific rule.
- Check important assets and rendering paths before disallowing CSS, JavaScript, images, or other resources needed to understand a page.
- Keep the sitemap URL accurate and verify the live response after deployment, not only the draft text.
- Use a separate indexability audit when the question is canonical, meta robots, X-Robots-Tag, status code, or search inclusion.
Audit canonical and indexability signals when crawl access is only one part of the diagnosis.
robots.txt is not a security or noindex control
Robots.txt asks compliant crawlers not to fetch matching paths; it does not protect private information, and a blocked URL can still be discovered or appear without a snippet. Use authentication for private content. When a crawlable response must stay out of search, use a valid noindex directive or another appropriate removal method.
Read the full robots.txt versus noindex guide for the decision and verification steps.
Common robots.txt mistakes
- Testing the wrong host: a subdomain has its own robots.txt scope.
- Using a folder-level file instead of the required top-level
/robots.txt. - Blocking CSS or JavaScript that a crawler needs to understand the page.
- Expecting a disallow rule to remove an already-known URL from search.
Robots.txt checker questions
Can robots.txt remove a page from Google?
No. Robots.txt controls whether a compliant crawler may request a URL; it is not a reliable removal or security mechanism. Keep a page crawlable when Google must see a noindex directive, or use authentication for private content.
Do robots.txt rules apply to other subdomains?
No. Test the host that serves the page; a file is needed only when that host requires its own crawl rules. Rules at example.com/robots.txt do not automatically define the crawl rules for shop.example.com or another host.
Why test an exact URL path?
Rules are evaluated for a specific crawler and URL path. A file can allow one path while blocking a more specific path, or contain a conflict that is easy to miss when reading the text without a result.
Method
Check robots.txt rules for Googlebot and other crawlers, test allow/disallow matches, inspect the matching rule, and build draft rules—free, private, and no account.
Privacy
Runs in the browser unless you explicitly load a public robots.txt file.
Limitations
Robots rules control crawling, not authorization, and do not guarantee deindexing.
Primary sources
- Robots Exclusion Protocol — IETF; 2022-09; 2026-08-14.
- Introduction to robots.txt — Google Search Central; current 2026-08-19; 2026-08-19.
- Robots meta tag, data-nosnippet, and X-Robots-Tag specifications — Google Search Central; current 2026-08-19; 2026-08-19.