SEO Tools

Robots.txt Tester

Find out which pages a crawler may visit. Enter a website and the URLs you care about, choose a crawler such as Googlebot, Bingbot or GPTBot, and see for each URL whether it is allowed or blocked and which line of robots.txt decides it. You can also paste a draft robots.txt to test changes before you publish them.

  • Encrypted connection
  • No sign-up
  • Free to use
Test pasted robots.txt instead of the live file

Useful for trying changes before you publish them. The website above is then only used to resolve the URLs.

How to use Robots.txt Tester

  1. Enter the website address.
  2. List the URLs or paths to test and pick a crawler.
  3. Optionally paste a robots.txt draft instead of the live file.
  4. Read the result and the deciding rule for each URL.

Robots.txt Tester features

RFC 9309 matching

Longest match wins, Allow wins ties, * and $ wildcards.

Crawler groups

Picks the most specific user-agent group, else *.

Up to 100 URLs

Full URLs or paths in one test.

Draft testing

Check pasted rules without publishing them.

File notes

Unknown directives, rules before user-agent, noindex, crawl-delay.

Error handling

Explains how 4xx and 5xx responses affect crawling.

When to use Robots.txt Tester

  • Checking why a page is not crawled.
  • Testing new rules before deployment.
  • Blocking or allowing AI crawlers deliberately.
  • Auditing robots.txt after a site migration.

Robots.txt Tester FAQ

How are conflicting rules resolved?

The rule with the longest matching path wins. If an Allow and a Disallow rule match with the same length, Allow wins. This is the behaviour defined in RFC 9309 and used by Google.

Which group applies to Googlebot-Image?

The most specific matching group: a “googlebot-image” group if it exists, otherwise “googlebot”, otherwise “*”. Rules from other groups are not combined.

Does blocking a page remove it from Google?

No. robots.txt controls crawling, not indexing. A blocked URL can still be indexed from links. Use a noindex meta tag on a crawlable page to remove it.

What if robots.txt returns an error?

A 4xx response means no restrictions. A 5xx response or timeout makes Google treat the whole site as disallowed until the file loads again.

Is crawl-delay supported?

Google ignores it; Bing and Yandex respect it. The tool notes it but does not use it for decisions.

Is my draft stored?

No. Pasted rules are only used for this test.

How robots.txt rules are matched

robots.txt is a plain-text file at the root of a host that tells crawlers which paths they may request. Since 2022 its rules are standardised in RFC 9309, which describes how crawlers pick a group of rules and how they match paths.

A file consists of groups. Each group starts with one or more User-agent lines and lists Allow and Disallow rules. A crawler uses the group that names it most specifically; only if none does, it uses the * group. Groups that name the same crawler are combined.

Within the chosen group, every rule whose path matches the start of the URL path is a candidate, and the longest one wins. An asterisk matches any sequence of characters and a dollar sign anchors the end of the URL. This means a short Disallow: / can be overridden by a longer Allow: /public/.

Mistakes in robots.txt can hide an entire site from search engines or expose areas you meant to keep out of crawlers’ way. Testing the important URLs after every change, and before publishing, is the safest routine.

Other useful tools