AI Crawler Checker: Is Your Content Accessible?
Check your website’s robots.txt rules for supported AI and search crawler tokens. Lantern shows homepage policy results and the rules behind them.

On this page
On this page
Follow Lantern on Google
Add Lantern as a preferred source to find more of our latest research in Google Search.
Your website’s robots.txt file communicates which paths automated crawlers are permitted to access. If those rules do not match your intentions, a compliant crawler may be restricted from content you want it to reach or permitted to access paths you intended to restrict.
Lantern’s AI Crawler Checker helps you inspect those policies. Enter a public domain or URL to see homepage policy results for supported AI and search crawler tokens, together with the rule evidence behind each decision.
The checker evaluates robots.txt. It does not score page content, measure citation readiness, or guarantee that a crawler will visit or cite your website.
This blog explains what the checker is, what it measures, how to use it, and what crawlers need to find and extract content effectively.
What Is an AI Crawler Checker?
Lantern’s AI Crawler Checker fetches a website’s root robots.txt file and evaluates its rules for the homepage path /.
For example, if you submit:
https://www.example.com/products/exampleThe checker fetches:
https://www.example.com/robots.txt
What the Checker Shows
Policy results for supported tokens
The checker returns results for these supported tokens:
- OAI-SearchBot
- GPTBot
- ChatGPT-User
- Claude-SearchBot
- ClaudeBot
- Claude-User
- PerplexityBot
- Perplexity-User
- Googlebot
- Google-Extended
- bingbot
- CCBot
- Applebot
- Applebot-Extended
These entries do not all represent the same activity. The catalog includes search crawlers, training crawlers, user-triggered retrieval agents, and data-use policy tokens.
Google-Extended and Applebot-Extended are policy tokens rather than independent crawling agents. Some user-triggered retrieval systems may not apply robots.txt in the same way as automated crawlers. A policy result should therefore not be interpreted as proof of how every request will behave.
Matched-rule evidence
When a rule determines the result, the checker shows:
- The applicable user-agent token or wildcard group.
- The matched Allow or Disallow directive.
- The rule pattern.
- Its source line number.
It also reports the number of applicable rules and makes the fetched robots.txt source available when the file was successfully returned.
If no applicable rule matches the homepage, access is treated as allowed. A result without a matched rule is not evidence of an explicit Allow directive.
Whether the policy could be fetched
The checker distinguishes three states:
| State | Meaning | Reported policy assumption |
|---|---|---|
| Found | A successful response returned policy content. | Evaluate the returned rules. |
| Not found | The target returned a 4xx response. | Treat compliant crawler access as allowed. |
| Unreachable | The policy could not be retrieved successfully, such as after a network failure or server error. | Use a conservative blocked assumption. |
These describe the checker’s handling of the fetch result. An unreachable result does not prove that robots.txt contains a blocking rule. Actual crawler behavior can also depend on implementation and previously cached policies.
How to Use the AI Crawler Checker
- Open Lantern’s AI Crawler Checker.
- Enter a public website domain or HTTP/HTTPS URL.
- Review the policy-fetch state and homepage results.
- Inspect the matched user-agent, directive, pattern, and line where available.
- Compare those rules with your intended access policy.
- If you update robots.txt, run another check to inspect the new result.
Response time depends on the target website and whether redirects or network delays occur. A failed fetch may require troubleshooting before a policy decision can be confirmed.
How to Interpret Allowed and Blocked Results
An allowed result means the evaluated robots.txt policy does not block the homepage for that token. It does not mean the crawler has visited, indexed the page, or selected it for an AI answer.
A blocked result can indicate a matching Disallow rule or, when the policy is unreachable, the checker’s conservative access assumption. Check the fetch state and evidence to distinguish the two.
For policy tokens, the result concerns the applicable data-use policy. For retrieval systems that may ignore robots.txt, it describes the policy rather than guaranteed enforcement.
There is no universal policy that every website should adopt. Decide which activities you intend to permit, consult the relevant operators’ documentation, and check that your rules express those choices.
What This Tool Does Not Check
The AI Crawler Checker does not:
- Analyze page copy, factual evidence, headings, or readability.
- Validate structured data or assess content freshness.
- Test JavaScript rendering or content extraction.
- Confirm firewall, authentication, or bot-management behavior.
- Inspect actual crawler visits or request logs.
- Check every path on a website.
- Measure rankings, AI citations, or brand mentions.
robots.txt is an access-policy mechanism, not a security boundary. Do not use it to protect sensitive information. Content requiring restricted access needs appropriate authentication and authorization.
FAQs
What does the AI Crawler Checker check?
It fetches a public website’s robots.txt file and evaluates homepage access policies for supported crawler and policy tokens. It returns rule evidence when a matching rule is available.
Does a result guarantee citations in ChatGPT or Perplexity?
No. An allowed result only means the evaluated robots.txt policy does not block the homepage for that token. Citation and visibility depend on additional factors, including whether the content is discovered, retrieved, selected, and used as a source.
Is the AI Crawler Checker free?
Yes. The AI Crawler Checker is free to use.
How often should pages be re-checked?
Quarterly for evergreen content, and after significant edits. Pages with frequently changing information may need more frequent reviews.
Does a blocked result mean every AI system is prevented from accessing the site?
No. robots.txt relies on compliance, and different types of retrieval can behave differently. It is not an enforcement mechanism like authentication or a firewall.
What should I do if robots.txt is unreachable?
Check whether the website and policy URL are accessible and whether server errors or redirects are involved. The tool’s blocked result in this state is an assumption, not evidence of a Disallow directive.
Check Your AI Crawler Policies
Understanding your access rules is one part of managing how automated systems interact with your website.
Run Lantern’s AI Crawler Checker to inspect homepage policies and see the rules behind the results.
Share this article