How to Run a Responsible AI Crawler Study
Original data can earn useful citations when the sample, definitions and limitations are transparent. A credible crawler study starts with a reproducible protocol, not a headline percentage.
Pre-register the domains, collection date, URL paths and fields you will record.
Define the sample and fields
Pre-register the domains, collection date, URL paths and fields you will record.
Useful fields include the presence of robots.txt, rules for documented tokens, llms.txt availability, sitemap status, structured data and the observed HTTP response. Do not collect private content or personal data.
Separate declared and observed results
Report policy directives separately from effective HTTP reachability and from search visibility.
A site can allow a token while returning a WAF challenge. Conversely, an accessible page does not prove that a crawler visited, indexed, trained on or cited it. Publish anonymized rows and a clear codebook when feasible.
Make updates auditable
Date every run, preserve the methodology and explain sampling bias and missing data.
Monthly updates are useful only when the collection process stays comparable. Avoid inflated claims such as universal adoption rates when the sample is small or convenience-based.
Practical checklist
Use these steps to turn the article’s principle into a repeatable publishing and measurement habit.
- Record the exact URL, crawler token and date before changing a rule.
- Separate a published directive from an observed HTTP response and from search visibility.
- Link to the relevant official documentation and state what the test cannot prove.
Frequently asked questions
How many sites are required?
There is no magic number; disclose the sample design, confidence limits and sources of bias.
Can a study prove crawler training?
No. Public directives and HTTP observations cannot prove how a model was trained.