What you may point the scanner at.
This tool causes real HTTP requests to real infrastructure that may not be yours. That is the whole reason this page exists.
Working draft. Behaviourally accurate; not reviewed by counsel.
The one rule that matters
Only submit sites you are entitled to have crawled. Your own, your client's, or a competitor you are monitoring within the published limits. Submitting a URL here makes us fetch it, from our addresses, on your behalf.
Not permitted
- Using the scanner as a traffic source. Pointing repeated scans at a third party to generate load. The per-domain cooldown exists partly to bound this, and circumventing it is a ban.
- Circumventing the abuse controls. Rotating addresses past the cooldown, or automating around the verification challenge.
- Scanning to find vulnerabilities in someone else's site. This is a content and structure rubric, not a security scanner, and using it as reconnaissance is a misuse.
- Reselling raw scan output as your own product without adding anything. Agency use, running client sites and reporting on them, is expected and fine.
- Submitting URLs that resolve into private address space. The SSRF guard rejects these; attempting it repeatedly is treated as an attack.
Competitor monitoring, specifically
Monitoring a competitor is legitimate and is a product feature. It is bounded to three sites per brand, twelve pages each, one request at a time, once a week, with robots.txt obeyed. Those limits are not configurable, and a site that disallows CrawldBot is not crawled at all.
If you are on the receiving end
Disallow CrawldBot in your robots.txt and monitoring stops on the next scheduled run. You
need no permission and no account. The crawler page has the exact syntax.
Enforcement
Abuse is met with a block. Requester addresses are stored as truncated one-way hashes: enough to recognise a repeat offender, not enough to be a log of who visited.