CCBot
Common Crawl documents CCBot as the crawler that builds an open web corpus. That corpus is widely reused as model feedstock. A CCBot hit is not a reader.
- Operator
- Common Crawl
- Motif family
- other_ai
- Motif class
- Model-related crawler
- In Motif classifier
- Yes
- robots.txt
- Vendor documents that it honors robots.txt
- Executes JavaScript
- No
- Verified in Motif
- Common Crawl publishes crawler docs. Motif does not check source IPs
What Motif will and will not claim
Motif classifies CCBot as family other_ai, class model, on tracker or collector hits. HTML-only Common Crawl fetches of the origin are out of scope.
Revenue
Not a visitor and not a channel. Policy for CCBot does not move Stripe.
Vendor documentation
Start with the vendor page at https://commoncrawl.org/ccbot. Motif stores the controlled token from a self-reported User-Agent.
CCBot User-Agent example
CCBot/2.0 (+https://commoncrawl.org/faq/)
Allow CCBot in robots.txt
User-agent: CCBot
Allow: /
Block CCBot in robots.txt
User-agent: CCBot
Disallow: /
Does CCBot count as a visit in Motif?
No. If CCBot hits the Motif tracker or collector it increments AI crawlers. It does not create a visitor and cannot win a recommendation.
Will Motif show CCBot on my HTML pages?
Only if that client requested /m.js or POST /v1/events. HTML-only crawls of the origin never reach Motif. Saying otherwise would be a product lie.
Should I block CCBot?
Model-related crawlers are a data-policy call. Allowing or blocking them does not create Stripe customers. Motif will not pretend a crawler hit is demand.
Related