AI crawler directory
18 classified HTTP tokens and 2 robots.txt product tokens. Each page states the vendor's documented job, whether Motif matches the token, and why a hit is never a Stripe customer. User-Agent strings prove nothing about source IP.
Three classes, three stakes
Search-index bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer, Amzn-SearchBot) are closer to Googlebot than to a reader. User-requested fetchers (ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher, Amzn-User, Google-CloudVertexBot) fire because a person asked. Model-related crawlers (GPTBot, ClaudeBot, Meta-ExternalAgent, Amazonbot, Bytespider, CCBot) collect corpus. Motif will not call that last group 'training' because it never sees the prompt.
The Motif limit, stated once
Motif observes these tokens on GET /m.js and POST /v1/events. HTML-only crawls of your origin never arrive. Competitors who show GPTBot on blog posts are reading logs Motif does not have. Empty crawler Insights means no tracker hits, not a quiet internet.
Policy tokens are not bots
Google-Extended and Applebot-Extended belong in robots.txt. They are not User-Agents. Motif's classifier omits them. Directories that put them in the same table as GPTBot without that sentence are teaching a false model of HTTP.
Tokens Motif does not classify today
Some public directories profile GoogleOther and DuckAssistBot. Motif's classifier does not match those strings. If they hit /m.js they are not labeled. We will not invent a profile for a token the product does not name.
Should I block AI crawlers?
Split the class. Blocking search-index bots opts you out of some assistant citations. Blocking model crawlers is a data-policy call with no Motif ranking effect. User-requested fetchers often ignore robots.txt. Motif will not rank any of them as channels.
How do I know a GPTBot hit is real?
You do not, from Motif. Check the source IP against OpenAI's published list. Motif stores the self-reported token when the request hit the tracker. Anyone can send User-Agent: GPTBot.
Do AI crawlers show up in Google Analytics?
Almost never. They do not run JavaScript, so tag-based analytics misses them. Motif also misses HTML-only crawls. Server or CDN logs are the only complete crawler record.
Which crawlers send traffic back?
None of them are the traffic. Search-index and user-fetch bots may precede a human click. The click is a person with an assistant referrer. Motif ranks that person if they pay.
Related