Directory / AI crawlers
AI crawlers 30
AI2Bot
Listed onlyAllen Institute for AI · AI crawlers
The Allen Institute for AI's crawler, used to build open training datasets such as Dolma. AI2 publishes no IP ranges, ASN, or reverse-DNS pattern.
Amazonbot
Listed onlyAmazon · AI crawlers
Amazon's web crawler used to improve Amazon's products and services, including training AI models. Amazon publishes a human-readable IP list but no machine-parseable feed in a format this project's schema supports.
Bytespider
Listed onlyByteDance · AI crawlers
ByteDance's web crawler, used to gather training data for its AI models. ByteDance publishes no IP ranges, ASN, or reverse-DNS pattern for it.
CCBot
Fully verifiableCommon Crawl Foundation · AI crawlers
Common Crawl's crawler that builds the freely available Common Crawl web archive, widely reused as AI training data by third parties.
Claude-SearchBot
Fully verifiableAnthropic · AI crawlers
Anthropic's crawler that indexes pages to improve search-style answer quality in Claude products, distinct from ClaudeBot's training-data crawl.
ClaudeBot
Fully verifiableAnthropic · AI crawlers
Anthropic's web crawler collecting publicly available web data for training Claude models. Anthropic publishes its crawler source IPs as a machine-readable feed.
Cloudflare AI Search
Listed onlyCloudflare · AI crawlers
The crawler behind Cloudflare AI Search, which indexes website content so it can be searched. Cloudflare documents that it only crawls a website the customer owns — the domain must exist in the same Cloudflare account and be selected as an AI Search data source.
Cloudflare Browser Run Crawler
Fully verifiableCloudflare · AI crawlers
The crawler behind the /crawl endpoint of Cloudflare's Browser Run (Browser Rendering) developer product, which crawls third-party websites on behalf of Cloudflare customers building applications on the platform. Its user agent is not configurable, and every request is signed with Web Bot Auth HTTP message signatures that site owners can verify against Cloudflare's published key directory.
Diffbot
Listed onlyDiffbot · AI crawlers
Diffbot's general-purpose web crawler, used to build its Knowledge Graph and power its structured-extraction APIs. Diffbot documents robots.txt behavior but publishes no IP list, ASN, or reverse-DNS pattern.
FirecrawlAgent
Listed onlyFirecrawl · AI crawlers
The fetcher behind Firecrawl's scraping and crawling API, which retrieves web pages on behalf of the developers and AI agents calling that API. Firecrawl's crawl endpoint honours robots.txt by default; ignoring it is an enterprise option that must be enabled by the operator. Firecrawl documents that it does not use a fixed set of outbound IP addresses.
Google-CloudVertexBot
Fully verifiableGoogle · AI crawlers
Google Cloud crawler that fetches sites on the site owners' request when building Vertex AI Agents. Crawling preferences addressed to its user agent have no effect on Google Search or other Google products.
GPTBot
Fully verifiableOpenAI · AI crawlers
OpenAI's web crawler that gathers publicly available data used to train OpenAI's models.
ImagesiftBot
Listed onlyHive · AI crawlers
ImageSift's crawler, operated by Hive. It scrapes publicly available images across the web to support Hive's ImageSift reverse-image-search and web intelligence products. Standard robots.txt directives are respected.
Kimi-SearchBot
Fully verifiableMoonshot AI · AI crawlers
Moonshot AI's search crawler for Kimi, which analyses pages for relevance and builds the index behind Kimi's search features. Blocking it removes a site from Kimi search results. Listed separately from KimiBot because Moonshot documents it as an independently configurable crawler with its own robots.txt token and its own published address list.
KimiBot
Fully verifiableMoonshot AI · AI crawlers
Moonshot AI's crawler for the Kimi assistant, gathering content that may be used to train Kimi's foundation models. Moonshot documents robots.txt control per crawler and publishes a machine-readable list of the addresses its crawlers use, while noting the ranges are dynamic and may change.
Leipzig Corpora Collection Crawler
Listed onlyLeipzig University Natural Language Processing Group · AI crawlers
The LCC crawler is operated by Leipzig University to collect web text for the Leipzig Corpora Collection, a set of linguistic corpora used in natural language processing research.
LinkupBot
Fully verifiableLinkup · AI crawlers
Linkup's web crawler. It discovers and retrieves publicly accessible pages to build and maintain the Linkup search index, which powers search and web grounding for AI applications. Linkup publishes the CIDR ranges its crawler originates from and asks site owners to check both the source IP and the user-agent token, since user agents can be spoofed.
MaCoCu
Listed onlyMaCoCu project (Jožef Stefan Institute) · AI crawlers
Crawler for the CEF-funded MaCoCu project, which collects, curates and enriches monolingual and parallel web text to build language corpora for under-resourced languages. The project publishes no IP ranges, so requests cannot be verified.
Meta External Agent
VerifiableMeta · AI crawlers
Meta's crawler for AI training data and content indexing across Meta products. Verified by ASN lookup (AS32934); Meta publishes no IP feed.
Meta Web Indexer
VerifiableMeta · AI crawlers
Meta's crawler that navigates the web to improve the quality of Meta AI search results, analysing page content for relevance and accuracy in Meta AI responses. Verified by ASN lookup (AS32934); Meta publishes no IP feed.
MistralAI-Index
Fully verifiableMistral AI · AI crawlers
Mistral AI's indexing crawler. It crawls the web automatically to build the index behind Mistral search, which answers user questions in Vibe. Mistral documents that content collected by this crawler is not used for generative AI training of any kind, and publishes the crawler's addresses as a JSON feed.
MistralAI-Training
Listed onlyMistral AI · AI crawlers
Mistral AI's training crawler. It collects web content to help build the datasets used to train Mistral's generative AI models, and is documented as separate from both the search-index crawler and the user-triggered fetcher. Mistral states webmasters can disallow this user agent in robots.txt.
ICC-Crawler
VerifiableNational Institute of Information and Communications Technology (NICT) · AI crawlers
ICC-Crawler is a web crawler operated by Japan's NICT that collects web pages across the internet to build datasets for information and language processing research.
OAI-SearchBot
Fully verifiableOpenAI · AI crawlers
OpenAI's crawler that indexes pages to power search results inside ChatGPT. It is distinct from GPTBot and is not used to gather training data.
omgili
Listed onlyWebz.io · AI crawlers
Webz.io's legacy crawler user agent (formerly "Omgilibot"), used to collect web, forum, and news content for its data feeds. Webz.io's current public documentation describes successor crawlers ("webzio" / "webzio-extended") and no longer documents this UA string or any IP verification method for it.
PerplexityBot
Fully verifiablePerplexity · AI crawlers
Perplexity's crawler that indexes pages to surface and link websites in Perplexity search results. Perplexity states it is not used to train models.
SBIntuitionsBot
Listed onlySB Intuitions Corp. · AI crawlers
Crawler operated by SB Intuitions (a SoftBank AI subsidiary) that collects web pages for AI development and information analysis, including training of its Sarashina language models. SB Intuitions publishes no IP ranges, so requests cannot be verified.
ShapBot
Fully verifiableParallel Web Systems · AI crawlers
Parallel's indexing crawler, which collects publicly available web content to build and maintain the search index behind Parallel's web APIs. The operator documents the robots.txt token ShapBot and publishes the crawler's addresses as a JSON feed.
Webzio-extended
Listed onlyWebz.io Ltd. · AI crawlers
Second crawler in Webz.io's crawler pair, which performs ethical validation on the data collected by Webzio and tags it as usable or not usable for AI and machine-learning training. Webz.io publishes no IP ranges.
YouBot
Fully verifiableYou.com · AI crawlers
You.com's crawler that indexes pages for its AI-powered search product. You.com documents a dedicated IP range, a reverse-DNS pattern, and support for signed-request verification via Web Bot Auth.