Directory / AI crawlers

AI crawlers 30

AI2Bot

Listed only

Allen Institute for AI · AI crawlers

The Allen Institute for AI's crawler, used to build open training datasets such as Dolma. AI2 publishes no IP ranges, ASN, or reverse-DNS pattern.

Amazonbot

Listed only

Amazon · AI crawlers

Amazon's web crawler used to improve Amazon's products and services, including training AI models. Amazon publishes a human-readable IP list but no machine-parseable feed in a format this project's schema supports.

Bytespider

Listed only

ByteDance · AI crawlers

ByteDance's web crawler, used to gather training data for its AI models. ByteDance publishes no IP ranges, ASN, or reverse-DNS pattern for it.

CCBot

Fully verifiable

Common Crawl Foundation · AI crawlers

Common Crawl's crawler that builds the freely available Common Crawl web archive, widely reused as AI training data by third parties.

Claude-SearchBot

Fully verifiable

Anthropic · AI crawlers

Anthropic's crawler that indexes pages to improve search-style answer quality in Claude products, distinct from ClaudeBot's training-data crawl.

ClaudeBot

Fully verifiable

Anthropic · AI crawlers

Anthropic's web crawler collecting publicly available web data for training Claude models. Anthropic publishes its crawler source IPs as a machine-readable feed.

Cloudflare AI Search

Listed only

Cloudflare · AI crawlers

The crawler behind Cloudflare AI Search, which indexes website content so it can be searched. Cloudflare documents that it only crawls a website the customer owns — the domain must exist in the same Cloudflare account and be selected as an AI Search data source.

Cloudflare Browser Run Crawler

Fully verifiable

Cloudflare · AI crawlers

The crawler behind the /crawl endpoint of Cloudflare's Browser Run (Browser Rendering) developer product, which crawls third-party websites on behalf of Cloudflare customers building applications on the platform. Its user agent is not configurable, and every request is signed with Web Bot Auth HTTP message signatures that site owners can verify against Cloudflare's published key directory.

Diffbot

Listed only

Diffbot · AI crawlers

Diffbot's general-purpose web crawler, used to build its Knowledge Graph and power its structured-extraction APIs. Diffbot documents robots.txt behavior but publishes no IP list, ASN, or reverse-DNS pattern.

FirecrawlAgent

Listed only

Firecrawl · AI crawlers

The fetcher behind Firecrawl's scraping and crawling API, which retrieves web pages on behalf of the developers and AI agents calling that API. Firecrawl's crawl endpoint honours robots.txt by default; ignoring it is an enterprise option that must be enabled by the operator. Firecrawl documents that it does not use a fixed set of outbound IP addresses.

Google-CloudVertexBot

Fully verifiable

Google · AI crawlers

Google Cloud crawler that fetches sites on the site owners' request when building Vertex AI Agents. Crawling preferences addressed to its user agent have no effect on Google Search or other Google products.

GPTBot

Fully verifiable

OpenAI · AI crawlers

OpenAI's web crawler that gathers publicly available data used to train OpenAI's models.

ImagesiftBot

Listed only

Hive · AI crawlers

ImageSift's crawler, operated by Hive. It scrapes publicly available images across the web to support Hive's ImageSift reverse-image-search and web intelligence products. Standard robots.txt directives are respected.

Kimi-SearchBot

Fully verifiable

Moonshot AI · AI crawlers

Moonshot AI's search crawler for Kimi, which analyses pages for relevance and builds the index behind Kimi's search features. Blocking it removes a site from Kimi search results. Listed separately from KimiBot because Moonshot documents it as an independently configurable crawler with its own robots.txt token and its own published address list.

KimiBot

Fully verifiable

Moonshot AI · AI crawlers

Moonshot AI's crawler for the Kimi assistant, gathering content that may be used to train Kimi's foundation models. Moonshot documents robots.txt control per crawler and publishes a machine-readable list of the addresses its crawlers use, while noting the ranges are dynamic and may change.

Leipzig Corpora Collection Crawler

Listed only

Leipzig University Natural Language Processing Group · AI crawlers

The LCC crawler is operated by Leipzig University to collect web text for the Leipzig Corpora Collection, a set of linguistic corpora used in natural language processing research.

LinkupBot

Fully verifiable

Linkup · AI crawlers

Linkup's web crawler. It discovers and retrieves publicly accessible pages to build and maintain the Linkup search index, which powers search and web grounding for AI applications. Linkup publishes the CIDR ranges its crawler originates from and asks site owners to check both the source IP and the user-agent token, since user agents can be spoofed.

MaCoCu

Listed only

MaCoCu project (Jožef Stefan Institute) · AI crawlers

Crawler for the CEF-funded MaCoCu project, which collects, curates and enriches monolingual and parallel web text to build language corpora for under-resourced languages. The project publishes no IP ranges, so requests cannot be verified.

Meta External Agent

Verifiable

Meta · AI crawlers

Meta's crawler for AI training data and content indexing across Meta products. Verified by ASN lookup (AS32934); Meta publishes no IP feed.

Meta Web Indexer

Verifiable

Meta · AI crawlers

Meta's crawler that navigates the web to improve the quality of Meta AI search results, analysing page content for relevance and accuracy in Meta AI responses. Verified by ASN lookup (AS32934); Meta publishes no IP feed.

MistralAI-Index

Fully verifiable

Mistral AI · AI crawlers

Mistral AI's indexing crawler. It crawls the web automatically to build the index behind Mistral search, which answers user questions in Vibe. Mistral documents that content collected by this crawler is not used for generative AI training of any kind, and publishes the crawler's addresses as a JSON feed.

MistralAI-Training

Listed only

Mistral AI · AI crawlers

Mistral AI's training crawler. It collects web content to help build the datasets used to train Mistral's generative AI models, and is documented as separate from both the search-index crawler and the user-triggered fetcher. Mistral states webmasters can disallow this user agent in robots.txt.

ICC-Crawler

Verifiable

National Institute of Information and Communications Technology (NICT) · AI crawlers

ICC-Crawler is a web crawler operated by Japan's NICT that collects web pages across the internet to build datasets for information and language processing research.

OAI-SearchBot

Fully verifiable

OpenAI · AI crawlers

OpenAI's crawler that indexes pages to power search results inside ChatGPT. It is distinct from GPTBot and is not used to gather training data.

omgili

Listed only

Webz.io · AI crawlers

Webz.io's legacy crawler user agent (formerly "Omgilibot"), used to collect web, forum, and news content for its data feeds. Webz.io's current public documentation describes successor crawlers ("webzio" / "webzio-extended") and no longer documents this UA string or any IP verification method for it.

PerplexityBot

Fully verifiable

Perplexity · AI crawlers

Perplexity's crawler that indexes pages to surface and link websites in Perplexity search results. Perplexity states it is not used to train models.

SBIntuitionsBot

Listed only

SB Intuitions Corp. · AI crawlers

Crawler operated by SB Intuitions (a SoftBank AI subsidiary) that collects web pages for AI development and information analysis, including training of its Sarashina language models. SB Intuitions publishes no IP ranges, so requests cannot be verified.

ShapBot

Fully verifiable

Parallel Web Systems · AI crawlers

Parallel's indexing crawler, which collects publicly available web content to build and maintain the search index behind Parallel's web APIs. The operator documents the robots.txt token ShapBot and publishes the crawler's addresses as a JSON feed.

Webzio-extended

Listed only

Webz.io Ltd. · AI crawlers

Second crawler in Webz.io's crawler pair, which performs ethical validation on the data collected by Webzio and tags it as usable or not usable for AI and machine-learning training. Webz.io publishes no IP ranges.

YouBot

Fully verifiable

You.com · AI crawlers

You.com's crawler that indexes pages for its AI-powered search product. You.com documents a dedicated IP range, a reverse-DNS pattern, and support for signed-request verification via Web Bot Auth.