Directory / Search engines

Search engines 59

AddSearchBot

Verifiable

AddSearch · Search engines

Crawler for AddSearch's hosted site-search service. It indexes the pages of customer sites so AddSearch can serve search results for them, obeys robots.txt rules written for the AddSearchBot token, and AddSearch publishes the fixed addresses it crawls from for allowlisting.

AdIdxBot

Verifiable

Microsoft · Search engines

Microsoft's crawler for Bing Ads / Microsoft Advertising. It crawls ads and follows through to the advertised landing pages for quality control, with both desktop and mobile variants.

Algolia Crawler

Verifiable

Algolia · Search engines

Algolia's Crawler, which visits a customer's pages, extracts search-relevant content, and pushes it to Algolia search indices. Algolia documents a fixed User-Agent token and a single static egress IP address for allowlisting.

Amazon AdBot

Verifiable

Amazon · Search engines

Amazon's advertising crawler. It scans web pages that request ads from Amazon's advertising systems, collecting page content for Amazon's classification systems to maintain brand safety and improve ad relevance.

AmazonProductDiscoverybot

Listed only

Amazon · Search engines

Amazon's product crawler, which collects publicly available product details from Amazon selling partner, brand and retailer websites to improve the accuracy and completeness of product information in Amazon's store. Amazon documents that it honours robots.txt user-agent and disallow directives, with changes taking up to 24 hours to apply, and that it does not support crawl-delay, nofollow or noindex.

Amzn-SearchBot

Listed only

Amazon · Search engines

Amazon's crawler used to improve search experiences in Amazon products and services; Amazon states it does not crawl content for generative AI model training. Amazon publishes a human-readable IP list but no machine-parseable feed in a format this project's schema supports.

Anomura

Verifiable

Direqt · Search engines

Direqt's search crawler, which discovers links and metadata to surface in Direqt's search features. The operator states it is not used to crawl content for model training, documents the robots.txt token Anomura, and publishes the two addresses the crawler requests come from.

Applebot

Fully verifiable

Apple · Search engines

Apple's web crawler, used to index content for Siri, Spotlight Suggestions, and Safari's search features.

Baiduspider

Verifiable

Baidu · Search engines

Baidu's primary web crawler, fetching pages for Baidu Search indexing.

BingVideoPreview

Verifiable

Microsoft · Search engines

Microsoft crawler, listed among Bing's crawlers, that fetches pages to generate video previews shown in Bing. It runs desktop and mobile variants and is verifiable through the same reverse-DNS check as Bing's other crawlers.

Bingbot

Fully verifiable

Microsoft · Search engines

Microsoft's primary web crawler, fetching pages for Bing Search indexing.

BingPreview

Verifiable

Microsoft · Search engines

Microsoft fetcher used to generate page snapshots for previews in Bing search results. Runs both desktop and mobile variants.

Bublup Bot

Verifiable

Bublup · Search engines

Bublup's content-discovery bot. It fetches pages to build the database behind Bublup's suggestion engine, reading page title, description and related images. Bublup states its crawling IPs are not fixed and documents reverse-DNS verification against the bublup.com domain instead.

Channel3Bot

Verifiable

Channel3 · Search engines

Channel3's product crawler. It visits publicly accessible product detail pages and collects images, titles, descriptions, prices, availability and variants for Channel3's product catalogue, which AI apps and agents query to route shoppers back to the original site. The operator tells site owners not to hard-code its addresses and to verify it by forward-confirmed reverse DNS instead.

coccocbot

Verifiable

Cốc Cốc · Search engines

Cốc Cốc's web crawler, fetching pages for the Vietnamese Cốc Cốc search engine. Separate web and image sub-bots share the same verification domain.

deepnoc

Listed only

deepnoc GmbH · Search engines

Research crawler operated by deepnoc GmbH that parses and stores public web page content so that new search engines can query it without running their own crawling infrastructure. No IP ranges are published.

DuckDuckBot

Fully verifiable

DuckDuckGo · Search engines

DuckDuckGo's web crawler, used to improve DuckDuckGo's search results.

ExaSearchBot

Fully verifiable

Exa · Search engines

Exa's search crawler, which discovers and indexes pages on the public web so that people and applications can find, retrieve and cite them through Exa. Every request is signed with HTTP Message Signatures (RFC 9421) under the Web Bot Auth scheme, so a request can be verified cryptographically regardless of its source IP address.

FreespokeCrawler

Listed only

Freespoke · Search engines

Freespoke's search crawler, which fetches pages to build the index behind the Freespoke search engine. The operator documents the crawler's user agent, the robots.txt token FreespokeCrawler and support for the Crawl-delay directive.

Geedo Product Search

Fully verifiable

Geedo · Search engines

GeedoShopProductFinder is the crawler for Geedo, a product-search engine that indexes publicly accessible product pages from online stores and follows robots.txt rules.

Storebot-Google

Fully verifiable

Google · Search engines

Google's crawler for Google Shopping surfaces, including the Shopping tab in Google Search and shopping.google.com. It crawls product and store pages to populate shopping results.

Google-InspectionTool

Fully verifiable

Google · Search engines

Google crawler used by Search testing tools such as the URL Inspection tool in Search Console and the Rich Result Test. Fetches pages on demand to show how Google Search renders them; it has no effect on ranking.

Google special-case crawlers

Fully verifiable

Google · Search engines

A group of Google crawlers that operate under separate agreements between the crawled site and a specific Google product (AdsBot, AdSense, Google Safety scanning, and similar). Several members of this group ignore the wildcard robots.txt rule and must be targeted explicitly.

Googlebot

Fully verifiable

Google · Search engines

Google's primary web crawler, fetching pages for Google Search indexing.

Googlebot Image

Fully verifiable

Google · Search engines

Google's image crawler, which fetches images for Google Images and image features in Search. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.

Googlebot Video

Fully verifiable

Google · Search engines

Google's video crawler, which fetches video content for Google Search video features. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.

GoogleOther

Fully verifiable

Google · Search engines

Google's generic crawler family (GoogleOther, GoogleOther-Image, GoogleOther-Video), used by various Google product teams for fetching publicly accessible content. Robots.txt rules addressed to it do not affect Google Search.

360Spider

Listed only

360 Search (Qihoo 360) · Search engines

360Spider is the web-search crawler of 360 Search (so.com), the search engine operated by Chinese internet company Qihoo 360. The operator also runs 360Spider-Image and 360Spider-Video for image and video search.

IONOS Crawler

Verifiable

IONOS · Search engines

IONOS Crawler (IonCrawl) continuously crawls publicly accessible domains to generate insights into how they are used, which IONOS uses to improve and expand its hosting products. IONOS documents reverse-DNS verification and deletes crawled data after 60 days.

Jooblebot

Listed only

Jooble · Search engines

The crawler for Jooble, a job-search aggregator that indexes job listings published across the web. Jooble's bot page states that it uses a web crawler which identifies itself as JoobleBot. No IP ranges are published.

Kagibot

Verifiable

Kagi · Search engines

Kagi's web crawler, fetching pages for the paid Kagi search engine. Kagi publishes a fixed, small set of source IPs rather than a CIDR feed, each paired with a kagibot.org hostname.

Linespider

Listed only

LINE · Search engines

LINE's crawler. It fetches web pages to support LINE's search features and content services, and adheres to the Robots Exclusion Protocol outlined in robots.txt.

Marginalia

Fully verifiable

Marginalia Search · Search engines

The crawler behind Marginalia Search, an independent search engine focused on older, text-heavy, and non-commercial web pages. Identifies itself by a bare URL rather than a conventional bot token.

MetaJobBot

Verifiable

METAJob · Search engines

MetaJobBot is the focused web crawler operated by the METAJob job meta-search engine. It searches websites for job listings to include in the METAJob index.

MojeekBot

Fully verifiable

Mojeek · Search engines

Mojeek's web crawler, fetching pages for the independent Mojeek search engine index.

Yeti

Verifiable

Naver · Search engines

Naver's web crawler, fetching pages for Naver Search indexing across South Korea's largest search engine.

Openindex Spider

Listed only

Openindex · Search engines

Openindex operates a web-crawling cluster (Apache Nutch on an Apache Hadoop cluster) for research and development of universal and focused search engines. The crawler respects robots.txt and the Crawl-delay directive and identifies itself in the User-Agent.

PetalBot

Verifiable

Huawei · Search engines

Huawei's web crawler, fetching pages for the Petal Search engine used in Huawei's mobile services and for Huawei Assistant content recommendations. Huawei documents reverse-DNS verification against aspiegel.com and petalsearch.com hostnames.

PiplBot

Listed only

Pipl · Search engines

PiplBot is the web-indexing crawler operated by Pipl, an identity and people-search service. It retrieves and indexes publicly available information, including content behind searchable databases, to build Pipl's people-search index.

Quantcastbot

Verifiable

Quantcast · Search engines

Quantcast's advertising crawler. It crawls websites to extract page content for interest-based audience categorization and to run quality-assurance checks on advertisement landing pages.

Qwantbot

Fully verifiable

Qwant · Search engines

Qwant's web crawler, fetching pages for the privacy-focused Qwant search engine. Older crawls identify as Qwantify; current documentation uses the Qwantbot name.

SeekportBot

Fully verifiable

SISTRIX · Search engines

The crawler for the Seekport search engine, operated by SISTRIX (Bonn, Germany). It fetches pages to build Seekport's search index and follows the Disallow directives in robots.txt.

SemanticScholarBot

Listed only

Allen Institute for AI (Ai2) · Search engines

SemanticScholarBot is the crawler for Semantic Scholar, an academic search engine built by the Allen Institute for AI. It crawls the web to find scholarly PDFs and metadata for indexing.

SeznamBot

Verifiable

Seznam.cz · Search engines

Seznam's web crawler, fetching pages for the Seznam.cz search engine used primarily in the Czech Republic.

Sogou Spider

Listed only

Sogou · Search engines

Sogou Spider is the web crawler for Sogou, a major Chinese search engine (owned by Tencent). It crawls and indexes web pages, news, and images to build Sogou's search index.

stepstoneCrawlBot

Listed only

The Stepstone Group · Search engines

The Stepstone Group's job-listing crawler. Its crawler page states that it processes only publicly available information, observes robots.txt directives and keeps intervals between requests to avoid loading servers, and publishes the user agent it sends for transparency.

StractBot

Listed only

Stract · Search engines

The crawler for Stract, an open-source search engine. Stract documents the crawler's user agent and robots.txt behavior but publishes no IP list or reverse-DNS verification method.

TinEye

Listed only

TinEye · Search engines

TinEye-bot is the web crawler operated by TinEye, the reverse image search engine built by Idee Inc. It crawls the web to discover and index images for TinEye's image-matching search service.

TrovitBot

Listed only

Trovit · Search engines

Trovit's web crawler, which discovers new and updated pages to add to the Trovit classifieds search index. Trovit documents that it does not fetch most sites more than once per second and that it can be blocked with the trovitBot robots.txt token.

Webzio

Listed only

Webz.io Ltd. · Search engines

Crawler for Webz.io (formerly Webhose.io), a web-data provider that collects publicly available content to build structured data feeds. Webz.io publishes no IP ranges, so requests cannot be verified.

Yahoo! Slurp

Listed only

Yahoo · Search engines

Slurp is Yahoo Search's web crawler. It crawls and indexes pages for Yahoo Search results and also gathers content for Yahoo News, Finance, and Sports.

Yahoo! JAPAN Crawler

Listed only

LY Corporation (Yahoo! JAPAN) · Search engines

Web crawler operated by Yahoo! JAPAN (LY Corporation) to collect pages for its Japanese search index and related services. It identifies itself with Y!J-prefixed user-agent tokens and follows the Robots Exclusion Protocol.

YandexBlogs

Verifiable

Yandex · Search engines

Yandex's blog-search robot. Yandex's robot table documents it as the blog search robot that indexes post comments, and lists it as following robots.txt directives.

YandexRenderResourcesBot

Verifiable

Yandex · Search engines

Yandex robot that loads resources such as JavaScript and CSS needed to render pages during crawling for Yandex Search.

YandexFavicons

Verifiable

Yandex · Search engines

Yandex's favicon fetcher. Yandex's robot table documents it as downloading a site's favicon file for display in search results, and lists it as one of the Yandex robots that do not follow robots.txt directives.

YandexImages

Verifiable

Yandex · Search engines

Yandex's image-indexing robot, which fetches images so they can be displayed in Yandex Images. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.

YandexMedia

Verifiable

Yandex · Search engines

Yandex's multimedia-indexing robot. Yandex's robot table documents it as a separate robot from YandexBot, describes it as indexing multimedia data, and lists it as following robots.txt directives.

YandexVideo

Verifiable

Yandex · Search engines

Yandex's video-indexing robot, which fetches pages and video content for display in Yandex video search. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.

YandexBot

Verifiable

Yandex · Search engines

Yandex's primary web crawler, fetching pages for Yandex Search indexing.