Directory / Search engines
Search engines 59
AddSearchBot
VerifiableAddSearch · Search engines
Crawler for AddSearch's hosted site-search service. It indexes the pages of customer sites so AddSearch can serve search results for them, obeys robots.txt rules written for the AddSearchBot token, and AddSearch publishes the fixed addresses it crawls from for allowlisting.
AdIdxBot
VerifiableMicrosoft · Search engines
Microsoft's crawler for Bing Ads / Microsoft Advertising. It crawls ads and follows through to the advertised landing pages for quality control, with both desktop and mobile variants.
Algolia Crawler
VerifiableAlgolia · Search engines
Algolia's Crawler, which visits a customer's pages, extracts search-relevant content, and pushes it to Algolia search indices. Algolia documents a fixed User-Agent token and a single static egress IP address for allowlisting.
Amazon AdBot
VerifiableAmazon · Search engines
Amazon's advertising crawler. It scans web pages that request ads from Amazon's advertising systems, collecting page content for Amazon's classification systems to maintain brand safety and improve ad relevance.
AmazonProductDiscoverybot
Listed onlyAmazon · Search engines
Amazon's product crawler, which collects publicly available product details from Amazon selling partner, brand and retailer websites to improve the accuracy and completeness of product information in Amazon's store. Amazon documents that it honours robots.txt user-agent and disallow directives, with changes taking up to 24 hours to apply, and that it does not support crawl-delay, nofollow or noindex.
Amzn-SearchBot
Listed onlyAmazon · Search engines
Amazon's crawler used to improve search experiences in Amazon products and services; Amazon states it does not crawl content for generative AI model training. Amazon publishes a human-readable IP list but no machine-parseable feed in a format this project's schema supports.
Anomura
VerifiableDireqt · Search engines
Direqt's search crawler, which discovers links and metadata to surface in Direqt's search features. The operator states it is not used to crawl content for model training, documents the robots.txt token Anomura, and publishes the two addresses the crawler requests come from.
Applebot
Fully verifiableApple · Search engines
Apple's web crawler, used to index content for Siri, Spotlight Suggestions, and Safari's search features.
Baiduspider
VerifiableBaidu · Search engines
Baidu's primary web crawler, fetching pages for Baidu Search indexing.
BingVideoPreview
VerifiableMicrosoft · Search engines
Microsoft crawler, listed among Bing's crawlers, that fetches pages to generate video previews shown in Bing. It runs desktop and mobile variants and is verifiable through the same reverse-DNS check as Bing's other crawlers.
Bingbot
Fully verifiableMicrosoft · Search engines
Microsoft's primary web crawler, fetching pages for Bing Search indexing.
BingPreview
VerifiableMicrosoft · Search engines
Microsoft fetcher used to generate page snapshots for previews in Bing search results. Runs both desktop and mobile variants.
Bublup Bot
VerifiableBublup · Search engines
Bublup's content-discovery bot. It fetches pages to build the database behind Bublup's suggestion engine, reading page title, description and related images. Bublup states its crawling IPs are not fixed and documents reverse-DNS verification against the bublup.com domain instead.
Channel3Bot
VerifiableChannel3 · Search engines
Channel3's product crawler. It visits publicly accessible product detail pages and collects images, titles, descriptions, prices, availability and variants for Channel3's product catalogue, which AI apps and agents query to route shoppers back to the original site. The operator tells site owners not to hard-code its addresses and to verify it by forward-confirmed reverse DNS instead.
coccocbot
VerifiableCốc Cốc · Search engines
Cốc Cốc's web crawler, fetching pages for the Vietnamese Cốc Cốc search engine. Separate web and image sub-bots share the same verification domain.
deepnoc
Listed onlydeepnoc GmbH · Search engines
Research crawler operated by deepnoc GmbH that parses and stores public web page content so that new search engines can query it without running their own crawling infrastructure. No IP ranges are published.
DuckDuckBot
Fully verifiableDuckDuckGo · Search engines
DuckDuckGo's web crawler, used to improve DuckDuckGo's search results.
ExaSearchBot
Fully verifiableExa · Search engines
Exa's search crawler, which discovers and indexes pages on the public web so that people and applications can find, retrieve and cite them through Exa. Every request is signed with HTTP Message Signatures (RFC 9421) under the Web Bot Auth scheme, so a request can be verified cryptographically regardless of its source IP address.
FreespokeCrawler
Listed onlyFreespoke · Search engines
Freespoke's search crawler, which fetches pages to build the index behind the Freespoke search engine. The operator documents the crawler's user agent, the robots.txt token FreespokeCrawler and support for the Crawl-delay directive.
Geedo Product Search
Fully verifiableGeedo · Search engines
GeedoShopProductFinder is the crawler for Geedo, a product-search engine that indexes publicly accessible product pages from online stores and follows robots.txt rules.
Storebot-Google
Fully verifiableGoogle · Search engines
Google's crawler for Google Shopping surfaces, including the Shopping tab in Google Search and shopping.google.com. It crawls product and store pages to populate shopping results.
Google-InspectionTool
Fully verifiableGoogle · Search engines
Google crawler used by Search testing tools such as the URL Inspection tool in Search Console and the Rich Result Test. Fetches pages on demand to show how Google Search renders them; it has no effect on ranking.
Google special-case crawlers
Fully verifiableGoogle · Search engines
A group of Google crawlers that operate under separate agreements between the crawled site and a specific Google product (AdsBot, AdSense, Google Safety scanning, and similar). Several members of this group ignore the wildcard robots.txt rule and must be targeted explicitly.
Googlebot
Fully verifiableGoogle · Search engines
Google's primary web crawler, fetching pages for Google Search indexing.
Googlebot Image
Fully verifiableGoogle · Search engines
Google's image crawler, which fetches images for Google Images and image features in Search. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.
Googlebot Video
Fully verifiableGoogle · Search engines
Google's video crawler, which fetches video content for Google Search video features. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.
GoogleOther
Fully verifiableGoogle · Search engines
Google's generic crawler family (GoogleOther, GoogleOther-Image, GoogleOther-Video), used by various Google product teams for fetching publicly accessible content. Robots.txt rules addressed to it do not affect Google Search.
360Spider
Listed only360 Search (Qihoo 360) · Search engines
360Spider is the web-search crawler of 360 Search (so.com), the search engine operated by Chinese internet company Qihoo 360. The operator also runs 360Spider-Image and 360Spider-Video for image and video search.
IONOS Crawler
VerifiableIONOS · Search engines
IONOS Crawler (IonCrawl) continuously crawls publicly accessible domains to generate insights into how they are used, which IONOS uses to improve and expand its hosting products. IONOS documents reverse-DNS verification and deletes crawled data after 60 days.
Jooblebot
Listed onlyJooble · Search engines
The crawler for Jooble, a job-search aggregator that indexes job listings published across the web. Jooble's bot page states that it uses a web crawler which identifies itself as JoobleBot. No IP ranges are published.
Kagibot
VerifiableKagi · Search engines
Kagi's web crawler, fetching pages for the paid Kagi search engine. Kagi publishes a fixed, small set of source IPs rather than a CIDR feed, each paired with a kagibot.org hostname.
Linespider
Listed onlyLINE · Search engines
LINE's crawler. It fetches web pages to support LINE's search features and content services, and adheres to the Robots Exclusion Protocol outlined in robots.txt.
Marginalia
Fully verifiableMarginalia Search · Search engines
The crawler behind Marginalia Search, an independent search engine focused on older, text-heavy, and non-commercial web pages. Identifies itself by a bare URL rather than a conventional bot token.
MetaJobBot
VerifiableMETAJob · Search engines
MetaJobBot is the focused web crawler operated by the METAJob job meta-search engine. It searches websites for job listings to include in the METAJob index.
MojeekBot
Fully verifiableMojeek · Search engines
Mojeek's web crawler, fetching pages for the independent Mojeek search engine index.
Yeti
VerifiableNaver · Search engines
Naver's web crawler, fetching pages for Naver Search indexing across South Korea's largest search engine.
Openindex Spider
Listed onlyOpenindex · Search engines
Openindex operates a web-crawling cluster (Apache Nutch on an Apache Hadoop cluster) for research and development of universal and focused search engines. The crawler respects robots.txt and the Crawl-delay directive and identifies itself in the User-Agent.
PetalBot
VerifiableHuawei · Search engines
Huawei's web crawler, fetching pages for the Petal Search engine used in Huawei's mobile services and for Huawei Assistant content recommendations. Huawei documents reverse-DNS verification against aspiegel.com and petalsearch.com hostnames.
PiplBot
Listed onlyPipl · Search engines
PiplBot is the web-indexing crawler operated by Pipl, an identity and people-search service. It retrieves and indexes publicly available information, including content behind searchable databases, to build Pipl's people-search index.
Quantcastbot
VerifiableQuantcast · Search engines
Quantcast's advertising crawler. It crawls websites to extract page content for interest-based audience categorization and to run quality-assurance checks on advertisement landing pages.
Qwantbot
Fully verifiableQwant · Search engines
Qwant's web crawler, fetching pages for the privacy-focused Qwant search engine. Older crawls identify as Qwantify; current documentation uses the Qwantbot name.
SeekportBot
Fully verifiableSISTRIX · Search engines
The crawler for the Seekport search engine, operated by SISTRIX (Bonn, Germany). It fetches pages to build Seekport's search index and follows the Disallow directives in robots.txt.
SemanticScholarBot
Listed onlyAllen Institute for AI (Ai2) · Search engines
SemanticScholarBot is the crawler for Semantic Scholar, an academic search engine built by the Allen Institute for AI. It crawls the web to find scholarly PDFs and metadata for indexing.
SeznamBot
VerifiableSeznam.cz · Search engines
Seznam's web crawler, fetching pages for the Seznam.cz search engine used primarily in the Czech Republic.
Sogou Spider
Listed onlySogou · Search engines
Sogou Spider is the web crawler for Sogou, a major Chinese search engine (owned by Tencent). It crawls and indexes web pages, news, and images to build Sogou's search index.
stepstoneCrawlBot
Listed onlyThe Stepstone Group · Search engines
The Stepstone Group's job-listing crawler. Its crawler page states that it processes only publicly available information, observes robots.txt directives and keeps intervals between requests to avoid loading servers, and publishes the user agent it sends for transparency.
StractBot
Listed onlyStract · Search engines
The crawler for Stract, an open-source search engine. Stract documents the crawler's user agent and robots.txt behavior but publishes no IP list or reverse-DNS verification method.
TinEye
Listed onlyTinEye · Search engines
TinEye-bot is the web crawler operated by TinEye, the reverse image search engine built by Idee Inc. It crawls the web to discover and index images for TinEye's image-matching search service.
TrovitBot
Listed onlyTrovit · Search engines
Trovit's web crawler, which discovers new and updated pages to add to the Trovit classifieds search index. Trovit documents that it does not fetch most sites more than once per second and that it can be blocked with the trovitBot robots.txt token.
Webzio
Listed onlyWebz.io Ltd. · Search engines
Crawler for Webz.io (formerly Webhose.io), a web-data provider that collects publicly available content to build structured data feeds. Webz.io publishes no IP ranges, so requests cannot be verified.
Yahoo! Slurp
Listed onlyYahoo · Search engines
Slurp is Yahoo Search's web crawler. It crawls and indexes pages for Yahoo Search results and also gathers content for Yahoo News, Finance, and Sports.
Yahoo! JAPAN Crawler
Listed onlyLY Corporation (Yahoo! JAPAN) · Search engines
Web crawler operated by Yahoo! JAPAN (LY Corporation) to collect pages for its Japanese search index and related services. It identifies itself with Y!J-prefixed user-agent tokens and follows the Robots Exclusion Protocol.
YandexBlogs
VerifiableYandex · Search engines
Yandex's blog-search robot. Yandex's robot table documents it as the blog search robot that indexes post comments, and lists it as following robots.txt directives.
YandexRenderResourcesBot
VerifiableYandex · Search engines
Yandex robot that loads resources such as JavaScript and CSS needed to render pages during crawling for Yandex Search.
YandexFavicons
VerifiableYandex · Search engines
Yandex's favicon fetcher. Yandex's robot table documents it as downloading a site's favicon file for display in search results, and lists it as one of the Yandex robots that do not follow robots.txt directives.
YandexImages
VerifiableYandex · Search engines
Yandex's image-indexing robot, which fetches images so they can be displayed in Yandex Images. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.
YandexMedia
VerifiableYandex · Search engines
Yandex's multimedia-indexing robot. Yandex's robot table documents it as a separate robot from YandexBot, describes it as indexing multimedia data, and lists it as following robots.txt directives.
YandexVideo
VerifiableYandex · Search engines
Yandex's video-indexing robot, which fetches pages and video content for display in Yandex video search. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.
YandexBot
VerifiableYandex · Search engines
Yandex's primary web crawler, fetching pages for Yandex Search indexing.