The web's bots,
verified.
The open, categorized directory of legitimate bots — their user agents, published IP ranges, and exactly how to verify them.
Search engines 53
AddSearchBot
VerifiableAddSearch · Search engines
Crawler for AddSearch's hosted site-search service. It indexes the pages of customer sites so AddSearch can serve search results for them, obeys robots.txt rules written for the AddSearchBot token, and AddSearch publishes the fixed addresses it crawls from for allowlisting.
AdIdxBot
VerifiableMicrosoft · Search engines
Microsoft's crawler for Bing Ads / Microsoft Advertising. It crawls ads and follows through to the advertised landing pages for quality control, with both desktop and mobile variants.
Algolia Crawler
VerifiableAlgolia · Search engines
Algolia's Crawler, which visits a customer's pages, extracts search-relevant content, and pushes it to Algolia search indices. Algolia documents a fixed User-Agent token and a single static egress IP address for allowlisting.
Amazon AdBot
VerifiableAmazon · Search engines
Amazon's advertising crawler. It scans web pages that request ads from Amazon's advertising systems, collecting page content for Amazon's classification systems to maintain brand safety and improve ad relevance.
Amzn-SearchBot
Listed onlyAmazon · Search engines
Amazon's crawler used to improve search experiences in Amazon products and services; Amazon states it does not crawl content for generative AI model training. Amazon publishes a human-readable IP list but no machine-parseable feed in a format this project's schema supports.
Applebot
Fully verifiableApple · Search engines
Apple's web crawler, used to index content for Siri, Spotlight Suggestions, and Safari's search features.
Baiduspider
VerifiableBaidu · Search engines
Baidu's primary web crawler, fetching pages for Baidu Search indexing.
BingVideoPreview
VerifiableMicrosoft · Search engines
Microsoft crawler, listed among Bing's crawlers, that fetches pages to generate video previews shown in Bing. It runs desktop and mobile variants and is verifiable through the same reverse-DNS check as Bing's other crawlers.
Bingbot
Fully verifiableMicrosoft · Search engines
Microsoft's primary web crawler, fetching pages for Bing Search indexing.
BingPreview
VerifiableMicrosoft · Search engines
Microsoft fetcher used to generate page snapshots for previews in Bing search results. Runs both desktop and mobile variants.
Bublup Bot
VerifiableBublup · Search engines
Bublup's content-discovery bot. It fetches pages to build the database behind Bublup's suggestion engine, reading page title, description and related images. Bublup states its crawling IPs are not fixed and documents reverse-DNS verification against the bublup.com domain instead.
coccocbot
VerifiableCốc Cốc · Search engines
Cốc Cốc's web crawler, fetching pages for the Vietnamese Cốc Cốc search engine. Separate web and image sub-bots share the same verification domain.
deepnoc
Listed onlydeepnoc GmbH · Search engines
Research crawler operated by deepnoc GmbH that parses and stores public web page content so that new search engines can query it without running their own crawling infrastructure. No IP ranges are published.
DuckDuckBot
Fully verifiableDuckDuckGo · Search engines
DuckDuckGo's web crawler, used to improve DuckDuckGo's search results.
Geedo Product Search
Fully verifiableGeedo · Search engines
GeedoShopProductFinder is the crawler for Geedo, a product-search engine that indexes publicly accessible product pages from online stores and follows robots.txt rules.
Storebot-Google
Fully verifiableGoogle · Search engines
Google's crawler for Google Shopping surfaces, including the Shopping tab in Google Search and shopping.google.com. It crawls product and store pages to populate shopping results.
Google-InspectionTool
Fully verifiableGoogle · Search engines
Google crawler used by Search testing tools such as the URL Inspection tool in Search Console and the Rich Result Test. Fetches pages on demand to show how Google Search renders them; it has no effect on ranking.
Google special-case crawlers
Fully verifiableGoogle · Search engines
A group of Google crawlers that operate under separate agreements between the crawled site and a specific Google product (AdsBot, AdSense, Google Safety scanning, and similar). Several members of this group ignore the wildcard robots.txt rule and must be targeted explicitly.
Googlebot
Fully verifiableGoogle · Search engines
Google's primary web crawler, fetching pages for Google Search indexing.
Googlebot Image
Fully verifiableGoogle · Search engines
Google's image crawler, which fetches images for Google Images and image features in Search. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.
Googlebot Video
Fully verifiableGoogle · Search engines
Google's video crawler, which fetches video content for Google Search video features. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.
GoogleOther
Fully verifiableGoogle · Search engines
Google's generic crawler family (GoogleOther, GoogleOther-Image, GoogleOther-Video), used by various Google product teams for fetching publicly accessible content. Robots.txt rules addressed to it do not affect Google Search.
360Spider
Listed only360 Search (Qihoo 360) · Search engines
360Spider is the web-search crawler of 360 Search (so.com), the search engine operated by Chinese internet company Qihoo 360. The operator also runs 360Spider-Image and 360Spider-Video for image and video search.
IONOS Crawler
VerifiableIONOS · Search engines
IONOS Crawler (IonCrawl) continuously crawls publicly accessible domains to generate insights into how they are used, which IONOS uses to improve and expand its hosting products. IONOS documents reverse-DNS verification and deletes crawled data after 60 days.
Jooblebot
Listed onlyJooble · Search engines
The crawler for Jooble, a job-search aggregator that indexes job listings published across the web. Jooble's bot page states that it uses a web crawler which identifies itself as JoobleBot. No IP ranges are published.
Kagibot
VerifiableKagi · Search engines
Kagi's web crawler, fetching pages for the paid Kagi search engine. Kagi publishes a fixed, small set of source IPs rather than a CIDR feed, each paired with a kagibot.org hostname.
Linespider
Listed onlyLINE · Search engines
LINE's crawler. It fetches web pages to support LINE's search features and content services, and adheres to the Robots Exclusion Protocol outlined in robots.txt.
Marginalia
Fully verifiableMarginalia Search · Search engines
The crawler behind Marginalia Search, an independent search engine focused on older, text-heavy, and non-commercial web pages. Identifies itself by a bare URL rather than a conventional bot token.
MetaJobBot
VerifiableMETAJob · Search engines
MetaJobBot is the focused web crawler operated by the METAJob job meta-search engine. It searches websites for job listings to include in the METAJob index.
MojeekBot
Fully verifiableMojeek · Search engines
Mojeek's web crawler, fetching pages for the independent Mojeek search engine index.
Yeti
VerifiableNaver · Search engines
Naver's web crawler, fetching pages for Naver Search indexing across South Korea's largest search engine.
Openindex Spider
Listed onlyOpenindex · Search engines
Openindex operates a web-crawling cluster (Apache Nutch on an Apache Hadoop cluster) for research and development of universal and focused search engines. The crawler respects robots.txt and the Crawl-delay directive and identifies itself in the User-Agent.
PetalBot
VerifiableHuawei · Search engines
Huawei's web crawler, fetching pages for the Petal Search engine used in Huawei's mobile services and for Huawei Assistant content recommendations. Huawei documents reverse-DNS verification against aspiegel.com and petalsearch.com hostnames.
PiplBot
Listed onlyPipl · Search engines
PiplBot is the web-indexing crawler operated by Pipl, an identity and people-search service. It retrieves and indexes publicly available information, including content behind searchable databases, to build Pipl's people-search index.
Quantcastbot
VerifiableQuantcast · Search engines
Quantcast's advertising crawler. It crawls websites to extract page content for interest-based audience categorization and to run quality-assurance checks on advertisement landing pages.
Qwantbot
Fully verifiableQwant · Search engines
Qwant's web crawler, fetching pages for the privacy-focused Qwant search engine. Older crawls identify as Qwantify; current documentation uses the Qwantbot name.
SeekportBot
Fully verifiableSISTRIX · Search engines
The crawler for the Seekport search engine, operated by SISTRIX (Bonn, Germany). It fetches pages to build Seekport's search index and follows the Disallow directives in robots.txt.
SemanticScholarBot
Listed onlyAllen Institute for AI (Ai2) · Search engines
SemanticScholarBot is the crawler for Semantic Scholar, an academic search engine built by the Allen Institute for AI. It crawls the web to find scholarly PDFs and metadata for indexing.
SeznamBot
VerifiableSeznam.cz · Search engines
Seznam's web crawler, fetching pages for the Seznam.cz search engine used primarily in the Czech Republic.
Sogou Spider
Listed onlySogou · Search engines
Sogou Spider is the web crawler for Sogou, a major Chinese search engine (owned by Tencent). It crawls and indexes web pages, news, and images to build Sogou's search index.
StractBot
Listed onlyStract · Search engines
The crawler for Stract, an open-source search engine. Stract documents the crawler's user agent and robots.txt behavior but publishes no IP list or reverse-DNS verification method.
TinEye
Listed onlyTinEye · Search engines
TinEye-bot is the web crawler operated by TinEye, the reverse image search engine built by Idee Inc. It crawls the web to discover and index images for TinEye's image-matching search service.
TrovitBot
Listed onlyTrovit · Search engines
Trovit's web crawler, which discovers new and updated pages to add to the Trovit classifieds search index. Trovit documents that it does not fetch most sites more than once per second and that it can be blocked with the trovitBot robots.txt token.
Webzio
Listed onlyWebz.io Ltd. · Search engines
Crawler for Webz.io (formerly Webhose.io), a web-data provider that collects publicly available content to build structured data feeds. Webz.io publishes no IP ranges, so requests cannot be verified.
Yahoo! Slurp
Listed onlyYahoo · Search engines
Slurp is Yahoo Search's web crawler. It crawls and indexes pages for Yahoo Search results and also gathers content for Yahoo News, Finance, and Sports.
Yahoo! JAPAN Crawler
Listed onlyLY Corporation (Yahoo! JAPAN) · Search engines
Web crawler operated by Yahoo! JAPAN (LY Corporation) to collect pages for its Japanese search index and related services. It identifies itself with Y!J-prefixed user-agent tokens and follows the Robots Exclusion Protocol.
YandexBlogs
VerifiableYandex · Search engines
Yandex's blog-search robot. Yandex's robot table documents it as the blog search robot that indexes post comments, and lists it as following robots.txt directives.
YandexRenderResourcesBot
VerifiableYandex · Search engines
Yandex robot that loads resources such as JavaScript and CSS needed to render pages during crawling for Yandex Search.
YandexFavicons
VerifiableYandex · Search engines
Yandex's favicon fetcher. Yandex's robot table documents it as downloading a site's favicon file for display in search results, and lists it as one of the Yandex robots that do not follow robots.txt directives.
YandexImages
VerifiableYandex · Search engines
Yandex's image-indexing robot, which fetches images so they can be displayed in Yandex Images. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.
YandexMedia
VerifiableYandex · Search engines
Yandex's multimedia-indexing robot. Yandex's robot table documents it as a separate robot from YandexBot, describes it as indexing multimedia data, and lists it as following robots.txt directives.
YandexVideo
VerifiableYandex · Search engines
Yandex's video-indexing robot, which fetches pages and video content for display in Yandex video search. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.
YandexBot
VerifiableYandex · Search engines
Yandex's primary web crawler, fetching pages for Yandex Search indexing.
AI crawlers 23
AI2Bot
Listed onlyAllen Institute for AI · AI crawlers
The Allen Institute for AI's crawler, used to build open training datasets such as Dolma. AI2 publishes no IP ranges, ASN, or reverse-DNS pattern.
Amazonbot
Listed onlyAmazon · AI crawlers
Amazon's web crawler used to improve Amazon's products and services, including training AI models. Amazon publishes a human-readable IP list but no machine-parseable feed in a format this project's schema supports.
Bytespider
Listed onlyByteDance · AI crawlers
ByteDance's web crawler, used to gather training data for its AI models. ByteDance publishes no IP ranges, ASN, or reverse-DNS pattern for it.
CCBot
Fully verifiableCommon Crawl Foundation · AI crawlers
Common Crawl's crawler that builds the freely available Common Crawl web archive, widely reused as AI training data by third parties.
Claude-SearchBot
Fully verifiableAnthropic · AI crawlers
Anthropic's crawler that indexes pages to improve search-style answer quality in Claude products, distinct from ClaudeBot's training-data crawl.
ClaudeBot
Fully verifiableAnthropic · AI crawlers
Anthropic's web crawler collecting publicly available web data for training Claude models. Anthropic publishes its crawler source IPs as a machine-readable feed.
Cloudflare AI Search
Listed onlyCloudflare · AI crawlers
The crawler behind Cloudflare AI Search, which indexes website content so it can be searched. Cloudflare documents that it only crawls a website the customer owns — the domain must exist in the same Cloudflare account and be selected as an AI Search data source.
Cloudflare Browser Run Crawler
Fully verifiableCloudflare · AI crawlers
The crawler behind the /crawl endpoint of Cloudflare's Browser Run (Browser Rendering) developer product, which crawls third-party websites on behalf of Cloudflare customers building applications on the platform. Its user agent is not configurable, and every request is signed with Web Bot Auth HTTP message signatures that site owners can verify against Cloudflare's published key directory.
Diffbot
Listed onlyDiffbot · AI crawlers
Diffbot's general-purpose web crawler, used to build its Knowledge Graph and power its structured-extraction APIs. Diffbot documents robots.txt behavior but publishes no IP list, ASN, or reverse-DNS pattern.
Google-CloudVertexBot
Fully verifiableGoogle · AI crawlers
Google Cloud crawler that fetches sites on the site owners' request when building Vertex AI Agents. Crawling preferences addressed to its user agent have no effect on Google Search or other Google products.
GPTBot
Fully verifiableOpenAI · AI crawlers
OpenAI's web crawler that gathers publicly available data used to train OpenAI's models.
ImagesiftBot
Listed onlyHive · AI crawlers
ImageSift's crawler, operated by Hive. It scrapes publicly available images across the web to support Hive's ImageSift reverse-image-search and web intelligence products. Standard robots.txt directives are respected.
Leipzig Corpora Collection Crawler
Listed onlyLeipzig University Natural Language Processing Group · AI crawlers
The LCC crawler is operated by Leipzig University to collect web text for the Leipzig Corpora Collection, a set of linguistic corpora used in natural language processing research.
MaCoCu
Listed onlyMaCoCu project (Jožef Stefan Institute) · AI crawlers
Crawler for the CEF-funded MaCoCu project, which collects, curates and enriches monolingual and parallel web text to build language corpora for under-resourced languages. The project publishes no IP ranges, so requests cannot be verified.
Meta External Agent
VerifiableMeta · AI crawlers
Meta's crawler for AI training data and content indexing across Meta products. Verified by ASN lookup (AS32934); Meta publishes no IP feed.
Meta Web Indexer
VerifiableMeta · AI crawlers
Meta's crawler that navigates the web to improve the quality of Meta AI search results, analysing page content for relevance and accuracy in Meta AI responses. Verified by ASN lookup (AS32934); Meta publishes no IP feed.
ICC-Crawler
VerifiableNational Institute of Information and Communications Technology (NICT) · AI crawlers
ICC-Crawler is a web crawler operated by Japan's NICT that collects web pages across the internet to build datasets for information and language processing research.
OAI-SearchBot
Fully verifiableOpenAI · AI crawlers
OpenAI's crawler that indexes pages to power search results inside ChatGPT. It is distinct from GPTBot and is not used to gather training data.
omgili
Listed onlyWebz.io · AI crawlers
Webz.io's legacy crawler user agent (formerly "Omgilibot"), used to collect web, forum, and news content for its data feeds. Webz.io's current public documentation describes successor crawlers ("webzio" / "webzio-extended") and no longer documents this UA string or any IP verification method for it.
PerplexityBot
Fully verifiablePerplexity · AI crawlers
Perplexity's crawler that indexes pages to surface and link websites in Perplexity search results. Perplexity states it is not used to train models.
SBIntuitionsBot
Listed onlySB Intuitions Corp. · AI crawlers
Crawler operated by SB Intuitions (a SoftBank AI subsidiary) that collects web pages for AI development and information analysis, including training of its Sarashina language models. SB Intuitions publishes no IP ranges, so requests cannot be verified.
Webzio-extended
Listed onlyWebz.io Ltd. · AI crawlers
Second crawler in Webz.io's crawler pair, which performs ethical validation on the data collected by Webzio and tags it as usable or not usable for AI and machine-learning training. Webz.io publishes no IP ranges.
YouBot
Fully verifiableYou.com · AI crawlers
You.com's crawler that indexes pages for its AI-powered search product. You.com documents a dedicated IP range, a reverse-DNS pattern, and support for signed-request verification via Web Bot Auth.
AI assistants 9
Amzn-User
Listed onlyAmazon · AI assistants
Amazon's on-demand fetcher supporting user actions, such as responding to Alexa queries that need up-to-date information; Amazon states it does not crawl content for generative AI model training. Amazon publishes a human-readable live-crawl IP list but no machine-parseable feed.
ChatGPT-User
Fully verifiableOpenAI · AI assistants
OpenAI's user-triggered fetcher, used when a ChatGPT user or a GPT Action asks the assistant to visit a specific page. It does not crawl autonomously.
Claude-User
Fully verifiableAnthropic · AI assistants
Anthropic's user-triggered fetcher, used when a person asks Claude to visit a specific web page (e.g. via tool use in a conversation).
DuckAssistBot
Fully verifiableDuckDuckGo · AI assistants
DuckDuckGo's real-time fetcher for DuckDuckGo Search's AI-assisted answers, which cite their sources. DuckDuckGo states the data is not used to train AI models, and publishes a machine-readable list of the bot's IP addresses.
Google-Agent
Fully verifiableGoogle · AI assistants
The fetcher used by AI agents hosted on Google infrastructure to navigate the web and perform actions on behalf of a user who asked for them. Google publishes a dedicated IP range list for it and signs a subset of its requests with Web Bot Auth under the agent.bot.goog identity. As a user-triggered fetcher it generally ignores robots.txt rules.
Google user-triggered fetchers
Fully verifiableGoogle · AI assistants
Google tools and product features that fetch a specific page because an end user asked for it (Google Read Aloud, Site Verifier, Gemini Notebook, Chrome Web Store, Google Messages, Pinpoint, Publisher Center, and similar), rather than autonomous crawling for search indexing. They generally ignore robots.txt because a human requested the fetch.
Meta External Fetcher
VerifiableMeta · AI assistants
Meta's on-demand fetcher that retrieves a single link at a user's request to support agentic AI features (e.g. an AI assistant navigating a page a user asked about), rather than broad indexing. Verified by ASN lookup (AS32934).
MistralAI-User
Fully verifiableMistral AI · AI assistants
Mistral AI's user-triggered fetcher, used when a user of a Mistral product (e.g. Le Chat) asks it to visit a specific web page to answer a question.
Perplexity-User
Fully verifiablePerplexity · AI assistants
Perplexity's user-triggered fetcher, used when a user's question requires visiting a specific web page to produce an accurate answer.
Social previews 22
Discordbot
Listed onlyDiscord · Social previews
Discord's crawler that fetches shared links to render embeds in chat. Discord's own docs only specify the User-Agent format required of API clients, not this embed fetcher; no IP ranges, ASN, or rDNS are published.
Embedly
VerifiableEmbedly · Social previews
Embedly's fetcher, used to generate rich embeds and previews of shared URLs for its customers' apps and sites. Embedly's FAQ publishes a fixed, small set of source IPs rather than a live CIDR feed.
Facebook External Hit
VerifiableMeta · Social previews
Meta's crawler that fetches shared links from Facebook, Instagram, and Messenger to generate rich link previews. Verified by ASN lookup (AS32934); Meta's own docs describe IP-based allowlisting but publish no IP feed.
Hatena service fetchers
Listed onlyHatena Co., Ltd. · Social previews
The fetchers Hatena's services use to collect information from pages its users link to. Hatena-Favicon identifies a page's favicon, Hatena::Scissors retrieves image thumbnails, HatenaBookmark fetches article information for Hatena Bookmark, Hatena Star associates stars with pages, and Hatena Antenna fetches update differences for the pages a user is monitoring.
HatenaBlog-bot
Listed onlyHatena Co., Ltd. · Social previews
Hatena Blog's fetcher. It retrieves a linked page when a blog author asks for it: the :title option of Hatena's URL notation pulls the page title, the :embed option collects the title, summary and favicon used to render a blog card, and the blog import feature fetches images so they can be re-uploaded to Hatena Fotolife.
Iframely
Fully verifiableIframely · Social previews
Iframely's fetcher, used to generate embeds and link previews of shared URLs for its customers' apps and sites. Iframely publishes live IPv4/IPv6 address lists and documents reverse-DNS resolution to its own domain.
LinkedInBot
Listed onlyLinkedIn · Social previews
LinkedIn's crawler that fetches shared links to build post link previews. LinkedIn's own robots.txt names the "LinkedInBot" token, but LinkedIn publishes no dedicated bot-documentation page, IP list, ASN, or reverse-DNS pattern for this crawler.
LivelapBot
Listed onlyLivelap · Social previews
LivelapBot is the crawler for Livelap, a content discovery app. It fetches pages shared on social media and crawls RSS feeds on a schedule, indexing HTML and media meta tags to build link previews shown in the Livelap app.
Mastodon (link preview fetcher)
Listed onlyMastodon gGmbH · Social previews
The generic link-preview fetcher built into Mastodon server software (FetchLinkCardService, using the http.rb gem), run independently by every self-hosted instance. There is no single operator, IP range, or ASN to verify against — traffic can originate from any of thousands of instances.
MicrosoftPreview
Fully verifiableMicrosoft · Social previews
Microsoft fetcher, listed among Bing's crawlers, that retrieves pages to generate link previews for Microsoft products and services.
PagePeeker
Listed onlyPagePeeker SRL · Social previews
PagePeeker's thumbnailing robot. It captures a screenshot of a page when one of PagePeeker's customers requests a thumbnail for it, refetching a given site at most once every five to seven days.
Pinterestbot
VerifiablePinterest · Social previews
Pinterest's crawler, used both for general web indexing and for validating Pin metadata/broken links. Pinterest documents a reverse-DNS verification method rather than a fixed IP list, since only its US-based traffic uses a stable range.
Redditbot
Listed onlyReddit · Social previews
Reddit's on-demand fetcher that retrieves a shared URL's metadata to build the link preview shown on a post. Reddit's robots.txt now disallows all crawling and documents no user agent, IP list, or verification method for this fetcher.
Slack-ImgProxy
Listed onlySlack · Social previews
Slack's image-proxying fetcher, which retrieves and caches images posted in channels while stripping referrer data. Slack documents the User-Agent string but publishes no IP ranges or other verification method.
Slackbot (Link Expanding)
Listed onlySlack · Social previews
Slack's crawler that expands links posted in channels into rich previews by reading oEmbed, Twitter Card, and Open Graph metadata. Slack documents the User-Agent strings but publishes no IP ranges or other verification method.
Snap URL Preview Service
Listed onlySnap Inc. · Social previews
Snapchat's link-preview fetcher, which scans the HTML of URLs shared in Snapchat chats to build preview cards from Open Graph or Twitter Card tags. Snap documents that responses are cached for 30 minutes to reduce traffic; no IP ranges or robots.txt behavior are documented.
StartmeBot
Listed onlyStart.me · Social previews
Start.me's fetcher. When a user adds a link or an RSS widget to their Start.me start page, this bot retrieves three things from the target site: the page title, the site's favicon, and the contents of any RSS feed the site offers.
TelegramBot (link preview)
Listed onlyTelegram · Social previews
Telegram's fetcher for generating link previews in chats. Telegram's docs document static CIDR ranges only for inbound webhook delivery to bot servers, not for this outbound preview fetcher, so no IP verification is attached here.
Twitterbot
Listed onlyX · Social previews
X (formerly Twitter)'s crawler that fetches shared links to render Card previews. X's own developer docs give no IP ranges, ASN, or reverse-DNS verification method; the commonly cited "*.twttr.com" rDNS check and aggregate IP ranges trace only to community write-ups, not official docs.
vkShare
Listed onlyVK · Social previews
vkShare is VK's link-preview fetcher. When a user shares a URL on the VK social network, it retrieves the target page to build a preview (title, description, and image) for the shared post.
WhatsApp Link Preview Fetcher
VerifiableMeta · Social previews
WhatsApp's fetcher that requests a shared URL to build the link preview shown in chat. Meta documents the request's User-Agent format but no IP feed; verified by ASN lookup (AS32934), same as Meta's other crawlers.
Yahoo Link Preview
Listed onlyYahoo · Social previews
Yahoo Link Preview fetches a page when a Yahoo Mail user includes its URL in an email message, generating a thumbnail, title and description preview. It fetches only user-referenced pages rather than crawling the web.
Monitoring 37
AwarioBot
Listed onlyAwario · Monitoring
Awario's social-listening and media-monitoring crawlers, which fetch web pages and RSS feeds to find mentions of the brands and keywords its customers track. Awario states the bots respect robots.txt and honour Crawl-delay, and explicitly asks site owners not to identify them by IP because it uses no consecutive IP blocks.
Azure Application Insights Availability
Listed onlyMicrosoft · Monitoring
Synthetic availability-test agents from Microsoft Azure Application Insights that send recurring HTTP requests to user-configured URLs to monitor uptime and responsiveness from multiple Azure regions.
Better Stack Uptime
Fully verifiableBetter Stack · Monitoring
Better Stack's uptime-monitoring probes, which request customer URLs on a schedule from a published, frequently-changing list of monitoring IPs.
BrandVerity
Listed onlyBrandVerity · Monitoring
BrandVerity is a brand-protection service that crawls websites and paid search landing pages to detect trademark and affiliate-compliance violations on behalf of its clients.
magpie-crawler
Listed onlyBrandwatch · Monitoring
magpie-crawler is Brandwatch's web crawler. It downloads publicly available blog, forum, news, and social-media pages to be indexed and analysed by Brandwatch's social-listening platform. It respects robots.txt.
Checkly
Fully verifiableCheckly · Monitoring
Synthetic-monitoring service that runs API and browser checks against customer endpoints from a published, fixed set of static outbound IPs.
Cloudflare Diagnostics
Listed onlyCloudflare · Monitoring
Cloudflare's support-diagnostics fetcher. Cloudflare documents that requests with this user agent are triggered when Cloudflare Support Engineers perform error checks, and by the continuous monitoring that raises alerts in the Cloudflare dashboard.
Cloudflare Health Checks
Listed onlyCloudflare · Monitoring
Cloudflare's Health Checks probes, which monitor customer origin servers from Cloudflare data centers in regions the customer selects. Cloudflare documents a fixed User-Agent format embedding the first 16 characters of the health check ID, and recommends matching on it; its health-checks docs do not tie probe traffic to the published Cloudflare IP lists, so no CIDR recipe is attached.
Cloudflare Prefetch
Listed onlyCloudflare · Monitoring
Cloudflare's prefetch bot, an Enterprise speed-optimization feature that pre-populates the CDN cache with content a visitor is likely to request next, using a site-owner-supplied manifest of URLs. Requests carry the CloudFlare-Prefetch user-agent.
Cloudflare Traffic Manager
Listed onlyCloudflare · Monitoring
Cloudflare's Load Balancing monitor, which sends health-check requests to customer origin pools at regular intervals to evaluate endpoint health. It embeds the load-balancer pool id in a fixed User-Agent.
Cookiebot Scanner
VerifiableCookiebot (Usercentrics) · Monitoring
Cookiebot's cookie-consent compliance scanner. It crawls the domains its customers have registered, on a roughly monthly schedule, to detect the cookies and tracking technologies in use and generate a cookie declaration. Cookiebot documents that scans run only from a fixed pool of addresses.
CookieHub Scanner
Fully verifiableCookieHub · Monitoring
CookieHub's cookie-consent compliance scanner. It crawls customer domains on a roughly monthly schedule to detect cookies and tracking technologies in use and generate a compliance report.
Datadog Synthetics
Fully verifiableDatadog · Monitoring
Datadog Synthetic Monitoring probes that run customer-configured API and browser tests against endpoints. API tests send a "Datadog/Synthetics" User-Agent; browser tests append "DatadogSynthetics" to a browser UA. Datadog publishes probe IP ranges as a machine-readable feed at ip-ranges.datadoghq.com/synthetics.json.
Dead Link Checker
VerifiableDLC Websites · Monitoring
Dead Link Checker is a hosted service that crawls a submitted website to find broken links and report them, with paid tiers that run automated periodic scans and email the results.
Dubbotbot
VerifiableDubBot · Monitoring
The crawler behind DubBot's web-governance platform. It inventories a customer's own website by following every link from a supplied URL and checks the pages for accessibility, broken links, spelling and content-policy problems. DubBot runs it from AWS on static IP addresses that the operator publishes for allowlisting.
Dynatrace Synthetic Monitoring
Listed onlyDynatrace · Monitoring
Dynatrace's synthetic monitoring runs browser and HTTP checks against sites its customers configure. Dynatrace always appends a RuxitSynthetic token to the user agent — even when the customer sets a custom one — so that synthetic traffic can be identified in server logs. Checks are user-configured, not crawling, so robots.txt is not part of the documented behaviour.
FreeWebMonitoring SiteChecker
VerifiableGreenWave Online Inc. · Monitoring
The website-monitoring robot of the FreeWebMonitoring service. It only checks URLs that registered members have submitted, and the operator documents that all checks originate from a single server address. GreenWave Online also warns that an older 0.1 agent name is forged by an unrelated scanner.
Freshping
VerifiableFreshworks · Monitoring
Freshworks' free uptime-monitoring service, which checks customer sites from a small, fixed set of documented monitoring-location IPs rather than a live feed.
Gnowit Newsbot
Listed onlyGnowit · Monitoring
Gnowit is an Ottawa-based media and government monitoring company whose crawler continuously fetches news, government, and other public web sources to power real-time monitoring, summarization, and analytics for its clients.
Grafana Synthetic Monitoring
Fully verifiableGrafana Labs · Monitoring
Grafana Cloud Synthetic Monitoring public probes, which run customer-configured checks (HTTP, browser, and protocol) from Grafana-run locations. HTTP checks send a documented synthetic-monitoring-agent User-Agent by default (customers can override it). Grafana publishes probe source IPs as a machine-readable feed at allowlists.grafana.com/synthetics.
HetrixTools
Fully verifiableHetrixTools · Monitoring
Uptime and blacklist monitoring service that probes customer sites from a published list of monitoring-node IPs.
Hydrozen
Fully verifiableHydrozen · Monitoring
Hydrozen.io is an uptime and website monitoring service; its checker fetches monitored endpoints from a documented, published set of IP addresses.
Buck
Listed onlyHypefactors · Monitoring
Buck is Hypefactors' media-monitoring crawler. It discovers and indexes web pages by following links, gathering publicly available content for the company's media-monitoring platform. It identifies itself and respects robots.txt.
Mediatoolkitbot
Listed onlyDeterm · Monitoring
Mediatoolkitbot is the web crawler operated by Determ (formerly Mediatoolkit), a media monitoring platform. It fetches publicly available web content to match brand mentions and topics that Determ's customers track.
Neticle Crawler
Listed onlyNeticle Technologies · Monitoring
Neticle's in-house crawler for its social and online media-monitoring service. It continuously fetches newly published web content so Neticle can score sentiment around the keywords, brands and products its customers track.
New Relic Synthetics
VerifiableNew Relic · Monitoring
New Relic's synthetic-monitoring minions, which probe customer-configured URLs from public minion IPs. New Relic documents an X-Abuse-Info request header identifying the monitor and account; logs also record a NewRelicbot User-Agent. New Relic states the published ranges "are reserved for use by New Relic and cannot be used by anyone else". The published IP list is a JSON object keyed by location, which this directory's feed formats cannot consume, so the ranges are recorded statically instead.
Pingdom
Fully verifiableSolarWinds · Monitoring
Uptime and page-speed monitoring service that probes customer sites from a published list of probe-server IPs. Owned by SolarWinds.
Amazon Route 53 Health Checks
Listed onlyAmazon Web Services · Monitoring
Amazon Route 53 health checkers, which probe customer-configured endpoints from AWS data centers worldwide to drive DNS failover. AWS publishes checker source ranges only inside the general ip-ranges.amazonaws.com/ip-ranges.json feed (entries tagged ROUTE53_HEALTHCHECKS), a shape this directory's feed formats cannot consume, so no CIDR recipe is attached.
SentiBot
Fully verifiableSentiOne · Monitoring
SentiOne's social-listening crawler. It indexes user-generated content for the SentiOne Listen platform, which its operator says analyses over 300,000 domains daily. Robots.txt rules written for "sentibot" are honoured, and Yandex-style reverse-DNS plus a published IP list allow verification.
Sentry Uptime Monitoring
Fully verifiableSentry · Monitoring
Sentry's uptime monitoring bot, which probes customer-configured URLs on a schedule from Sentry's uptime-check infrastructure and raises alerts on failures. Sentry documents a fixed User-Agent and publishes the current uptime-check IP addresses at a machine-readable endpoint.
Site24x7
Fully verifiableZoho · Monitoring
Zoho's infrastructure and website monitoring service, which probes customer sites from 130+ global monitoring locations published as a machine-readable IP list.
StatusCake
Fully verifiableStatusCake · Monitoring
Uptime and page-speed testing service that requests customer URLs from a global network of test-location servers, published as a machine-readable IP list.
Testomatobot
Fully verifiableTestomato · Monitoring
Testomato's website-monitoring agent. It downloads pages and resources and submits web forms for the checks its customers configure (uptime, error, SSL, response-time, meta-tag and JSON-LD monitoring), on a schedule the customer sets. Testomato publishes a plain-text list of the addresses its monitoring nodes request from.
Trendiction Bot
Listed onlyTrendiction · Monitoring
Trendiction operates a crawler that fetches public news sites, message boards, and blogs to build a search index and feed media-monitoring and market-research analytics for its clients.
updown.io
Fully verifiableupdown.io · Monitoring
Website uptime-monitoring service that checks customer sites from a published JSON array of node IPs. Documents "updown.io" as the common substring of its requests rather than one fixed User-Agent.
Uptime.com
Listed onlyUptime.com · Monitoring
Uptime.com's synthetic-monitoring probes. They request the URLs its customers configure as checks, from probe servers in many locations, including headless-browser transaction checks that execute JavaScript. Uptime.com publishes its probe IP list only inside the authenticated dashboard, so the user agent is the public identifier.
UptimeRobot
Fully verifiableUptimeRobot · Monitoring
Uptime monitoring service that probes customer sites on a schedule from a published list of monitoring IPs.
SEO tools 42
AhrefsSiteAudit
Fully verifiableAhrefs · SEO tools
Ahrefs' on-demand site-auditing crawler, distinct from AhrefsBot, used when Ahrefs customers run the Site Audit tool against their own or a competitor's domain. Shares Ahrefs' published IP-range feed.
AhrefsBot
Fully verifiableAhrefs · SEO tools
Ahrefs' primary web crawler, powering the backlink and keyword database behind the Ahrefs SEO platform and the Yep search engine. Crawls from publicly published IP ranges with a matching reverse-DNS suffix.
Audisto Crawler
Fully verifiableAudisto GmbH · SEO tools
Crawler for Audisto's hosted technical SEO and site-audit platform, fetching pages of the sites its customers analyse. Audisto publishes its crawler addresses as JSON and documents reverse-DNS verification.
Barkrowler
Fully verifiableBabbar · SEO tools
Barkrowler is the web crawler operated by Babbar (formerly Exensa). It builds and updates Babbar's graph of the web, which powers the company's SEO and link-analysis tools, and applies a politeness delay between requests.
BomboraBot
Listed onlyBombora, Inc. · SEO tools
BomboraBot is Bombora's web crawler. It classifies the content and topics of web pages that carry Bombora's tags so the company can model B2B purchase intent, visiting each tagged page at most once every 30 days.
Botify
Listed onlyBotify · SEO tools
Botify's SEO crawler, used by its SiteCrawler product to analyze enterprise websites for search-indexing insights. Botify does not publish a fixed IP range or CIDR list for site owners to allowlist.
BuiltWith
Listed onlyBuiltWith Pty Ltd · SEO tools
Crawler for BuiltWith's technology-profiling service, which visits sites and analyses publicly visible markup to determine which web technologies they use. BuiltWith publishes no IP ranges, so requests cannot be verified.
Caliperbot
VerifiableConductor · SEO tools
Conductor's single web crawler. It reads the HTML of pages on sites its customers track, recording on-page elements such as title tags, header tags and other metadata for Conductor's SEO and search-visibility reporting. Conductor publishes the address range it crawls from and will lower the crawl rate on request.
Cincraw
Listed onlyCINC Corp. (株式会社CINC) · SEO tools
The web crawler operated by CINC, a Japanese data-solutions company, to collect the page data behind its marketing and SEO analytics products. Its documented policy is to fetch page body content, header and HTTP status information and the JS/CSS needed to render a page, then store a rendered screen capture. CINC states that it does not follow advertising links, deletes all cookies between requests, and does not load analytics or ad-measurement tags. No robots.txt policy and no IP ranges are published.
Cocolyzebot
Listed onlyCocolyze · SEO tools
Crawler for Cocolyze's SEO analysis platform, fetching pages of sites its users analyse. Cocolyze publishes no IP ranges, so requests cannot be verified beyond the user agent.
cognitiveSEO
Listed onlycognitiveSEO · SEO tools
James BOT is the web crawler operated by cognitiveSEO, an SEO toolset. It crawls the web and analyzes links to power the backlink and SEO analysis offered by the cognitiveSEO platform.
CriteoBot
VerifiableCriteo · SEO tools
Criteo's advertising crawler. It fetches merchant and publisher pages to extract product and content data used for Criteo's commerce and retargeting ads, and respects robots.txt and crawl-delay directives.
DataForSeoBot
VerifiableDataForSEO · SEO tools
DataForSEO's crawler. It fetches pages to build the backlink and SEO datasets that power DataForSEO's marketing-data APIs, and honours robots.txt and crawl-delay directives.
Dataproviderbot
VerifiableDataprovider.com · SEO tools
Dataprovider.com's in-house crawler. It indexes more than 400 million domains each month and structures what it finds into the company's web dataset (business information, technology detection, classifications and risk signals). The operator documents that it follows the robot exclusion protocol and that its crawlers can be identified by a reverse DNS lookup.
DotBot
Listed onlyMoz · SEO tools
Moz's general-purpose web crawler, distinct from rogerbot, that gathers link data powering the Moz Link Index and Link Explorer. Moz's own help pages document no fixed IP range for it.
Dragonbot
Listed onlyDragon Metrics · SEO tools
Dragon Metrics' SEO crawler, which collects data for the platform's Site Audit and Site Explorer features. Its operator documents that it respects robots.txt using Google's open-source parser, and that it crawls from dynamic IP addresses so it can only be identified by user agent.
HubSpot Crawler
VerifiableHubSpot · SEO tools
HubSpot's crawler, which fetches customer and external pages to power the SEO recommendations and link analysis in HubSpot's marketing tools. HubSpot publishes its egress ranges tagged by service, including web crawling.
IAS Crawler
Listed onlyIntegral Ad Science · SEO tools
Content-rating and ad-verification crawler operated by Integral Ad Science. It visits web pages to assess content quality and brand safety and to support invalid-traffic detection for advertisers.
Linkdexbot
Listed onlyAuthoritas (Analytics SEO Limited) · SEO tools
Linkdexbot is the web crawler for Linkdex, an SEO and search-marketing analytics platform now operated by Authoritas. It gathers link and page data used to power the platform's SEO reporting tools.
MegaIndex Crawler
Listed onlyMegaIndex · SEO tools
MegaIndex is an SEO and web-analytics platform whose crawler indexes links across the web to power backlink analysis, keyword tracking, and site audits for its subscribers.
Meta External Ads
VerifiableMeta · SEO tools
Meta's crawler that fetches pages for advertising and other business-related products and services, separate from the AI-training and link-preview crawlers. Verified by ASN lookup (AS32934); Meta publishes no IP feed.
MTRobot
Listed onlyMetrics Tools (Andreas Knatz) · SEO tools
Crawler for Metrics Tools, a German SEO analytics service, collecting page data for its visibility and ranking analyses. The operator publishes no IP ranges, so requests cannot be verified beyond the user agent.
MJ12bot
Listed onlyMajestic-12 · SEO tools
The crawler behind Majestic's backlink index. Majestic explicitly states it is a community-based distributed crawler with no fixed IP allocation, so requests cannot be verified by IP, ASN, or reverse DNS.
OnCrawl
Listed onlyOnCrawl · SEO tools
OnCrawl's SEO crawler, used to analyze a customer's own site structure and content for technical SEO reporting. OnCrawl's help docs describe no fixed IP range; the bot's identity is user-configurable per crawl.
Outbrain crawler
VerifiableOutbrain · SEO tools
Outbrain's content-recommendation crawler. It fetches advertiser landing pages so Outbrain's system can pull the correct image and headline for a promoted-content unit, and rejects submitted URLs it cannot reach.
Panscient Crawler
Listed onlyPanscient Inc. · SEO tools
Panscient's large-scale crawler, which traverses public websites so that Panscient can build structured company and professional data feeds licensed to enterprise customers. The operator documents a full-corpus refresh each quarter, a rate limit of at most one request per second to any single domain, and compliance with the Robot Exclusion Standard. A separate "pantest" agent is used for testing. No IP ranges are published.
Proximic (Comscore Crawler)
Listed onlyComscore, Inc. · SEO tools
Proximic is Comscore's web crawler. It downloads the static textual content of pages to perform contextual analysis (content language, rating, and IAB categories) so advertising partners can match campaigns to page content. It identifies itself and honors robots.txt.
Rogerbot
Listed onlyMoz · SEO tools
Moz's site-audit crawler for Moz Pro Campaigns, distinct from DotBot. Moz's own FAQ states plainly that Rogerbot has no IP range: "we do not use a static IP address or range of IP addresses."
RyteBot
Listed onlySemrush · SEO tools
The crawler behind the Ryte.com tools, which analyse on-page SEO, technical and usability issues. Ryte was absorbed by Semrush, and RyteBot is now documented as a member of the Semrush bot family with its own robots.txt user agent. No IP ranges are published for it.
Search Atlas Bot
Listed onlySearch Atlas · SEO tools
The crawler behind Search Atlas's SEO platform, which fetches pages for its Site Auditor and monitoring features. The operator publishes the bot's user agent for allowlisting and states that the crawler does not use static IP addresses, so it can only be identified by its user agent.
SiteAuditBot
VerifiableSemrush · SEO tools
Semrush's site-auditing crawler, distinct from SemrushBot: it crawls a domain on demand when a Semrush customer runs the Site Audit tool, looking for SEO and technical issues. Unlike the backlink crawler, which Semrush says cannot be identified by IP, Site Audit is documented as running from a single dedicated subnet.
SemrushBot
Listed onlySemrush · SEO tools
Semrush's web crawler, feeding the backlink and site-audit data behind the Semrush SEO platform. Semrush's own bot page explicitly states it does not use consecutive IP blocks, so no CIDR list can be sourced.
SemrushBot-SI
VerifiableSemrush · SEO tools
The crawler behind Semrush's On Page SEO Checker and related on-page tools, run against a domain when a Semrush customer sets up a campaign for it. It is a separate robots.txt user agent from SemrushBot, and Semrush documents its own addresses to allowlist for it.
SeobilityBot
Fully verifiableSeobility GmbH · SEO tools
Crawler for Seobility's hosted SEO analysis and site-audit tooling, fetching pages of sites its customers analyse. Seobility publishes a machine-readable list of the addresses its bots crawl from.
SEOkicks
Listed onlyJobkicks SLU · SEO tools
SEOkicks operates a web crawler that builds a backlink database powering its SEO tools. The crawler visits sites to collect link data for analysis.
serpstatbot
Fully verifiableSerpstat · SEO tools
Serpstat's backlink crawler. It continuously crawls the web to add new links and track changes in Serpstat's link database, honouring robots.txt and Crawl-delay directives, and publishes the full list of addresses it crawls from.
SISTRIX Crawler
VerifiableSISTRIX · SEO tools
The crawler behind the SISTRIX Toolbox, a German SEO visibility platform. SISTRIX documents that every crawler IP resolves via reverse DNS to the "sistrix.net" domain rather than publishing a static CIDR list.
Siteimprove Crawler
VerifiableSiteimprove · SEO tools
Siteimprove's content-suite crawler, which fetches pages of sites its customers have configured in their account to run quality-assurance, accessibility, policy and SEO checks. Companion agents (LinkCheck, Image size, Probe) fetch links and resources for the same checks and crawl from the same published address list.
t3versionsBot
Listed onlyTorben Hansen (t3versions) · SEO tools
Private-project crawler that makes single GET requests to sites and looks for TYPO3 fingerprints, collecting statistics on the worldwide usage and development of the open-source TYPO3 CMS. No IP ranges are published, and the operator documents no robots.txt support (exclusion is by email request).
VelenPublicWebCrawler
Listed onlyHunter · SEO tools
Hunter's public web crawler, written in Go. It analyses millions of publicly accessible pages every month to build the business datasets and machine learning models behind Hunter's products, and never fetches anything behind a login. The operator documents a deliberate rate limit of one page at a time and one page every two seconds per site.
XoviBot
Listed onlyXovi GmbH · SEO tools
XoviBot is the web crawler for XOVI, an SEO and online-marketing analytics suite. It crawls sites to gather backlink and ranking data for the platform's SEO tools.
Zoominfobot
Listed onlyZoomInfo Technologies · SEO tools
ZoomInfo's indexing robot, which scans corporate websites, press releases, news services and SEC filings to build ZoomInfo's search index of businesses and business professionals. The operator documents that it obeys robots.txt, spaces out requests on larger sites and never opens more than one connection to a site at a time. No IP ranges are published.
Feed fetchers 13
Apple Podcasts (iTMS)
VerifiableApple · Feed fetchers
Apple's podcast feed fetcher, which crawls only URLs associated with content registered on Apple Podcasts. Apple documents that iTMS traffic may come from applebot.apple.com hosts and that it does not follow robots.txt because it is not a general search crawler.
BazQux Fetcher
Listed onlyBazQux Reader · Feed fetchers
The feed fetcher of the BazQux Reader hosted RSS service. It retrieves and periodically refreshes the RSS/Atom and comment feeds that users have subscribed to, typically no more than once an hour per feed. BazQux documents that the fetcher acts as an agent of those users and therefore ignores robots.txt.
Facebook Catalog
Listed onlyMeta · Feed fetchers
Meta's product-catalog fetcher, identified by the facebookcatalog user-agent. It retrieves merchant product-data feeds used to build and refresh commerce catalogs surfaced across Facebook and Instagram.
Feedbin
VerifiableFeedbin · Feed fetchers
Feedbin's feed fetcher, which retrieves RSS/Atom feeds that users have subscribed to. Its user agent includes the internal feed id and current subscriber count, and Feedbin documents forward-confirmed reverse DNS in *.bot.feedbin.com as the way to verify its requests.
Feedly Fetcher
Listed onlyFeedly · Feed fetchers
Feedly's fetcher, which retrieves RSS/Atom feed URLs after a user has explicitly added them to their Feedly. Feedly documents that it behaves as a direct agent of the user rather than a robot, and does not publish a fixed IP list because its source IPs change over time.
Feedspot
Listed onlyFeedspot · Feed fetchers
Feedspot is a hosted content reader and feed aggregation service. Its bot fetches RSS and Atom feeds and web content on behalf of Feedspot users.
Flipboard Proxy
Listed onlyFlipboard, Inc. · Feed fetchers
Flipboard's proxy service, which fetches and prepares elements of a page (e.g. a social feed a user asked Flipboard to scan) for presentation in the Flipboard app. Flipboard's own docs say these requests currently originate from an Amazon EC2 cluster but publish no fixed IP list.
Feedfetcher-Google
Fully verifiableGoogle · Feed fetchers
Google's feed retrieval agent for RSS and Atom feeds used by Google News and WebSub. It fetches and periodically refreshes feeds that users of an app or service have explicitly subscribed to.
Hatena::Russia::Crawler
Listed onlyHatena Co., Ltd. (Hatelabo) · Feed fetchers
The fetcher behind Daichecker, the antenna service run on Hatelabo, Hatena's experimental-services lab. It checks the pages and feeds that users have registered for updates. Hatena documents that it parses only the robots.txt groups that name this user agent directly and does not apply the User-agent: * group, so a wildcard rule will not stop it. Hatena notes the name comes from an internal code name for RSS-reader development and has no connection to the country.
Inoreader Fetcher
Fully verifiableInnologica · Feed fetchers
Inoreader's feed fetcher, which retrieves RSS/Atom feeds that Inoreader users have subscribed to. Its own docs state it does not read robots.txt because it fetches specific, user-requested feed URLs rather than crawling a site, and it publishes a live list of its backend fetcher IPs.
Miniflux
Listed onlyMiniflux · Feed fetchers
Miniflux is a minimalist, open-source, self-hosted feed reader. User-run instances fetch the RSS and Atom feeds their subscribers add, identifying themselves with a Miniflux User-Agent.
NewsBlur Feed Fetcher
Listed onlyNewsBlur · Feed fetchers
NewsBlur's open-source feed fetcher, which polls RSS/Atom feeds on behalf of subscribed users. Its user agent embeds the live subscriber count and the feed's permalink; NewsBlur publishes no fixed IP range for it.
Superfeedr
VerifiableSuperfeedr · Feed fetchers
Superfeedr's PubSubHubbub feed-polling infrastructure, which fetches feed URLs that publishers or subscribers have registered with the service. Its docs publish a list of current node IPs but warn it changes as they add or remove cloud capacity.
Archivers 7
AcademicBotRTU
Listed onlyRiga Technical University (Institute of Applied Computer Systems) · Archivers
Crawler run by Riga Technical University that indexes websites and documents to compare against student and researcher works for plagiarism detection. The operator publishes no IP ranges, so requests cannot be verified.
Arquivo.pt Web Crawler
Listed onlyArquivo.pt (FCCN/FCT) · Archivers
The crawler behind Arquivo.pt, Portugal's public web archive, which captures full page renders (HTML, CSS, JS, images) for long-term preservation. Built on Heritrix; the operator documents no fixed IP range.
BnF Web Archiving Robot
Listed onlyBibliothèque nationale de France · Archivers
The web crawler of the Bibliothèque nationale de France, which harvests French websites for the legal deposit of the web to preserve the national documentary heritage. It runs on Heritrix and applies request delays to avoid overloading servers.
Cloudflare Always Online
Listed onlyCloudflare · Archivers
Cloudflare's Always Online crawler, which fetches pages from sites that have the feature enabled so a cached copy can be served to visitors when the origin server is unreachable. Cloudflare's crawler reference documents the CloudFlare-AlwaysOnline user agent for this product.
archive.org_bot
Listed onlyInternet Archive · Archivers
The Internet Archive's Heritrix-based crawler used for its wide crawl of the web, feeding the Wayback Machine. The Archive says it crawls slowly to avoid disrupting sites and publishes no IP ranges. Its help pages note that robots exclusions may prevent archiving, but the operator does not document a commitment to obey robots.txt across its crawls.
ArchiveBot
Listed onlyArchive Team · Archivers
ArchiveBot is an IRC-controlled archiving bot run by Archive Team that crawls websites on request, writes WARC files, and uploads the captures to the Internet Archive.
TurnitinBot
Listed onlyTurnitin · Archivers
TurnitinBot is the web crawler operated by Turnitin. It collects publicly available web pages to build the content database used by Turnitin's academic-integrity and plagiarism-detection services.
Security scanners 13
Censys Inspect
VerifiableCensys · Security scanners
Censys' internet-wide scanner, which probes public IP addresses to build the host/service data behind Censys Search. Censys publishes a fixed set of scanner subnets and ASNs for opt-out purposes.
Cloudflare Custom Hostname Verification
Listed onlyCloudflare · Security scanners
Cloudflare's custom-hostname ownership checker. Cloudflare documents that requests carrying this user agent are triggered when a customer chooses to validate a custom hostname with an HTTP ownership token, which requires fetching the token from the hostname being claimed.
Cloudflare SSL Detector
Listed onlyCloudflare · Security scanners
Cloudflare's SSL/TLS Recommender crawler, which fetches a customer origin over both HTTP and HTTPS to determine whether the site can be served fully over HTTPS and to recommend TLS settings.
Cortex Xpanse
VerifiablePalo Alto Networks · Security scanners
Palo Alto Networks' attack-surface-management scanner. It continuously scans the global internet from a published set of ranges to map its customers' internet-facing assets and discover emerging threats, and its requests carry a plain-English user agent naming the company and an opt-out contact address.
Detectify
VerifiableDetectify · Security scanners
Detectify's external attack-surface and vulnerability scanner, which probes customer-configured assets from a documented set of AWS-hosted source IPs. The operator publishes both the source IPs and the scanner's user-agent strings.
Driftnet Internet Measurement
VerifiableDriftnet · Security scanners
Driftnet's internet-measurement scanner, run from the internet-measurement.com domain, which probes publicly exposed services to give network owners an external view of their infrastructure. The operator states this traffic never attempts to log in to systems and publishes its scanner IP ranges.
Jugendschutzprogramm-Crawler
Listed onlyJusProg e.V. · Security scanners
Web crawler operated by JusProg e.V., a German non-profit youth-protection association, that fetches and rates web pages to build the age-classification database used by its parental-control filtering software.
LeakIX (l9explore)
VerifiableLeakIX · Security scanners
LeakIX's internet-wide recon scanner (l9explore), which probes exposed services to populate the LeakIX misconfiguration/vulnerability search engine. Its probe fleet publishes a live, per-host list of source IPs at scan.leakix.net whose hostnames resolve under scan.leakix.org, and its scanning tool sets an identifiable user-agent.
Qualys SSL Labs
VerifiableQualys · Security scanners
Qualys' SSL Labs server test, which runs visitor-initiated and monthly (SSL Pulse) TLS-configuration assessments of public HTTPS sites. Qualys documents the assessments as slow and non-intrusive and publishes the scanner's source IP ranges in its support knowledge base.
Streamline3Bot
VerifiableUBT (EU) Ltd · Security scanners
UBT's web crawler, which powers a classification service that categorises public websites by their content. It re-crawls a given site roughly once every three days. UBT publishes no IP list, and instead documents reverse-DNS verification against the ubtsupport.com domain.
Stripebot
VerifiableStripe · Security scanners
Stripe's automated web crawler. Collects data from Stripe users' websites so Stripe can provide its services and comply with financial regulations. Distinct from Stripe's webhook delivery traffic.
SurdotlyBot
Listed onlySur.ly · Security scanners
Sur.ly's crawler. The operator runs a spam-fighting link-safety service and uses this bot to fetch third-party sites and build a short security profile for each one, querying metadata and favicons and taking a screenshot of the homepage. The operator states it never harvests e-mail addresses or content unrelated to security.
W3C Markup Validator
VerifiableWorld Wide Web Consortium (W3C) · Security scanners
The W3C Markup Validation Service's fetcher, used when a user submits a URL to be checked for HTML conformance. W3C documents a fixed source address for its validation services alongside the user-agent string.
Webhooks 8
Adyen Webhooks
VerifiableAdyen · Webhooks
Adyen's payment-event webhook delivery service. Not a crawler; POSTs notifications to merchant endpoints. Adyen documents its egress domain for DNS-based allowlisting rather than a static IP list, since its outbound IPs change over time.
GitHub Webhooks
Fully verifiableGitHub · Webhooks
GitHub's webhook delivery service. Not a crawler; sends repository and organization event notifications to configured endpoints from IP ranges published in GitHub's meta API.
APIs-Google
Fully verifiableGoogle · Webhooks
Google's agent that delivers push notification messages sent through Google APIs (such as Pub/Sub and WebSub push subscriptions) to subscriber endpoints.
PayPal IPN
VerifiablePayPal · Webhooks
PayPal's Instant Payment Notification service. Not a crawler; POSTs payment-event notifications to merchant listener endpoints from a documented, static set of server CIDR ranges (shared with other PayPal server traffic).
Stripe Webhooks
Fully verifiableStripe · Webhooks
Stripe's webhook delivery service. Not a crawler; sends event notifications (payments, subscriptions) to merchant endpoints from published IPs.
Svix Webhooks
VerifiableSvix · Webhooks
Webhook-sending infrastructure used by Svix's customers to deliver events. Not a crawler; Pro/Enterprise plans get a documented, static set of per-region source IPs, and requests include a Svix sender identifier in the User-Agent.
Telegram Bot Webhooks
VerifiableTelegram · Webhooks
Telegram's webhook delivery to bot servers. Not a crawler; POSTs update events to a bot's registered HTTPS endpoint from two documented, static CIDR ranges. Telegram documents no User-Agent for these POSTs; the UA listed here is Telegram's documented fetcher token and is unconfirmed for webhook traffic — verify by source IP, not UA.
Twilio Webhooks
Listed onlyTwilio · Webhooks
Twilio's webhook delivery service, which POSTs event callbacks (incoming messages and calls, status updates) to customer-configured endpoints. Not a crawler. Twilio states there is no fixed range of source IPs — requests come from a dynamic pool — so recipients are told to validate the X-Twilio-Signature request signature instead.
No bots match that filter.