The web's bots,
verified.

The open, categorized directory of legitimate bots — their user agents, published IP ranges, and exactly how to verify them.

227 bots 123 verifiable 10 categories 9298 IP ranges

Search engines 53

AddSearchBot

Verifiable

AddSearch · Search engines

Crawler for AddSearch's hosted site-search service. It indexes the pages of customer sites so AddSearch can serve search results for them, obeys robots.txt rules written for the AddSearchBot token, and AddSearch publishes the fixed addresses it crawls from for allowlisting.

AdIdxBot

Verifiable

Microsoft · Search engines

Microsoft's crawler for Bing Ads / Microsoft Advertising. It crawls ads and follows through to the advertised landing pages for quality control, with both desktop and mobile variants.

Algolia Crawler

Verifiable

Algolia · Search engines

Algolia's Crawler, which visits a customer's pages, extracts search-relevant content, and pushes it to Algolia search indices. Algolia documents a fixed User-Agent token and a single static egress IP address for allowlisting.

Amazon AdBot

Verifiable

Amazon · Search engines

Amazon's advertising crawler. It scans web pages that request ads from Amazon's advertising systems, collecting page content for Amazon's classification systems to maintain brand safety and improve ad relevance.

Amzn-SearchBot

Listed only

Amazon · Search engines

Amazon's crawler used to improve search experiences in Amazon products and services; Amazon states it does not crawl content for generative AI model training. Amazon publishes a human-readable IP list but no machine-parseable feed in a format this project's schema supports.

Applebot

Fully verifiable

Apple · Search engines

Apple's web crawler, used to index content for Siri, Spotlight Suggestions, and Safari's search features.

Baiduspider

Verifiable

Baidu · Search engines

Baidu's primary web crawler, fetching pages for Baidu Search indexing.

BingVideoPreview

Verifiable

Microsoft · Search engines

Microsoft crawler, listed among Bing's crawlers, that fetches pages to generate video previews shown in Bing. It runs desktop and mobile variants and is verifiable through the same reverse-DNS check as Bing's other crawlers.

Bingbot

Fully verifiable

Microsoft · Search engines

Microsoft's primary web crawler, fetching pages for Bing Search indexing.

BingPreview

Verifiable

Microsoft · Search engines

Microsoft fetcher used to generate page snapshots for previews in Bing search results. Runs both desktop and mobile variants.

Bublup Bot

Verifiable

Bublup · Search engines

Bublup's content-discovery bot. It fetches pages to build the database behind Bublup's suggestion engine, reading page title, description and related images. Bublup states its crawling IPs are not fixed and documents reverse-DNS verification against the bublup.com domain instead.

coccocbot

Verifiable

Cốc Cốc · Search engines

Cốc Cốc's web crawler, fetching pages for the Vietnamese Cốc Cốc search engine. Separate web and image sub-bots share the same verification domain.

deepnoc

Listed only

deepnoc GmbH · Search engines

Research crawler operated by deepnoc GmbH that parses and stores public web page content so that new search engines can query it without running their own crawling infrastructure. No IP ranges are published.

DuckDuckBot

Fully verifiable

DuckDuckGo · Search engines

DuckDuckGo's web crawler, used to improve DuckDuckGo's search results.

Geedo Product Search

Fully verifiable

Geedo · Search engines

GeedoShopProductFinder is the crawler for Geedo, a product-search engine that indexes publicly accessible product pages from online stores and follows robots.txt rules.

Storebot-Google

Fully verifiable

Google · Search engines

Google's crawler for Google Shopping surfaces, including the Shopping tab in Google Search and shopping.google.com. It crawls product and store pages to populate shopping results.

Google-InspectionTool

Fully verifiable

Google · Search engines

Google crawler used by Search testing tools such as the URL Inspection tool in Search Console and the Rich Result Test. Fetches pages on demand to show how Google Search renders them; it has no effect on ranking.

Google special-case crawlers

Fully verifiable

Google · Search engines

A group of Google crawlers that operate under separate agreements between the crawled site and a specific Google product (AdsBot, AdSense, Google Safety scanning, and similar). Several members of this group ignore the wildcard robots.txt rule and must be targeted explicitly.

Googlebot

Fully verifiable

Google · Search engines

Google's primary web crawler, fetching pages for Google Search indexing.

Googlebot Image

Fully verifiable

Google · Search engines

Google's image crawler, which fetches images for Google Images and image features in Search. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.

Googlebot Video

Fully verifiable

Google · Search engines

Google's video crawler, which fetches video content for Google Search video features. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.

GoogleOther

Fully verifiable

Google · Search engines

Google's generic crawler family (GoogleOther, GoogleOther-Image, GoogleOther-Video), used by various Google product teams for fetching publicly accessible content. Robots.txt rules addressed to it do not affect Google Search.

360Spider

Listed only

360 Search (Qihoo 360) · Search engines

360Spider is the web-search crawler of 360 Search (so.com), the search engine operated by Chinese internet company Qihoo 360. The operator also runs 360Spider-Image and 360Spider-Video for image and video search.

IONOS Crawler

Verifiable

IONOS · Search engines

IONOS Crawler (IonCrawl) continuously crawls publicly accessible domains to generate insights into how they are used, which IONOS uses to improve and expand its hosting products. IONOS documents reverse-DNS verification and deletes crawled data after 60 days.

Jooblebot

Listed only

Jooble · Search engines

The crawler for Jooble, a job-search aggregator that indexes job listings published across the web. Jooble's bot page states that it uses a web crawler which identifies itself as JoobleBot. No IP ranges are published.

Kagibot

Verifiable

Kagi · Search engines

Kagi's web crawler, fetching pages for the paid Kagi search engine. Kagi publishes a fixed, small set of source IPs rather than a CIDR feed, each paired with a kagibot.org hostname.

Linespider

Listed only

LINE · Search engines

LINE's crawler. It fetches web pages to support LINE's search features and content services, and adheres to the Robots Exclusion Protocol outlined in robots.txt.

Marginalia

Fully verifiable

Marginalia Search · Search engines

The crawler behind Marginalia Search, an independent search engine focused on older, text-heavy, and non-commercial web pages. Identifies itself by a bare URL rather than a conventional bot token.

MetaJobBot

Verifiable

METAJob · Search engines

MetaJobBot is the focused web crawler operated by the METAJob job meta-search engine. It searches websites for job listings to include in the METAJob index.

MojeekBot

Fully verifiable

Mojeek · Search engines

Mojeek's web crawler, fetching pages for the independent Mojeek search engine index.

Yeti

Verifiable

Naver · Search engines

Naver's web crawler, fetching pages for Naver Search indexing across South Korea's largest search engine.

Openindex Spider

Listed only

Openindex · Search engines

Openindex operates a web-crawling cluster (Apache Nutch on an Apache Hadoop cluster) for research and development of universal and focused search engines. The crawler respects robots.txt and the Crawl-delay directive and identifies itself in the User-Agent.

PetalBot

Verifiable

Huawei · Search engines

Huawei's web crawler, fetching pages for the Petal Search engine used in Huawei's mobile services and for Huawei Assistant content recommendations. Huawei documents reverse-DNS verification against aspiegel.com and petalsearch.com hostnames.

PiplBot

Listed only

Pipl · Search engines

PiplBot is the web-indexing crawler operated by Pipl, an identity and people-search service. It retrieves and indexes publicly available information, including content behind searchable databases, to build Pipl's people-search index.

Quantcastbot

Verifiable

Quantcast · Search engines

Quantcast's advertising crawler. It crawls websites to extract page content for interest-based audience categorization and to run quality-assurance checks on advertisement landing pages.

Qwantbot

Fully verifiable

Qwant · Search engines

Qwant's web crawler, fetching pages for the privacy-focused Qwant search engine. Older crawls identify as Qwantify; current documentation uses the Qwantbot name.

SeekportBot

Fully verifiable

SISTRIX · Search engines

The crawler for the Seekport search engine, operated by SISTRIX (Bonn, Germany). It fetches pages to build Seekport's search index and follows the Disallow directives in robots.txt.

SemanticScholarBot

Listed only

Allen Institute for AI (Ai2) · Search engines

SemanticScholarBot is the crawler for Semantic Scholar, an academic search engine built by the Allen Institute for AI. It crawls the web to find scholarly PDFs and metadata for indexing.

SeznamBot

Verifiable

Seznam.cz · Search engines

Seznam's web crawler, fetching pages for the Seznam.cz search engine used primarily in the Czech Republic.

Sogou Spider

Listed only

Sogou · Search engines

Sogou Spider is the web crawler for Sogou, a major Chinese search engine (owned by Tencent). It crawls and indexes web pages, news, and images to build Sogou's search index.

StractBot

Listed only

Stract · Search engines

The crawler for Stract, an open-source search engine. Stract documents the crawler's user agent and robots.txt behavior but publishes no IP list or reverse-DNS verification method.

TinEye

Listed only

TinEye · Search engines

TinEye-bot is the web crawler operated by TinEye, the reverse image search engine built by Idee Inc. It crawls the web to discover and index images for TinEye's image-matching search service.

TrovitBot

Listed only

Trovit · Search engines

Trovit's web crawler, which discovers new and updated pages to add to the Trovit classifieds search index. Trovit documents that it does not fetch most sites more than once per second and that it can be blocked with the trovitBot robots.txt token.

Webzio

Listed only

Webz.io Ltd. · Search engines

Crawler for Webz.io (formerly Webhose.io), a web-data provider that collects publicly available content to build structured data feeds. Webz.io publishes no IP ranges, so requests cannot be verified.

Yahoo! Slurp

Listed only

Yahoo · Search engines

Slurp is Yahoo Search's web crawler. It crawls and indexes pages for Yahoo Search results and also gathers content for Yahoo News, Finance, and Sports.

Yahoo! JAPAN Crawler

Listed only

LY Corporation (Yahoo! JAPAN) · Search engines

Web crawler operated by Yahoo! JAPAN (LY Corporation) to collect pages for its Japanese search index and related services. It identifies itself with Y!J-prefixed user-agent tokens and follows the Robots Exclusion Protocol.

YandexBlogs

Verifiable

Yandex · Search engines

Yandex's blog-search robot. Yandex's robot table documents it as the blog search robot that indexes post comments, and lists it as following robots.txt directives.

YandexRenderResourcesBot

Verifiable

Yandex · Search engines

Yandex robot that loads resources such as JavaScript and CSS needed to render pages during crawling for Yandex Search.

YandexFavicons

Verifiable

Yandex · Search engines

Yandex's favicon fetcher. Yandex's robot table documents it as downloading a site's favicon file for display in search results, and lists it as one of the Yandex robots that do not follow robots.txt directives.

YandexImages

Verifiable

Yandex · Search engines

Yandex's image-indexing robot, which fetches images so they can be displayed in Yandex Images. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.

YandexMedia

Verifiable

Yandex · Search engines

Yandex's multimedia-indexing robot. Yandex's robot table documents it as a separate robot from YandexBot, describes it as indexing multimedia data, and lists it as following robots.txt directives.

YandexVideo

Verifiable

Yandex · Search engines

Yandex's video-indexing robot, which fetches pages and video content for display in Yandex video search. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.

YandexBot

Verifiable

Yandex · Search engines

Yandex's primary web crawler, fetching pages for Yandex Search indexing.

AI crawlers 23

AI2Bot

Listed only

Allen Institute for AI · AI crawlers

The Allen Institute for AI's crawler, used to build open training datasets such as Dolma. AI2 publishes no IP ranges, ASN, or reverse-DNS pattern.

Amazonbot

Listed only

Amazon · AI crawlers

Amazon's web crawler used to improve Amazon's products and services, including training AI models. Amazon publishes a human-readable IP list but no machine-parseable feed in a format this project's schema supports.

Bytespider

Listed only

ByteDance · AI crawlers

ByteDance's web crawler, used to gather training data for its AI models. ByteDance publishes no IP ranges, ASN, or reverse-DNS pattern for it.

CCBot

Fully verifiable

Common Crawl Foundation · AI crawlers

Common Crawl's crawler that builds the freely available Common Crawl web archive, widely reused as AI training data by third parties.

Claude-SearchBot

Fully verifiable

Anthropic · AI crawlers

Anthropic's crawler that indexes pages to improve search-style answer quality in Claude products, distinct from ClaudeBot's training-data crawl.

ClaudeBot

Fully verifiable

Anthropic · AI crawlers

Anthropic's web crawler collecting publicly available web data for training Claude models. Anthropic publishes its crawler source IPs as a machine-readable feed.

Cloudflare AI Search

Listed only

Cloudflare · AI crawlers

The crawler behind Cloudflare AI Search, which indexes website content so it can be searched. Cloudflare documents that it only crawls a website the customer owns — the domain must exist in the same Cloudflare account and be selected as an AI Search data source.

Cloudflare Browser Run Crawler

Fully verifiable

Cloudflare · AI crawlers

The crawler behind the /crawl endpoint of Cloudflare's Browser Run (Browser Rendering) developer product, which crawls third-party websites on behalf of Cloudflare customers building applications on the platform. Its user agent is not configurable, and every request is signed with Web Bot Auth HTTP message signatures that site owners can verify against Cloudflare's published key directory.

Diffbot

Listed only

Diffbot · AI crawlers

Diffbot's general-purpose web crawler, used to build its Knowledge Graph and power its structured-extraction APIs. Diffbot documents robots.txt behavior but publishes no IP list, ASN, or reverse-DNS pattern.

Google-CloudVertexBot

Fully verifiable

Google · AI crawlers

Google Cloud crawler that fetches sites on the site owners' request when building Vertex AI Agents. Crawling preferences addressed to its user agent have no effect on Google Search or other Google products.

GPTBot

Fully verifiable

OpenAI · AI crawlers

OpenAI's web crawler that gathers publicly available data used to train OpenAI's models.

ImagesiftBot

Listed only

Hive · AI crawlers

ImageSift's crawler, operated by Hive. It scrapes publicly available images across the web to support Hive's ImageSift reverse-image-search and web intelligence products. Standard robots.txt directives are respected.

Leipzig Corpora Collection Crawler

Listed only

Leipzig University Natural Language Processing Group · AI crawlers

The LCC crawler is operated by Leipzig University to collect web text for the Leipzig Corpora Collection, a set of linguistic corpora used in natural language processing research.

MaCoCu

Listed only

MaCoCu project (Jožef Stefan Institute) · AI crawlers

Crawler for the CEF-funded MaCoCu project, which collects, curates and enriches monolingual and parallel web text to build language corpora for under-resourced languages. The project publishes no IP ranges, so requests cannot be verified.

Meta External Agent

Verifiable

Meta · AI crawlers

Meta's crawler for AI training data and content indexing across Meta products. Verified by ASN lookup (AS32934); Meta publishes no IP feed.

Meta Web Indexer

Verifiable

Meta · AI crawlers

Meta's crawler that navigates the web to improve the quality of Meta AI search results, analysing page content for relevance and accuracy in Meta AI responses. Verified by ASN lookup (AS32934); Meta publishes no IP feed.

ICC-Crawler

Verifiable

National Institute of Information and Communications Technology (NICT) · AI crawlers

ICC-Crawler is a web crawler operated by Japan's NICT that collects web pages across the internet to build datasets for information and language processing research.

OAI-SearchBot

Fully verifiable

OpenAI · AI crawlers

OpenAI's crawler that indexes pages to power search results inside ChatGPT. It is distinct from GPTBot and is not used to gather training data.

omgili

Listed only

Webz.io · AI crawlers

Webz.io's legacy crawler user agent (formerly "Omgilibot"), used to collect web, forum, and news content for its data feeds. Webz.io's current public documentation describes successor crawlers ("webzio" / "webzio-extended") and no longer documents this UA string or any IP verification method for it.

PerplexityBot

Fully verifiable

Perplexity · AI crawlers

Perplexity's crawler that indexes pages to surface and link websites in Perplexity search results. Perplexity states it is not used to train models.

SBIntuitionsBot

Listed only

SB Intuitions Corp. · AI crawlers

Crawler operated by SB Intuitions (a SoftBank AI subsidiary) that collects web pages for AI development and information analysis, including training of its Sarashina language models. SB Intuitions publishes no IP ranges, so requests cannot be verified.

Webzio-extended

Listed only

Webz.io Ltd. · AI crawlers

Second crawler in Webz.io's crawler pair, which performs ethical validation on the data collected by Webzio and tags it as usable or not usable for AI and machine-learning training. Webz.io publishes no IP ranges.

YouBot

Fully verifiable

You.com · AI crawlers

You.com's crawler that indexes pages for its AI-powered search product. You.com documents a dedicated IP range, a reverse-DNS pattern, and support for signed-request verification via Web Bot Auth.

AI assistants 9

Amzn-User

Listed only

Amazon · AI assistants

Amazon's on-demand fetcher supporting user actions, such as responding to Alexa queries that need up-to-date information; Amazon states it does not crawl content for generative AI model training. Amazon publishes a human-readable live-crawl IP list but no machine-parseable feed.

ChatGPT-User

Fully verifiable

OpenAI · AI assistants

OpenAI's user-triggered fetcher, used when a ChatGPT user or a GPT Action asks the assistant to visit a specific page. It does not crawl autonomously.

Claude-User

Fully verifiable

Anthropic · AI assistants

Anthropic's user-triggered fetcher, used when a person asks Claude to visit a specific web page (e.g. via tool use in a conversation).

DuckAssistBot

Fully verifiable

DuckDuckGo · AI assistants

DuckDuckGo's real-time fetcher for DuckDuckGo Search's AI-assisted answers, which cite their sources. DuckDuckGo states the data is not used to train AI models, and publishes a machine-readable list of the bot's IP addresses.

Google-Agent

Fully verifiable

Google · AI assistants

The fetcher used by AI agents hosted on Google infrastructure to navigate the web and perform actions on behalf of a user who asked for them. Google publishes a dedicated IP range list for it and signs a subset of its requests with Web Bot Auth under the agent.bot.goog identity. As a user-triggered fetcher it generally ignores robots.txt rules.

Google user-triggered fetchers

Fully verifiable

Google · AI assistants

Google tools and product features that fetch a specific page because an end user asked for it (Google Read Aloud, Site Verifier, Gemini Notebook, Chrome Web Store, Google Messages, Pinpoint, Publisher Center, and similar), rather than autonomous crawling for search indexing. They generally ignore robots.txt because a human requested the fetch.

Meta External Fetcher

Verifiable

Meta · AI assistants

Meta's on-demand fetcher that retrieves a single link at a user's request to support agentic AI features (e.g. an AI assistant navigating a page a user asked about), rather than broad indexing. Verified by ASN lookup (AS32934).

MistralAI-User

Fully verifiable

Mistral AI · AI assistants

Mistral AI's user-triggered fetcher, used when a user of a Mistral product (e.g. Le Chat) asks it to visit a specific web page to answer a question.

Perplexity-User

Fully verifiable

Perplexity · AI assistants

Perplexity's user-triggered fetcher, used when a user's question requires visiting a specific web page to produce an accurate answer.

Social previews 22

Discordbot

Listed only

Discord · Social previews

Discord's crawler that fetches shared links to render embeds in chat. Discord's own docs only specify the User-Agent format required of API clients, not this embed fetcher; no IP ranges, ASN, or rDNS are published.

Embedly

Verifiable

Embedly · Social previews

Embedly's fetcher, used to generate rich embeds and previews of shared URLs for its customers' apps and sites. Embedly's FAQ publishes a fixed, small set of source IPs rather than a live CIDR feed.

Facebook External Hit

Verifiable

Meta · Social previews

Meta's crawler that fetches shared links from Facebook, Instagram, and Messenger to generate rich link previews. Verified by ASN lookup (AS32934); Meta's own docs describe IP-based allowlisting but publish no IP feed.

Hatena service fetchers

Listed only

Hatena Co., Ltd. · Social previews

The fetchers Hatena's services use to collect information from pages its users link to. Hatena-Favicon identifies a page's favicon, Hatena::Scissors retrieves image thumbnails, HatenaBookmark fetches article information for Hatena Bookmark, Hatena Star associates stars with pages, and Hatena Antenna fetches update differences for the pages a user is monitoring.

HatenaBlog-bot

Listed only

Hatena Co., Ltd. · Social previews

Hatena Blog's fetcher. It retrieves a linked page when a blog author asks for it: the :title option of Hatena's URL notation pulls the page title, the :embed option collects the title, summary and favicon used to render a blog card, and the blog import feature fetches images so they can be re-uploaded to Hatena Fotolife.

Iframely

Fully verifiable

Iframely · Social previews

Iframely's fetcher, used to generate embeds and link previews of shared URLs for its customers' apps and sites. Iframely publishes live IPv4/IPv6 address lists and documents reverse-DNS resolution to its own domain.

LinkedInBot

Listed only

LinkedIn · Social previews

LinkedIn's crawler that fetches shared links to build post link previews. LinkedIn's own robots.txt names the "LinkedInBot" token, but LinkedIn publishes no dedicated bot-documentation page, IP list, ASN, or reverse-DNS pattern for this crawler.

LivelapBot

Listed only

Livelap · Social previews

LivelapBot is the crawler for Livelap, a content discovery app. It fetches pages shared on social media and crawls RSS feeds on a schedule, indexing HTML and media meta tags to build link previews shown in the Livelap app.

Mastodon (link preview fetcher)

Listed only

Mastodon gGmbH · Social previews

The generic link-preview fetcher built into Mastodon server software (FetchLinkCardService, using the http.rb gem), run independently by every self-hosted instance. There is no single operator, IP range, or ASN to verify against — traffic can originate from any of thousands of instances.

MicrosoftPreview

Fully verifiable

Microsoft · Social previews

Microsoft fetcher, listed among Bing's crawlers, that retrieves pages to generate link previews for Microsoft products and services.

PagePeeker

Listed only

PagePeeker SRL · Social previews

PagePeeker's thumbnailing robot. It captures a screenshot of a page when one of PagePeeker's customers requests a thumbnail for it, refetching a given site at most once every five to seven days.

Pinterestbot

Verifiable

Pinterest · Social previews

Pinterest's crawler, used both for general web indexing and for validating Pin metadata/broken links. Pinterest documents a reverse-DNS verification method rather than a fixed IP list, since only its US-based traffic uses a stable range.

Redditbot

Listed only

Reddit · Social previews

Reddit's on-demand fetcher that retrieves a shared URL's metadata to build the link preview shown on a post. Reddit's robots.txt now disallows all crawling and documents no user agent, IP list, or verification method for this fetcher.

Slack-ImgProxy

Listed only

Slack · Social previews

Slack's image-proxying fetcher, which retrieves and caches images posted in channels while stripping referrer data. Slack documents the User-Agent string but publishes no IP ranges or other verification method.

Slackbot (Link Expanding)

Listed only

Slack · Social previews

Slack's crawler that expands links posted in channels into rich previews by reading oEmbed, Twitter Card, and Open Graph metadata. Slack documents the User-Agent strings but publishes no IP ranges or other verification method.

Snap URL Preview Service

Listed only

Snap Inc. · Social previews

Snapchat's link-preview fetcher, which scans the HTML of URLs shared in Snapchat chats to build preview cards from Open Graph or Twitter Card tags. Snap documents that responses are cached for 30 minutes to reduce traffic; no IP ranges or robots.txt behavior are documented.

StartmeBot

Listed only

Start.me · Social previews

Start.me's fetcher. When a user adds a link or an RSS widget to their Start.me start page, this bot retrieves three things from the target site: the page title, the site's favicon, and the contents of any RSS feed the site offers.

TelegramBot (link preview)

Listed only

Telegram · Social previews

Telegram's fetcher for generating link previews in chats. Telegram's docs document static CIDR ranges only for inbound webhook delivery to bot servers, not for this outbound preview fetcher, so no IP verification is attached here.

Twitterbot

Listed only

X · Social previews

X (formerly Twitter)'s crawler that fetches shared links to render Card previews. X's own developer docs give no IP ranges, ASN, or reverse-DNS verification method; the commonly cited "*.twttr.com" rDNS check and aggregate IP ranges trace only to community write-ups, not official docs.

vkShare

Listed only

VK · Social previews

vkShare is VK's link-preview fetcher. When a user shares a URL on the VK social network, it retrieves the target page to build a preview (title, description, and image) for the shared post.

WhatsApp Link Preview Fetcher

Verifiable

Meta · Social previews

WhatsApp's fetcher that requests a shared URL to build the link preview shown in chat. Meta documents the request's User-Agent format but no IP feed; verified by ASN lookup (AS32934), same as Meta's other crawlers.

Yahoo Link Preview

Listed only

Yahoo · Social previews

Yahoo Link Preview fetches a page when a Yahoo Mail user includes its URL in an email message, generating a thumbnail, title and description preview. It fetches only user-referenced pages rather than crawling the web.

Monitoring 37

AwarioBot

Listed only

Awario · Monitoring

Awario's social-listening and media-monitoring crawlers, which fetch web pages and RSS feeds to find mentions of the brands and keywords its customers track. Awario states the bots respect robots.txt and honour Crawl-delay, and explicitly asks site owners not to identify them by IP because it uses no consecutive IP blocks.

Azure Application Insights Availability

Listed only

Microsoft · Monitoring

Synthetic availability-test agents from Microsoft Azure Application Insights that send recurring HTTP requests to user-configured URLs to monitor uptime and responsiveness from multiple Azure regions.

Better Stack Uptime

Fully verifiable

Better Stack · Monitoring

Better Stack's uptime-monitoring probes, which request customer URLs on a schedule from a published, frequently-changing list of monitoring IPs.

BrandVerity

Listed only

BrandVerity · Monitoring

BrandVerity is a brand-protection service that crawls websites and paid search landing pages to detect trademark and affiliate-compliance violations on behalf of its clients.

magpie-crawler

Listed only

Brandwatch · Monitoring

magpie-crawler is Brandwatch's web crawler. It downloads publicly available blog, forum, news, and social-media pages to be indexed and analysed by Brandwatch's social-listening platform. It respects robots.txt.

Checkly

Fully verifiable

Checkly · Monitoring

Synthetic-monitoring service that runs API and browser checks against customer endpoints from a published, fixed set of static outbound IPs.

Cloudflare Diagnostics

Listed only

Cloudflare · Monitoring

Cloudflare's support-diagnostics fetcher. Cloudflare documents that requests with this user agent are triggered when Cloudflare Support Engineers perform error checks, and by the continuous monitoring that raises alerts in the Cloudflare dashboard.

Cloudflare Health Checks

Listed only

Cloudflare · Monitoring

Cloudflare's Health Checks probes, which monitor customer origin servers from Cloudflare data centers in regions the customer selects. Cloudflare documents a fixed User-Agent format embedding the first 16 characters of the health check ID, and recommends matching on it; its health-checks docs do not tie probe traffic to the published Cloudflare IP lists, so no CIDR recipe is attached.

Cloudflare Prefetch

Listed only

Cloudflare · Monitoring

Cloudflare's prefetch bot, an Enterprise speed-optimization feature that pre-populates the CDN cache with content a visitor is likely to request next, using a site-owner-supplied manifest of URLs. Requests carry the CloudFlare-Prefetch user-agent.

Cloudflare Traffic Manager

Listed only

Cloudflare · Monitoring

Cloudflare's Load Balancing monitor, which sends health-check requests to customer origin pools at regular intervals to evaluate endpoint health. It embeds the load-balancer pool id in a fixed User-Agent.

Cookiebot Scanner

Verifiable

Cookiebot (Usercentrics) · Monitoring

Cookiebot's cookie-consent compliance scanner. It crawls the domains its customers have registered, on a roughly monthly schedule, to detect the cookies and tracking technologies in use and generate a cookie declaration. Cookiebot documents that scans run only from a fixed pool of addresses.

CookieHub Scanner

Fully verifiable

CookieHub · Monitoring

CookieHub's cookie-consent compliance scanner. It crawls customer domains on a roughly monthly schedule to detect cookies and tracking technologies in use and generate a compliance report.

Datadog Synthetics

Fully verifiable

Datadog · Monitoring

Datadog Synthetic Monitoring probes that run customer-configured API and browser tests against endpoints. API tests send a "Datadog/Synthetics" User-Agent; browser tests append "DatadogSynthetics" to a browser UA. Datadog publishes probe IP ranges as a machine-readable feed at ip-ranges.datadoghq.com/synthetics.json.

Dead Link Checker

Verifiable

DLC Websites · Monitoring

Dead Link Checker is a hosted service that crawls a submitted website to find broken links and report them, with paid tiers that run automated periodic scans and email the results.

Dubbotbot

Verifiable

DubBot · Monitoring

The crawler behind DubBot's web-governance platform. It inventories a customer's own website by following every link from a supplied URL and checks the pages for accessibility, broken links, spelling and content-policy problems. DubBot runs it from AWS on static IP addresses that the operator publishes for allowlisting.

Dynatrace Synthetic Monitoring

Listed only

Dynatrace · Monitoring

Dynatrace's synthetic monitoring runs browser and HTTP checks against sites its customers configure. Dynatrace always appends a RuxitSynthetic token to the user agent — even when the customer sets a custom one — so that synthetic traffic can be identified in server logs. Checks are user-configured, not crawling, so robots.txt is not part of the documented behaviour.

FreeWebMonitoring SiteChecker

Verifiable

GreenWave Online Inc. · Monitoring

The website-monitoring robot of the FreeWebMonitoring service. It only checks URLs that registered members have submitted, and the operator documents that all checks originate from a single server address. GreenWave Online also warns that an older 0.1 agent name is forged by an unrelated scanner.

Freshping

Verifiable

Freshworks · Monitoring

Freshworks' free uptime-monitoring service, which checks customer sites from a small, fixed set of documented monitoring-location IPs rather than a live feed.

Gnowit Newsbot

Listed only

Gnowit · Monitoring

Gnowit is an Ottawa-based media and government monitoring company whose crawler continuously fetches news, government, and other public web sources to power real-time monitoring, summarization, and analytics for its clients.

Grafana Synthetic Monitoring

Fully verifiable

Grafana Labs · Monitoring

Grafana Cloud Synthetic Monitoring public probes, which run customer-configured checks (HTTP, browser, and protocol) from Grafana-run locations. HTTP checks send a documented synthetic-monitoring-agent User-Agent by default (customers can override it). Grafana publishes probe source IPs as a machine-readable feed at allowlists.grafana.com/synthetics.

HetrixTools

Fully verifiable

HetrixTools · Monitoring

Uptime and blacklist monitoring service that probes customer sites from a published list of monitoring-node IPs.

Hydrozen

Fully verifiable

Hydrozen · Monitoring

Hydrozen.io is an uptime and website monitoring service; its checker fetches monitored endpoints from a documented, published set of IP addresses.

Buck

Listed only

Hypefactors · Monitoring

Buck is Hypefactors' media-monitoring crawler. It discovers and indexes web pages by following links, gathering publicly available content for the company's media-monitoring platform. It identifies itself and respects robots.txt.

Mediatoolkitbot

Listed only

Determ · Monitoring

Mediatoolkitbot is the web crawler operated by Determ (formerly Mediatoolkit), a media monitoring platform. It fetches publicly available web content to match brand mentions and topics that Determ's customers track.

Neticle Crawler

Listed only

Neticle Technologies · Monitoring

Neticle's in-house crawler for its social and online media-monitoring service. It continuously fetches newly published web content so Neticle can score sentiment around the keywords, brands and products its customers track.

New Relic Synthetics

Verifiable

New Relic · Monitoring

New Relic's synthetic-monitoring minions, which probe customer-configured URLs from public minion IPs. New Relic documents an X-Abuse-Info request header identifying the monitor and account; logs also record a NewRelicbot User-Agent. New Relic states the published ranges "are reserved for use by New Relic and cannot be used by anyone else". The published IP list is a JSON object keyed by location, which this directory's feed formats cannot consume, so the ranges are recorded statically instead.

Pingdom

Fully verifiable

SolarWinds · Monitoring

Uptime and page-speed monitoring service that probes customer sites from a published list of probe-server IPs. Owned by SolarWinds.

Amazon Route 53 Health Checks

Listed only

Amazon Web Services · Monitoring

Amazon Route 53 health checkers, which probe customer-configured endpoints from AWS data centers worldwide to drive DNS failover. AWS publishes checker source ranges only inside the general ip-ranges.amazonaws.com/ip-ranges.json feed (entries tagged ROUTE53_HEALTHCHECKS), a shape this directory's feed formats cannot consume, so no CIDR recipe is attached.

SentiBot

Fully verifiable

SentiOne · Monitoring

SentiOne's social-listening crawler. It indexes user-generated content for the SentiOne Listen platform, which its operator says analyses over 300,000 domains daily. Robots.txt rules written for "sentibot" are honoured, and Yandex-style reverse-DNS plus a published IP list allow verification.

Sentry Uptime Monitoring

Fully verifiable

Sentry · Monitoring

Sentry's uptime monitoring bot, which probes customer-configured URLs on a schedule from Sentry's uptime-check infrastructure and raises alerts on failures. Sentry documents a fixed User-Agent and publishes the current uptime-check IP addresses at a machine-readable endpoint.

Site24x7

Fully verifiable

Zoho · Monitoring

Zoho's infrastructure and website monitoring service, which probes customer sites from 130+ global monitoring locations published as a machine-readable IP list.

StatusCake

Fully verifiable

StatusCake · Monitoring

Uptime and page-speed testing service that requests customer URLs from a global network of test-location servers, published as a machine-readable IP list.

Testomatobot

Fully verifiable

Testomato · Monitoring

Testomato's website-monitoring agent. It downloads pages and resources and submits web forms for the checks its customers configure (uptime, error, SSL, response-time, meta-tag and JSON-LD monitoring), on a schedule the customer sets. Testomato publishes a plain-text list of the addresses its monitoring nodes request from.

Trendiction Bot

Listed only

Trendiction · Monitoring

Trendiction operates a crawler that fetches public news sites, message boards, and blogs to build a search index and feed media-monitoring and market-research analytics for its clients.

updown.io

Fully verifiable

updown.io · Monitoring

Website uptime-monitoring service that checks customer sites from a published JSON array of node IPs. Documents "updown.io" as the common substring of its requests rather than one fixed User-Agent.

Uptime.com

Listed only

Uptime.com · Monitoring

Uptime.com's synthetic-monitoring probes. They request the URLs its customers configure as checks, from probe servers in many locations, including headless-browser transaction checks that execute JavaScript. Uptime.com publishes its probe IP list only inside the authenticated dashboard, so the user agent is the public identifier.

UptimeRobot

Fully verifiable

UptimeRobot · Monitoring

Uptime monitoring service that probes customer sites on a schedule from a published list of monitoring IPs.

SEO tools 42

AhrefsSiteAudit

Fully verifiable

Ahrefs · SEO tools

Ahrefs' on-demand site-auditing crawler, distinct from AhrefsBot, used when Ahrefs customers run the Site Audit tool against their own or a competitor's domain. Shares Ahrefs' published IP-range feed.

AhrefsBot

Fully verifiable

Ahrefs · SEO tools

Ahrefs' primary web crawler, powering the backlink and keyword database behind the Ahrefs SEO platform and the Yep search engine. Crawls from publicly published IP ranges with a matching reverse-DNS suffix.

Audisto Crawler

Fully verifiable

Audisto GmbH · SEO tools

Crawler for Audisto's hosted technical SEO and site-audit platform, fetching pages of the sites its customers analyse. Audisto publishes its crawler addresses as JSON and documents reverse-DNS verification.

Barkrowler

Fully verifiable

Babbar · SEO tools

Barkrowler is the web crawler operated by Babbar (formerly Exensa). It builds and updates Babbar's graph of the web, which powers the company's SEO and link-analysis tools, and applies a politeness delay between requests.

BomboraBot

Listed only

Bombora, Inc. · SEO tools

BomboraBot is Bombora's web crawler. It classifies the content and topics of web pages that carry Bombora's tags so the company can model B2B purchase intent, visiting each tagged page at most once every 30 days.

Botify

Listed only

Botify · SEO tools

Botify's SEO crawler, used by its SiteCrawler product to analyze enterprise websites for search-indexing insights. Botify does not publish a fixed IP range or CIDR list for site owners to allowlist.

BuiltWith

Listed only

BuiltWith Pty Ltd · SEO tools

Crawler for BuiltWith's technology-profiling service, which visits sites and analyses publicly visible markup to determine which web technologies they use. BuiltWith publishes no IP ranges, so requests cannot be verified.

Caliperbot

Verifiable

Conductor · SEO tools

Conductor's single web crawler. It reads the HTML of pages on sites its customers track, recording on-page elements such as title tags, header tags and other metadata for Conductor's SEO and search-visibility reporting. Conductor publishes the address range it crawls from and will lower the crawl rate on request.

Cincraw

Listed only

CINC Corp. (株式会社CINC) · SEO tools

The web crawler operated by CINC, a Japanese data-solutions company, to collect the page data behind its marketing and SEO analytics products. Its documented policy is to fetch page body content, header and HTTP status information and the JS/CSS needed to render a page, then store a rendered screen capture. CINC states that it does not follow advertising links, deletes all cookies between requests, and does not load analytics or ad-measurement tags. No robots.txt policy and no IP ranges are published.

Cocolyzebot

Listed only

Cocolyze · SEO tools

Crawler for Cocolyze's SEO analysis platform, fetching pages of sites its users analyse. Cocolyze publishes no IP ranges, so requests cannot be verified beyond the user agent.

cognitiveSEO

Listed only

cognitiveSEO · SEO tools

James BOT is the web crawler operated by cognitiveSEO, an SEO toolset. It crawls the web and analyzes links to power the backlink and SEO analysis offered by the cognitiveSEO platform.

CriteoBot

Verifiable

Criteo · SEO tools

Criteo's advertising crawler. It fetches merchant and publisher pages to extract product and content data used for Criteo's commerce and retargeting ads, and respects robots.txt and crawl-delay directives.

DataForSeoBot

Verifiable

DataForSEO · SEO tools

DataForSEO's crawler. It fetches pages to build the backlink and SEO datasets that power DataForSEO's marketing-data APIs, and honours robots.txt and crawl-delay directives.

Dataproviderbot

Verifiable

Dataprovider.com · SEO tools

Dataprovider.com's in-house crawler. It indexes more than 400 million domains each month and structures what it finds into the company's web dataset (business information, technology detection, classifications and risk signals). The operator documents that it follows the robot exclusion protocol and that its crawlers can be identified by a reverse DNS lookup.

DotBot

Listed only

Moz · SEO tools

Moz's general-purpose web crawler, distinct from rogerbot, that gathers link data powering the Moz Link Index and Link Explorer. Moz's own help pages document no fixed IP range for it.

Dragonbot

Listed only

Dragon Metrics · SEO tools

Dragon Metrics' SEO crawler, which collects data for the platform's Site Audit and Site Explorer features. Its operator documents that it respects robots.txt using Google's open-source parser, and that it crawls from dynamic IP addresses so it can only be identified by user agent.

HubSpot Crawler

Verifiable

HubSpot · SEO tools

HubSpot's crawler, which fetches customer and external pages to power the SEO recommendations and link analysis in HubSpot's marketing tools. HubSpot publishes its egress ranges tagged by service, including web crawling.

IAS Crawler

Listed only

Integral Ad Science · SEO tools

Content-rating and ad-verification crawler operated by Integral Ad Science. It visits web pages to assess content quality and brand safety and to support invalid-traffic detection for advertisers.

Linkdexbot

Listed only

Authoritas (Analytics SEO Limited) · SEO tools

Linkdexbot is the web crawler for Linkdex, an SEO and search-marketing analytics platform now operated by Authoritas. It gathers link and page data used to power the platform's SEO reporting tools.

MegaIndex Crawler

Listed only

MegaIndex · SEO tools

MegaIndex is an SEO and web-analytics platform whose crawler indexes links across the web to power backlink analysis, keyword tracking, and site audits for its subscribers.

Meta External Ads

Verifiable

Meta · SEO tools

Meta's crawler that fetches pages for advertising and other business-related products and services, separate from the AI-training and link-preview crawlers. Verified by ASN lookup (AS32934); Meta publishes no IP feed.

MTRobot

Listed only

Metrics Tools (Andreas Knatz) · SEO tools

Crawler for Metrics Tools, a German SEO analytics service, collecting page data for its visibility and ranking analyses. The operator publishes no IP ranges, so requests cannot be verified beyond the user agent.

MJ12bot

Listed only

Majestic-12 · SEO tools

The crawler behind Majestic's backlink index. Majestic explicitly states it is a community-based distributed crawler with no fixed IP allocation, so requests cannot be verified by IP, ASN, or reverse DNS.

OnCrawl

Listed only

OnCrawl · SEO tools

OnCrawl's SEO crawler, used to analyze a customer's own site structure and content for technical SEO reporting. OnCrawl's help docs describe no fixed IP range; the bot's identity is user-configurable per crawl.

Outbrain crawler

Verifiable

Outbrain · SEO tools

Outbrain's content-recommendation crawler. It fetches advertiser landing pages so Outbrain's system can pull the correct image and headline for a promoted-content unit, and rejects submitted URLs it cannot reach.

Panscient Crawler

Listed only

Panscient Inc. · SEO tools

Panscient's large-scale crawler, which traverses public websites so that Panscient can build structured company and professional data feeds licensed to enterprise customers. The operator documents a full-corpus refresh each quarter, a rate limit of at most one request per second to any single domain, and compliance with the Robot Exclusion Standard. A separate "pantest" agent is used for testing. No IP ranges are published.

Proximic (Comscore Crawler)

Listed only

Comscore, Inc. · SEO tools

Proximic is Comscore's web crawler. It downloads the static textual content of pages to perform contextual analysis (content language, rating, and IAB categories) so advertising partners can match campaigns to page content. It identifies itself and honors robots.txt.

Rogerbot

Listed only

Moz · SEO tools

Moz's site-audit crawler for Moz Pro Campaigns, distinct from DotBot. Moz's own FAQ states plainly that Rogerbot has no IP range: "we do not use a static IP address or range of IP addresses."

RyteBot

Listed only

Semrush · SEO tools

The crawler behind the Ryte.com tools, which analyse on-page SEO, technical and usability issues. Ryte was absorbed by Semrush, and RyteBot is now documented as a member of the Semrush bot family with its own robots.txt user agent. No IP ranges are published for it.

Search Atlas Bot

Listed only

Search Atlas · SEO tools

The crawler behind Search Atlas's SEO platform, which fetches pages for its Site Auditor and monitoring features. The operator publishes the bot's user agent for allowlisting and states that the crawler does not use static IP addresses, so it can only be identified by its user agent.

SiteAuditBot

Verifiable

Semrush · SEO tools

Semrush's site-auditing crawler, distinct from SemrushBot: it crawls a domain on demand when a Semrush customer runs the Site Audit tool, looking for SEO and technical issues. Unlike the backlink crawler, which Semrush says cannot be identified by IP, Site Audit is documented as running from a single dedicated subnet.

SemrushBot

Listed only

Semrush · SEO tools

Semrush's web crawler, feeding the backlink and site-audit data behind the Semrush SEO platform. Semrush's own bot page explicitly states it does not use consecutive IP blocks, so no CIDR list can be sourced.

SemrushBot-SI

Verifiable

Semrush · SEO tools

The crawler behind Semrush's On Page SEO Checker and related on-page tools, run against a domain when a Semrush customer sets up a campaign for it. It is a separate robots.txt user agent from SemrushBot, and Semrush documents its own addresses to allowlist for it.

SeobilityBot

Fully verifiable

Seobility GmbH · SEO tools

Crawler for Seobility's hosted SEO analysis and site-audit tooling, fetching pages of sites its customers analyse. Seobility publishes a machine-readable list of the addresses its bots crawl from.

SEOkicks

Listed only

Jobkicks SLU · SEO tools

SEOkicks operates a web crawler that builds a backlink database powering its SEO tools. The crawler visits sites to collect link data for analysis.

serpstatbot

Fully verifiable

Serpstat · SEO tools

Serpstat's backlink crawler. It continuously crawls the web to add new links and track changes in Serpstat's link database, honouring robots.txt and Crawl-delay directives, and publishes the full list of addresses it crawls from.

SISTRIX Crawler

Verifiable

SISTRIX · SEO tools

The crawler behind the SISTRIX Toolbox, a German SEO visibility platform. SISTRIX documents that every crawler IP resolves via reverse DNS to the "sistrix.net" domain rather than publishing a static CIDR list.

Siteimprove Crawler

Verifiable

Siteimprove · SEO tools

Siteimprove's content-suite crawler, which fetches pages of sites its customers have configured in their account to run quality-assurance, accessibility, policy and SEO checks. Companion agents (LinkCheck, Image size, Probe) fetch links and resources for the same checks and crawl from the same published address list.

t3versionsBot

Listed only

Torben Hansen (t3versions) · SEO tools

Private-project crawler that makes single GET requests to sites and looks for TYPO3 fingerprints, collecting statistics on the worldwide usage and development of the open-source TYPO3 CMS. No IP ranges are published, and the operator documents no robots.txt support (exclusion is by email request).

VelenPublicWebCrawler

Listed only

Hunter · SEO tools

Hunter's public web crawler, written in Go. It analyses millions of publicly accessible pages every month to build the business datasets and machine learning models behind Hunter's products, and never fetches anything behind a login. The operator documents a deliberate rate limit of one page at a time and one page every two seconds per site.

XoviBot

Listed only

Xovi GmbH · SEO tools

XoviBot is the web crawler for XOVI, an SEO and online-marketing analytics suite. It crawls sites to gather backlink and ranking data for the platform's SEO tools.

Zoominfobot

Listed only

ZoomInfo Technologies · SEO tools

ZoomInfo's indexing robot, which scans corporate websites, press releases, news services and SEC filings to build ZoomInfo's search index of businesses and business professionals. The operator documents that it obeys robots.txt, spaces out requests on larger sites and never opens more than one connection to a site at a time. No IP ranges are published.

Feed fetchers 13

Apple Podcasts (iTMS)

Verifiable

Apple · Feed fetchers

Apple's podcast feed fetcher, which crawls only URLs associated with content registered on Apple Podcasts. Apple documents that iTMS traffic may come from applebot.apple.com hosts and that it does not follow robots.txt because it is not a general search crawler.

BazQux Fetcher

Listed only

BazQux Reader · Feed fetchers

The feed fetcher of the BazQux Reader hosted RSS service. It retrieves and periodically refreshes the RSS/Atom and comment feeds that users have subscribed to, typically no more than once an hour per feed. BazQux documents that the fetcher acts as an agent of those users and therefore ignores robots.txt.

Facebook Catalog

Listed only

Meta · Feed fetchers

Meta's product-catalog fetcher, identified by the facebookcatalog user-agent. It retrieves merchant product-data feeds used to build and refresh commerce catalogs surfaced across Facebook and Instagram.

Feedbin

Verifiable

Feedbin · Feed fetchers

Feedbin's feed fetcher, which retrieves RSS/Atom feeds that users have subscribed to. Its user agent includes the internal feed id and current subscriber count, and Feedbin documents forward-confirmed reverse DNS in *.bot.feedbin.com as the way to verify its requests.

Feedly Fetcher

Listed only

Feedly · Feed fetchers

Feedly's fetcher, which retrieves RSS/Atom feed URLs after a user has explicitly added them to their Feedly. Feedly documents that it behaves as a direct agent of the user rather than a robot, and does not publish a fixed IP list because its source IPs change over time.

Feedspot

Listed only

Feedspot · Feed fetchers

Feedspot is a hosted content reader and feed aggregation service. Its bot fetches RSS and Atom feeds and web content on behalf of Feedspot users.

Flipboard Proxy

Listed only

Flipboard, Inc. · Feed fetchers

Flipboard's proxy service, which fetches and prepares elements of a page (e.g. a social feed a user asked Flipboard to scan) for presentation in the Flipboard app. Flipboard's own docs say these requests currently originate from an Amazon EC2 cluster but publish no fixed IP list.

Feedfetcher-Google

Fully verifiable

Google · Feed fetchers

Google's feed retrieval agent for RSS and Atom feeds used by Google News and WebSub. It fetches and periodically refreshes feeds that users of an app or service have explicitly subscribed to.

Hatena::Russia::Crawler

Listed only

Hatena Co., Ltd. (Hatelabo) · Feed fetchers

The fetcher behind Daichecker, the antenna service run on Hatelabo, Hatena's experimental-services lab. It checks the pages and feeds that users have registered for updates. Hatena documents that it parses only the robots.txt groups that name this user agent directly and does not apply the User-agent: * group, so a wildcard rule will not stop it. Hatena notes the name comes from an internal code name for RSS-reader development and has no connection to the country.

Inoreader Fetcher

Fully verifiable

Innologica · Feed fetchers

Inoreader's feed fetcher, which retrieves RSS/Atom feeds that Inoreader users have subscribed to. Its own docs state it does not read robots.txt because it fetches specific, user-requested feed URLs rather than crawling a site, and it publishes a live list of its backend fetcher IPs.

Miniflux

Listed only

Miniflux · Feed fetchers

Miniflux is a minimalist, open-source, self-hosted feed reader. User-run instances fetch the RSS and Atom feeds their subscribers add, identifying themselves with a Miniflux User-Agent.

NewsBlur Feed Fetcher

Listed only

NewsBlur · Feed fetchers

NewsBlur's open-source feed fetcher, which polls RSS/Atom feeds on behalf of subscribed users. Its user agent embeds the live subscriber count and the feed's permalink; NewsBlur publishes no fixed IP range for it.

Superfeedr

Verifiable

Superfeedr · Feed fetchers

Superfeedr's PubSubHubbub feed-polling infrastructure, which fetches feed URLs that publishers or subscribers have registered with the service. Its docs publish a list of current node IPs but warn it changes as they add or remove cloud capacity.

Archivers 7

AcademicBotRTU

Listed only

Riga Technical University (Institute of Applied Computer Systems) · Archivers

Crawler run by Riga Technical University that indexes websites and documents to compare against student and researcher works for plagiarism detection. The operator publishes no IP ranges, so requests cannot be verified.

Arquivo.pt Web Crawler

Listed only

Arquivo.pt (FCCN/FCT) · Archivers

The crawler behind Arquivo.pt, Portugal's public web archive, which captures full page renders (HTML, CSS, JS, images) for long-term preservation. Built on Heritrix; the operator documents no fixed IP range.

BnF Web Archiving Robot

Listed only

Bibliothèque nationale de France · Archivers

The web crawler of the Bibliothèque nationale de France, which harvests French websites for the legal deposit of the web to preserve the national documentary heritage. It runs on Heritrix and applies request delays to avoid overloading servers.

Cloudflare Always Online

Listed only

Cloudflare · Archivers

Cloudflare's Always Online crawler, which fetches pages from sites that have the feature enabled so a cached copy can be served to visitors when the origin server is unreachable. Cloudflare's crawler reference documents the CloudFlare-AlwaysOnline user agent for this product.

archive.org_bot

Listed only

Internet Archive · Archivers

The Internet Archive's Heritrix-based crawler used for its wide crawl of the web, feeding the Wayback Machine. The Archive says it crawls slowly to avoid disrupting sites and publishes no IP ranges. Its help pages note that robots exclusions may prevent archiving, but the operator does not document a commitment to obey robots.txt across its crawls.

ArchiveBot

Listed only

Archive Team · Archivers

ArchiveBot is an IRC-controlled archiving bot run by Archive Team that crawls websites on request, writes WARC files, and uploads the captures to the Internet Archive.

TurnitinBot

Listed only

Turnitin · Archivers

TurnitinBot is the web crawler operated by Turnitin. It collects publicly available web pages to build the content database used by Turnitin's academic-integrity and plagiarism-detection services.

Security scanners 13

Censys Inspect

Verifiable

Censys · Security scanners

Censys' internet-wide scanner, which probes public IP addresses to build the host/service data behind Censys Search. Censys publishes a fixed set of scanner subnets and ASNs for opt-out purposes.

Cloudflare Custom Hostname Verification

Listed only

Cloudflare · Security scanners

Cloudflare's custom-hostname ownership checker. Cloudflare documents that requests carrying this user agent are triggered when a customer chooses to validate a custom hostname with an HTTP ownership token, which requires fetching the token from the hostname being claimed.

Cloudflare SSL Detector

Listed only

Cloudflare · Security scanners

Cloudflare's SSL/TLS Recommender crawler, which fetches a customer origin over both HTTP and HTTPS to determine whether the site can be served fully over HTTPS and to recommend TLS settings.

Cortex Xpanse

Verifiable

Palo Alto Networks · Security scanners

Palo Alto Networks' attack-surface-management scanner. It continuously scans the global internet from a published set of ranges to map its customers' internet-facing assets and discover emerging threats, and its requests carry a plain-English user agent naming the company and an opt-out contact address.

Detectify

Verifiable

Detectify · Security scanners

Detectify's external attack-surface and vulnerability scanner, which probes customer-configured assets from a documented set of AWS-hosted source IPs. The operator publishes both the source IPs and the scanner's user-agent strings.

Driftnet Internet Measurement

Verifiable

Driftnet · Security scanners

Driftnet's internet-measurement scanner, run from the internet-measurement.com domain, which probes publicly exposed services to give network owners an external view of their infrastructure. The operator states this traffic never attempts to log in to systems and publishes its scanner IP ranges.

Jugendschutzprogramm-Crawler

Listed only

JusProg e.V. · Security scanners

Web crawler operated by JusProg e.V., a German non-profit youth-protection association, that fetches and rates web pages to build the age-classification database used by its parental-control filtering software.

LeakIX (l9explore)

Verifiable

LeakIX · Security scanners

LeakIX's internet-wide recon scanner (l9explore), which probes exposed services to populate the LeakIX misconfiguration/vulnerability search engine. Its probe fleet publishes a live, per-host list of source IPs at scan.leakix.net whose hostnames resolve under scan.leakix.org, and its scanning tool sets an identifiable user-agent.

Qualys SSL Labs

Verifiable

Qualys · Security scanners

Qualys' SSL Labs server test, which runs visitor-initiated and monthly (SSL Pulse) TLS-configuration assessments of public HTTPS sites. Qualys documents the assessments as slow and non-intrusive and publishes the scanner's source IP ranges in its support knowledge base.

Streamline3Bot

Verifiable

UBT (EU) Ltd · Security scanners

UBT's web crawler, which powers a classification service that categorises public websites by their content. It re-crawls a given site roughly once every three days. UBT publishes no IP list, and instead documents reverse-DNS verification against the ubtsupport.com domain.

Stripebot

Verifiable

Stripe · Security scanners

Stripe's automated web crawler. Collects data from Stripe users' websites so Stripe can provide its services and comply with financial regulations. Distinct from Stripe's webhook delivery traffic.

SurdotlyBot

Listed only

Sur.ly · Security scanners

Sur.ly's crawler. The operator runs a spam-fighting link-safety service and uses this bot to fetch third-party sites and build a short security profile for each one, querying metadata and favicons and taking a screenshot of the homepage. The operator states it never harvests e-mail addresses or content unrelated to security.

W3C Markup Validator

Verifiable

World Wide Web Consortium (W3C) · Security scanners

The W3C Markup Validation Service's fetcher, used when a user submits a URL to be checked for HTML conformance. W3C documents a fixed source address for its validation services alongside the user-agent string.

Webhooks 8

Adyen Webhooks

Verifiable

Adyen · Webhooks

Adyen's payment-event webhook delivery service. Not a crawler; POSTs notifications to merchant endpoints. Adyen documents its egress domain for DNS-based allowlisting rather than a static IP list, since its outbound IPs change over time.

GitHub Webhooks

Fully verifiable

GitHub · Webhooks

GitHub's webhook delivery service. Not a crawler; sends repository and organization event notifications to configured endpoints from IP ranges published in GitHub's meta API.

APIs-Google

Fully verifiable

Google · Webhooks

Google's agent that delivers push notification messages sent through Google APIs (such as Pub/Sub and WebSub push subscriptions) to subscriber endpoints.

PayPal IPN

Verifiable

PayPal · Webhooks

PayPal's Instant Payment Notification service. Not a crawler; POSTs payment-event notifications to merchant listener endpoints from a documented, static set of server CIDR ranges (shared with other PayPal server traffic).

Stripe Webhooks

Fully verifiable

Stripe · Webhooks

Stripe's webhook delivery service. Not a crawler; sends event notifications (payments, subscriptions) to merchant endpoints from published IPs.

Svix Webhooks

Verifiable

Svix · Webhooks

Webhook-sending infrastructure used by Svix's customers to deliver events. Not a crawler; Pro/Enterprise plans get a documented, static set of per-region source IPs, and requests include a Svix sender identifier in the User-Agent.

Telegram Bot Webhooks

Verifiable

Telegram · Webhooks

Telegram's webhook delivery to bot servers. Not a crawler; POSTs update events to a bot's registered HTTPS endpoint from two documented, static CIDR ranges. Telegram documents no User-Agent for these POSTs; the UA listed here is Telegram's documented fetcher token and is unconfirmed for webhook traffic — verify by source IP, not UA.

Twilio Webhooks

Listed only

Twilio · Webhooks

Twilio's webhook delivery service, which POSTs event callbacks (incoming messages and calls, status updates) to customer-configured endpoints. Not a crawler. Twilio states there is no fixed range of source IPs — requests come from a dynamic pool — so recipients are told to validate the X-Twilio-Signature request signature instead.