Directory / SEO tools
SEO tools 51
AhrefsSiteAudit
Fully verifiableAhrefs · SEO tools
Ahrefs' on-demand site-auditing crawler, distinct from AhrefsBot, used when Ahrefs customers run the Site Audit tool against their own or a competitor's domain. Shares Ahrefs' published IP-range feed.
AhrefsBot
Fully verifiableAhrefs · SEO tools
Ahrefs' primary web crawler, powering the backlink and keyword database behind the Ahrefs SEO platform and the Yep search engine. Crawls from publicly published IP ranges with a matching reverse-DNS suffix.
AudigentAdBot
Listed onlyAudigent · SEO tools
Audigent's advertising crawler. The operator documents that it collects only the metadata in the header of an HTML page and does not scrape page body content, and that it must be named explicitly in robots.txt because a wildcard rule does not block it.
Audisto Crawler
Fully verifiableAudisto GmbH · SEO tools
Crawler for Audisto's hosted technical SEO and site-audit platform, fetching pages of the sites its customers analyse. Audisto publishes its crawler addresses as JSON and documents reverse-DNS verification.
Barkrowler
Fully verifiableBabbar · SEO tools
Barkrowler is the web crawler operated by Babbar (formerly Exensa). It builds and updates Babbar's graph of the web, which powers the company's SEO and link-analysis tools, and applies a politeness delay between requests.
BomboraBot
Listed onlyBombora, Inc. · SEO tools
BomboraBot is Bombora's web crawler. It classifies the content and topics of web pages that carry Bombora's tags so the company can model B2B purchase intent, visiting each tagged page at most once every 30 days.
Botify
Listed onlyBotify · SEO tools
Botify's SEO crawler, used by its SiteCrawler product to analyze enterprise websites for search-indexing insights. Botify does not publish a fixed IP range or CIDR list for site owners to allowlist.
Brightbot
VerifiableBright Data · SEO tools
Bright Data's data-collection crawler, documented as the main collection pipeline for its products, with a 24-hour cache layer to avoid re-downloading the same page. The operator states Brightbot is deliberately transparent — a unique user agent plus a single published source subnet — so its traffic can be separated from user traffic. Its documented control mechanism is Bright Data's own collectors.txt file and Web Master console; the page states no robots.txt behaviour.
BuiltWith
Listed onlyBuiltWith Pty Ltd · SEO tools
Crawler for BuiltWith's technology-profiling service, which visits sites and analyses publicly visible markup to determine which web technologies they use. BuiltWith publishes no IP ranges, so requests cannot be verified.
Caliperbot
VerifiableConductor · SEO tools
Conductor's single web crawler. It reads the HTML of pages on sites its customers track, recording on-page elements such as title tags, header tags and other metadata for Conductor's SEO and search-visibility reporting. Conductor publishes the address range it crawls from and will lower the crawl rate on request.
Cincraw
Listed onlyCINC Corp. (株式会社CINC) · SEO tools
The web crawler operated by CINC, a Japanese data-solutions company, to collect the page data behind its marketing and SEO analytics products. Its documented policy is to fetch page body content, header and HTTP status information and the JS/CSS needed to render a page, then store a rendered screen capture. CINC states that it does not follow advertising links, deletes all cookies between requests, and does not load analytics or ad-measurement tags. No robots.txt policy and no IP ranges are published.
Claritybot
Listed onlyseoClarity · SEO tools
seoClarity's page and site audit crawler. Crawls are triggered on demand by seoClarity clients to analyse pages for technical and content issues, and clients can also schedule daily managed page crawls. The operator documents that it obeys robots.txt and Crawl-delay, and that its addresses are dynamic.
Cocolyzebot
Listed onlyCocolyze · SEO tools
Crawler for Cocolyze's SEO analysis platform, fetching pages of sites its users analyse. Cocolyze publishes no IP ranges, so requests cannot be verified beyond the user agent.
cognitiveSEO
Listed onlycognitiveSEO · SEO tools
James BOT is the web crawler operated by cognitiveSEO, an SEO toolset. It crawls the web and analyzes links to power the backlink and SEO analysis offered by the cognitiveSEO platform.
CriteoBot
VerifiableCriteo · SEO tools
Criteo's advertising crawler. It fetches merchant and publisher pages to extract product and content data used for Criteo's commerce and retargeting ads, and respects robots.txt and crawl-delay directives.
DataForSeoBot
VerifiableDataForSEO · SEO tools
DataForSEO's crawler. It fetches pages to build the backlink and SEO datasets that power DataForSEO's marketing-data APIs, and honours robots.txt and crawl-delay directives.
Dataproviderbot
VerifiableDataprovider.com · SEO tools
Dataprovider.com's in-house crawler. It indexes more than 400 million domains each month and structures what it finds into the company's web dataset (business information, technology detection, classifications and risk signals). The operator documents that it follows the robot exclusion protocol and that its crawlers can be identified by a reverse DNS lookup.
DomCopBot
Listed onlyDomCop · SEO tools
DomCop's availability crawler, used by the domain-research service of the same name. The operator documents that it accesses only the robots.txt file, once per domain, and uses that request to establish whether a website is live on the domain.
DotBot
Listed onlyMoz · SEO tools
Moz's general-purpose web crawler, distinct from rogerbot, that gathers link data powering the Moz Link Index and Link Explorer. Moz's own help pages document no fixed IP range for it.
Dragonbot
Listed onlyDragon Metrics · SEO tools
Dragon Metrics' SEO crawler, which collects data for the platform's Site Audit and Site Explorer features. Its operator documents that it respects robots.txt using Google's open-source parser, and that it crawls from dynamic IP addresses so it can only be identified by user agent.
EzoicBot
Listed onlyEzoic · SEO tools
Ezoic's crawler family, run by the digital-publisher technology platform of the same name. A desktop and a mobile variant crawl pages to study how sites, search engines and content interact, alongside named subtypes for Core Web Vitals measurement, uptime checks, ads.txt verification, integration checks and page-topic analysis. All variants share the single robots.txt token EzoicBot.
HubSpot Crawler
VerifiableHubSpot · SEO tools
HubSpot's crawler, which fetches customer and external pages to power the SEO recommendations and link analysis in HubSpot's marketing tools. HubSpot publishes its egress ranges tagged by service, including web crawling.
IAS Crawler
Listed onlyIntegral Ad Science · SEO tools
Content-rating and ad-verification crawler operated by Integral Ad Science. It visits web pages to assess content quality and brand safety and to support invalid-traffic detection for advertisers.
Linkdexbot
Listed onlyAuthoritas (Analytics SEO Limited) · SEO tools
Linkdexbot is the web crawler for Linkdex, an SEO and search-marketing analytics platform now operated by Authoritas. It gathers link and page data used to power the platform's SEO reporting tools.
MegaIndex Crawler
Listed onlyMegaIndex · SEO tools
MegaIndex is an SEO and web-analytics platform whose crawler indexes links across the web to power backlink analysis, keyword tracking, and site audits for its subscribers.
Meta External Ads
VerifiableMeta · SEO tools
Meta's crawler that fetches pages for advertising and other business-related products and services, separate from the AI-training and link-preview crawlers. Verified by ASN lookup (AS32934); Meta publishes no IP feed.
MTRobot
Listed onlyMetrics Tools (Andreas Knatz) · SEO tools
Crawler for Metrics Tools, a German SEO analytics service, collecting page data for its visibility and ranking analyses. The operator publishes no IP ranges, so requests cannot be verified beyond the user agent.
MJ12bot
Listed onlyMajestic-12 · SEO tools
The crawler behind Majestic's backlink index. Majestic explicitly states it is a community-based distributed crawler with no fixed IP allocation, so requests cannot be verified by IP, ASN, or reverse DNS.
Monsidobot
VerifiableAcquia · SEO tools
Acquia Web Governance (formerly Monsido) crawler, which scans the public websites its customers have configured to run accessibility, quality assurance and policy checks. It also issues link-status checks against third-party sites that customers have linked to, preferring HEAD requests for those. Acquia documents a fallback user agent — a plain Chrome string with no identifying token — used only when the primary one fails, so a share of its traffic is identifiable by address rather than by user agent.
Nano Interactive Crawler
VerifiableNano Interactive · SEO tools
Nano Interactive's contextual-advertising crawler. Its published crawler policy lists four desktop and mobile user agents, all carrying the NanoInteractive/1.0 token, and names the two addresses the crawler requests come from so site owners can allow it explicitly.
OnCrawl
Listed onlyOnCrawl · SEO tools
OnCrawl's SEO crawler, used to analyze a customer's own site structure and content for technical SEO reporting. OnCrawl's help docs describe no fixed IP range; the bot's identity is user-configurable per crawl.
Outbrain crawler
VerifiableOutbrain · SEO tools
Outbrain's content-recommendation crawler. It fetches advertiser landing pages so Outbrain's system can pull the correct image and headline for a promoted-content unit, and rejects submitted URLs it cannot reach.
Panscient Crawler
Listed onlyPanscient Inc. · SEO tools
Panscient's large-scale crawler, which traverses public websites so that Panscient can build structured company and professional data feeds licensed to enterprise customers. The operator documents a full-corpus refresh each quarter, a rate limit of at most one request per second to any single domain, and compliance with the Robot Exclusion Standard. A separate "pantest" agent is used for testing. No IP ranges are published.
Proximic (Comscore Crawler)
Listed onlyComscore, Inc. · SEO tools
Proximic is Comscore's web crawler. It downloads the static textual content of pages to perform contextual analysis (content language, rating, and IAB categories) so advertising partners can match campaigns to page content. It identifies itself and honors robots.txt.
Rogerbot
Listed onlyMoz · SEO tools
Moz's site-audit crawler for Moz Pro Campaigns, distinct from DotBot. Moz's own FAQ states plainly that Rogerbot has no IP range: "we do not use a static IP address or range of IP addresses."
RyteBot
Listed onlySemrush · SEO tools
The crawler behind the Ryte.com tools, which analyse on-page SEO, technical and usability issues. Ryte was absorbed by Semrush, and RyteBot is now documented as a member of the Semrush bot family with its own robots.txt user agent. No IP ranges are published for it.
Scope3 Crawler
VerifiableScope3 · SEO tools
Scope3's crawler, which indexes publicly available web content (and paywalled media where the publisher has granted access) to produce content classification and brand-safety assessments for advertising. Scope3 documents adaptive rate limiting of five pages per minute per domain and publishes the single address it crawls from.
Search Atlas Bot
Listed onlySearch Atlas · SEO tools
The crawler behind Search Atlas's SEO platform, which fetches pages for its Site Auditor and monitoring features. The operator publishes the bot's user agent for allowlisting and states that the crawler does not use static IP addresses, so it can only be identified by its user agent.
SiteAuditBot
VerifiableSemrush · SEO tools
Semrush's site-auditing crawler, distinct from SemrushBot: it crawls a domain on demand when a Semrush customer runs the Site Audit tool, looking for SEO and technical issues. Unlike the backlink crawler, which Semrush says cannot be identified by IP, Site Audit is documented as running from a single dedicated subnet.
SemrushBot
Listed onlySemrush · SEO tools
Semrush's web crawler, feeding the backlink and site-audit data behind the Semrush SEO platform. Semrush's own bot page explicitly states it does not use consecutive IP blocks, so no CIDR list can be sourced.
SemrushBot-SI
VerifiableSemrush · SEO tools
The crawler behind Semrush's On Page SEO Checker and related on-page tools, run against a domain when a Semrush customer sets up a campaign for it. It is a separate robots.txt user agent from SemrushBot, and Semrush documents its own addresses to allowlist for it.
SeobilityBot
Fully verifiableSeobility GmbH · SEO tools
Crawler for Seobility's hosted SEO analysis and site-audit tooling, fetching pages of sites its customers analyse. Seobility publishes a machine-readable list of the addresses its bots crawl from.
SEOkicks
Listed onlyJobkicks SLU · SEO tools
SEOkicks operates a web crawler that builds a backlink database powering its SEO tools. The crawler visits sites to collect link data for analysis.
serpstatbot
Fully verifiableSerpstat · SEO tools
Serpstat's backlink crawler. It continuously crawls the web to add new links and track changes in Serpstat's link database, honouring robots.txt and Crawl-delay directives, and publishes the full list of addresses it crawls from.
SISTRIX Crawler
VerifiableSISTRIX · SEO tools
The crawler behind the SISTRIX Toolbox, a German SEO visibility platform. SISTRIX documents that every crawler IP resolves via reverse DNS to the "sistrix.net" domain rather than publishing a static CIDR list.
Siteimprove Crawler
VerifiableSiteimprove · SEO tools
Siteimprove's content-suite crawler, which fetches pages of sites its customers have configured in their account to run quality-assurance, accessibility, policy and SEO checks. Companion agents (LinkCheck, Image size, Probe) fetch links and resources for the same checks and crawl from the same published address list.
t3versionsBot
Listed onlyTorben Hansen (t3versions) · SEO tools
Private-project crawler that makes single GET requests to sites and looks for TYPO3 fingerprints, collecting statistics on the worldwide usage and development of the open-source TYPO3 CMS. No IP ranges are published, and the operator documents no robots.txt support (exclusion is by email request).
TTD-Content
Listed onlyThe Trade Desk · SEO tools
The Trade Desk's content scraper. When a page sends an ad request to The Trade Desk, this crawler scans the page to determine the context in which the ads were displayed, and caches that contextual data for ad serving. The operator publishes a plain-text list of the addresses it crawls from at ttd-content.adsrvr.org/ips; that list currently holds 2,640 individual addresses, which is beyond this directory's per-feed range cap, so it is not recorded as a machine-readable recipe here.
VelenPublicWebCrawler
Listed onlyHunter · SEO tools
Hunter's public web crawler, written in Go. It analyses millions of publicly accessible pages every month to build the business datasets and machine learning models behind Hunter's products, and never fetches anything behind a login. The operator documents a deliberate rate limit of one page at a time and one page every two seconds per site.
XoviBot
Listed onlyXovi GmbH · SEO tools
XoviBot is the web crawler for XOVI, an SEO and online-marketing analytics suite. It crawls sites to gather backlink and ranking data for the platform's SEO tools.
Zoominfobot
Listed onlyZoomInfo Technologies · SEO tools
ZoomInfo's indexing robot, which scans corporate websites, press releases, news services and SEC filings to build ZoomInfo's search index of businesses and business professionals. The operator documents that it obeys robots.txt, spaces out requests on larger sites and never opens more than one connection to a site at a time. No IP ranges are published.