Directory / SEO tools
SEO tools 66
AhrefsSiteAudit
Fully verifiableAhrefs · SEO tools
Ahrefs' on-demand site-auditing crawler, distinct from AhrefsBot, used when Ahrefs customers run the Site Audit tool against their own or a competitor's domain. Shares Ahrefs' published IP-range feed.
AhrefsBot
Fully verifiableAhrefs · SEO tools
Ahrefs' primary web crawler, powering the backlink and keyword database behind the Ahrefs SEO platform and the Yep search engine. Crawls from publicly published IP ranges with a matching reverse-DNS suffix.
AudigentAdBot
Listed onlyAudigent · SEO tools
Audigent's advertising crawler. The operator documents that it collects only the metadata in the header of an HTML page and does not scrape page body content, and that it must be named explicitly in robots.txt because a wildcard rule does not block it.
Audisto Crawler
Fully verifiableAudisto GmbH · SEO tools
Crawler for Audisto's hosted technical SEO and site-audit platform, fetching pages of the sites its customers analyse. Audisto publishes its crawler addresses as JSON and documents reverse-DNS verification.
Barkrowler
Fully verifiableBabbar · SEO tools
Barkrowler is the web crawler operated by Babbar (formerly Exensa). It builds and updates Babbar's graph of the web, which powers the company's SEO and link-analysis tools, and applies a politeness delay between requests.
BomboraBot
Listed onlyBombora, Inc. · SEO tools
BomboraBot is Bombora's web crawler. It classifies the content and topics of web pages that carry Bombora's tags so the company can model B2B purchase intent, visiting each tagged page at most once every 30 days.
Botify
Listed onlyBotify · SEO tools
Botify's SEO crawler, used by its SiteCrawler product to analyze enterprise websites for search-indexing insights. Botify does not publish a fixed IP range or CIDR list for site owners to allowlist.
Brightbot
VerifiableBright Data · SEO tools
Bright Data's data-collection crawler, documented as the main collection pipeline for its products, with a 24-hour cache layer to avoid re-downloading the same page. The operator states Brightbot is deliberately transparent — a unique user agent plus a single published source subnet — so its traffic can be separated from user traffic. Its documented control mechanism is Bright Data's own collectors.txt file and Web Master console; the page states no robots.txt behaviour.
BuiltWith
Listed onlyBuiltWith Pty Ltd · SEO tools
Crawler for BuiltWith's technology-profiling service, which visits sites and analyses publicly visible markup to determine which web technologies they use. BuiltWith publishes no IP ranges, so requests cannot be verified.
Caliperbot
VerifiableConductor · SEO tools
Conductor's single web crawler. It reads the HTML of pages on sites its customers track, recording on-page elements such as title tags, header tags and other metadata for Conductor's SEO and search-visibility reporting. Conductor publishes the address range it crawls from and will lower the crawl rate on request.
Cincraw
Listed onlyCINC Corp. (株式会社CINC) · SEO tools
The web crawler operated by CINC, a Japanese data-solutions company, to collect the page data behind its marketing and SEO analytics products. Its documented policy is to fetch page body content, header and HTTP status information and the JS/CSS needed to render a page, then store a rendered screen capture. CINC states that it does not follow advertising links, deletes all cookies between requests, and does not load analytics or ad-measurement tags. No robots.txt policy and no IP ranges are published.
Claritybot
Listed onlyseoClarity · SEO tools
seoClarity's page and site audit crawler. Crawls are triggered on demand by seoClarity clients to analyse pages for technical and content issues, and clients can also schedule daily managed page crawls. The operator documents that it obeys robots.txt and Crawl-delay, and that its addresses are dynamic.
Cocolyzebot
Listed onlyCocolyze · SEO tools
Crawler for Cocolyze's SEO analysis platform, fetching pages of sites its users analyse. Cocolyze publishes no IP ranges, so requests cannot be verified beyond the user agent.
cognitiveSEO
Listed onlycognitiveSEO · SEO tools
James BOT is the web crawler operated by cognitiveSEO, an SEO toolset. It crawls the web and analyzes links to power the backlink and SEO analysis offered by the cognitiveSEO platform.
CriteoBot
VerifiableCriteo · SEO tools
Criteo's advertising crawler. It fetches merchant and publisher pages to extract product and content data used for Criteo's commerce and retargeting ads, and respects robots.txt and crawl-delay directives.
DataForSeoBot
VerifiableDataForSEO · SEO tools
DataForSEO's crawler. It fetches pages to build the backlink and SEO datasets that power DataForSEO's marketing-data APIs, and honours robots.txt and crawl-delay directives.
Dataproviderbot
VerifiableDataprovider.com · SEO tools
Dataprovider.com's in-house crawler. It indexes more than 400 million domains each month and structures what it finds into the company's web dataset (business information, technology detection, classifications and risk signals). The operator documents that it follows the robot exclusion protocol and that its crawlers can be identified by a reverse DNS lookup.
DomCopBot
Listed onlyDomCop · SEO tools
DomCop's availability crawler, used by the domain-research service of the same name. The operator documents that it accesses only the robots.txt file, once per domain, and uses that request to establish whether a website is live on the domain.
DotBot
Listed onlyMoz · SEO tools
Moz's general-purpose web crawler, distinct from rogerbot, that gathers link data powering the Moz Link Index and Link Explorer. Moz's own help pages document no fixed IP range for it.
Dragonbot
Listed onlyDragon Metrics · SEO tools
Dragon Metrics' SEO crawler, which collects data for the platform's Site Audit and Site Explorer features. Its operator documents that it respects robots.txt using Google's open-source parser, and that it crawls from dynamic IP addresses so it can only be identified by user agent.
EzoicBot
Listed onlyEzoic · SEO tools
Ezoic's crawler family, run by the digital-publisher technology platform of the same name. A desktop and a mobile variant crawl pages to study how sites, search engines and content interact, alongside named subtypes for Core Web Vitals measurement, uptime checks, ads.txt verification, integration checks and page-topic analysis. All variants share the single robots.txt token EzoicBot.
HubSpot Crawler
VerifiableHubSpot · SEO tools
HubSpot's crawler, which fetches customer and external pages to power the SEO recommendations and link analysis in HubSpot's marketing tools. HubSpot publishes its egress ranges tagged by service, including web crawling.
IAS Crawler
Listed onlyIntegral Ad Science · SEO tools
Content-rating and ad-verification crawler operated by Integral Ad Science. It visits web pages to assess content quality and brand safety and to support invalid-traffic detection for advertisers.
LinkCheckerBot
Fully verifiableLocal Profy LLC (LinkChecker.pro) · SEO tools
LinkChecker.pro's backlink-monitoring crawler. It revisits the backlinks that the service's customers track, checking whether each link is still present and how it is marked. The operator documents that the crawler strictly respects robots.txt and Crawl-delay, and publishes the addresses it crawls from.
Linkdexbot
Listed onlyAuthoritas (Analytics SEO Limited) · SEO tools
Linkdexbot is the web crawler for Linkdex, an SEO and search-marketing analytics platform now operated by Authoritas. It gathers link and page data used to power the platform's SEO reporting tools.
LinksIndexerBot
VerifiableLinks Indexer (Kalpraj Solutions (OPC) Private Limited) · SEO tools
The crawler behind the Links Indexer URL indexing service. It parses third-party sites to verify their URLs and status for sitemap campaigns, querying pages for metadata and favicons and taking a homepage screenshot. The operator states it never harvests e-mail addresses or content unrelated to sitemap campaigns, runs at most four parallel requests, and publishes the single address it crawls from.
McontextualBot
Listed onlyMContextual · SEO tools
MContextual's crawler for contextual advertising. The operator documents that it reads page content so advertisers can build cookieless contextual audiences from a site's subject matter, that it can be blocked with standard robots.txt directives, and that its requests originate from AWS and GCP addresses rather than a published range.
MegaIndex Crawler
Listed onlyMegaIndex · SEO tools
MegaIndex is an SEO and web-analytics platform whose crawler indexes links across the web to power backlink analysis, keyword tracking, and site audits for its subscribers.
Meta External Ads
VerifiableMeta · SEO tools
Meta's crawler that fetches pages for advertising and other business-related products and services, separate from the AI-training and link-preview crawlers. Verified by ASN lookup (AS32934); Meta publishes no IP feed.
MTRobot
Listed onlyMetrics Tools (Andreas Knatz) · SEO tools
Crawler for Metrics Tools, a German SEO analytics service, collecting page data for its visibility and ranking analyses. The operator publishes no IP ranges, so requests cannot be verified beyond the user agent.
MJ12bot
Listed onlyMajestic-12 · SEO tools
The crawler behind Majestic's backlink index. Majestic explicitly states it is a community-based distributed crawler with no fixed IP allocation, so requests cannot be verified by IP, ASN, or reverse DNS.
Monsidobot
VerifiableAcquia · SEO tools
Acquia Web Governance (formerly Monsido) crawler, which scans the public websites its customers have configured to run accessibility, quality assurance and policy checks. It also issues link-status checks against third-party sites that customers have linked to, preferring HEAD requests for those. Acquia documents a fallback user agent — a plain Chrome string with no identifying token — used only when the primary one fails, so a share of its traffic is identifiable by address rather than by user agent.
Nano Interactive Crawler
VerifiableNano Interactive · SEO tools
Nano Interactive's contextual-advertising crawler. Its published crawler policy lists four desktop and mobile user agents, all carrying the NanoInteractive/1.0 token, and names the two addresses the crawler requests come from so site owners can allow it explicitly.
OAI-AdsBot
Fully verifiableOpenAI · SEO tools
OpenAI's fetcher for advertising on ChatGPT. When an advertiser submits an ad, OAI-AdsBot visits the submitted landing page to check that it complies with OpenAI's ad policies and to judge when the ad is relevant to show. OpenAI states it only visits pages submitted as ads and that what it collects is not used to train generative AI foundation models. It has its own published IP range list, separate from GPTBot, OAI-SearchBot and ChatGPT-User.
OnCrawl
Listed onlyOnCrawl · SEO tools
OnCrawl's SEO crawler, used to analyze a customer's own site structure and content for technical SEO reporting. OnCrawl's help docs describe no fixed IP range; the bot's identity is user-configurable per crawl.
Outbrain crawler
VerifiableOutbrain · SEO tools
Outbrain's content-recommendation crawler. It fetches advertiser landing pages so Outbrain's system can pull the correct image and headline for a promoted-content unit, and rejects submitted URLs it cannot reach.
Panscient Crawler
Listed onlyPanscient Inc. · SEO tools
Panscient's large-scale crawler, which traverses public websites so that Panscient can build structured company and professional data feeds licensed to enterprise customers. The operator documents a full-corpus refresh each quarter, a rate limit of at most one request per second to any single domain, and compliance with the Robot Exclusion Standard. A separate "pantest" agent is used for testing. No IP ranges are published.
Proximic (Comscore Crawler)
Listed onlyComscore, Inc. · SEO tools
Proximic is Comscore's web crawler. It downloads the static textual content of pages to perform contextual analysis (content language, rating, and IAB categories) so advertising partners can match campaigns to page content. It identifies itself and honors robots.txt.
Rogerbot
Listed onlyMoz · SEO tools
Moz's site-audit crawler for Moz Pro Campaigns, distinct from DotBot. Moz's own FAQ states plainly that Rogerbot has no IP range: "we do not use a static IP address or range of IP addresses."
RyteBot
Listed onlySemrush · SEO tools
The crawler behind the Ryte.com tools, which analyse on-page SEO, technical and usability issues. Ryte was absorbed by Semrush, and RyteBot is now documented as a member of the Semrush bot family with its own robots.txt user agent. No IP ranges are published for it.
Scope3 Crawler
VerifiableScope3 · SEO tools
Scope3's crawler, which indexes publicly available web content (and paywalled media where the publisher has granted access) to produce content classification and brand-safety assessments for advertising. Scope3 documents adaptive rate limiting of five pages per minute per domain and publishes the single address it crawls from.
Search Atlas Bot
Listed onlySearch Atlas · SEO tools
The crawler behind Search Atlas's SEO platform, which fetches pages for its Site Auditor and monitoring features. The operator publishes the bot's user agent for allowlisting and states that the crawler does not use static IP addresses, so it can only be identified by its user agent.
SiteAuditBot
VerifiableSemrush · SEO tools
Semrush's site-auditing crawler, distinct from SemrushBot: it crawls a domain on demand when a Semrush customer runs the Site Audit tool, looking for SEO and technical issues. Unlike the backlink crawler, which Semrush says cannot be identified by IP, Site Audit is documented as running from a single dedicated subnet.
SemrushBot
Listed onlySemrush · SEO tools
Semrush's web crawler, feeding the backlink and site-audit data behind the Semrush SEO platform. Semrush's own bot page explicitly states it does not use consecutive IP blocks, so no CIDR list can be sourced.
SemrushBot-SI
VerifiableSemrush · SEO tools
The crawler behind Semrush's On Page SEO Checker and related on-page tools, run against a domain when a Semrush customer sets up a campaign for it. It is a separate robots.txt user agent from SemrushBot, and Semrush documents its own addresses to allowlist for it.
SeobilityBot
Fully verifiableSeobility GmbH · SEO tools
Crawler for Seobility's hosted SEO analysis and site-audit tooling, fetching pages of sites its customers analyse. Seobility publishes a machine-readable list of the addresses its bots crawl from.
SEOkicks
Listed onlyJobkicks SLU · SEO tools
SEOkicks operates a web crawler that builds a backlink database powering its SEO tools. The crawler visits sites to collect link data for analysis.
serpstatbot
Fully verifiableSerpstat · SEO tools
Serpstat's backlink crawler. It continuously crawls the web to add new links and track changes in Serpstat's link database, honouring robots.txt and Crawl-delay directives, and publishes the full list of addresses it crawls from.
SirdataBot
Listed onlySirdata · SEO tools
Crawler for Sirdata's contextual advertising API. It fetches and categorises page content when Sirdata's on-page script cannot read the document from the parent frame. Sirdata publishes the addresses it crawls from at proxies-list.sirdata.fr, but that pool rotates continuously — it recommends re-fetching every ten minutes — so no snapshot of it stays accurate.
SISTRIX Crawler
VerifiableSISTRIX · SEO tools
The crawler behind the SISTRIX Toolbox, a German SEO visibility platform. SISTRIX documents that every crawler IP resolves via reverse DNS to the "sistrix.net" domain rather than publishing a static CIDR list.
Siteimprove Crawler
VerifiableSiteimprove · SEO tools
Siteimprove's content-suite crawler, which fetches pages of sites its customers have configured in their account to run quality-assurance, accessibility, policy and SEO checks. Companion agents (LinkCheck, Image size, Probe) fetch links and resources for the same checks and crawl from the same published address list.
SmartologyBot
VerifiableSmartology · SEO tools
Smartology's SmartMatch contextual advertising crawler. It reads pages on domains where an advertiser is buying advertising space, so that ads can be matched to page context; the operator states the bot is only active on a domain while suitable ad slots and advertiser budget exist. Smartology publishes its current address list and user-agent match pattern as a JSON document.
StatusNest Backlink Spider
Listed onlyStatusNest · SEO tools
StatusNest's link-graph crawler. It visits public web pages to record their outbound and inbound links and maintain a graph of how sites connect to each other, for use by SEO and web analysts. The operator documents a maximum rate of one request per second per domain with automatic backoff, states that it collects only public link data and basic page metadata, and that it re-reads robots.txt before every crawl session. No addresses are published.
t3versionsBot
Listed onlyTorben Hansen (t3versions) · SEO tools
Private-project crawler that makes single GET requests to sites and looks for TYPO3 fingerprints, collecting statistics on the worldwide usage and development of the open-source TYPO3 CMS. No IP ranges are published, and the operator documents no robots.txt support (exclusion is by email request).
TTD-Content
Listed onlyThe Trade Desk · SEO tools
The Trade Desk's content scraper. When a page sends an ad request to The Trade Desk, this crawler scans the page to determine the context in which the ads were displayed, and caches that contextual data for ad serving. The operator publishes a plain-text list of the addresses it crawls from at ttd-content.adsrvr.org/ips; that list currently holds 2,640 individual addresses, which is beyond this directory's per-feed range cap, so it is not recorded as a machine-readable recipe here.
VelenPublicWebCrawler
Listed onlyHunter · SEO tools
Hunter's public web crawler, written in Go. It analyses millions of publicly accessible pages every month to build the business datasets and machine learning models behind Hunter's products, and never fetches anything behind a login. The operator documents a deliberate rate limit of one page at a time and one page every two seconds per site.
W3C Link Checker
VerifiableWorld Wide Web Consortium (W3C) · SEO tools
W3C's link checker, which follows the links on a user-submitted page to report broken and redirected references. W3C's validation services page marks it as the one validator that, being a crawling service, honors robots.txt directives, and gives the same published source addresses as the other W3C validators.
W3C CSS Validation Service
VerifiableWorld Wide Web Consortium (W3C) · SEO tools
W3C's CSS validator, which fetches a user-submitted page and its stylesheets to check them against the CSS specifications. It runs on W3C's Jigsaw server and identifies itself with a Jigsaw prefix followed by its own W3C_CSS_Validator_JFouffa token, from the same published validator addresses.
W3C Internationalization Checker
VerifiableWorld Wide Web Consortium (W3C) · SEO tools
The fetcher behind W3C's Internationalization Checker, which retrieves a user-submitted page and reports on its language, encoding and other i18n markup. W3C's validation services page lists its user agent alongside the single IPv4 address and IPv6 prefix all W3C validators fetch from.
W3C Markup Validator
VerifiableWorld Wide Web Consortium (W3C) · SEO tools
The W3C Markup Validation Service's fetcher, used when a user submits a URL to be checked for HTML conformance. W3C documents a fixed source address for its validation services alongside the user-agent string.
Validator.nu (W3C Nu HTML Checker)
VerifiableWorld Wide Web Consortium (W3C) · SEO tools
The fetcher behind the Nu HTML Checker, the W3C validation service that checks a user-submitted URL against current HTML conformance rules. W3C's validation services page lists its user agent and states that all W3C validator traffic comes from one published IPv4 address and one IPv6 prefix.
XoviBot
Listed onlyXovi GmbH · SEO tools
XoviBot is the web crawler for XOVI, an SEO and online-marketing analytics suite. It crawls sites to gather backlink and ranking data for the platform's SEO tools.
Yahoo Ad Monitoring
Listed onlyYahoo · SEO tools
Page-fetch client that retrieves the landing pages of URLs listed with Yahoo advertising services. Yahoo documents that it fetches each advertiser-supplied landing page to check policy compliance and to improve the accuracy of the ad listing, and that it uses separate desktop and mobile user agents.
YandexAccessibilityBot
VerifiableYandex · SEO tools
Yandex's accessibility checker, which downloads pages to assess how accessible they are to users. Yandex documents that it sends up to three requests per second, ignores the crawl-rate setting in Yandex Webmaster, and does not follow the general robots.txt rules written for arbitrary robots.
YandexPagechecker
VerifiableYandex · SEO tools
The fetcher behind Yandex's structured data validator, which retrieves a page so its markup can be validated. Yandex's robot table records it as taking the general robots.txt rules into account.
Zoominfobot
Listed onlyZoomInfo Technologies · SEO tools
ZoomInfo's indexing robot, which scans corporate websites, press releases, news services and SEC filings to build ZoomInfo's search index of businesses and business professionals. The operator documents that it obeys robots.txt, spaces out requests on larger sites and never opens more than one connection to a site at a time. No IP ranges are published.