Directory / Search engines

Search engines 73

AddSearchBot

Verifiable

AddSearch · Search engines

Crawler for AddSearch's hosted site-search service. It indexes the pages of customer sites so AddSearch can serve search results for them, obeys robots.txt rules written for the AddSearchBot token, and AddSearch publishes the fixed addresses it crawls from for allowlisting.

AdIdxBot

Verifiable

Microsoft · Search engines

Microsoft's crawler for Bing Ads / Microsoft Advertising. It crawls ads and follows through to the advertised landing pages for quality control, with both desktop and mobile variants.

Algolia Crawler

Verifiable

Algolia · Search engines

Algolia's Crawler, which visits a customer's pages, extracts search-relevant content, and pushes it to Algolia search indices. Algolia documents a fixed User-Agent token and a single static egress IP address for allowlisting.

Amazon AdBot

Verifiable

Amazon · Search engines

Amazon's advertising crawler. It scans web pages that request ads from Amazon's advertising systems, collecting page content for Amazon's classification systems to maintain brand safety and improve ad relevance.

AmazonProductDiscoverybot

Listed only

Amazon · Search engines

Amazon's product crawler, which collects publicly available product details from Amazon selling partner, brand and retailer websites to improve the accuracy and completeness of product information in Amazon's store. Amazon documents that it honours robots.txt user-agent and disallow directives, with changes taking up to 24 hours to apply, and that it does not support crawl-delay, nofollow or noindex.

Amzn-SearchBot

Verifiable

Amazon · Search engines

Amazon's crawler used to improve search experiences in Amazon products and services; Amazon states it does not crawl content for generative AI model training. Amazon publishes its source addresses on a dedicated page as a dated snapshot; they are recorded here verbatim.

Anomura

Verifiable

Direqt · Search engines

Direqt's search crawler, which discovers links and metadata to surface in Direqt's search features. The operator states it is not used to crawl content for model training, documents the robots.txt token Anomura, and publishes the two addresses the crawler requests come from.

Applebot

Fully verifiable

Apple · Search engines

Apple's web crawler, used to index content for Siri, Spotlight Suggestions, and Safari's search features.

atlassian-bot

Verifiable

Atlassian · Search engines

Crawler behind the Teamwork Graph "custom website" connector, which indexes a site an Atlassian customer has connected so its content is searchable in Rovo. Atlassian documents the robots.txt token and publishes the connector's egress ranges under the rovo-crawler product in its IP-ranges feed.

Baiduspider

Verifiable

Baidu · Search engines

Baidu's primary web crawler, fetching pages for Baidu Search indexing.

BingVideoPreview

Verifiable

Microsoft · Search engines

Microsoft crawler, listed among Bing's crawlers, that fetches pages to generate video previews shown in Bing. It runs desktop and mobile variants and is verifiable through the same reverse-DNS check as Bing's other crawlers.

Bingbot

Fully verifiable

Microsoft · Search engines

Microsoft's primary web crawler, fetching pages for Bing Search indexing.

BingPreview

Verifiable

Microsoft · Search engines

Microsoft fetcher used to generate page snapshots for previews in Bing search results. Runs both desktop and mobile variants.

Bublup Bot

Verifiable

Bublup · Search engines

Bublup's content-discovery bot. It fetches pages to build the database behind Bublup's suggestion engine, reading page title, description and related images. Bublup states its crawling IPs are not fixed and documents reverse-DNS verification against the bublup.com domain instead.

Channel3Bot

Verifiable

Channel3 · Search engines

Channel3's product crawler. It visits publicly accessible product detail pages and collects images, titles, descriptions, prices, availability and variants for Channel3's product catalogue, which AI apps and agents query to route shoppers back to the original site. The operator tells site owners not to hard-code its addresses and to verify it by forward-confirmed reverse DNS instead.

coccocbot

Verifiable

Cốc Cốc · Search engines

Cốc Cốc's web crawler, fetching pages for the Vietnamese Cốc Cốc search engine. Separate web and image sub-bots share the same verification domain.

deepnoc

Listed only

deepnoc GmbH · Search engines

Research crawler operated by deepnoc GmbH that parses and stores public web page content so that new search engines can query it without running their own crawling infrastructure. No IP ranges are published.

DuckDuckBot

Fully verifiable

DuckDuckGo · Search engines

DuckDuckGo's web crawler, used to improve DuckDuckGo's search results.

ExaSearchBot

Fully verifiable

Exa · Search engines

Exa's search crawler, which discovers and indexes pages on the public web so that people and applications can find, retrieve and cite them through Exa. Every request is signed with HTTP Message Signatures (RFC 9421) under the Web Bot Auth scheme, so a request can be verified cryptographically regardless of its source IP address.

FindFilesBot

Verifiable

FindFiles.net · Search engines

The crawler behind the FindFiles.net file search engine. It checks that publicly accessible files users search for are still available, using multiplexed HTTP/2 HEAD requests with a minimum ten-second interval per server, and may download images, videos and executables for classification and safety checks. Purpose-specific user agents are published for its link checking, favicon fetching, virus scanning and image classification agents, which all share one crawler host and address.

FleebsBot

Listed only

iontic GmbH (fleebs.com) · Search engines

Crawler for fleebs.com, a German real-time search engine. The operator's bot page documents two identities: one that searches for new information and one that analyses the pages it finds. It documents robots.txt blocking with the FleebsBot token and notes that pages indexed before a block may take time to drop out of the index.

FreespokeCrawler

Listed only

Freespoke · Search engines

Freespoke's search crawler, which fetches pages to build the index behind the Freespoke search engine. The operator documents the crawler's user agent, the robots.txt token FreespokeCrawler and support for the Crawl-delay directive.

Geedo Product Search

Fully verifiable

Geedo · Search engines

GeedoShopProductFinder is the crawler for Geedo, a product-search engine that indexes publicly accessible product pages from online stores and follows robots.txt rules.

Storebot-Google

Fully verifiable

Google · Search engines

Google's crawler for Google Shopping surfaces, including the Shopping tab in Google Search and shopping.google.com. It crawls product and store pages to populate shopping results.

Google-InspectionTool

Fully verifiable

Google · Search engines

Google crawler used by Search testing tools such as the URL Inspection tool in Search Console and the Rich Result Test. Fetches pages on demand to show how Google Search renders them; it has no effect on ranking.

Google special-case crawlers

Fully verifiable

Google · Search engines

A group of Google crawlers that operate under separate agreements between the crawled site and a specific Google product (AdsBot, AdSense, Google Safety scanning, and similar). Several members of this group ignore the wildcard robots.txt rule and must be targeted explicitly.

Googlebot

Fully verifiable

Google · Search engines

Google's primary web crawler, fetching pages for Google Search indexing.

Googlebot Image

Fully verifiable

Google · Search engines

Google's image crawler, which fetches images for Google Images and image features in Search. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.

Googlebot Video

Fully verifiable

Google · Search engines

Google's video crawler, which fetches video content for Google Search video features. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.

GoogleOther

Fully verifiable

Google · Search engines

Google's generic crawler family (GoogleOther, GoogleOther-Image, GoogleOther-Video), used by various Google product teams for fetching publicly accessible content. Robots.txt rules addressed to it do not affect Google Search.

360Spider

Listed only

360 Search (Qihoo 360) · Search engines

360Spider is the web-search crawler of 360 Search (so.com), the search engine operated by Chinese internet company Qihoo 360. The operator also runs 360Spider-Image and 360Spider-Video for image and video search.

IbouBot

Fully verifiable

Babbar (Ibou) · Search engines

The crawler behind Ibou, a French conversational search engine. It discovers publicly accessible pages and indexes them so Ibou can cite and link back to the sites it answers from. The operator documents a 5-second politeness delay per host, publishes its address ranges as JSON, and documents forward-confirmed reverse DNS as the reliable way to tell a genuine request from a spoofed user agent.

IONOS Crawler

Verifiable

IONOS · Search engines

IONOS Crawler (IonCrawl) continuously crawls publicly accessible domains to generate insights into how they are used, which IONOS uses to improve and expand its hosting products. IONOS documents reverse-DNS verification and deletes crawled data after 60 days.

Jooblebot

Listed only

Jooble · Search engines

The crawler for Jooble, a job-search aggregator that indexes job listings published across the web. Jooble's bot page states that it uses a web crawler which identifies itself as JoobleBot. No IP ranges are published.

Kagibot

Verifiable

Kagi · Search engines

Kagi's web crawler, fetching pages for the paid Kagi search engine. Kagi publishes a fixed, small set of source IPs rather than a CIDR feed, each paired with a kagibot.org hostname.

Linespider

Listed only

LINE · Search engines

LINE's crawler. It fetches web pages to support LINE's search features and content services, and adheres to the Robots Exclusion Protocol outlined in robots.txt.

LyonlBot

Listed only

Lyonl Search · Search engines

The crawler used by Lyonl Search to discover and refresh public web documents for its index, with separate identities for web, image and news crawling so site owners can control each class independently. The operator documents robots.txt and crawl-delay support, conditional fetches with ETag and Last-Modified, and that the crawler does not submit forms or bypass authentication. It states that crawler address ranges are not currently published, so no verification recipe is recorded.

MagnetmeBot

Fully verifiable

Magnet.me · Search engines

Magnet.me's web crawler. It discovers and refreshes career-related pages — vacancies, events and employer information — for the Magnet.me careers platform. The operator documents that it may render JavaScript, follows the canonical version of a page, does not submit forms and adjusts its crawl speed automatically, and it publishes the addresses it crawls from as both a JSON and a plain-text list.

Marginalia

Fully verifiable

Marginalia Search · Search engines

The crawler behind Marginalia Search, an independent search engine focused on older, text-heavy, and non-commercial web pages. Identifies itself by a bare URL rather than a conventional bot token.

MetaJobBot

Verifiable

METAJob · Search engines

MetaJobBot is the focused web crawler operated by the METAJob job meta-search engine. It searches websites for job listings to include in the METAJob index.

MojeekBot

Fully verifiable

Mojeek · Search engines

Mojeek's web crawler, fetching pages for the independent Mojeek search engine index.

MotoMinerBot

Verifiable

MotoMiner · Search engines

MotoMiner's vehicle-listing crawler. It primarily crawls automotive dealership websites and adds their vehicle detail pages to the MotoMiner index. The operator documents robots.txt and crawl-delay support, honours bot-specific noindex meta directives, throttles outbound requests against defined traffic thresholds, and publishes the address the crawler currently requests from.

Yeti

Verifiable

Naver · Search engines

Naver's web crawler, fetching pages for Naver Search indexing across South Korea's largest search engine.

NestDaddyBot

Listed only

NestDaddy Technologies · Search engines

Web crawler behind the NestDaddy search engine. The operator's webmaster page documents that it extracts page title, meta description, headings, language and link structure to build the index, filters adult and abusive content out of results, and fully respects robots.txt including the Crawl-delay directive. It states that the crawler runs from a distributed network whose current ranges are available only on request.

Openindex Spider

Listed only

Openindex · Search engines

Openindex operates a web-crawling cluster (Apache Nutch on an Apache Hadoop cluster) for research and development of universal and focused search engines. The crawler respects robots.txt and the Crawl-delay directive and identifies itself in the User-Agent.

Paqlebot

Verifiable

Paqle · Search engines

Paqle's news crawler. It indexes news articles to discover who is mentioned in the media, crawling at most one page per minute per site and tuning its rate to how often a site publishes. Paqle documents that the crawler falls back to a site's Googlebot robots.txt rules when no Paqlebot rules are present, and that its addresses reverse-resolve under paqle.net.

PetalBot

Verifiable

Huawei · Search engines

Huawei's web crawler, fetching pages for the Petal Search engine used in Huawei's mobile services and for Huawei Assistant content recommendations. Huawei documents reverse-DNS verification against aspiegel.com and petalsearch.com hostnames.

PiplBot

Listed only

Pipl · Search engines

PiplBot is the web-indexing crawler operated by Pipl, an identity and people-search service. It retrieves and indexes publicly available information, including content behind searchable databases, to build Pipl's people-search index.

Quantcastbot

Verifiable

Quantcast · Search engines

Quantcast's advertising crawler. It crawls websites to extract page content for interest-based audience categorization and to run quality-assurance checks on advertisement landing pages.

Qwantbot

Fully verifiable

Qwant · Search engines

Qwant's web crawler, fetching pages for the privacy-focused Qwant search engine. Older crawls identify as Qwantify; current documentation uses the Qwantbot name.

SeekportBot

Fully verifiable

SISTRIX · Search engines

The crawler for the Seekport search engine, operated by SISTRIX (Bonn, Germany). It fetches pages to build Seekport's search index and follows the Disallow directives in robots.txt.

SemanticScholarBot

Listed only

Allen Institute for AI (Ai2) · Search engines

SemanticScholarBot is the crawler for Semantic Scholar, an academic search engine built by the Allen Institute for AI. It crawls the web to find scholarly PDFs and metadata for indexing.

SeznamBot

Verifiable

Seznam.cz · Search engines

Seznam's web crawler, fetching pages for the Seznam.cz search engine used primarily in the Czech Republic.

Sogou Spider

Listed only

Sogou · Search engines

Sogou Spider is the web crawler for Sogou, a major Chinese search engine (owned by Tencent). It crawls and indexes web pages, news, and images to build Sogou's search index.

stepstoneCrawlBot

Listed only

The Stepstone Group · Search engines

The Stepstone Group's job-listing crawler. Its crawler page states that it processes only publicly available information, observes robots.txt directives and keeps intervals between requests to avoid loading servers, and publishes the user agent it sends for transparency.

StractBot

Listed only

Stract · Search engines

The crawler for Stract, an open-source search engine. Stract documents the crawler's user agent and robots.txt behavior but publishes no IP list or reverse-DNS verification method.

Swiftbot

Listed only

Elasticsearch B.V. (Swiftype) · Search engines

The web crawler behind Swiftype's site-search product, now operated by Elasticsearch B.V. Unlike a general search-engine crawler it only visits sites that Swiftype customers have asked it to crawl, in order to build a search index for those sites. The operator documents that Swiftbot obeys every restriction it finds in robots.txt, including a full-site disallow, and that it looks for the token Swiftbot.

TinEye

Listed only

TinEye · Search engines

TinEye-bot is the web crawler operated by TinEye, the reverse image search engine built by Idee Inc. It crawls the web to discover and index images for TinEye's image-matching search service.

TrovitBot

Listed only

Trovit · Search engines

Trovit's web crawler, which discovers new and updated pages to add to the Trovit classifieds search index. Trovit documents that it does not fetch most sites more than once per second and that it can be blocked with the trovitBot robots.txt token.

urlsuma

Listed only

urlsuma.de (Gerhard Stöbe) · Search engines

Crawler for urlsuma.de, a German general-purpose web search engine under construction. The operator documents that it re-reads robots.txt without caching before every single fetch, makes about two requests per visit (robots.txt plus one page), executes no JavaScript, and treats a 4xx response as a permanent block. Current crawler addresses are shared on request only.

Webzio

Listed only

Webz.io Ltd. · Search engines

Crawler for Webz.io (formerly Webhose.io), a web-data provider that collects publicly available content to build structured data feeds. Webz.io publishes no IP ranges, so requests cannot be verified.

WepchSearchEngine

Listed only

Wepch · Search engines

The crawler for Wepch, an independent privacy-focused search engine still in development. Its operator page states the project's purpose, gives a contact address for crawl complaints, and documents the robots.txt token site owners can use to crawl-delay or disallow it.

Yahoo! Slurp

Listed only

Yahoo · Search engines

Slurp is Yahoo Search's web crawler. It crawls and indexes pages for Yahoo Search results and also gathers content for Yahoo News, Finance, and Sports.

Yahoo! JAPAN Crawler

Listed only

LY Corporation (Yahoo! JAPAN) · Search engines

Web crawler operated by Yahoo! JAPAN (LY Corporation) to collect pages for its Japanese search index and related services. It identifies itself with Y!J-prefixed user-agent tokens and follows the Robots Exclusion Protocol.

YandexBlogs

Verifiable

Yandex · Search engines

Yandex's blog-search robot. Yandex's robot table documents it as the blog search robot that indexes post comments, and lists it as following robots.txt directives.

YandexComBot

Verifiable

Yandex · Search engines

Yandex's indexing robot for content in languages other than Russian, feeding the international yandex.com search index. Yandex's robot table documents that it can index content when there is no explicit robot-specific restriction, and lists it among the robots that do not follow the general robots.txt rules written for arbitrary robots.

YandexRenderResourcesBot

Verifiable

Yandex · Search engines

Yandex robot that loads resources such as JavaScript and CSS needed to render pages during crawling for Yandex Search.

YandexFavicons

Verifiable

Yandex · Search engines

Yandex's favicon fetcher. Yandex's robot table documents it as downloading a site's favicon file for display in search results, and lists it as one of the Yandex robots that do not follow robots.txt directives.

YandexImages

Verifiable

Yandex · Search engines

Yandex's image-indexing robot, which fetches images so they can be displayed in Yandex Images. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.

YandexMedia

Verifiable

Yandex · Search engines

Yandex's multimedia-indexing robot. Yandex's robot table documents it as a separate robot from YandexBot, describes it as indexing multimedia data, and lists it as following robots.txt directives.

YandexMobileBot

Verifiable

Yandex · Search engines

Yandex's mobile-layout robot, which fetches pages to determine whether their layout is suitable for mobile devices. It advertises an iPhone Safari prefix followed by its own YandexMobileBot token, and Yandex lists it among the robots that do not follow the general robots.txt rules written for arbitrary robots.

YandexVideo

Verifiable

Yandex · Search engines

Yandex's video-indexing robot, which fetches pages and video content for display in Yandex video search. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.

YandexBot

Verifiable

Yandex · Search engines

Yandex's primary web crawler, fetching pages for Yandex Search indexing.