Directory / Search engines
Search engines 73
AddSearchBot
VerifiableAddSearch · Search engines
Crawler for AddSearch's hosted site-search service. It indexes the pages of customer sites so AddSearch can serve search results for them, obeys robots.txt rules written for the AddSearchBot token, and AddSearch publishes the fixed addresses it crawls from for allowlisting.
AdIdxBot
VerifiableMicrosoft · Search engines
Microsoft's crawler for Bing Ads / Microsoft Advertising. It crawls ads and follows through to the advertised landing pages for quality control, with both desktop and mobile variants.
Algolia Crawler
VerifiableAlgolia · Search engines
Algolia's Crawler, which visits a customer's pages, extracts search-relevant content, and pushes it to Algolia search indices. Algolia documents a fixed User-Agent token and a single static egress IP address for allowlisting.
Amazon AdBot
VerifiableAmazon · Search engines
Amazon's advertising crawler. It scans web pages that request ads from Amazon's advertising systems, collecting page content for Amazon's classification systems to maintain brand safety and improve ad relevance.
AmazonProductDiscoverybot
Listed onlyAmazon · Search engines
Amazon's product crawler, which collects publicly available product details from Amazon selling partner, brand and retailer websites to improve the accuracy and completeness of product information in Amazon's store. Amazon documents that it honours robots.txt user-agent and disallow directives, with changes taking up to 24 hours to apply, and that it does not support crawl-delay, nofollow or noindex.
Amzn-SearchBot
VerifiableAmazon · Search engines
Amazon's crawler used to improve search experiences in Amazon products and services; Amazon states it does not crawl content for generative AI model training. Amazon publishes its source addresses on a dedicated page as a dated snapshot; they are recorded here verbatim.
Anomura
VerifiableDireqt · Search engines
Direqt's search crawler, which discovers links and metadata to surface in Direqt's search features. The operator states it is not used to crawl content for model training, documents the robots.txt token Anomura, and publishes the two addresses the crawler requests come from.
Applebot
Fully verifiableApple · Search engines
Apple's web crawler, used to index content for Siri, Spotlight Suggestions, and Safari's search features.
atlassian-bot
VerifiableAtlassian · Search engines
Crawler behind the Teamwork Graph "custom website" connector, which indexes a site an Atlassian customer has connected so its content is searchable in Rovo. Atlassian documents the robots.txt token and publishes the connector's egress ranges under the rovo-crawler product in its IP-ranges feed.
Baiduspider
VerifiableBaidu · Search engines
Baidu's primary web crawler, fetching pages for Baidu Search indexing.
BingVideoPreview
VerifiableMicrosoft · Search engines
Microsoft crawler, listed among Bing's crawlers, that fetches pages to generate video previews shown in Bing. It runs desktop and mobile variants and is verifiable through the same reverse-DNS check as Bing's other crawlers.
Bingbot
Fully verifiableMicrosoft · Search engines
Microsoft's primary web crawler, fetching pages for Bing Search indexing.
BingPreview
VerifiableMicrosoft · Search engines
Microsoft fetcher used to generate page snapshots for previews in Bing search results. Runs both desktop and mobile variants.
Bublup Bot
VerifiableBublup · Search engines
Bublup's content-discovery bot. It fetches pages to build the database behind Bublup's suggestion engine, reading page title, description and related images. Bublup states its crawling IPs are not fixed and documents reverse-DNS verification against the bublup.com domain instead.
Channel3Bot
VerifiableChannel3 · Search engines
Channel3's product crawler. It visits publicly accessible product detail pages and collects images, titles, descriptions, prices, availability and variants for Channel3's product catalogue, which AI apps and agents query to route shoppers back to the original site. The operator tells site owners not to hard-code its addresses and to verify it by forward-confirmed reverse DNS instead.
coccocbot
VerifiableCốc Cốc · Search engines
Cốc Cốc's web crawler, fetching pages for the Vietnamese Cốc Cốc search engine. Separate web and image sub-bots share the same verification domain.
deepnoc
Listed onlydeepnoc GmbH · Search engines
Research crawler operated by deepnoc GmbH that parses and stores public web page content so that new search engines can query it without running their own crawling infrastructure. No IP ranges are published.
DuckDuckBot
Fully verifiableDuckDuckGo · Search engines
DuckDuckGo's web crawler, used to improve DuckDuckGo's search results.
ExaSearchBot
Fully verifiableExa · Search engines
Exa's search crawler, which discovers and indexes pages on the public web so that people and applications can find, retrieve and cite them through Exa. Every request is signed with HTTP Message Signatures (RFC 9421) under the Web Bot Auth scheme, so a request can be verified cryptographically regardless of its source IP address.
FindFilesBot
VerifiableFindFiles.net · Search engines
The crawler behind the FindFiles.net file search engine. It checks that publicly accessible files users search for are still available, using multiplexed HTTP/2 HEAD requests with a minimum ten-second interval per server, and may download images, videos and executables for classification and safety checks. Purpose-specific user agents are published for its link checking, favicon fetching, virus scanning and image classification agents, which all share one crawler host and address.
FleebsBot
Listed onlyiontic GmbH (fleebs.com) · Search engines
Crawler for fleebs.com, a German real-time search engine. The operator's bot page documents two identities: one that searches for new information and one that analyses the pages it finds. It documents robots.txt blocking with the FleebsBot token and notes that pages indexed before a block may take time to drop out of the index.
FreespokeCrawler
Listed onlyFreespoke · Search engines
Freespoke's search crawler, which fetches pages to build the index behind the Freespoke search engine. The operator documents the crawler's user agent, the robots.txt token FreespokeCrawler and support for the Crawl-delay directive.
Geedo Product Search
Fully verifiableGeedo · Search engines
GeedoShopProductFinder is the crawler for Geedo, a product-search engine that indexes publicly accessible product pages from online stores and follows robots.txt rules.
Storebot-Google
Fully verifiableGoogle · Search engines
Google's crawler for Google Shopping surfaces, including the Shopping tab in Google Search and shopping.google.com. It crawls product and store pages to populate shopping results.
Google-InspectionTool
Fully verifiableGoogle · Search engines
Google crawler used by Search testing tools such as the URL Inspection tool in Search Console and the Rich Result Test. Fetches pages on demand to show how Google Search renders them; it has no effect on ranking.
Google special-case crawlers
Fully verifiableGoogle · Search engines
A group of Google crawlers that operate under separate agreements between the crawled site and a specific Google product (AdsBot, AdSense, Google Safety scanning, and similar). Several members of this group ignore the wildcard robots.txt rule and must be targeted explicitly.
Googlebot
Fully verifiableGoogle · Search engines
Google's primary web crawler, fetching pages for Google Search indexing.
Googlebot Image
Fully verifiableGoogle · Search engines
Google's image crawler, which fetches images for Google Images and image features in Search. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.
Googlebot Video
Fully verifiableGoogle · Search engines
Google's video crawler, which fetches video content for Google Search video features. It is a distinct crawler from Googlebot with its own robots.txt token, though rules for Googlebot also apply to it.
GoogleOther
Fully verifiableGoogle · Search engines
Google's generic crawler family (GoogleOther, GoogleOther-Image, GoogleOther-Video), used by various Google product teams for fetching publicly accessible content. Robots.txt rules addressed to it do not affect Google Search.
360Spider
Listed only360 Search (Qihoo 360) · Search engines
360Spider is the web-search crawler of 360 Search (so.com), the search engine operated by Chinese internet company Qihoo 360. The operator also runs 360Spider-Image and 360Spider-Video for image and video search.
IbouBot
Fully verifiableBabbar (Ibou) · Search engines
The crawler behind Ibou, a French conversational search engine. It discovers publicly accessible pages and indexes them so Ibou can cite and link back to the sites it answers from. The operator documents a 5-second politeness delay per host, publishes its address ranges as JSON, and documents forward-confirmed reverse DNS as the reliable way to tell a genuine request from a spoofed user agent.
IONOS Crawler
VerifiableIONOS · Search engines
IONOS Crawler (IonCrawl) continuously crawls publicly accessible domains to generate insights into how they are used, which IONOS uses to improve and expand its hosting products. IONOS documents reverse-DNS verification and deletes crawled data after 60 days.
Jooblebot
Listed onlyJooble · Search engines
The crawler for Jooble, a job-search aggregator that indexes job listings published across the web. Jooble's bot page states that it uses a web crawler which identifies itself as JoobleBot. No IP ranges are published.
Kagibot
VerifiableKagi · Search engines
Kagi's web crawler, fetching pages for the paid Kagi search engine. Kagi publishes a fixed, small set of source IPs rather than a CIDR feed, each paired with a kagibot.org hostname.
Linespider
Listed onlyLINE · Search engines
LINE's crawler. It fetches web pages to support LINE's search features and content services, and adheres to the Robots Exclusion Protocol outlined in robots.txt.
LyonlBot
Listed onlyLyonl Search · Search engines
The crawler used by Lyonl Search to discover and refresh public web documents for its index, with separate identities for web, image and news crawling so site owners can control each class independently. The operator documents robots.txt and crawl-delay support, conditional fetches with ETag and Last-Modified, and that the crawler does not submit forms or bypass authentication. It states that crawler address ranges are not currently published, so no verification recipe is recorded.
MagnetmeBot
Fully verifiableMagnet.me · Search engines
Magnet.me's web crawler. It discovers and refreshes career-related pages — vacancies, events and employer information — for the Magnet.me careers platform. The operator documents that it may render JavaScript, follows the canonical version of a page, does not submit forms and adjusts its crawl speed automatically, and it publishes the addresses it crawls from as both a JSON and a plain-text list.
Marginalia
Fully verifiableMarginalia Search · Search engines
The crawler behind Marginalia Search, an independent search engine focused on older, text-heavy, and non-commercial web pages. Identifies itself by a bare URL rather than a conventional bot token.
MetaJobBot
VerifiableMETAJob · Search engines
MetaJobBot is the focused web crawler operated by the METAJob job meta-search engine. It searches websites for job listings to include in the METAJob index.
MojeekBot
Fully verifiableMojeek · Search engines
Mojeek's web crawler, fetching pages for the independent Mojeek search engine index.
MotoMinerBot
VerifiableMotoMiner · Search engines
MotoMiner's vehicle-listing crawler. It primarily crawls automotive dealership websites and adds their vehicle detail pages to the MotoMiner index. The operator documents robots.txt and crawl-delay support, honours bot-specific noindex meta directives, throttles outbound requests against defined traffic thresholds, and publishes the address the crawler currently requests from.
Yeti
VerifiableNaver · Search engines
Naver's web crawler, fetching pages for Naver Search indexing across South Korea's largest search engine.
NestDaddyBot
Listed onlyNestDaddy Technologies · Search engines
Web crawler behind the NestDaddy search engine. The operator's webmaster page documents that it extracts page title, meta description, headings, language and link structure to build the index, filters adult and abusive content out of results, and fully respects robots.txt including the Crawl-delay directive. It states that the crawler runs from a distributed network whose current ranges are available only on request.
Openindex Spider
Listed onlyOpenindex · Search engines
Openindex operates a web-crawling cluster (Apache Nutch on an Apache Hadoop cluster) for research and development of universal and focused search engines. The crawler respects robots.txt and the Crawl-delay directive and identifies itself in the User-Agent.
Paqlebot
VerifiablePaqle · Search engines
Paqle's news crawler. It indexes news articles to discover who is mentioned in the media, crawling at most one page per minute per site and tuning its rate to how often a site publishes. Paqle documents that the crawler falls back to a site's Googlebot robots.txt rules when no Paqlebot rules are present, and that its addresses reverse-resolve under paqle.net.
PetalBot
VerifiableHuawei · Search engines
Huawei's web crawler, fetching pages for the Petal Search engine used in Huawei's mobile services and for Huawei Assistant content recommendations. Huawei documents reverse-DNS verification against aspiegel.com and petalsearch.com hostnames.
PiplBot
Listed onlyPipl · Search engines
PiplBot is the web-indexing crawler operated by Pipl, an identity and people-search service. It retrieves and indexes publicly available information, including content behind searchable databases, to build Pipl's people-search index.
Quantcastbot
VerifiableQuantcast · Search engines
Quantcast's advertising crawler. It crawls websites to extract page content for interest-based audience categorization and to run quality-assurance checks on advertisement landing pages.
Qwantbot
Fully verifiableQwant · Search engines
Qwant's web crawler, fetching pages for the privacy-focused Qwant search engine. Older crawls identify as Qwantify; current documentation uses the Qwantbot name.
SeekportBot
Fully verifiableSISTRIX · Search engines
The crawler for the Seekport search engine, operated by SISTRIX (Bonn, Germany). It fetches pages to build Seekport's search index and follows the Disallow directives in robots.txt.
SemanticScholarBot
Listed onlyAllen Institute for AI (Ai2) · Search engines
SemanticScholarBot is the crawler for Semantic Scholar, an academic search engine built by the Allen Institute for AI. It crawls the web to find scholarly PDFs and metadata for indexing.
SeznamBot
VerifiableSeznam.cz · Search engines
Seznam's web crawler, fetching pages for the Seznam.cz search engine used primarily in the Czech Republic.
Sogou Spider
Listed onlySogou · Search engines
Sogou Spider is the web crawler for Sogou, a major Chinese search engine (owned by Tencent). It crawls and indexes web pages, news, and images to build Sogou's search index.
stepstoneCrawlBot
Listed onlyThe Stepstone Group · Search engines
The Stepstone Group's job-listing crawler. Its crawler page states that it processes only publicly available information, observes robots.txt directives and keeps intervals between requests to avoid loading servers, and publishes the user agent it sends for transparency.
StractBot
Listed onlyStract · Search engines
The crawler for Stract, an open-source search engine. Stract documents the crawler's user agent and robots.txt behavior but publishes no IP list or reverse-DNS verification method.
Swiftbot
Listed onlyElasticsearch B.V. (Swiftype) · Search engines
The web crawler behind Swiftype's site-search product, now operated by Elasticsearch B.V. Unlike a general search-engine crawler it only visits sites that Swiftype customers have asked it to crawl, in order to build a search index for those sites. The operator documents that Swiftbot obeys every restriction it finds in robots.txt, including a full-site disallow, and that it looks for the token Swiftbot.
TinEye
Listed onlyTinEye · Search engines
TinEye-bot is the web crawler operated by TinEye, the reverse image search engine built by Idee Inc. It crawls the web to discover and index images for TinEye's image-matching search service.
TrovitBot
Listed onlyTrovit · Search engines
Trovit's web crawler, which discovers new and updated pages to add to the Trovit classifieds search index. Trovit documents that it does not fetch most sites more than once per second and that it can be blocked with the trovitBot robots.txt token.
urlsuma
Listed onlyurlsuma.de (Gerhard Stöbe) · Search engines
Crawler for urlsuma.de, a German general-purpose web search engine under construction. The operator documents that it re-reads robots.txt without caching before every single fetch, makes about two requests per visit (robots.txt plus one page), executes no JavaScript, and treats a 4xx response as a permanent block. Current crawler addresses are shared on request only.
Webzio
Listed onlyWebz.io Ltd. · Search engines
Crawler for Webz.io (formerly Webhose.io), a web-data provider that collects publicly available content to build structured data feeds. Webz.io publishes no IP ranges, so requests cannot be verified.
WepchSearchEngine
Listed onlyWepch · Search engines
The crawler for Wepch, an independent privacy-focused search engine still in development. Its operator page states the project's purpose, gives a contact address for crawl complaints, and documents the robots.txt token site owners can use to crawl-delay or disallow it.
Yahoo! Slurp
Listed onlyYahoo · Search engines
Slurp is Yahoo Search's web crawler. It crawls and indexes pages for Yahoo Search results and also gathers content for Yahoo News, Finance, and Sports.
Yahoo! JAPAN Crawler
Listed onlyLY Corporation (Yahoo! JAPAN) · Search engines
Web crawler operated by Yahoo! JAPAN (LY Corporation) to collect pages for its Japanese search index and related services. It identifies itself with Y!J-prefixed user-agent tokens and follows the Robots Exclusion Protocol.
YandexBlogs
VerifiableYandex · Search engines
Yandex's blog-search robot. Yandex's robot table documents it as the blog search robot that indexes post comments, and lists it as following robots.txt directives.
YandexComBot
VerifiableYandex · Search engines
Yandex's indexing robot for content in languages other than Russian, feeding the international yandex.com search index. Yandex's robot table documents that it can index content when there is no explicit robot-specific restriction, and lists it among the robots that do not follow the general robots.txt rules written for arbitrary robots.
YandexRenderResourcesBot
VerifiableYandex · Search engines
Yandex robot that loads resources such as JavaScript and CSS needed to render pages during crawling for Yandex Search.
YandexFavicons
VerifiableYandex · Search engines
Yandex's favicon fetcher. Yandex's robot table documents it as downloading a site's favicon file for display in search results, and lists it as one of the Yandex robots that do not follow robots.txt directives.
YandexImages
VerifiableYandex · Search engines
Yandex's image-indexing robot, which fetches images so they can be displayed in Yandex Images. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.
YandexMedia
VerifiableYandex · Search engines
Yandex's multimedia-indexing robot. Yandex's robot table documents it as a separate robot from YandexBot, describes it as indexing multimedia data, and lists it as following robots.txt directives.
YandexMobileBot
VerifiableYandex · Search engines
Yandex's mobile-layout robot, which fetches pages to determine whether their layout is suitable for mobile devices. It advertises an iPhone Safari prefix followed by its own YandexMobileBot token, and Yandex lists it among the robots that do not follow the general robots.txt rules written for arbitrary robots.
YandexVideo
VerifiableYandex · Search engines
Yandex's video-indexing robot, which fetches pages and video content for display in Yandex video search. Yandex documents it as a separate robot from YandexBot and lists it as following robots.txt directives.
YandexBot
VerifiableYandex · Search engines
Yandex's primary web crawler, fetching pages for Yandex Search indexing.