Directory / Archivers
Archivers 11
AcademicBotRTU
Listed onlyRiga Technical University (Institute of Applied Computer Systems) · Archivers
Crawler run by Riga Technical University that indexes websites and documents to compare against student and researcher works for plagiarism detection. The operator publishes no IP ranges, so requests cannot be verified.
Arquivo.pt Web Crawler
Listed onlyArquivo.pt (FCCN/FCT) · Archivers
The crawler behind Arquivo.pt, Portugal's public web archive, which captures full page renders (HTML, CSS, JS, images) for long-term preservation. Built on Heritrix; the operator documents no fixed IP range.
BnF Web Archiving Robot
Listed onlyBibliothèque nationale de France · Archivers
The web crawler of the Bibliothèque nationale de France, which harvests French websites for the legal deposit of the web to preserve the national documentary heritage. It runs on Heritrix and applies request delays to avoid overloading servers.
CLASSLA-web
Listed onlyCLARIN.SI (CLASSLA Knowledge Centre) · Archivers
The CLASSLA-web crawler collects web corpora for South Slavic and other languages under the CLARIN Knowledge Centre for South Slavic languages, continuing the CEF-funded MaCoCu project's crawling infrastructure. It runs the SpiderLing crawler developed at Masaryk University. The operator documents that it adheres to the robots exclusion standard and reads robots.txt on the first access of each crawl run.
Cloudflare Always Online
Listed onlyCloudflare · Archivers
Cloudflare's Always Online crawler, which fetches pages from sites that have the feature enabled so a cached copy can be served to visitors when the origin server is unreachable. Cloudflare's crawler reference documents the CloudFlare-AlwaysOnline user agent for this product.
Cotoyogi
VerifiableROIS-DS Data Lake Research and Development Center · Archivers
A research web crawler run by the Data Lake Research and Development Center of the Joint Support-Center for Data Science Research (ROIS-DS), Research Organization of Information and Systems, Japan. It collects Japanese-language web data for research use. The operator page documents RFC 9309 robots.txt handling with Crawl-delay support, robots meta tag support, and the address range the crawler requests from.
archive.org_bot
Listed onlyInternet Archive · Archivers
The Internet Archive's Heritrix-based crawler used for its wide crawl of the web, feeding the Wayback Machine. The Archive says it crawls slowly to avoid disrupting sites and publishes no IP ranges. Its help pages note that robots exclusions may prevent archiving, but the operator does not document a commitment to obey robots.txt across its crawls.
ArchiveBot
Listed onlyArchive Team · Archivers
ArchiveBot is an IRC-controlled archiving bot run by Archive Team that crawls websites on request, writes WARC files, and uploads the captures to the Internet Archive.
PlagAwareBot
Listed onlyPlagAware · Archivers
PlagAware's fetcher, used by the German plagiarism-checking service of the same name. Text sections of a submitted document are passed to search engines, and this bot then reads only the individual pages those searches flagged as possible sources, caching them for 48 hours. The operator documents that it does not crawl whole sites and that it adheres to robots.txt directives.
TurnitinBot
Listed onlyTurnitin · Archivers
TurnitinBot is the web crawler operated by Turnitin. It collects publicly available web pages to build the content database used by Turnitin's academic-integrity and plagiarism-detection services.
XY-Archive-Compliance
VerifiableXY Planning Network · Archivers
Website archiver for XY Archive, XYPN's compliance-archiving product for financial advisory firms. It crawls a subscribing firm's site to decide which pages to capture and then screenshots each page for the archive; XYPN documents the single address it archives from.