Directory / Archivers
Archivers 8
AcademicBotRTU
Listed onlyRiga Technical University (Institute of Applied Computer Systems) · Archivers
Crawler run by Riga Technical University that indexes websites and documents to compare against student and researcher works for plagiarism detection. The operator publishes no IP ranges, so requests cannot be verified.
Arquivo.pt Web Crawler
Listed onlyArquivo.pt (FCCN/FCT) · Archivers
The crawler behind Arquivo.pt, Portugal's public web archive, which captures full page renders (HTML, CSS, JS, images) for long-term preservation. Built on Heritrix; the operator documents no fixed IP range.
BnF Web Archiving Robot
Listed onlyBibliothèque nationale de France · Archivers
The web crawler of the Bibliothèque nationale de France, which harvests French websites for the legal deposit of the web to preserve the national documentary heritage. It runs on Heritrix and applies request delays to avoid overloading servers.
Cloudflare Always Online
Listed onlyCloudflare · Archivers
Cloudflare's Always Online crawler, which fetches pages from sites that have the feature enabled so a cached copy can be served to visitors when the origin server is unreachable. Cloudflare's crawler reference documents the CloudFlare-AlwaysOnline user agent for this product.
archive.org_bot
Listed onlyInternet Archive · Archivers
The Internet Archive's Heritrix-based crawler used for its wide crawl of the web, feeding the Wayback Machine. The Archive says it crawls slowly to avoid disrupting sites and publishes no IP ranges. Its help pages note that robots exclusions may prevent archiving, but the operator does not document a commitment to obey robots.txt across its crawls.
ArchiveBot
Listed onlyArchive Team · Archivers
ArchiveBot is an IRC-controlled archiving bot run by Archive Team that crawls websites on request, writes WARC files, and uploads the captures to the Internet Archive.
PlagAwareBot
Listed onlyPlagAware · Archivers
PlagAware's fetcher, used by the German plagiarism-checking service of the same name. Text sections of a submitted document are passed to search engines, and this bot then reads only the individual pages those searches flagged as possible sources, caching them for 48 hours. The operator documents that it does not crawl whole sites and that it adheres to robots.txt directives.
TurnitinBot
Listed onlyTurnitin · Archivers
TurnitinBot is the web crawler operated by Turnitin. It collects publicly available web pages to build the content database used by Turnitin's academic-integrity and plagiarism-detection services.