archive.org_bot
Listed onlyThe Internet Archive's Heritrix-based crawler used for its wide crawl of the web, feeding the Wayback Machine. The Archive says it crawls slowly to avoid disrupting sites and publishes no IP ranges. Its help pages note that robots exclusions may prevent archiving, but the operator does not document a commitment to obey robots.txt across its crawls.
Operated by Internet Archive · Official documentation
User agents
Patterns this directory matches, with real observed strings.
- regexarchive\.org_bot
- regexia_archiver
- observedMozilla/5.0 (compatible; archive.org_bot +http://archive.org/details/archive.org_bot)
- observedMozilla/5.0 (compatible; archive.org_bot +http://www.archive.org/details/archive.org_bot)
- observedMozilla/5.0 (compatible; special_archiver/3.1.1 +http://www.archive.org/details/archive.org_bot)
- observedia_archiver-web.archive.org
How to verify
No verification recipe published by the operator.
Good Bot Practices scorecard
- Identifies honestly
Stable UA token documented (2 patterns)
- Verifiable
Operator publishes no verification path
- Respects robots.txt
Does not honor robots.txt
- Behaves
Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset
- Reachable operator
Operator and documentation published
Measured against the Good Bot Practices.