Directory / Archivers

archive.org_bot

Listed only

The Internet Archive's Heritrix-based crawler used for its wide crawl of the web, feeding the Wayback Machine. The Archive says it crawls slowly to avoid disrupting sites and publishes no IP ranges. Its help pages note that robots exclusions may prevent archiving, but the operator does not document a commitment to obey robots.txt across its crawls.

Operated by Internet Archive · Official documentation

User agents

Patterns this directory matches, with real observed strings.

  • regexarchive\.org_bot
  • regexia_archiver
  • observedMozilla/5.0 (compatible; archive.org_bot +http://archive.org/details/archive.org_bot)
  • observedMozilla/5.0 (compatible; archive.org_bot +http://www.archive.org/details/archive.org_bot)
  • observedMozilla/5.0 (compatible; special_archiver/3.1.1 +http://www.archive.org/details/archive.org_bot)
  • observedia_archiver-web.archive.org

How to verify

No verification recipe published by the operator.

Good Bot Practices scorecard

  • Identifies honestly

    Stable UA token documented (2 patterns)

  • Verifiable

    Operator publishes no verification path

  • Respects robots.txt

    Does not honor robots.txt

  • Behaves

    Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset

  • Reachable operator

    Operator and documentation published

Measured against the Good Bot Practices.