Directory / Archivers

Arquivo.pt Web Crawler

Listed only

The crawler behind Arquivo.pt, Portugal's public web archive, which captures full page renders (HTML, CSS, JS, images) for long-term preservation. Built on Heritrix; the operator documents no fixed IP range.

Operated by Arquivo.pt (FCCN/FCT) · Official documentation

User agents

Patterns this directory matches, with real observed strings.

  • regexArquivo-web-crawler
  • observedArquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling)

How to verify

No verification recipe published by the operator.

Good Bot Practices scorecard

  • Identifies honestly

    Stable UA token documented (1 pattern)

  • Verifiable

    Operator publishes no verification path

  • Respects robots.txt

    Honors robots.txt (token: Arquivo-web-crawler)

  • Behaves

    Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset

  • Reachable operator

    Operator and documentation published

Measured against the Good Bot Practices.