Directory / Archivers

Arquivo.pt Web Crawler

Fully verifiable

The crawler behind Arquivo.pt, Portugal's public web archive, which captures full page renders (HTML, CSS, JS, images) for long-term preservation. Built on Heritrix. Arquivo.pt publishes the ranges its crawlers operate from as a machine-readable JSON file in the JAFAR format, and recommends combining that IP check with the User-Agent check.

Operated by Arquivo.pt (FCCN/FCT) · Official documentation

User agents

Patterns this directory matches, with real observed strings.

  • regexArquivo-web-crawler
  • observedArquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling)

How to verify

Fetch the official IP range feed

curl -s https://arquivo.pt/crawlerips.json

IP ranges (2)

Refreshed daily from the operator's feed. Also available as /data/ips/arquivo-pt.ips and /data/bots/arquivo-pt.json.

194.210.235.0/26
2001:690:a00:1039::/64

Good Bot Practices scorecard

  • Identifies honestly

    Stable UA token documented (1 pattern)

  • Verifiable

    Fully verifiable

  • Respects robots.txt

    Honors robots.txt (token: Arquivo-web-crawler)

  • Behaves

    Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset

  • Reachable operator

    Operator and documentation published

Measured against the Good Bot Practices.