Arquivo.pt Web Crawler
Fully verifiableThe crawler behind Arquivo.pt, Portugal's public web archive, which captures full page renders (HTML, CSS, JS, images) for long-term preservation. Built on Heritrix. Arquivo.pt publishes the ranges its crawlers operate from as a machine-readable JSON file in the JAFAR format, and recommends combining that IP check with the User-Agent check.
Operated by Arquivo.pt (FCCN/FCT) · Official documentation
User agents
Patterns this directory matches, with real observed strings.
- regexArquivo-web-crawler
- observedArquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling)
How to verify
Fetch the official IP range feed
curl -s https://arquivo.pt/crawlerips.json IP ranges (2)
Refreshed daily from the operator's feed. Also available as /data/ips/arquivo-pt.ips and /data/bots/arquivo-pt.json.
194.210.235.0/26 2001:690:a00:1039::/64
Good Bot Practices scorecard
- Identifies honestly
Stable UA token documented (1 pattern)
- Verifiable
Fully verifiable
- Respects robots.txt
Honors robots.txt (token: Arquivo-web-crawler)
- Behaves
Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset
- Reachable operator
Operator and documentation published
Measured against the Good Bot Practices.