Arquivo.pt Web Crawler
Listed onlyThe crawler behind Arquivo.pt, Portugal's public web archive, which captures full page renders (HTML, CSS, JS, images) for long-term preservation. Built on Heritrix; the operator documents no fixed IP range.
Operated by Arquivo.pt (FCCN/FCT) · Official documentation
User agents
Patterns this directory matches, with real observed strings.
- regexArquivo-web-crawler
- observedArquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling)
How to verify
No verification recipe published by the operator.
Good Bot Practices scorecard
- Identifies honestly
Stable UA token documented (1 pattern)
- Verifiable
Operator publishes no verification path
- Respects robots.txt
Honors robots.txt (token: Arquivo-web-crawler)
- Behaves
Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset
- Reachable operator
Operator and documentation published
Measured against the Good Bot Practices.