CLASSLA-web
Listed onlyThe CLASSLA-web crawler collects web corpora for South Slavic and other languages under the CLARIN Knowledge Centre for South Slavic languages, continuing the CEF-funded MaCoCu project's crawling infrastructure. It runs the SpiderLing crawler developed at Masaryk University. The operator documents that it adheres to the robots exclusion standard and reads robots.txt on the first access of each crawl run.
Operated by CLARIN.SI (CLASSLA Knowledge Centre) · Official documentation
User agents
Patterns this directory matches, with real observed strings.
- regexCLASSLA-web
- observedMozilla/5.0 (compatible; CLASSLA-web; +https://www.clarin.si/info/classla-web-crawler/)
How to verify
No verification recipe published by the operator.
Good Bot Practices scorecard
- Identifies honestly
Stable UA token documented (1 pattern)
- Verifiable
Operator publishes no verification path
- Respects robots.txt
Honors robots.txt (token: CLASSLA-web)
- Behaves
Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset
- Reachable operator
Operator and documentation published
Measured against the Good Bot Practices.