Leipzig Corpora Collection Crawler
Listed onlyThe LCC crawler is operated by Leipzig University to collect web text for the Leipzig Corpora Collection, a set of linguistic corpora used in natural language processing research.
Operated by Leipzig University Natural Language Processing Group · Official documentation
User agents
Patterns this directory matches, with real observed strings.
- regex^LCC
- observedLCC (+http://corpora.informatik.uni-leipzig.de/crawler_faq.html)
How to verify
No verification recipe published by the operator.
Good Bot Practices scorecard
- Identifies honestly
Stable UA token documented (1 pattern)
- Verifiable
Operator publishes no verification path
- Respects robots.txt
Does not honor robots.txt
- Behaves
Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset
- Reachable operator
Operator and documentation published
Measured against the Good Bot Practices.