Directory / Archivers

CLASSLA-web

Listed only

The CLASSLA-web crawler collects web corpora for South Slavic and other languages under the CLARIN Knowledge Centre for South Slavic languages, continuing the CEF-funded MaCoCu project's crawling infrastructure. It runs the SpiderLing crawler developed at Masaryk University. The operator documents that it adheres to the robots exclusion standard and reads robots.txt on the first access of each crawl run.

Operated by CLARIN.SI (CLASSLA Knowledge Centre) · Official documentation

User agents

Patterns this directory matches, with real observed strings.

  • regexCLASSLA-web
  • observedMozilla/5.0 (compatible; CLASSLA-web; +https://www.clarin.si/info/classla-web-crawler/)

How to verify

No verification recipe published by the operator.

Good Bot Practices scorecard

  • Identifies honestly

    Stable UA token documented (1 pattern)

  • Verifiable

    Operator publishes no verification path

  • Respects robots.txt

    Honors robots.txt (token: CLASSLA-web)

  • Behaves

    Crawl-rate behavior is operator-declared; not machine-verifiable from this dataset

  • Reachable operator

    Operator and documentation published

Measured against the Good Bot Practices.