Published December 7, 2018 | Version v1

csTenTen17, a Recent Czech Web Corpus

Authors/Creators

  • 1. Lexical Computing

Description

This article introduces a very large Czech text corpus for language research – csTenTen17 compiled from texts downloaded in 2015, 2016 and 2017. The corpus is consisting of 10.5 billion words reaching double the size of its predecessor from 2012. A brief comparison with other recent Czech corpora follows.

Files

paper10-Suchomel.pdf

Files (439.6 kB)

Name Size Download all
md5:70a24a61fb6c7d35321a9e7aac46b4dc
439.6 kB Preview Download

Additional details

Funding

European Commission
ELEXIS - European Lexicographic Infrastructure 731015