Published October 10, 2023 | Version v1

Italian FastText models

  • 1. University of Modena and Reggio Emilia

Description

Italian FastText models trained from scratch on a dataset composed of:

- wiki: a dump of Italian Wikipedia (as of December 15, 2022), comprising 25,548,651 sentences and 526,640,982 words (3.2 GB of raw text);
- webz: a dataset of Italian news (159,226 documents) from the webz.io platform, crawled in October 2015, containing 44,041,823 sentences and 44,544,385 words (244 MB);
- a dataset of 5,510 Italian news articles from the newspaper ModenaToday (MT) or 15,115 documents from the Italian version of Reuters (RCV2).

ft_wiki_wbz_mt_20_epochs: FastText model trained on the dataset consisting of wiki, webz, and MT for 20 epochs

ft_wiki_wbz_mt_50_epochs: FastText model trained on the dataset consisting of wiki, webz, and MT for 50 epochs

ft_wiki_wbz_reut_20_epochs: FastText model trained on the dataset consisting of wiki, webz, and RCV2 for 20 epochs

ft_wiki_wbz_reut_50_epochs: FastText model trained on the dataset consisting of wiki, webz, and RCV2 for 50 epochs

Files

ft_wiki_wbz_mt_20_epochs.zip

Files (13.1 GB)

Name Size
md5:54e2161839492d26b3dbd8ec579595e6
3.3 GB Preview Download
md5:505a5d26441ed7550c3dc81cdb54855b
3.3 GB Preview Download
md5:97c8a9ec384d29dc89411549bc423b64
3.3 GB Preview Download
md5:fbada47e419664ae8b7c3a4ec4ebfc45
3.3 GB Preview Download