Published October 10, 2023 | Version v1

Italian Word2Vec models

  • 1. University of Modena and Reggio Emilia

Description

Italian Word2Vec models trained from scratch on a dataset composed of:

- wiki: a dump of Italian Wikipedia (as of December 15, 2022), comprising 25,548,651 sentences and 526,640,982 words (3.2 GB of raw text);
- webz: a dataset of Italian news (159,226 documents) from the webz.io platform, crawled in October 2015, containing 44,041,823 sentences and 44,544,385 words (244 MB);
- a dataset of 5,510 Italian news articles from the newspaper ModenaToday (MT) or 15,115 documents from the Italian version of Reuters (RCV2).

w2v_wiki_wbz_mt_20_epochs.zip: Word2Vec model trained on the dataset consisting of wiki, webz, and MT for 20 epochs

w2v_wiki_wbz_mt_50_epochs.zip: Word2Vec model trained on the dataset consisting of wiki, webz, and MT for 50 epochs

w2v_wiki_wbz_reut_20_epochs.zip: Word2Vec model trained on the dataset consisting of wiki, webz, and RCV2 for 20 epochs

w2v_wiki_wbz_reut_50_epochs.zip: Word2Vec model trained on the dataset consisting of wiki, webz, and RCV2 for 50 epochs

Files

w2v_wiki_wbz_mt_20_epochs.zip

Files (4.1 GB)

Name Size
md5:c3741381a24ce2345f0223be1c6d6a28
1.0 GB Preview Download
md5:34eb36ef4d692a6d73f5f2132dc84a38
1.0 GB Preview Download
md5:d6c7727f1bea9b475e86d48401976431
1.0 GB Preview Download
md5:e112539e26bd6e813ccb28728451821c
1.0 GB Preview Download