Published September 8, 2026 | Version v2

Croatian word2vec embeddings trained on OpenSubtitles Part 2

  • 1. ROR icon Harrisburg University of Science and Technology

Description

This dataset contains the subs2vec embeddings for Croatian, as presented in https://zenodo.org/records/17243814. The embeddings were trained on large-scale subtitle corpora and represent semantic vector spaces derived from naturalistic language use in films and television from the OpenSubtitles 2018 datasets: https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles

For this language, we provide all embedding variants explored in the study. Specifically, the dataset includes vectors generated under different combinations of:

  • Dimensionality: multiple vector sizes (e.g., 100, 200, 300, …)
  • Window size: varying context windows (e.g., 2, 5, 10, …)
  • Each file corresponds to a unique configuration (dimension × window size). 

Each file contains the vocabulary for that language (column 1) and then the embedding values (columns 2 through dimension size + 1). 

If you use this dataset, please cite:

Files

Files (1.4 GB)

Name Size
md5:f09c97a06e6b63c316174e41d39b4753
138.9 MB Download
md5:1772b262defe7dfc1b4a7de8d0a5ad4b
138.7 MB Download
md5:ee33e81e7033bd8696fac7b32faa57f2
139.0 MB Download
md5:1a9c25501e3e3175449e9c4469b0e81e
138.6 MB Download
md5:6b0077f74028589d5f39cbadfa779b20
139.0 MB Download
md5:222f452d15540a4b41150e4b5751f772
138.7 MB Download
md5:298c54385de5a96cd2750465a1745c51
139.0 MB Download
md5:22671c6c36ac0b3e4c043587c62f19ba
138.6 MB Download
md5:3645c5c806ddcb799aa0e5dae177f886
139.0 MB Download
md5:e0568b725d92824c4a21bd3adfc6cf34
138.6 MB Download

Additional details

Related works

Is supplement to
Standard: 10.5281/zenodo.17243812 (DOI)

Software