Czech word2vec embeddings trained on OpenSubtitles Part 2
Description
This dataset contains the subs2vec embeddings for Czech, as presented in https://zenodo.org/records/17243814. The embeddings were trained on large-scale subtitle corpora and represent semantic vector spaces derived from naturalistic language use in films and television from the OpenSubtitles 2018 datasets: https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles.
For this language, we provide all embedding variants explored in the study. Specifically, the dataset includes vectors generated under different combinations of:
- Dimensionality: multiple vector sizes (e.g., 100, 200, 300, …)
- Window size: varying context windows (e.g., 2, 5, 10, …)
- Each file corresponds to a unique configuration (dimension × window size).
Each file contains the vocabulary for that language (column 1) and then the embedding values (columns 2 through dimension size + 1).
If you use this dataset, please cite:
- Manuscript: https://doi.org/10.5281/zenodo.17243812
- Data: This Zenodo dataset (using the DOI provided here)
Files
README.md
Files
(14.2 GB)
| Name | Size | |
|---|---|---|
|
md5:cab9b3d8a1df19239ec9a336f5ed9afd
|
500.0 MB | Download |
|
md5:8a0e3ff82281367644421a86ee7389f0
|
500.0 MB | Download |
|
md5:72ecc84853d55a3cd8dbf6955d8e576a
|
500.0 MB | Download |
|
md5:eb2fb498f04921b7506dcb9c0b957e3b
|
462.0 MB | Download |
|
md5:cbd7888ddb0151289f1e1c76debcd336
|
500.0 MB | Download |
|
md5:9dc8b1eea20477ea239b8989466e8405
|
500.0 MB | Download |
|
md5:bfbd74503dbea589a6fd1b3a5983685b
|
500.0 MB | Download |
|
md5:ff46d57b321ed41d90924259e38ef944
|
460.3 MB | Download |
|
md5:27a19f1525ba9ee5ab7179a2a8005d88
|
500.0 MB | Download |
|
md5:74a14658569f828bc37bf8098bbf5b57
|
500.0 MB | Download |
|
md5:383fc3baf28238420350b1c4e07c5634
|
500.0 MB | Download |
|
md5:b0515c93582785da6dadc53dfc7a5bab
|
460.9 MB | Download |
|
md5:1cdcf3c8c274d5080d775c849512f120
|
500.0 MB | Download |
|
md5:3ebe3b0afea7ebc8635f957171352ce3
|
500.0 MB | Download |
|
md5:b85b168deb796b9c9baf7e31f3061a69
|
500.0 MB | Download |
|
md5:22ce099bfb711925b67e3a1c9cfe172c
|
459.7 MB | Download |
|
md5:2225324b2305aa6ed89127ab5301d5e4
|
500.0 MB | Download |
|
md5:e3a560a22d9eb49e76315684da29434c
|
500.0 MB | Download |
|
md5:f40af889c852ab4069c399a462f0db5d
|
500.0 MB | Download |
|
md5:e90cb6f5a84a2ecfcbcf05c703a6917e
|
458.4 MB | Download |
|
md5:83e5d79e752186a93ffd85415654900c
|
500.0 MB | Download |
|
md5:0fa3754b36fd11d70803c3b049856fde
|
500.0 MB | Download |
|
md5:a9f162556c0987a92005df928ac67d5c
|
500.0 MB | Download |
|
md5:07cdb015142b1975a797afca28bf435d
|
459.3 MB | Download |
|
md5:b8cd890ba8c08d66cf0d2de7d16f86f0
|
200.1 MB | Download |
|
md5:bf2b42c7e2a851de04c9f46d5458a866
|
199.7 MB | Download |
|
md5:43a7c1fd5420f4fcc68ae5f0086050b1
|
200.2 MB | Download |
|
md5:ab9da00c19c764d36042db08d313a364
|
199.7 MB | Download |
|
md5:4e4b39ce08f4e5a906eb7fbd5d0c3ccb
|
200.3 MB | Download |
|
md5:8fbd7d2926db4997096623676d04506c
|
199.6 MB | Download |
|
md5:fcbefba7f88cbdd2de42ca0bc8278efe
|
200.2 MB | Download |
|
md5:3a9e8d20af5c50f489c7e219fa71d22b
|
199.6 MB | Download |
|
md5:0076667117a84b15b0c1864db72c35bc
|
200.2 MB | Download |
|
md5:854247be7f2a730cd1e14c39f2c0e604
|
199.7 MB | Download |
|
md5:8da7c0502fd3a70c0f3f7f9c2d74ca74
|
200.2 MB | Download |
|
md5:e0164c934d0524b35194674743770915
|
199.6 MB | Download |
|
md5:826f5465e694cf140b7a48209d422620
|
7.1 kB | Download |
|
md5:9771a8c82ee0f40ffbb9a4a2262e0357
|
3.2 kB | Preview Download |
Additional details
Related works
- Is supplement to
- Standard: 10.5281/zenodo.17243812 (DOI)
Software
- Repository URL
- https://github.com/SemanticPriming/word2manylanguages
- Programming language
- Python , R