Published January 18, 2022
| Version v1
Dataset
Open
Korean embedding files using the different morphological segmentation granularity of the word
Description
Embedding files using the following segmentation:
- wordUD
- morphUD
- +morphUD
Based on wordUD there are 9,692,938 sentences and 157,653,628 words (tokenized) including all articles published in The Hankyoreh during 2016 (1.2M sentences), Sejong morphologically analyzed corpus (3M), and Korean Wiki (20201101) (5.3M):
./fasttext skipgram -input input -output embedding -dim 300
Files
Files
(17.3 GB)
| Name | Size | |
|---|---|---|
|
md5:5f1519031614a874b31f61b8b8532d15
|
3.5 GB | Download |
|
md5:c479edb1d723658dff0a0771d4ad4165
|
1.2 GB | Download |
|
md5:6eb44d10923f850623b4ab808fce77c0
|
3.5 GB | Download |
|
md5:09af129267bbc72427d1fc3c4c85314b
|
1.3 GB | Download |
|
md5:f7af7c58eb00080e591e3fe243210e7f
|
5.1 GB | Download |
|
md5:7e1d356d4680f88e8887f1fc3a3f7d83
|
2.9 GB | Download |