Published January 18, 2022 | Version v1

Korean embedding files using the different morphological segmentation granularity of the word

Authors/Creators

  • 1. UBC

Description

Embedding files using the following segmentation:

  1. wordUD
  2. morphUD
  3. +morphUD 

Based on wordUD there are 9,692,938 sentences and 157,653,628 words (tokenized) including all articles published in The Hankyoreh during 2016 (1.2M sentences), Sejong morphologically analyzed corpus (3M), and Korean Wiki (20201101) (5.3M):

 

./fasttext skipgram -input input -output embedding -dim 300

Files

Files (17.3 GB)

Name Size
md5:5f1519031614a874b31f61b8b8532d15
3.5 GB Download
md5:c479edb1d723658dff0a0771d4ad4165
1.2 GB Download
md5:6eb44d10923f850623b4ab808fce77c0
3.5 GB Download
md5:09af129267bbc72427d1fc3c4c85314b
1.3 GB Download
md5:f7af7c58eb00080e591e3fe243210e7f
5.1 GB Download
md5:7e1d356d4680f88e8887f1fc3a3f7d83
2.9 GB Download