Published July 12, 2021
| Version V1.0
Dataset
Open
COMPRISE_Data10_MENYO-20K_V1.0
Description
A multi-domain parallel corpus for English-Yoruba language pair that can be used to benchmark machine translation systems. The dataset has 20,100 parallel sentences split into 10,070 train, 3,397 dev, and 6,633 test sentences.
Files
COMPRISE_Data10_MENYO-20K_V1.0.zip
Files
(2.8 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:77cbd001e22339e03cd20ea98e5661a1
|
2.8 MB | Preview Download |