Published November 24, 2021 | Version 1

Latin-to-Ethiopic Script Transliteration Dataset

Authors/Creators

  • 1. Addis Ababa University

Description

The dataset contains social media user comments written in Latin scripts but having Amharic meaning. For the purpose of hate speech detection and sentiment analysis, we have developed a Latin-to-Ethiopic script transliteration model so as to incorporate users' comments written in Latin scripts. For the experiment, we have used these datasets. The dataset contains social media user comments and named entities (which include lists of students' names). 

Files

namedEntity_v1.txt

Files (30.0 kB)

Name Size Download all
md5:436ea3662ae3025e0f7cb5a9b09cafa9
5.5 kB Preview Download
md5:dc03b276a581da76ea54d9b6d1c70d39
9.3 kB Preview Download
md5:d2357262f9b35fba78a87956c6cd98d8
15.2 kB Preview Download