Tunizi: Tunisian Arabizi Sentiment Analysis Dataset

doi:10.5281/zenodo.4275240

African Natural Language Processing (AfricaNLP)

Published November 13, 2020 | Version 01

Dataset Open

Tunizi: Tunisian Arabizi Sentiment Analysis Dataset

Chayma Fourati¹

1. iCompass

Tunizi is the first 100% Tunisian Arabizi sentiment analysis dataset. Tunisian Arabizi is the representation of the tunisian dialect written in Latin characters and numbers rather than Arabic letters.We gathered comments from social media platforms that express sentiment about popular topics. For this purpose, we extracted 100k comments using public streaming APIs. Tunizi was preprocessed by removing links, emoji symbols, and punctuations.

The collected comments were manually annotated using an overall polarity: positive (1), negative (-1) and neutral (0) class. We divided the dataset into separate training, validation and test sets, with a ratio of 7:1:2 with a balanced split where the number of comments from positive class and negative class are almost the same.

Notes

This dataset contains comments written in Tunisian Arabizi which represents the Tunisian dialect written in Latin letters and numbers.

Files

Files (4.7 MB)

Name	Size	Download all
tunizi_train md5:5be37f4fa75cc3b5af5860f4504cdcc9	4.7 MB	Download

Views

235

Downloads

Show more details

	All versions	This version
Views	1,547	1,513
Downloads	235	233
Data volume	1.2 GB	1.2 GB

More info on how stats are collected....

DOI

Resource type

Dataset

Publisher

Zenodo

Languages

Tunisian Arabic

Creative Commons Attribution 4.0 International

The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited. Read more

Technical metadata

Created: November 16, 2020
Modified: December 9, 2020

Tunizi: Tunisian Arabizi Sentiment Analysis Dataset

Creators

Description

Notes

Files

Files (4.7 MB)