Chilean Twitter Hate Speech Dataset
Authors/Creators
Description
The dataset comprises a total of 4,547 Tweets ID’s, Authors ID’s about Hate Speech related to Chilean dialect or news that were posted on Twitter from 2020 to July 2022 and, 6542 Tweets ID’s related to the classified tweet's context.
-
tweets_data.csv - Whole dataset.
-
referenced_tweets_data.csv - Referenced Tweets data
-
tweets_data.parquet.gzip - Whole dataset.
-
referenced_tweets_data.parquet.gzip - Referenced Tweets data
The training set includes 4,547 examples labeled in 5 clases: "Odio", "Mujeres", "Comunidad LGBTQ+", "Comunidades Migrantes", "Pueblos Originarios" with values from 0 to 3 indicating the amount of annotators that indicated the tweet belonged in that class.
The Whole dataset includes the following columns:
-
tweet_id: Tweet identifier. (Anonymized)
-
author_id: Author identifier. (Anonymized)
-
conversation_id: Tuple which contains the tweets_id (from the file referenced_tweeets_data.csv) to which the labeled tweet references. this Id’s are in such order that in the first position is the tweet referenced in the labeled tweet, then the id in the second position is referenced by the tweet in the first position and so on and so on…
-
text: full text.
- Odio: Hate classification votes.
- Mujeres: Women classification votes.
- Comunidades Migrantes: Immigrant communities classification votes.
- Pueblos Originarios: Native americans classification votes.
The Referenced Tweets data has the following columns:
-
tweet_id: Tweet identifier. (Anonymized)
-
author_id: Author identifier. (Anonymized)
-
conversation_id: Contains either an ID number indicating the next referenced tweet in Referenced Tweets data or a 0 indicative of that tweet being the last referenced tweet.
- text: full text.
Abstract
In the last few years, several organizations have manifested their concern over the increase in use of Hateful Speech or Hate Speech for short, this concept refers to forms of expression or audiovisual content that encourage discrimination or violence against individuals or groups solely based on their gender, sexual orientation, ethnicity, religion, or nationality. Being able to monitor this phenomenon in a timely manner can help societies and their governments to prevent tensions, crimes, and conflicts that endangers not only the most fundamental democratic values but also order stability and social peace.
The fast massification of social platforms has transformed them into one of the main mediums used by people today for creating and sharing information. Consequently, social media platforms such as Twitter, Instagram, or Facebook are the staging in which Hate Speech is mostly propagated today. Sadly the great reach of these platforms, their public nature, the social dynamics that are perpetuated in them and the absence of an explicit regulatory framework, only worsen and increase the magnitude of this phenomena. Mining such conversations, such as Tweets, to develop a dataset can serve as a data resource for interdisciplinary research related to the analysis of interest, views, opinions and help us in the creation of tools to further our understanding of social dynamics related to Hate Speech propagation and analysis.
Files
tweets_data.csv
Additional details
Additional titles
- Subtitle (Spanish)
- CL2
Dates
- Available
-
2025-01-15