There is a newer version of this record available.

Dataset Restricted Access

Restricted Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior"

Antigoni-Maria Founta; Constantinos Djouvas; Despoina Chatzakou; Ilias Leontiadis; Jeremy Blackburn; Gianluca Stringhini; Athena Vakali; Michael Sirivianos; Nicolas Kourtellis

Restricted Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found here.  The Public version of the dataset can be found here

  • hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K raws with tweet text, the associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). 

  • retweets.csv: contains ~2K rows, where every row consists of the row number in the  hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the  first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences.

Please cite the paper in any published work that uses any of these resources. 

@inproceedings{founta2018large, 
    title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior}, 
    author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas}, 
    booktitle={11th International Conference on Web and Social Media, ICWSM 2018}, 
    year={2018}, 
    organization={AAAI Press} 

For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy  

Restricted Access

You may request access to the files in this upload, provided that you fulfil the conditions below. The decision whether to grant/deny access is solely under the responsibility of the record owner.


  1. You will not attempt to use this data to de-anonymize, in any way, any users in this or any other dataset.
  2. You will not re-share the dataset with anyone not included in this request.
  3. You will appropriately cite the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" ICWSM 2018 paper in any publication, of any form and kind, using this data

306
68
views
downloads
All versions This version
Views 306174
Downloads 6837
Data volume 788.0 MB470.0 MB
Unique views 199137
Unique downloads 4428

Share

Cite as