Journal article Open Access

Automatic Detection of Online Abuse and Analysis of Problematic Users in Wikipedia

Rawat, Charu; Sarkar, Arnab; Singh, Sameer; Alvarado, Rafael; Rasberry, Lane

Today’s digital landscape is characterized by the pervasive presence of online communities. One of the persistent challenges to the ideal of free-flowing discourse in these communities has been online abuse. Wikipedia is a case in point, as it’s large community of contributors have experienced the perils of online abuse ranging from hateful speech to personal attacks to spam. Currently, Wikipedia has a human-driven process in place to identify online abuse. In this paper, we propose a framework to understand and detect such abuse in the English Wikipedia community. We analyze the publicly available data sources provided by Wikipedia. We discover that Wikipedia’s XML dumps require extensive computing power to be used for temporal textual analysis, and, as an alternative, we propose a web scraping methodology to extract user-level data and perform extensive exploratory data analysis to understand the characteristics of users who have been blocked for abusive behavior in the past. With these data, we develop an abuse detection model that leverages Natural Language Processing techniques, such as character and word n-grams, sentiment analysis and topic modeling, and generates features that are used as inputs in a model based on machine learning algorithms to predict abusive behavior. Our best abuse detection model, using XGBoost Classifier, gives us an AUC of ~84%.

Files (416.3 kB)
Name Size
Automatic Detection of Online Abuse and Analysis of Problematic Users in Wikipedia.pdf
md5:a17cc1eadcd0183a7673ef93277da7dc
416.3 kB Download
78
33
views
downloads
All versions This version
Views 7879
Downloads 3333
Data volume 13.7 MB13.7 MB
Unique views 6768
Unique downloads 3131

Share

Cite as