Journal article Open Access

Automatic Detection of Online Abuse and Analysis of Problematic Users in Wikipedia

Rawat, Charu; Sarkar, Arnab; Singh, Sameer; Alvarado, Rafael; Rasberry, Lane

Today’s digital landscape is characterized by the pervasive presence of online communities. One of the persistent challenges to the ideal of free-flowing discourse in these communities has been online abuse. Wikipedia is a case in point, as it’s large community of contributors have experienced the perils of online abuse ranging from hateful speech to personal attacks to spam. Currently, Wikipedia has a human-driven process in place to identify online abuse. In this paper, we propose a framework to understand and detect such abuse in the English Wikipedia community. We analyze the publicly available data sources provided by Wikipedia. We discover that Wikipedia’s XML dumps require extensive computing power to be used for temporal textual analysis, and, as an alternative, we propose a web scraping methodology to extract user-level data and perform extensive exploratory data analysis to understand the characteristics of users who have been blocked for abusive behavior in the past. With these data, we develop an abuse detection model that leverages Natural Language Processing techniques, such as character and word n-grams, sentiment analysis and topic modeling, and generates features that are used as inputs in a model based on machine learning algorithms to predict abusive behavior. Our best abuse detection model, using XGBoost Classifier, gives us an AUC of ~84%.

Files (416.3 kB)
Name Size
Automatic Detection of Online Abuse and Analysis of Problematic Users in Wikipedia.pdf
md5:a17cc1eadcd0183a7673ef93277da7dc
416.3 kB Download
132
47
views
downloads
All versions This version
Views 132133
Downloads 4747
Data volume 19.6 MB19.6 MB
Unique views 120121
Unique downloads 4545

Share

Cite as