Wikipedia Talk Page 'Climate Change' Sentiment and Toxicity Dataset
Description
The given dataset was prepared as part of a master's research project under the Master's program in Computational Social Systems at RWTH Aachen University.
The talk page was parsed using the GraWiTas tool in JSON format. The file Climate_change.comment_list.json is raw export of discussions that needs to be cleaned before using it for calculating sentiment and toxicity scores.
The sentiment scores were calculated using VADER and toxicity scores using the Perspective API by Google.
Date & time of dump is: 27-06-2022 12:12 UTC+02:00
Files
Climate_change.comment_list.json
Files
(36.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:de73187d1ca12f10b4862653b2583a41
|
18.3 MB | Preview Download |
|
md5:f8b2412510f775f9ddf4bb3374d55be6
|
17.9 MB | Preview Download |
Additional details
References
- Benjamin Cabrera, Laura Steinert, and Björn Ross. 2017. GraWiTas: a Grammar-based Wikipedia Talk Page Parser. In Proceedings of the Software Demonstrations of the 15th Conference of the European Chapter of the Association for Computational Linguistics, pages 21–24, Valencia, Spain. Association for Computational Linguistics.
- https://perspectiveapi.com/
- VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text (by C.J. Hutto and Eric Gilbert) Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.