Published October 26, 2019 | Version Version 1.0

Datasets from Costa Rican news sources for fake news detection

Description

Today, technology has changed the way information is propagated and how the message is received. The interpretation of the news may have different angles depending on the source of origin. Because of this, there has been an increase in misinformation, in the way of influencing public opinion and in how we perceive or estimate reality.
The objective of this beta dataset is to be used for the evaluation of data mining models that allow the classification of true or potentially fake news that are generated by Costa Rican news sites only.  This is intended to assess the level of reliability of the models and extend the scope of this research in future work.

The dataset has been pre-processed (standarized using lower cases, lemmatized and removed any possible noise from it) and analyzed using LIWC dictionaries. One version has the news text in Spanish and was processed using LIWC2007 dictionary in Spanish. The second version was processed using LIWC2015 English dictionary and has the news text in English. The reason to having two versions is to be able to test using the newer LIWC dictionary which includes more Summary Language Variables that the Spanish version doesn't have and analyze how this and other variables can contribute to different results when creating models. 

The file "DescripcionVariables" provides a description of all variables used.

 

Notes

There is a third file included "datasource_clasificado_webhose" which has over 6000 news articles from Costa Rica. This is in raw format and hasn't been pre-processed nor analyzed with LIWC dictionaries and was obtained using Webohose.io API.

Files

Files (45.1 MB)

Name Size Download all
md5:945cf9e10841e2e3183b4c14b5c0a330
27.7 MB Download
md5:7ed090af6516833a15591547d7c12e05
15.9 MB Download
md5:ca12a593f82a54d658bc0da06f4f86b9
802.5 kB Download
md5:d2913a055b9660022dc5f3ce59a53b4f
813.9 kB Download