There is a newer version of the record available.

Published April 26, 2023 | Version v1.0.0

FactNews: Sentence-Level Annotated Dataset To Predict Factually and Media Bias

  • 1. University of São Paulo
  • 2. National University of Singapore
  • 3. Federal University of Minas Gerais

Description

Automated news credibility and fact-checking at scale require accurate prediction of news factuality and media bias. Here, we introduce a large sentence-level dataset, titled "FactNews", composed of 6,191 sentences expertly annotated according to factuality and media bias definitions proposed by AllSides. We use "FactNews" to assess the overall reliability of news sources by formulating two text classification problems for predicting sentence-level factuality of news reporting and bias of media outlets. Our experiments demonstrate that biased sentences present a higher number of words compared to factual sentences, besides having a predominance of emotions. Hence, the fine-grained analysis of subjectivity and impartiality of news articles showed promising results for predicting the reliability of the entire media outlet. Finally, due to the severity of fake news and political polarization in Brazil, and the lack of research for Portuguese, both dataset and baseline were proposed for Brazilian Portuguese.

Files

factnews_dataset.csv

Files (2.2 MB)

Name Size Download all
md5:2f73a13104085a58c94604646138ca0f
1.1 MB Preview Download
md5:377fe8e40b1a165fa026bfa84334802e
1.1 MB Preview Download
md5:54473da7cca74a6b005e485c77240ad1
86 Bytes Preview Download

Additional details

References

  • This paper was published at RANLP 2023