Published June 10, 2022 | Version 1.0

Fibvid dataset with multiple extracted features (both sparse and dense)

Authors/Creators

  • 1. Laboratoire L3i, Université de La Rochelle

Description

This is a publication of the FibVid dataset originaly dedicated to fake news detection. We changed here the purpose of this dataset in order to use it in the context of event tracking in press documents.

Kim, Jisu, Jihwan Aum, SangEun Lee, Yeonju Jang, Eunil Park, et Daejin Choi. 2021. « FibVID: Comprehensive Fake News Diffusion Dataset during the COVID-19 Period ». Telematics and Informatics 64 (novembre): 101688. https://doi.org/10.1016/j.tele.2021.101688.

In this dataset, we provide multiple features extracted from the text itself. Please note the text is missing from the dataset published in the CSV format for copyright reasons. You can download the original datasets and manually add the missing texts from the original publications.

Features are extracted using:

- A corpus of reference articles in multiple languages languages for TF-IDF weighting. (features_news) [1]

- A corpus of tweets reporting news for TF-IDF weighting. (features_tweets) [1]

- A S-BERT model [2] that uses distiluse-base-multilingual-cased-v1 (called features_use) [3]

- A S-BERT model [2] that uses paraphrase-multilingual-mpnet-base-v2 (called features_mpnet) [4]

References:

[1]: Guillaume Bernard. (2022). Resources to compute TF-IDF weightings on press articles and tweets (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6610406

[2]: Reimers, Nils, et Iryna Gurevych. 2019. « Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks ». In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3982‑92. Hong Kong, China: Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410.

[3]: https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v1

[4]: https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2

Files

fibvid_test.csv

Files (30.7 MB)

Name Size Download all
md5:9a6e1ca3a8b7ae6a8b07a23f5d81ca8c
25.2 kB Preview Download
md5:be6d9a0f6957e32a097ac1d944df9b93
4.9 MB Preview Download
md5:a1478d567e02ffba8f23f63026598c2f
361.6 kB Preview Download
md5:584777482b22d79724ade294bceb0be7
343.1 kB Preview Download
md5:3d3e99ca68ff7b1c4dcf33bd346420b4
3.2 MB Preview Download
md5:b45b5127977e3ade2373a46cedd4bb3f
62.0 kB Preview Download
md5:dd57de03a80354a1a66cd5ca7c182521
12.2 MB Preview Download
md5:7b034f04f4845b6f6efb32a4f1614d69
931.0 kB Preview Download
md5:ac8bf9ae6a5b0cbe0dd423ed4da340ea
888.2 kB Preview Download
md5:78a9377ec842bcdcdca3a7dc76ae3716
7.8 MB Preview Download

Additional details

Funding

European Commission
NewsEye - NewsEye: A Digital Investigator for Historical Newspapers 770299