Published June 3, 2022 | Version 1.0

Resources to compute TF-IDF weightings on press articles and tweets

Authors/Creators

  • 1. Laboratoire L3i, Université de La Rochelle

Description

These two datasets of features are used in order to compute TF-IDF weightings of documents. It is meant to be used with the compute-tf-idf-vectors program written in Python and available on Pypi.org.

- features_tweets.csv contains features (tokens, lemmas and entities) extracted from Tweets published by press agencies in french, german, spanish and english.

- features_news.csv contains features (tokens, lemmas and entities) extracted from articles published by Deutsche Welle in the same languages.

Files

Files (194.6 MB)

Name Size Download all
md5:df4571e74dff57d4995d5d90abe9fca4
188.1 MB Download
md5:dc138406caf9361bc58d71953352b779
6.5 MB Download

Additional details

Funding

European Commission
NewsEye - NewsEye: A Digital Investigator for Historical Newspapers 770299

References

  • Miranda, Sebastião, Artūrs Znotiņš, Shay B. Cohen, et Guntis Barzdins. 2018. « Multilingual Clustering of Streaming News ». In 2018 Conference on Empirical Methods in Natural Language Processing, 4535‑44. Brussels, Belgium: Association for Computational Linguistics. https://www.aclweb.org/anthology/D18-1483/.