Resources to compute TF-IDF weightings on press articles and tweets
Description
These two datasets of features are used in order to compute TF-IDF weightings of documents. It is meant to be used with the compute-tf-idf-vectors program written in Python and available on Pypi.org.
- features_tweets.csv contains features (tokens, lemmas and entities) extracted from Tweets published by press agencies in french, german, spanish and english.
- features_news.csv contains features (tokens, lemmas and entities) extracted from articles published by Deutsche Welle in the same languages.
Files
Files
(194.6 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:df4571e74dff57d4995d5d90abe9fca4
|
188.1 MB | Download |
|
md5:dc138406caf9361bc58d71953352b779
|
6.5 MB | Download |
Additional details
Funding
References
- Miranda, Sebastião, Artūrs Znotiņš, Shay B. Cohen, et Guntis Barzdins. 2018. « Multilingual Clustering of Streaming News ». In 2018 Conference on Empirical Methods in Natural Language Processing, 4535‑44. Brussels, Belgium: Association for Computational Linguistics. https://www.aclweb.org/anthology/D18-1483/.