Tracking Knowledge Propagation Across Wikipedia Languages

Valentim, Rodolfo; Comarela, Giovanni; Park, Souneil; Saez-Trumper, Diego

doi:10.5281/zenodo.4433137

Published March 15, 2021 | Version v1

Dataset Open

Tracking Knowledge Propagation Across Wikipedia Languages

1. Politecnico di Torino
2. Federal University of Espírito Santo
3. Telefonica Research
4. Wikimedia Foundation

We present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow follow up research on building predictive models of them. For this purpose, we align all the Wikipedia articles in a language-agnostic manner according to the concept they cover, which results in 13M propagation instances. To the best of our knowledge, this dataset is the first to explore the full inter-language propagation at a large scale. Together with the dataset, a holistic overview of the propagation and key insights about the underlying structural factors are provided to aid future research. For example, we find that although long cascades are unusual, the propagation tends to continue further once it reaches more than four language editions. We also find that the size of language editions are associated with the speed of propagation. We believe the dataset not only contributes to the prior literature on Wikipedia growth but also enables new use cases such as edit recommendation for addressing knowledge gaps, detection of disinformation, and cultural relationship analysis.

Files

dataset.csv.zip

Files (673.8 MB)

Name	Size	Download all
dataset.csv.zip md5:49cf4224d6cd75423773c17656fab416	321.9 MB	Preview Download
dataset.jsonl.zip md5:3f7d416827fe0edc887933f785096560	351.9 MB	Preview Download

	All versions	This version
Views	2,298	2,289
Downloads	477	477
Data volume	195.1 GB	195.1 GB

Tracking Knowledge Propagation Across Wikipedia Languages

Authors/Creators

Description

Files

dataset.csv.zip

Files (673.8 MB)