TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia

Flöck, Fabian; Erdogan, Kenan; Acosta, Maribel

doi:10.5281/zenodo.345571

Published March 4, 2017 | Version v1

Dataset Open

TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia

1. GESIS - Leibniz Institute for the Social Sciences
2. Karlsruhe Institute of Technology

Please cite 10.5281/zenodo.789289 for all versions of this dataset, which will always resolve to the latest.

-----------------

This dataset contains every instance of all tokens (≈ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article revision it was originally created in, and (ii) lists with all the revisions in which the token was ever deleted and (potentially) re-added and re-deleted from its article, enabling a complete and straightforward tracking of its history.

This data would be exceedingly hard to create by an average potential user as it is (i) very expensive to compute and as (ii) accurately tracking the history of each token in revisioned documents is a non-trivial task.
Adapting a state-of-the-art algorithm, we have produced a dataset that allows for a range of analyses and metrics, already popular in research and going beyond, to be generated on complete-Wikipedia scale; ensuring quality and allowing researchers to forego expensive text-comparison computation, which so far has hindered scalable usage.

This dataset, its creation process and use cases are described in a dedicated dataset paper of the same name, published at the ICWSM 2017 conference. In this paper, we show how this data enables, on token level, computation of provenance, measuring survival of content over time, very detailed conflict metrics, and fine-grained interactions of editors like partial reverts, re-additions and other metrics.

Tokenization used: https://gist.github.com/faflo/3f5f30b1224c38b1836d63fa05d1ac94

Toy example for how the token metadata is generated:
https://gist.github.com/faflo/8bd212e81e594676f8d002b175b79de8

Be sure to read the ReadMe.txt or - even more detailed - the supporting paper which is referenced under "related identifiers".

Notes

Attention: In this current version we spotted an inconsistency with the 'str' column (token values ) in token csv files. Some of the tokens which contain regular quotes ( '"' ) inside were written into csv files without considering " as a quoting character. This will be fixed in an upcoming version. For example a token '"press' must be written into csv as '"""press', but it is sometimes written as '"press' in csv files. To overcome this while parsing csv files: 1. Iterate file in lines - 2. Split line with ',' - 3. Check if str value (4th item after split) starts and ends with '"'. If yes, remove them and replace '""' with '"'. Example python function: https://gist.github.com/faflo/19d3cf1768fbd7939f76ce3e9ee3b087

Files

ReadMe.txt

Files (68.8 GB)

Name	Size	Download all
20161101-current_content-parts-1-50-pageids-12-117215.7z md5:9e8f7a54c341c568e69e06a7754ea41d	1.6 GB	Download
20161101-current_content-parts-101-150-pageids-418317-1081580.7z md5:d308469bbf760a0e65b2cee8ab8c6935	2.1 GB	Download
20161101-current_content-parts-151-200-pageids-1081586-2203796.7z md5:0b70a450bac2e84823098284e7782936	2.2 GB	Download
20161101-current_content-parts-201-250-pageids-2203809-4051322.7z md5:44acc4c9a3c0aeb7abb0ec3949baaaad	2.3 GB	Download
20161101-current_content-parts-251-300-pageids-4051356-7027309.7z md5:3433d20fb707b9ba947983291870b0b2	2.4 GB	Download
20161101-current_content-parts-301-350-pageids-7027310-11781922.7z md5:87cdb9c81ea6dadc0de91fb76b2e66db	2.5 GB	Download
20161101-current_content-parts-351-400-pageids-11781924-17443368.7z md5:a5fc1ea8ec93ba314436066dd6000ea2	2.7 GB	Download
20161101-current_content-parts-401-450-pageids-17443414-23281466.7z md5:f49751a97ce48342f25e76b53e573995	2.7 GB	Download
20161101-current_content-parts-451-500-pageids-23281469-29590519.7z md5:508d889366d9d636e0ba7b91f931fcb6	2.9 GB	Download
20161101-current_content-parts-501-550-pageids-29590554-36522618.7z md5:949d81981bd51d1ea37f573ee973a2de	3.0 GB	Download
20161101-current_content-parts-51-100-pageids-117216-418311.7z md5:4cd4b1c263db0de13a15d0f5f3d55b2c	2.0 GB	Download
20161101-current_content-parts-551-600-pageids-36522655-43525178.7z md5:91c6ff5bd7a5bc48d6ddd9efcd053c19	3.1 GB	Download
20161101-current_content-parts-601-646-pageids-43525205-52158752.7z md5:03693a4affc80eca9cae488c6c23ec2c	3.0 GB	Download
20161101-deleted_content-parts-1-50-pageids-12-117215.7z md5:03b3fe3a13c8dc460786b7c4a8688b46	3.8 GB	Download
20161101-deleted_content-parts-101-150-pageids-418317-1081580.7z md5:3985763b3f1dab3fab53903e135cb96b	3.3 GB	Download
20161101-deleted_content-parts-151-200-pageids-1081586-2203796.7z md5:71681cdf933991d5b1efbc33b1a9d7fd	3.1 GB	Download
20161101-deleted_content-parts-201-250-pageids-2203809-4051322.7z md5:6c1faec136e9610cf79ab5b187aede86	2.9 GB	Download
20161101-deleted_content-parts-251-300-pageids-4051356-7027309.7z md5:b317542e7126c1427feebaae6ed3b579	2.7 GB	Download
20161101-deleted_content-parts-301-350-pageids-7027310-11781922.7z md5:7df7782c1f34423f7ded596965c9239d	2.5 GB	Download
20161101-deleted_content-parts-351-400-pageids-11781924-17443368.7z md5:51345b0dd5cbe54cdad28478972d6cc0	2.2 GB	Download
20161101-deleted_content-parts-401-450-pageids-17443414-23281466.7z md5:142011b6f99c0cb2f91f9ee7f817e576	2.1 GB	Download
20161101-deleted_content-parts-451-500-pageids-23281469-29590519.7z md5:febb9a9e66c27b6d33ee01f46bf86059	1.9 GB	Download
20161101-deleted_content-parts-501-550-pageids-29590554-36522618.7z md5:abf2fa6e951365facf69bc4e193009c1	1.7 GB	Download
20161101-deleted_content-parts-51-100-pageids-117216-418311.7z md5:020a64d61e6a4d9b952bab5304a748d2	3.4 GB	Download
20161101-deleted_content-parts-551-600-pageids-36522655-43525178.7z md5:634ac42e1216bf925245052f07de2bd1	1.4 GB	Download
20161101-deleted_content-parts-601-646-pageids-43525205-52158752.7z md5:d4ee1a0636827f555312f99fba8b1b3f	879.4 MB	Download
ReadMe.txt md5:2fa12b2142007c33f870e217452bdf9d	4.5 kB	Preview Download
revisions.7z md5:3ef8944885617ac0efa4738aab786f2b	4.6 GB	Download

Additional details

Is referenced by: https://arxiv.org/abs/1703.08244 (URL)
Is supplemented by: https://arxiv.org/abs/1703.08244 (URL); 10.5281/zenodo.439699 (DOI)

Flöck, Fabian, and Acosta, Maribel. "WikiWho: Precise and efficient attribution of authorship of revisioned content." Proceedings of the 23rd international conference on World Wide Web. ACM, 2014.
Fabian Flöck, Kenan Erdogan, Maribel Acosta. "TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia." Proceedings of ICWSM2017 (to appear). Preprint: https://arxiv.org/abs/1703.08244

	All versions	This version
Views	6,912	4,832
Downloads	6,550	2,797
Data volume	17.4 TB	6.9 TB

ReadMe.txt

Files (68.8 GB)

Related works

References

TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia

Authors/Creators

Description

Notes

Files

ReadMe.txt

Files (68.8 GB)

Additional details

Related works

References