Architectural styles of curiosity in global Wikipedia mobile app readership
Authors/Creators
Description
Description of the data and file structure
These directories contain the code, aggregated data, and preprocessing scripts to re-create the figures in "Architectural styles of curiosity in global Wikipedia readership"
Files and variables
File: Archive.zip
Figures: Publically usable illustrations are available here.
Description: Contains analysis, data, preprocessing, results, and utils folders. 14 directories, 25 files.
|-- analysis| |-- KNOT_analysis.R <- analyzes laboratory data| |-- analyze_1000-networks.ipynb <- analyzes naturalistic data| |-- analyze_1000-networks_comparison-knot-rw_clean.ipynb <- compares datasets and nulls| |-- analyze_KNOT_networks.ipynb <- analyzes laboratory data| |-- analyze_forward_flow.ipynb <- calculates forward flow| |-- forest_plots.R <- correlations wtih sociodemographic variables| |-- topic_analysis.R <- analysis of topic and information diversity| `-- worldmap.R <- visualization of geographical data sources|-- data| |-- laboratory_data <- variables for laboratory browsing and survey data| |-- mobile_app_data <- aggregated data for network structure and topic (rows are individuals)| |-- pretrained_embeddings <- fastText word embeddings| |-- spatial_navigation <- data from Sea Hero Quest| |-- surveys <- data from nationally aggregated sociodemographic surveys| `-- wikispeedia <- data from WikiSpeedia game|-- preprocessing| |-- data_knowledge-networks_generate-subsample_clean.ipynb <- processes mobile app data| |-- data_knowledge-networks_metrics-combined_clean.ipynb <- calculates network metrics| |-- data_knowledge-networks_rw_get-data.ipynb <- calculates null networks| `-- data_knowledge-networks_sessions-app_cleaned.ipynb <- processes individual browsing|-- requirements.txt|-- results| |-- UMAP <- data used to generate network embedding (rows are individuals)| `-- figs <- code for generated figure on forward flow `-- utils |-- plot_knowledge-networks_network-comparison.ipynb <- visualizations of network comparisons |-- plot_knowledge-networks_network-metrics_distance.ipynb <- visualizations of distance between datasets |-- plot_knowledge-networks_summary-stats.ipynb <- visualizations of summary stats |-- utils_embedding.py <- get word and document embeddings |-- utils_filtration_metrics.py <- higher-order topology functions (unused) |-- utils_gt.py <- graph-tool functions |-- utils_network.py <- functions to make networks from series of article IDs |-- utils_network_metrics.py <- network metrics |-- utils_networkx.py <- networkx functions |-- utils_rw.py <- functions to generate random walks and null models `-- utils_tokenizer.py <- functions for processing embeddings.
Code/software
See requirements.txt
Access information
Other publicly accessible locations of the data:
* https://gitlab.wikimedia.org/repos/research/curiosity
Additional data was derived from the following sources:
* [Human Development Index]
* [World Happiness Report]
* [WikiSpeedia]
* [FastText]
* [Sea Hero Quest]
* [Knowledge Networks Over Time Study]
Files
Archive.zip
Files
(71.8 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:a81ed1bb45696260ca5c89fe6c3aa3e6
|
71.8 MB | Preview Download |
Additional details
Identifiers
Related works
- Is supplement to
- 10.31234/osf.io/szuyj (DOI)