Published December 1, 2021 | Version v1

Topic Model for English Wikipedia's Biographies with list of all 1.8M articles linked to Wikidata

  • 1. City University of New York

Description

A Genism LDA Topic Model of English Wikipedia biographical articles with list of all 1.8M articles, and some associated Wikidata information

The model has 150 Topics.

This model was developed in the process of isolating a set of visual arts biographical articles, as described in "Clowns in the Visual Artists: Topic Modeling Wikipedia and Wikidata" in the Spring 2022 issue of Art Documentation - https://doi.org/10.1086/719999

Because names, nationalities, and birthdays are so prominent in biographies, the stopwords list removed 170,000 names, surnames, city names, place names, countries, days, months and other time related words (https://github.com/mandiberg/Names-Surnames-and-Countries-for-Stopwords).  We also directly removed each article subject’s given and surname, which were almost always the most frequently occurring words in any given article. Otherwise, the model just produced topics based on nationality, and common names and surnames.

Files:

all_enwiki_bios_from_wikidata.csv
The list of all Wikidata items for humans with an enwiki page (e.g biographical article) was extracted from Wikidata JSON dump; list includes gender, occupation, and nationality. This was joined with the converted plaintext from an English Wikipedia dump. This data was downloaded in March 2021.

Wikipedia Biographies LDA Topic Model human readable summary.csv
A human readable file with the 150 topics ranked by count of articles per topic from the 1.8M corpus. The most popular topics have categorical descriptions of the occupations of each cluster. Some are marked as not an occupation cluster. 

BoW_corpus.mm*
model_lda_full_Sep2_150Tv2*
These six files comprise the topic model. The code to load them is present in the python files. 

dict_full_Aug-28-2021
processed_docs_full_Aug-28-2021.txt
processed_docs_1000_Aug-18-2021.txt
These are the dictionary and processed corpuses required to build and implement the model using this code. The corpus with the first 1000 items is meant to be used for testing, as the full one is quite large and takes a long time to complete. 

topic-model-wikipedia-sept2021.zip
The code and settings used for creating and implementing this model are included in this zip and are also available here: https://github.com/mandiberg/topic-model-wikipedia

All-Wikipedia-Biographies-with-topic1.csv
All-Wikipedia-Biographies-with-topic1and2.csv
These are the list of 1.8M biographies matched to topics. The "topic1" file just includes the first topic, this is a slightly larger list. The "topic1and2" file is slightly smaller because about 2% articles do not match to a second topic.

Analysis-for-Clowns-Visual-Arts.zip
These are the raw data and final data produced for the "Clowns in the Visual Artists." Please see the article for context.

Files

All-Wikipedia-Biographies-with-topic1.csv

Files (5.6 GB)

Name Size
md5:3f0b854078ebfc1cf24653b6627517c7
87.2 MB Preview Download
md5:65998cc02cda830e7dc5bad657488a87
124.2 MB Preview Download
md5:7ea8250a4e53dcbad9a23ecd2994180c
131.7 MB Preview Download
md5:d86735f1bb70b634ade0ccc8ac4b2243
10.1 MB Preview Download
md5:bd1127590f9de936e591bb8397ecf85f
2.0 GB Download
md5:e8c727bef35d0c18b3a9ecfd90e0e63e
8.9 MB Download
md5:73cd83003cfb9842c0226a6b9a8f9762
77.8 MB Download
md5:e1353df6c90abb37d083934a484efa2c
505.7 kB Download
md5:2ceaff874087e79317543a3f6cd6c140
60.0 MB Download
md5:0d3cb90406df5d24ecc6f5c42f2466e3
4.1 MB Download
md5:560e9151e22888397ea10b80f1985a0f
500.4 kB Download
md5:5ef6015931cffcaede96e0e4393ecacd
60.0 MB Download
md5:63db704c7dfab1d857042a88a2c88834
5.5 MB Preview Download
md5:c6b1f25edaa2906a54bfad5fb7513213
2.9 GB Preview Download
md5:81969bbbce367514a66fd657bfe00ef7
72.4 MB Preview Download
md5:37319c248a279ac0ddc4c431ec06f9a8
13.5 kB Preview Download