Trends in gender homophily in scientific publications (data)
Authors/Creators
- 1. Universidad Carlos III de Madrid
Description
This dataset contains records of research articles extracted from the Web of Science (WoS) from 1980 to 2019---in total, 15,642 journals, 28,241,100 articles and 111,980,858 authorships across 153 research areas.
The main dataset (author_address_article_gend_v3.parquet), in Parquet format, contains all the authorships, where an authorship is defined as the tuple article-author. There are 12 variables per authorship (row):
- ut: unique article identifier.
- daisng_id: unique author identifier.
- author_no: author number, as listed in the article.
- country: author country (two-letter ISO code).
- date: publication date.
- gender: gender of the author ("male" or "female"), as provided by the Genderize.io API.
- probability: probability of the gender attribute, as provided by the Genderize.io API.
- count: number of entries for the author first name, as provided by the Genderize.io API.
- jsc: journal subject category.
- field: field of research.
- research_area: area of research.
- n_aut: number of authors in this publication.
- journal: journal name.
- alphabetical: whether the author list for this article is in alphabetical order.
With the previous dataset, a resampler was applied to generate null homophily values for each year. There are 4 datasets in R Data Serialization (RDS) format:
- null_field.rds: null homophily values per country, year and field of research.
- null_field_comp.rds: null homophily values per year and field of research (only for complete authorships).
- null_research.rds: null homophily values per year and area of research.
- null_research_comp.rds: null homophily values per year and area of research (only for complete authorships).
All these datasets have the same structure:
- country: country (two-letter ISO code).
- year: year.
- variable: either field or research area name.
- m: average homophily.
- s: homophily std. error.
Finally, some supplementary files used in the descriptive analysis and methods:
- File null_research_l2019.rds is an example of the output from the resampling algorithm for year 2019.
- File wos_category_to_field.csv is a mapping from WoS categories to more general fields.
- File jcr_if_2020.csv contains the percentiles of the journal impact factor for the JCR 2020.
Files
jcr_if_2020.csv
Files
(1.9 GB)
| Name | Size | |
|---|---|---|
|
md5:fec1cdccc0a66af984df140bbae0daa7
|
1.9 GB | Download |
|
md5:40cf486c8848f07e5e80c2ceb04e0507
|
453.9 kB | Preview Download |
|
md5:64a1ee26646c673e39bfc797b2ec4d44
|
244.5 kB | Download |
|
md5:a42d962d972e9cd0dc808f629bef4b75
|
208.0 kB | Download |
|
md5:6a8b3c2a1c381a4156574a0af489a912
|
2.5 MB | Download |
|
md5:00dfc5ee9f5f0c4b3931253f26c3500e
|
1.9 MB | Download |
|
md5:afc303887a613e026eb6fdd943ab06ec
|
78.3 MB | Download |
|
md5:cba51fbe5ae78a397fbf248189ffdeb6
|
6.5 kB | Preview Download |