Published 2024 | Version v2

Trends in gender homophily in scientific publications (data)

  • 1. Universidad Carlos III de Madrid

Description

This dataset contains records of research articles extracted from the Web of Science (WoS) from 1980 to 2019---in total, 15,642 journals, 28,241,100 articles and 111,980,858 authorships across 153 research areas.

The main dataset (author_address_article_gend_v3.parquet), in Parquet format, contains all the authorships, where an authorship is defined as the tuple article-author. There are 12 variables per authorship (row):

  • ut: unique article identifier.
  • daisng_id: unique author identifier.
  • author_no: author number, as listed in the article.
  • country: author country (two-letter ISO code).
  • date: publication date.
  • gender: gender of the author ("male" or "female"), as provided by the Genderize.io API.
  • probability: probability of the gender attribute, as provided by the Genderize.io API.
  • count: number of entries for the author first name, as provided by the Genderize.io API.
  • jsc: journal subject category.
  • field: field of research.
  • research_area: area of research.
  • n_aut: number of authors in this publication.
  • journal: journal name.
  • alphabetical: whether the author list for this article is in alphabetical order.

With the previous dataset, a resampler was applied to generate null homophily values for each year. There are 4 datasets in R Data Serialization (RDS) format:

  • null_field.rds: null homophily values per country, year and field of research.
  • null_field_comp.rds: null homophily values per year and field of research (only for complete authorships).
  • null_research.rds: null homophily values per year and area of research.
  • null_research_comp.rds: null homophily values per year and area of research (only for complete authorships).

All these datasets have the same structure:

  • country: country (two-letter ISO code).
  • year: year.
  • variable: either field or research area name.
  • m: average homophily.
  • s: homophily std. error.

Finally, some supplementary files used in the descriptive analysis and methods:

  • File null_research_l2019.rds is an example of the output from the resampling algorithm for year 2019.
  • File wos_category_to_field.csv is a mapping from WoS categories to more general fields.
  • File jcr_if_2020.csv contains the percentiles of the journal impact factor for the JCR 2020.

Files

jcr_if_2020.csv

Files (1.9 GB)

Name Size
md5:fec1cdccc0a66af984df140bbae0daa7
1.9 GB Download
md5:40cf486c8848f07e5e80c2ceb04e0507
453.9 kB Preview Download
md5:64a1ee26646c673e39bfc797b2ec4d44
244.5 kB Download
md5:a42d962d972e9cd0dc808f629bef4b75
208.0 kB Download
md5:6a8b3c2a1c381a4156574a0af489a912
2.5 MB Download
md5:00dfc5ee9f5f0c4b3931253f26c3500e
1.9 MB Download
md5:afc303887a613e026eb6fdd943ab06ec
78.3 MB Download
md5:cba51fbe5ae78a397fbf248189ffdeb6
6.5 kB Preview Download