Published September 1, 2023 | Version 1.0.0

Workflow for detecting biomedical articles with openly available underlying datasets - Datasets and extraction forms

  • 1. Berlin Institute of Health (QUEST Center for Responsible Research)

Description

The open data screening datasets contain both automatically detected (TRUE) Open Data statements by ODDPub, and its manual validation using Numbat extraction tool. Furthermore, extraction forms for both screenings – 2020 and 2021 – are included. The manually processed dataset for the calculation of the inter-rater reliability of manual validation can be also found here.  

(i) Data from articles published in 2020 (file ‘charite_open_data_2020.csv’) have been collected applying a slightly different sequence of questions in the extraction workflow than the articles published in 2021 (file ‘charite_open_data_2021.csv’). Both datasets were cleaned for any personal data or internal comments. Thus, they do not contain the default columns which in the raw export from Numbat contained commentaries regarding different question. Also, in another regard these files do not represent raw outputs of the Numbat extraction tool, but a processed version. This means that articles validated by more than two raters were first reconciled in Numbat, resulting in one final decision (output of extractions after reconciliation). Then from the output of extractions before reconciliation those articles validated by only 1 rater (and thus not part of the inter-rater reliability calculation) were selected, which were afterwards joined with the already reconciled dataset.  

The actual decision about Openness of validated dataset can be analysed in various ways: 

  1. Column ‘open_data_assessment’/’assessment’ shows a binary decision between Open Data TRUE and FALSE. 
  2. If that column indicates ‘NULL’, the dataset was classified into ‘non’-open category, and the result can be found on one of the following ways: 
    • Column ‘reference_to_data’ as ‘n_a’ for excluded articles, e.g. not producing any data.
    • Column ‘data_access’ as ‘restricted’. 
    • Column ‘own_or_reuse_data’ as ‘open_data_reuse’. 

The original extraction form contains an option ‘unsure_open_data’ besides ‘open_data’/’no_open_data’ which was resolved either during reconciliation between multiple raters or by case-related consultation with a second rater in case of doubt, and is not included here. 

(ii) The inter-rater reliability calculation was made on randomly selected 100 articles for 2 raters. The third rater screened 20 articles sample, which is part of 100 sample. The tables provided here include both article-level data, and dataset-level data. 

(iii) The Numbat extarction forms used for the screenings in 2020 and 2021 are included in two formats - JSON and Markdown.

(iv) ‘data_dictionary_open_data.csv’ table documents all variables of each data file containing here. 

Files

2023-08-17-Openness (en) 2021.json

Files (2.5 MB)

Name Size Download all
md5:204f3e9537030356e565f8eb87b5d0ac
61.6 kB Preview Download
md5:68d95cf4ab7c5bd20895995bb320a54f
28.1 kB Preview Download
md5:f1c89353640e91d100664a61d4b4f1f7
85.7 kB Preview Download
md5:f8256115126e31114f8c03da23ea988c
25.1 kB Preview Download
md5:aadba07857689e1348e4ff7e1a2ce7fd
562.4 kB Preview Download
md5:2193bbf68bf6d3b8314c042e1efe5e8b
1.6 MB Preview Download
md5:33de9efc1e7553d747cfc36b66e7fb39
25.1 kB Preview Download
md5:291e309d01d63dfead56e012598929d3
155.1 kB Preview Download
md5:555d30d1c4871356847918a10ac3a5a3
9.1 kB Preview Download
md5:379b963742d2ed6e8b6b2ca91b05d17f
26.5 kB Preview Download
md5:0d28a747dd2c115b04989877b3182c15
2.8 kB Preview Download
md5:b2850f25d9a44dc845c537d099d9477d
6.9 kB Preview Download

Additional details

Related works

Is described by
Preprint: 10.31222/osf.io/z4bkf (DOI)
Is documented by
Workflow: 10.17504/protocols.io.q26g74p39gwz/v1 (DOI)