Published June 23, 2021 | Version Second

Missing Citations in COCI: Publishers Analytics Result

Description

This dataset contains a JSON file containing the results retrieved through the software developed by the authors. We opted for JSON file format to store the obtained data, since this format allows the storage of heterogeneous information in a complex and structured way. The four main structures stored in the present file are: 

1) "publishers", a list of dictionaries representing each publisher encountered;

2) "citations", a dictionary containing two lists, the one storing the validated citational data and the other storing the still invalid citational data. Each processed citation is represented as a dictionary. 

3) "total_num_of_valid_citations", whose value is the number of citational data that could be validated throughout the process implemented by our software.

4) "external_data_for_unrecognized_prefixes": a dictionary of dictionaries representing the publishers we didn't find on Crossref, but that were identified through other online services. 

We used as input material open data from the dataset “Citations to invalid DOI-identified entities obtained from processing DOI-to-DOI citations to add in COCI”.  

Files

output.json

Files (274.4 MB)

Name Size Download all
md5:69941622da98e2ebe5ad3df20560dc2a
274.4 MB Preview Download