Published September 7, 2022 | Version v1

Coverage and biases in freely accessible academic bibliographic platforms: an exploratory analysis

  • 1. Universidad de Córdoba
  • 2. Cybermetrics Lab
  • 3. Universidad de Granada
  • 4. Universitat Politècnica de València

Description

There is a significant overlap between sources which suggests that Crossref could be a usual source to feed these platforms because it freely provides bibliographic information of the most important academic publishers. In this sense, the largest coverages are in Scilit (99.77%) and Lens.org (99.13%), two platforms that use Crossref as primary source. The high overlap among sources has also showed that there are 27,144 documents in our sample that were not found in at least one of the databases under study. Within this group, 77% were only missing from one or two of the six databases. This results shows that inclusion criteria and gathering methods are significantly different between databases (Gusenbauer, 2019), but also that searching in three of these databases it would be possible to find the entire collection (99.47%).

The analysis by document type has revealed that all the platforms have limited coverage of books and book chapters, while they perform better with journal articles. This evidences that the primary aim of these platforms is to exhaustively cover research articles, while other important document types such as book chapters or datasets are underrepresented. A possible cause is that certain formats might not be considered relevant for some platforms (i.e. datasets, others) or they would have problems to indexing and describing these typologies. In this sense, book chapters could not be separated from the book or that ISBN would be used as identifier instead DOI. Dimensions, Semantic Scholar and Microsoft Academic are the sources that have the most problems indexing these type of documents. This problem with book chapters could have special consequences in Social Sciences and Humanities where a great of their outputs are released in this format.

According to the language, the analysis has also detected possible biases in favour of papers written in English in all the platforms, being more important in Lens, Google Scholar and Semantic Scholar. Although, this is just an exploratory approach, these results suggest a possible language biases that could be due to different reasons. The most probable could be that some non-English journals are released by small publishers with little distribution, impairing their findability and influencing their impact.

The results provided by this exploratory work evidence that there is no single master source to collect evidence of scientific impact, due to the different coverage criteria and indexing technical solutions adopted by each source. These findings lead us to point out the need to design multiplatform indicators, in which bibliometric indicators are created based on data from different sources. The results provided are of interest both to experts in science studies (unravelling the bibliometric coverage, benefits and limitations of the sources) and to the agencies in charge of evaluating research results.

Files

102.pdf

Files (474.2 kB)

Name Size Download all
md5:cf3c25e9032ed14d2c9c7ed4a04d254b
474.2 kB Preview Download