Published November 27, 2024 | Version v1

Beyond the surface: Revealing researchers' behavoiur in public repositories (First version)

  • 1. Centre for Engineering Biology and School of Biological Science, University of Edinburgh
  • 2. School of Informatics, University of Edinburgh
  • 3. University of Dundee
  • 4. Divisions of Molecular Cell and Developmental Biology, and Computational Biology, University of Dundee
  • 5. Centre of Engineering Biology and School of Biological Science, University of Edinburgh

Description

This poster was presented in the Edinburgh Open Research Conference 2024 in May 29th at the University of Edinburgh.

Ensuring the availability and accessibility of data has become necessary in the pursuit of advancing knowledge. This information is vital for adhering to the FAIR (Findable, Accessible, Interoperable, and Reusable) principles in scientific data management. Accurate documentation of the studies, commonly known as metadata, is indispensable to achieve this fundamental goal. Regrettably, public records often fall short, providing inadequate, repetitive, and incomplete descriptions, hindering the seamless flow of knowledge. Our project addresses the metadata challenge by analysing current repositories and designing prompts for better metadata in future.

The rapid advancements in AI create an ideal environment for integrating these methods into everyday scientific research practices to publish datasets. Leading repositories secure the metadata quality by employing skilled data curators; nevertheless, this approach demands substantial expenses for database management.  Therefore, we aim to develop a user-friendly and cost-effective tool to enrich metadata, specifically targeting named entities within unstructured textual data using AI. 

Here, we will present our preliminary results from the original metadata assessment in our target repositories. This analysis, employing standard text mining metrics, serves to identify critical features, characteristics, similarities, and differences within the dataset.  In our initial steps, we analysed records sourced from The BioDare2 (https://biodare2.ed.ac.uk/), a domain-specific repository for biological time series data that stores over 15,000 datasets, and DataShare (https://datashare.ed.ac.uk/) a domain-agnostic, research data repository at the University of Edinburgh, with more than 6.500 data entries.

The BioDare2 database has no curation process apart from minimal length requirements. In contrast, DataShare is considered semi-curated because deposits are filtered for relevance to the repository's scope, valid layout and format, and the exclusion of spam. The repository's curators also offer suggestions for enhancing metadata quality.

The difference between the databases manifests itself in our analysis of metadata metrics and reinforces the necessity of efficient and automatic curation. They serve as the foundational groundwork for scoping and the development of our tools for metadata enhancement with these repositories in future. Tools that improve metadata quality will help to unlock the potential of research data, drive scientific advancement, and foster international collaboration among researchers, overcoming geographical constraints and facilitating cooperation among scientists who may have never interacted in person.

Files

Edinburgh_Open_Research_Conference_PosterMJRC_20240510_submitted_version.pdf

Additional details

Funding

UK Research and Innovation
EASTBIO DTP BB/J01446X/1

Software

Repository URL
https://github.com/mjrodriguezc/metadata_project
Programming language
Python
Development Status
Active