Monitoring data sharing and FAIR practices with LLMs: a synthesis of expert views
-
1.
Digital Science (United Kingdom)
- 2. OPIX
-
3.
OpenAIRE Non-Profit Civil Partnership
-
4.
Athena Research and Innovation Center In Information Communication & Knowledge Technologies
-
5.
Ministère de l'Enseignement Supérieur, de la Recherche et de l'Espace
-
6.
American Association For The Advancement of Science
-
7.
The Open University
-
8.
University of Maribor
-
9.
Fundação para a Ciência e Tecnologia
-
10.
University of Oxford
- 11. Stanford University
- 12. DataSeer
Description
This paper examines the potential of large language models (LLMs) to improve the monitoring of data sharing practices, Data Availability Statements (DAS), and adherence to the FAIR principles (Findable, Accessible, Interoperable, Reusable). It draws on expert consultation and independent analysis to assess both the opportunities and limitations of emerging AI-based approaches.
The central finding is that LLMs are likely to deliver meaningful improvements in monitoring performance, particularly through better interpretation of context, greater recognition of non-standard reporting practices, and enhanced ability to integrate information across publications, repositories, metadata and persistent identifier systems. These capabilities offer the potential for substantial gains in precision, recall and coverage compared with traditional rule-based or keyword-based methods.
However, the study concludes that LLMs are not a standalone solution. Monitoring should be understood as a sociotechnical system that depends on high-quality scholarly infrastructure, shared standards, metadata, repositories and reporting practices. While LLMs can improve interpretation and scale, they cannot – yet – independently verify whether data are genuinely available, accessible or reusable, nor can they resolve ambiguity in the definitions of data sharing and FAIR compliance.
The report identifies four enduring constraints on monitoring: verification gaps, inference gaps, definitional ambiguity and infrastructure limitations. It finds broad agreement that the most effective future systems will combine LLMs with structured metadata, repository checks, persistent identifiers, knowledge graphs and human oversight. Rigorous validation, transparency and reproducibility will be essential to establish trust in the resulting indicators.
The key policy implication is that investment in monitoring should not focus solely on analytical tools. Equal attention must be given to improving research workflows, metadata quality, reporting standards and infrastructure integration. Ultimately, the quality of monitoring reflects the quality of the underlying research system. LLMs can strengthen monitoring, but they cannot substitute for the foundations of a transparent, well-documented and reusable research ecosystem.
Files
UKRN_Working_Paper_14 - Monitoring data sharing with LLMs.pdf
Files
(315.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:1428daef287c8b4f94cfc03fe9cecb38
|
315.2 kB | Preview Download |