Can yesterday's data meet the needs of today's researchers? Crafting additional methods to improve fitness for use
Description
For 20 years, Persée has been the French operator for scientific literature heritage digitization and online dissemination. Persée currently has a 34 people team, involved in publishing, technical and scientific projects, and working over a fully-fledged paper-to-digital document processing chain. So far Persée has produced :
- a web portal (standard digital library) ;
- a triplestore (linked open metadata outlet) ;
- 5 ‘perséides’ (project-specific online corpora).
Alongside the perseides which are the result of a co-design process, Persée has been providing legacy (meta)data to digital humanists for whom data science and computational methods, not to mention AI/ML, have become standard practice. Since ‘Collections as data designed for everyone serve no one’ (The Santa Barbara Statement on Collections as Data, 2017), Persée has implemented various methods to ensure the fitness for use, including quality checks, user feedback analysis, co-design. Yet, the same questions eventually arise again: 'Can yesterday’s data fulfil today’s researchers’ needs? What are these needs? How can we ascertain quality?’.
Taking our part in the collaboration of library science, computer science and humanities working together with digital cultural heritage (Oberbichler et al., 2022), we saw fit to extend customer study. Drawing up explicitly the Persée FAIR Implementation Profile (Schultes et al., 2020), mapping visually user types and data dissemination channels, then applying graph analysis, makes it possible to highlight critical nodes, popular use cases and intensive use areas. Moreover, modelling a typical data reuse process makes it possible to (re)locate occurrences of data friction (Edwards et al., 2011). Besides, to our knowledge, few are the summary papers about OCR quality for so called ‘downstream tasks’ (van Strien et al., 2020) in an otherwise abundant literature since Holley's OCR accuracy categories (Holley, 2009). Setting up a dedicated process of literature review, we found that :
- some tasks, like part-of-speech tagging and NER, seem more affected by OCR quality than others ;
- the performance of a task might relate to specific ranges of OCR quality metrics, although the matter, still open to question, largely depends on the chosen task, language, metrics, and deserves to be explored further.
The introduction of graph study, process modelling and literature review has proved useful to qualify usage more accurately and to get a better scope of the corresponding quality dimensions and metrics. We argue that, for lack of a one-stop standard, such endeavours could help to :
- formulate project-specific data reuse strategies : using data ‘as-is’, counterbalancing defects with smarter processing, applying post-production corrections ;
- tackle the ubiquitous topic of data quality more effectively : finetuning supply and services, avoiding the ‘garbage in, garbage out’ effect in research projects, closing the data lifecycle by re-ingesting the research output.
Notes
Files
20230509_DARIAH2023_PerseePoster_abstract.pdf
Files
(534.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:6626885a726121d2c029ff584b70aa3c
|
242.5 kB | Preview Download |
|
md5:c690462fd7745854b116a57d024f3eae
|
291.8 kB | Preview Download |