Investigating the performance of GROBID and OUTCITE
Authors/Creators
Description
In a prior study (Cioffi & Peroni, 2022), we analysed the available reference extraction tools to understand their performances off-the-shelf – i.e. by using them as they have been configured, without prior training. We evaluated them against a corpus of 56 PDF articles (our gold standard) published in 27 subject areas (Computer Science, Arts and Humanities, Mathematics, etc.). From that analysis, we have identified the two most promising tools for bibliographic reference extraction and parsing, i.e. Anystyle and Grobid, which are CRF based. We have extended such study by training Grobid against an extended gold standard with various training configurations to understand how much the performances improve. As a result, we have also revised the code used for testing and comparing the reference extraction software to make it available also for others to be reused for similar analysis. Othe tests have been performed on OUTCITE and new conversions and evaluations softwares have been created for the purpose. The final aim of this work is to develop a reference extraction service which enables a user to provide a PDF of a scholarly article in input and to have, in return, citation data and bibliographic metadata from all the references that are cited by the given article in a format that enables their ingestion in OpenCitations (Peroni & Shotton, 2020). The demo of the service is available online.
This publication is part of my Thesis research for the Digital Humanities and Digital Knowledge Master's Course at University of Bologna
Some publications related to the research:
The gold standard can be found here, Pagnotta, O. (2024). CEX Project - Dataset and Gold Standard [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10535653.
The code can be found here Pagnotta, O. (2024). olgagolgan/CEX-Project: CEX Project Code (software). Zenodo. https://doi.org/10.5281/zenodo.10638757.
The output dataset of GROBID, Anystyle and OUTCITE can be found here Pagnotta, O. (2024). CEX Project - Output Dataset (Anystyle, GROBID, OUTCITE) (Version 1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10524898.
The training dataset of GROBID can be found here Pagnotta, O. (2024). CEX Project - GROBID annotation aligned Gold Standard (Version 1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10529646.
The trained GROBID citation models can be found here Pagnotta, O. (2024). CEX Project - trained GROBID citation models. Zenodo. https://doi.org/10.5281/zenodo.10529709.
The final service can be found here Pagnotta, O. and Paolini, L. (2024). opencitations/cec: alpha version (service). Zenodo. https://doi.org/10.5281/zenodo.10635630.
Files
WOOC-posterPagnotta.pdf
Files
(499.1 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:ea661c71e7e11822ce2a63725ed8eac8
|
499.1 kB | Preview Download |
Additional details
Identifiers
Dates
- Other
-
2023-10-24