Intermediate JSONL for Softcite Extractions from the Open Access Literature
Authors/Creators
Description
Check documentation at https://github.com/softcite/softcite-extractions-oa
This archive is an intermediate resource used in the pipeline that created the parquet files available at https://doi.org/10.5281/zenodo.15066399
While this archive is ~26 GiB the parquet files are much easier to handle at about 5 GiB and are better documented.
As I create this seems to me that the only reason that one would be working with these files is to understand conversion issues, to work with the extractions from non-PDF files, or to access additional metadata extracted by GROBID (although the only additional metadata not in the parquet files should be accessible from the article DOI via cross-ref). Note, though, that those extractions are sparse and some files incomplete. See additional details in the README inside the archive or at https://github.com/softcite/softcite-extractions-oa/blob/main/EXTRACTING_TABLES.md
This work used JetStream 2 at Indiana through allocation CIS220172 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296. The authors acknowledge the Texas Advanced Computing Center (TACC) at The University of Texas at Austin for providing computational resources that have contributed to the creation and processing of this research dataset. URL: http://www.tacc.utexas.edu
Version Notes
v 1.0.1 Removed files restrictions
v 1 Initial upload and beta testing
Files
json_dataset.zip
Additional details
Related works
- Is original form of
- Dataset: 10.5281/zenodo.15066399 (DOI)
Software
- Repository URL
- https://github.com/softcite/softcite-extractions-oa
- Development Status
- Active