Published April 4, 2025 | Version 1.0.1

Intermediate JSONL for Softcite Extractions from the Open Access Literature

  • 1. ROR icon The University of Texas at Austin
  • 2. ROR icon University of California, Berkeley

Description

Check documentation at https://github.com/softcite/softcite-extractions-oa

This archive is an intermediate resource used in the pipeline that created the parquet files available at https://doi.org/10.5281/zenodo.15066399

While this archive is ~26 GiB the parquet files are much easier to handle at about 5 GiB and are better documented.

As I create this seems to me that the only reason that one would be working with these files is to understand conversion issues, to work with the extractions from non-PDF files, or to access additional metadata extracted by GROBID (although the only additional metadata not in the parquet files should be accessible from the article DOI via cross-ref).  Note, though, that those extractions are sparse and some files incomplete. See additional details in the README inside the archive or at https://github.com/softcite/softcite-extractions-oa/blob/main/EXTRACTING_TABLES.md

This work used JetStream 2 at Indiana through allocation CIS220172 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296. The authors acknowledge the Texas Advanced Computing Center (TACC) at The University of Texas at Austin for providing computational resources that have contributed to the creation and processing of this research dataset. URL: http://www.tacc.utexas.edu

Version Notes

v 1.0.1 Removed files restrictions

v 1 Initial upload and beta testing

Files

json_dataset.zip

Files (26.6 GB)

Name Size
md5:ff2bb3f98fdccdab99074fa6295391c0
26.6 GB Preview Download
md5:5bcbe947c16e78eba3a02b546128ce54
6.9 kB Preview Download

Additional details

Related works

Is original form of
Dataset: 10.5281/zenodo.15066399 (DOI)

Software