PRESTO Corpus: Structured Rumour Claims from 19th-Century British Newspapers (Heritage Made Digital)
Description
The PRESTO Corpus is a silver-standard dataset of 7,460 automatically extracted rumour claims from 19th-century British newspapers, derived from the Heritage Made Digital (HMD) collection (biglam/hmd_newspapers, British Library). It is produced by PRESTO (Pattern-based Rumour Extraction with Semantic Tracking), a two-phase methodology for discovering and tracking rumoured events in historical OCR'd text.
Each entry contains the original matched sentence, the extracted rumour proposition, named entity annotations (PERSON, GPE, ORG), the dependency pattern type used for extraction (of, that, it-is-rumoured, to-the-effect, standalone), publication metadata (title, date, location), and OCR quality scores.
Extraction precision was estimated at 67.0% (95% CI: 60.2%–73.1%) via manual annotation of a random 200-entry sample (PRESTO_sample_200_annotated.csv, included). The corpus is a silver-standard resource suited for exploratory analysis and retrieval; users requiring high-precision structured claims for quantitative analysis should verify a sample before use.
This corpus accompanies the paper: Zhang, Wanshu. "PRESTO: A Two-Phase Methodology for Discovering and Tracking Rumoured Events in Historical Newspapers." Digital Humanities 2026 (DH2026), Daejeon.
Files
PRESTO_corpus_7460.csv
Files
(3.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:ff72472ce6659a7b31fceb632e81227d
|
3.1 MB | Preview Download |
|
md5:037da9af7c8be77b9d8eda947967a021
|
75.2 kB | Preview Download |