Published May 8, 2026 | Version 1.0.0

PRESTO Corpus: Structured Rumour Claims from 19th-Century British Newspapers (Heritage Made Digital)

  • 1. C2DH
  • 2. Luxembourg

Description

The PRESTO Corpus is a silver-standard dataset of 7,460 automatically extracted rumour claims from 19th-century British newspapers, derived from the Heritage Made Digital (HMD) collection (biglam/hmd_newspapers, British Library). It is produced by PRESTO (Pattern-based Rumour Extraction with Semantic Tracking), a two-phase methodology for discovering and tracking rumoured events in historical OCR'd text.

Each entry contains the original matched sentence, the extracted rumour proposition, named entity annotations (PERSON, GPE, ORG), the dependency pattern type used for extraction (of, that, it-is-rumoured, to-the-effect, standalone), publication metadata (title, date, location), and OCR quality scores.

Extraction precision was estimated at 67.0% (95% CI: 60.2%–73.1%) via manual annotation of a random 200-entry sample (PRESTO_sample_200_annotated.csv, included). The corpus is a silver-standard resource suited for exploratory analysis and retrieval; users requiring high-precision structured claims for quantitative analysis should verify a sample before use.

This corpus accompanies the paper: Zhang, Wanshu. "PRESTO: A Two-Phase Methodology for Discovering and Tracking Rumoured Events in Historical Newspapers." Digital Humanities 2026 (DH2026), Daejeon.

Files

PRESTO_corpus_7460.csv

Files (3.2 MB)

Name Size Download all
md5:ff72472ce6659a7b31fceb632e81227d
3.1 MB Preview Download
md5:037da9af7c8be77b9d8eda947967a021
75.2 kB Preview Download