There is a newer version of the record available.

Published May 16, 2026 | Version v1

TabulusBench: A Benchmark Dataset for Scientific Survey Table Extraction and Reference-Aware Table Processing

Description

TabulusBench is the accompanying benchmark dataset for the Tabulus pipeline, an OCR-driven framework for extracting and semantically processing comparison tables from scientific survey and review papers. The dataset focuses on related-work and comparison tables that summarize methods, datasets, metrics, and experimental findings across scientific literature.

The collection spans five scientific domains:

  • Biomedicine and Health

  • Agriculture, Food, and Environmental Systems

  • Computer Science, AI, and Data Science

  • Energy, Materials, and Chemical Sciences

  • Engineering, Robotics, and Built Infrastructure

To construct the benchmark, four OCR systems — Chandra (https://github.com/datalab-to/chandra), DeepSeek-OCR-2 (https://github.com/deepseek-ai/DeepSeek-OCR-2), PaddleOCR (https://github.com/PADDLEPADDLE/PADDLEOCR), and Kreuzberg (https://github.com/kreuzberg-dev/kreuzberg) — were applied to scientific PDF documents to extract tabular content. The extracted tables were subsequently manually verified and corrected against the original PDF tables to create high-quality gold-standard annotations.

The dataset includes:

  • cropped table images,

  • OCR-generated table reconstructions,

  • manually corrected gold-standard CSV tables,

  • bibliography extraction outputs,

  • DOI matching results,

  • runtime statistics,

  • RMS similarity metrics,

  • precision, recall, and F1-score evaluations.

The resource preserves both intermediate and final pipeline outputs to support reproducibility, benchmarking, and future comparison experiments for scientific document understanding workflows.

Dataset Structure

The dataset is organized hierarchically by:

  1. research domain,

  2. topic,

  3. paper.

Example structure:

tabulusbench/
├── Agriculture_Food_And_Environmental_Systems/
│   └── agroecology/
│       └── P51/
│           ├── Ref/
│           ├── Ref_Tables/
│           └── P51.pdf
│
├── Biomedicine_And_Health/
├── Computer_Science_AI_And_Data_Science/
├── Energy_Materials_And_Chemical_Sciences/
└── Engineering_Robotics_And_Built_Infrastructure/

Each paper directory contains:

  • Ref/ — bibliography extraction outputs, DOI matching files, and evaluation metrics.

  • Ref_Tables/ — cropped table images, OCR predictions, gold-standard tables, and table extraction benchmark results.

  • PXX.pdf — the original survey paper PDF (when redistribution is permitted).

The Ref_Tables/ directory contains:

  • OCR outputs generated using Chandra, DeepSeek-OCR-2, PaddleOCR, and Kreuzberg,

  • manually corrected gold-standard CSV tables,

  • MinerU table crops,

  • benchmark and evaluation outputs.

The Ref/ directory contains:

  • raw OCR bibliography text,

  • GROBID extraction outputs,

  • regex-based reference extraction results,

  • DOI matching files,

  • precision, recall, and F1-score evaluations.

TabulusBench supports research in:

  • scientific table extraction,

  • OCR robustness evaluation,

  • document understanding,

  • bibliography extraction,

  • reference matching,

  • scholarly knowledge graphs,

  • FAIR scientific information systems.

Within the released dataset dump, we do not redistribute the original survey paper PDFs from which the tables were extracted. However, the file papers_list.xlsx contains the list of scientific papers considered during dataset construction.

Files

table_dataset_characteristics.csv

Files (84.0 MB)

Name Size Download all
md5:bae5636d36e734b1771960e7854e09bc
31.4 kB Download
md5:93ed123018bdd3767cbc31ec5c4d9263
57.0 kB Preview Download
md5:b6b01fa6d137a49a9f3b5e20c4b7c6a3
83.9 MB Preview Download

Additional details

Software

Development Status
Active