Published March 29, 2015 | Version v1

Fuse

Description

The contributors have provided two related datasets, which together constitute the FUSE spreadsheet corpus2.

  + A Web Analysis dataset of 2,127,284 URLs that return spreadsheet content, along with the full HTTP web server response, formatted as JSON records. This dataset was obtained by filtering through 26.83 billion HTTP responses within the Common Crawl archive.
  + A Binary Analysis dataset of 249,376 spreadsheets, extracted from the 1.9 PB of raw data within the Common Crawl archive. For each spreadsheet, the authors provide JSON metadata containing their analysis, which includes NLP token extraction and spreadsheet metrics.

Files

fuse.zip

Files (9.4 GB)

Name Size
md5:13e955c44f0b77d1c36088c0bbb3366d
9.4 GB Preview Download