Dataset Open Access

Dataset for generating TL;DR

Syed, Shahbaz; Voelske, Michael; Potthast, Martin; Stein, Benno


Dublin Core Export

<?xml version='1.0' encoding='utf-8'?>
<oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:creator>Syed, Shahbaz</dc:creator>
  <dc:creator>Voelske, Michael</dc:creator>
  <dc:creator>Potthast, Martin</dc:creator>
  <dc:creator>Stein, Benno</dc:creator>
  <dc:date>2018-02-08</dc:date>
  <dc:description>This is the dataset for the TL;DR challenge containing posts from the Reddit corpus, suitable for abstractive summarization using deep learning. The format is a json file where each line is a JSON object representing a post. The schema of each post is shown below:


	author: string (nullable = true)
	body: string (nullable = true)
	normalizedBody: string (nullable = true)
	content: string (nullable = true)
	content_len: long (nullable = true)
	summary: string (nullable = true)
	summary_len: long (nullable = true)
	id: string (nullable = true)
	subreddit: string (nullable = true)
	subreddit_id: string (nullable = true)
	title: string (nullable = true)


Specifically, the content and summary fields can be directly used as inputs to a deep learning model (e.g. Sequence to Sequence model ). The dataset consists of 3,084,410 posts with an average length of 211 words for content, and 25 words for the summary.

Note : As this is the complete dataset for the challenge, it is up to the participants to split it into training and validation sets accordingly.</dc:description>
  <dc:identifier>https://zenodo.org/record/1168855</dc:identifier>
  <dc:identifier>10.5281/zenodo.1168855</dc:identifier>
  <dc:identifier>oai:zenodo.org:1168855</dc:identifier>
  <dc:language>eng</dc:language>
  <dc:relation>doi:10.5281/zenodo.1043504</dc:relation>
  <dc:relation>doi:10.5281/zenodo.1168854</dc:relation>
  <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
  <dc:rights>http://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
  <dc:subject>tl;dr challenge</dc:subject>
  <dc:subject>abstractive summarization</dc:subject>
  <dc:subject>social media</dc:subject>
  <dc:subject>user-generated content</dc:subject>
  <dc:title>Dataset for generating TL;DR</dc:title>
  <dc:type>info:eu-repo/semantics/other</dc:type>
  <dc:type>dataset</dc:type>
</oai_dc:dc>
733
540
views
downloads
All versions This version
Views 733734
Downloads 540540
Data volume 1.2 TB1.2 TB
Unique views 672673
Unique downloads 440440

Share

Cite as