Published September 2, 2023 | Version v1

Indonesia Text Simplification Dataset

  • 1. Institut Teknologi Bandung

Description

This is a dataset for Indonesian text simplification. Dataset input is provided by Liputan6 dataset (Koto et. al., 2020) and referenced by the id attribute as a document number, and the sentence_index attribute as the sentence position inside the corresponding document. The dataset consists of 95 sentences with four cases: relative clauses (30), appositions (14), conjoined clauses (in complex sentences; 30), and anaphora (26). Notice that these cases are not mutually exclusive. There are 33 sentences that don’t handle these cases.

Files

Indonesia Text Simplification Dataset - dataset.csv

Files (13.8 kB)