Indonesia Text Simplification Dataset
Authors/Creators
- 1. Institut Teknologi Bandung
Description
This is a dataset for Indonesian text simplification. Dataset input is provided by Liputan6 dataset (Koto et. al., 2020) and referenced by the id attribute as a document number, and the sentence_index attribute as the sentence position inside the corresponding document. The dataset consists of 95 sentences with four cases: relative clauses (30), appositions (14), conjoined clauses (in complex sentences; 30), and anaphora (26). Notice that these cases are not mutually exclusive. There are 33 sentences that don’t handle these cases.
Files
Indonesia Text Simplification Dataset - dataset.csv
Files
(13.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:d97ce2a5a17096531f180f5b62b7d84d
|
13.8 kB | Preview Download |