Published April 12, 2022
| Version 1.0
Dataset
Open
LARD: Large-scale Artificial Disfluency Generation
Description
This dataset contains 95,992 examples of utterances with 71,994 artificial inserted disfluencies using the LARD method. We use the Schema-Guided Dialogue (SGD) dataset as a base to construct the synthetic disfluencies. The LARD dataset contains three different types of disfluencies: repetitions, replacements, and restarts.
Files
README.md
Additional details
Related works
- Is compiled by
- https://github.com/tatianapassali/artificial-disfluency-generation (URL)
- arxiv.org/pdf/2201.05041.pdf (Handle)