Published April 12, 2022 | Version 1.0

LARD: Large-scale Artificial Disfluency Generation

Description

This dataset contains 95,992 examples of utterances with 71,994 artificial inserted disfluencies using the LARD method. We use the Schema-Guided Dialogue (SGD) dataset as a base to construct the synthetic disfluencies. The LARD dataset contains three different types of disfluencies: repetitions, replacements, and restarts. 

Files

README.md

Files (31.2 MB)

Name Size Download all
md5:f98ab14cb107ced9f3f41e98e9e527ff
1.7 kB Preview Download
md5:ccba0fd9f1e3fb973961527be6c58762
6.3 MB Preview Download
md5:ada1e11d565c04779940100cb9323ba8
18.8 MB Preview Download
md5:af36afcd23712d3ebc01064a5e65ede0
6.2 MB Preview Download

Additional details