There is a newer version of the record available.

Published June 8, 2023 | Version v1.1.0
Dataset Restricted

Charades-STA Speech Caption Dataset

Authors/Creators

  • 1. HFUT(Hefei University of Technology)

Contributors

Other:

  • 1. Anhui University

Description

This dataset is an extension of Charades-STA dataset, where audio is read from text using machine simulation method "microsoft/speecht5_tts"

Files

Restricted

The record is publicly accessible, but files are restricted. Log in to check if you have access.

Additional details

References

  • Gao, Jiyang, et al. "Tall: Temporal activity localization via language query." Proceedings of the IEEE international conference on computer vision. 2017
  • Ao, Junyi, et al. "Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing." arXiv preprint arXiv:2110.07205 (2021).