Published June 8, 2023
| Version v1.2.0
Dataset
Open
Charades-STA Speech Caption Dataset
Description
- Dataset introduction: This dataset is an extension of Charades-STA dataset, where audio is read from text using machine simulation method "microsoft/speecht5_tts"
- Associated Code: https://github.com/xian-sh/UniSDNet
- Associated Paper: https://arxiv.org/abs/2403.14174
- Disclaimer: This dataset is for academic research only, non-commercial use, if you use this dataset please cite the associated paper
- Cite:
@article{hu2024unified,
title={Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding},
author={Jingjing Hu and Dan Guo and Kun Li and Zhan Si and Xun Yang and Xiaojun Chang and Meng Wang},
year={2024},
Journal={CoRR},
volume={abs/2403.14174},
}
Files
raw_audios.zip
Additional details
References
- Gao, Jiyang, et al. "Tall: Temporal activity localization via language query." Proceedings of the IEEE international conference on computer vision. 2017
- Ao, Junyi, et al. "Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing." arXiv preprint arXiv:2110.07205 (2021).
- https://huggingface.co/datasets/Matthijs/cmu-arctic-xvectors