Published June 8, 2023 | Version v1.2.0

Charades-STA Speech Caption Dataset

Authors/Creators

  • 1. HFUT(Hefei University of Technology)

Contributors

Other:

  • 1. Anhui University

Description

@article{hu2024unified,
  title={Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding},
  author={Jingjing Hu and Dan Guo and Kun Li and Zhan Si and Xun Yang and Xiaojun Chang and Meng Wang},
  year={2024},
  Journal={CoRR},
  volume={abs/2403.14174},
}

Files

raw_audios.zip

Files (991.4 MB)

Name Size
md5:fefd7388a706ce2c2ed00e325d1081e0
987.2 MB Preview Download
md5:760a82a5a8699d5ab96c5ad89e5cbe13
910.2 kB Preview Download
md5:aeab1075390f747ce4b418586581ff35
3.3 MB Preview Download

Additional details

References

  • Gao, Jiyang, et al. "Tall: Temporal activity localization via language query." Proceedings of the IEEE international conference on computer vision. 2017
  • Ao, Junyi, et al. "Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing." arXiv preprint arXiv:2110.07205 (2021).
  • https://huggingface.co/datasets/Matthijs/cmu-arctic-xvectors