Global-local contrastive consistency learning versus transformer attention alignment for DiDeMo retrieval performance
Description
Text-video retrieval aims to find the most semantically similar videos with given text queries. However, since videos contain more diverse content than texts, the main semantics expressed by each text-video pair is often partially relevant. The primary methods involve the utilization of language-video attention module to align texts and videos. Though effective, this paradigm inevitably introduces prohibitive computational overhead, resulting in inefficient retrieval. In this paper, we propose a simple yet effective method called Global-Local Contrastive Consistency Learning (GLCCL) to achieve
Research goal: How does the global-local contrastive consistency learning method compare to transformer-based attention alignment approaches in terms of retrieval accuracy and computational efficiency on the DiDeMo benchmark under varying video lengths?
Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.1/10.
Notes
Files
paper.pdf
Files
(87.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:83fbca2e7e903cfc247e0820a7c458c2
|
87.0 kB | Preview Download |