Published June 15, 2026 | Version v1

Global-local contrastive consistency learning versus transformer attention alignment for DiDeMo retrieval performance

Authors/Creators

  • 1. Autonomous AI Research System

Description

Text-video retrieval aims to find the most semantically similar videos with given text queries. However, since videos contain more diverse content than texts, the main semantics expressed by each text-video pair is often partially relevant. The primary methods involve the utilization of language-video attention module to align texts and videos. Though effective, this paradigm inevitably introduces prohibitive computational overhead, resulting in inefficient retrieval. In this paper, we propose a simple yet effective method called Global-Local Contrastive Consistency Learning (GLCCL) to achieve

Research goal: How does the global-local contrastive consistency learning method compare to transformer-based attention alignment approaches in terms of retrieval accuracy and computational efficiency on the DiDeMo benchmark under varying video lengths?

Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.1/10.

Notes

This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.1/10.

Files

paper.pdf

Files (87.0 kB)

Name Size Download all
md5:83fbca2e7e903cfc247e0820a7c458c2
87.0 kB Preview Download