Impact of Contrastive Constraints in Cross-Modal Attention on Zero-Shot Retrieval Under Domain Shift
Description
Cross-modal attention mechanisms have been widely applied to the image-text matching task and have achieved remarkable improvements thanks to its capability of learning fine-grained relevance across different modalities. However, the cross-modal attention models of existing methods could be sub-optimal and inaccurate because there is no direct supervision provided during the training process. In this work, we propose two novel training strategies, namely Contrastive Content Re-sourcing (CCR) and Contrastive Content Swapping (CCS) constraints, to address such limitations. These constraints supe
Research goal: How does integrating contrastive constraints into cross-modal attention affect zero-shot retrieval accuracy on domain-shifted benchmarks like Flickr30k and MS-COCO?
Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 9.2/10.
Notes
Files
paper.pdf
Files
(76.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:e30efaa46ce38e04c81247111781f6d8
|
76.9 kB | Preview Download |