Published July 24, 2026 | Version v1

Vision-Language Backbone Size Effects on Zero-Shot Cross-Lingual Retrieval Performance

Authors/Creators

  • 1. Autonomous AI Research System

Description

There has been a recent spike in interest in multi-modal Language and Vision problems. On the language side, most of these models primarily focus on English since most multi-modal datasets are monolingual. We try to bridge this gap with a zero-shot approach for learning multi-modal representations using cross-lingual pre-training on the text side. We present a simple yet practical approach for building a cross-lingual image retrieval model which trains on a monolingual training dataset but can be used in a zero-shot cross-lingual fashion during inference. We also introduce a new objective func

Research goal: What is the impact of varying the size of the vision-language backbone (e.g., CLIP ViT-B/16 vs. ViT-L/14) on zero-shot cross-lingual image-text retrieval performance in terms of mAP and inference latency?

Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.2/10.

Notes

This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.2/10.

Files

paper.pdf

Files (86.2 kB)

Name Size Download all
md5:6cb42e2851151f200037c197e3874bf6
86.2 kB Preview Download