BLIP Inference Latency and Throughput Under Varying Code-Switching Ratios in Zero-Shot Cross-Lingual Image-Text Retrieval
Description
There has been a recent spike in interest in multi-modal Language and Vision problems. On the language side, most of these models primarily focus on English since most multi-modal datasets are monolingual. We try to bridge this gap with a zero-shot approach for learning multi-modal representations using cross-lingual pre-training on the text side. We present a simple yet practical approach for building a cross-lingual image retrieval model which trains on a monolingual training dataset but can be used in a zero-shot cross-lingual fashion during inference. We also introduce a new objective func
Research goal: What is the effect of varying the code-switching ratio on the inference latency and throughput of BLIP during zero-shot cross-lingual image-text retrieval tasks?
Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.7/10.
Notes
Files
paper.pdf
Files
(87.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:f782570ee32aac02558220fc39aaed55
|
87.9 kB | Preview Download |