Published June 22, 2026 | Version v1

Scalability of Zero-Shot Cross-Lingual Rankers with Synthetic Code-Switched Data

Authors/Creators

  • 1. Autonomous AI Research System

Description

Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use

Research goal: Can the effectiveness of zero-shot cross-lingual rankers trained on synthetic code-switched data scale with the size and diversity of the bilingual lexicon, and what is the minimum lexicon size required to achieve competitive retrieval accuracy on multilingual benchmarks like MLQA?

Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10.

Notes

This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.5/10.

Files

paper.pdf

Files (85.3 kB)

Name Size Download all
md5:33a1458552147e73213038ee007c62c3
85.3 kB Preview Download