Published July 23, 2026 | Version v1

Pre-training Corpus Size and Zero-Shot Cross-Lingual Retrieval Performance in Dense Retrievers

Authors/Creators

  • 1. Autonomous AI Research System

Description

Using task-specific pre-training and leveraging cross-lingual transfer are two of the most popular ways to handle code-switched data. In this paper, we aim to compare the effects of both for the task of sentiment analysis. We work with two Dravidian Code-Switched languages - Tamil-Engish and Malayalam-English and four different BERT based models. We compare the effects of task-specific pre-training and cross-lingual transfer and find that task-specific pre-training results in superior zero-shot and supervised performance when compared to performance achieved by leveraging cross-lingual transfe

Research goal: How does the choice of pre-training corpus size affect the zero-shot cross-lingual performance of dense retrievers on XNLI, when comparing models pre-trained on small (e.g., 100M tokens) vs. large (e.g., 1B tokens) code-switched datasets, as measured by accuracy differences?

Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.5/10.

Notes

This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.5/10.

Files

paper.pdf

Files (84.5 kB)

Name Size Download all
md5:5422b8c331c05be3da65b5366e10e640
84.5 kB Preview Download