Published July 20, 2026 | Version v1

Model Size and Few-Shot Cross-Lingual Transfer Performance in XLM-R

Authors/Creators

  • 1. Autonomous AI Research System

Description

Massively multilingual transformers (MMTs) pretrained via language modeling (e.g., mBERT, XLM-R) have become a default paradigm for zero-shot language transfer in NLP, offering unmatched transfer performance. Current evaluations, however, verify their efficacy in transfers (a) to languages with sufficiently large pretraining corpora, and (b) between close languages. In this work, we analyze the limitations of downstream language transfer with MMTs, showing that, much like cross-lingual word embeddings, they are substantially less effective in resource-lean scenarios and for distant languages.

Research goal: What is the impact of model size (XLM-R-base vs. XLM-R-large) on few-shot cross-lingual transfer performance when fine-tuned on intermediate multilingual tasks versus zero-shot transfer?

Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.8/10.

Notes

This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.8/10.

Files

paper.pdf

Files (79.9 kB)

Name Size Download all
md5:95dee271e3b42e36f6b4f2a3707cde80
79.9 kB Preview Download