Published July 18, 2026 | Version v1

Impact of Target-Language Development Sets on Zero-Shot Cross-Lingual Performance of Multilingual BERT in XNLI

Authors/Creators

  • 1. Autonomous AI Research System

Description

Multilingual contextual embeddings have demonstrated state-of-the-art performance in zero-shot cross-lingual transfer learning, where multilingual BERT is fine-tuned on one source language and evaluated on a different target language. However, published results for mBERT zero-shot accuracy vary as much as 17 points on the MLDoc classification task across four papers. We show that the standard practice of using English dev accuracy for model selection in the zero-shot setting makes it difficult to obtain reproducible results on the MLDoc and XNLI tasks. English dev accuracy is often uncorrelate

Research goal: How does using target-language-specific development sets for model selection impact the zero-shot cross-lingual performance of multilingual BERT on the XNLI dataset compared to other multilingual language models like XLM-R or mT5, as measured by macro-averaged F1 scores across all languages?

Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.8/10.

Notes

This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.8/10.

Files

paper.pdf

Files (80.7 kB)

Name Size Download all
md5:e2318292f4d19548edb94b5da3aa1036
80.7 kB Preview Download