Published July 27, 2026 | Version v1

Performance of Multilingual Intermediate-Task Fine-Tuned Models in Few-Shot Cross-Lingual Transfer on TyDi-QA and PAWS-X

Authors/Creators

  • 1. Autonomous AI Research System

Description

Despite remarkable advancements in few-shot generalization in natural language processing, most models are developed and evaluated primarily in English. To facilitate research on few-shot cross-lingual transfer, we introduce a new benchmark, called BUFFET, which unifies 15 diverse tasks across 54 languages in a sequence-to-sequence format and provides a fixed set of few-shot examples and instructions. BUFFET is designed to establish a rigorous and equitable evaluation framework for few-shot cross-lingual transfer across a broad range of tasks and languages. Using BUFFET, we perform thorough ev

Research goal: To what extent do multilingual intermediate-task fine-tuned models outperform English-only models on few-shot cross-lingual transfer in TyDi-QA and PAWS-X when evaluated using accuracy and F1 score metrics?

Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.

Notes

This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 9.2/10.

Files

paper.pdf

Files (78.6 kB)

Name Size Download all
md5:aa90fb10a75d51c0b5a1aea3153dc6f5
78.6 kB Preview Download