Code-Focused Intermediate-Task Training for Cross-Lingual Code Generation
Description
The use of Large Language Models (LLMs) for program code generation has gained substantial attention, but their biases and limitations with non-English prompts challenge global inclusivity. This paper investigates the complexities of multilingual prompt-based code generation. Our evaluations of LLMs, including CODELLAMA and CODEGEMMA, reveal significant disparities in code quality for non-English prompts; we also demonstrate the inadequacy of simple approaches like prompt translation, bootstrapped data augmentation, and fine-tuning. To address this, we propose a zero-shot cross-lingual approac
Research goal: Does intermediate-task training with code-focused benchmarks (e.g., HumanEval, MBPP) improve zero-shot cross-lingual code generation performance on benchmarks like XCodeEval or Multilingual CodeGen?
Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.8/10.
Notes
Files
paper.pdf
Files
(75.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:d80ff80e6a598b01ebf07902b54a380d
|
75.3 kB | Preview Download |