Impact of RLHF on Multilingual LLM Alignment in BUFFET Tasks
Description
While Reinforcement Learning from Human Feedback (RLHF) effectively aligns pretrained Large Language and Vision-Language Models (LLMs, and VLMs) with human preferences, its computational cost and complexity hamper its wider adoption. To alleviate some of the computational burden of fine-tuning, parameter efficient methods, like LoRA were introduced. In this work, we empirically evaluate the setup of Parameter Efficient Reinforcement Learning from Human Feedback (PE-RLHF) that leverages LoRA fine-tuning for Reward Modeling, and Reinforcement Learning. We benchmark the PE-RLHF setup on six diver
Research goal: How does reinforcement learning from human feedback (RLHF) influence the alignment of multilingual LLMs on BUFFET tasks, as evaluated by human preference scores versus baseline fine-tuning?
Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.2/10.
Notes
Files
paper.pdf
Files
(77.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:20cf9cea8bae632dcd1158edeaee0f71
|
77.4 kB | Preview Download |