Jul 05, 2024
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
We introduce a novel, scalable method for generating high-quality multilingual feedback data to balance data coverage. We establish the benefits of cross-lingual transfer and increased dataset size in preference training.

Authors
John Dang with Arash Ahmadian, Kelly Marchisio, Julia Kreutzer, Ahmet Üstün, Sara Hooker
Abstract
Related works

Research
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards
Read

Research
Unlocking Reasoning Capability on Machine Translation in Large Language Models
Read

Research
Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation
Read






