Cohere Labs - Stella Li, PhD Student

other
Cohere Labs - Stella Li - Spurious Rewards: Rethinking Training Signals in RLVR

Date: Jun 11, 2025

Time: 4:00 PM - 5:00 PM

Location: Online

This talk explores a surprising finding in reinforcement learning for mathematical reasoning: certain language models can achieve substantial performance gains even when trained with completely uninformative or incorrect reward signals. Through extensive experiments on mathematical benchmarks, we demonstrate that Qwen2.5-Math models improve significantly when trained with random rewards, format-only rewards, or even rewards that explicitly favor incorrect answers.

Stella a second year Ph.D. student in the Allen School of Computer Science and Engineering at the University of Washington, advised by Yulia Tsvetkov.

Stella received her B.S. and M.S.E. at Johns Hopkins with majors in Computer Science, Cognitive Science (linguistics focus), and Applied Mathematics (statistics focus). I worked as a research assistant at the Center for Language and Speech Processing advised by Philipp Koehn and Kenton Murray.

Add event to calendar

Apple Google Office 365 Outlook Outlook.com Yahoo