Cohere Labs - Feng Yao, PhD student in the CSE dept. UC San Diego

other
Cohere Labs - Feng Yao - DenseMixer: Improving MoE Post-Training with Precise Router Gradient

Date: Jul 23, 2025

Time: 4:00 PM - 5:00 PM

Location: Online

Training MoE (Mixture of Experts) models is significantly more challenging than training dense models. A key difficulty lies in the widely-used TopK router, which is non-differentiable and leads to challenges such as unstable optimization, difficulty in end-to-end training, and suboptimal solutions. To address this, we propose DenseMixer, which introduces a single forward pass through the non-activated experts (with controllable computational overhead) to provide the router with more accurate gradient signals, thereby improving post-training performance. We observe consistent improvements across three types of models — OLMoE, Qwen1.5-MoE, and Qwen3-MoE — with sizes 7B/14B/30B, on 10+ datasets. For example, on the GPQA-Diamond benchmark, Qwen3-30B-MoE achieves a 3.7 percentage point improvement over standard training when using just 1k SFT examples. In terms of efficiency, DenseMixer only adds one extra forward pass for non-activated experts, increasing theoretical FLOPs by 46% (1.46x the original). However, the actual increase in training time is lower than the theoretical estimate. When the model size or data volume is moderate, the additional time cost remains acceptable, achieving the goal of improving training quality with controllable computational overhead.

Feng Yao is currently a second-year PhD student in the CSE department at the University of California, San Diego (UCSD), advised by Prof. Jingbo Shang and Prof. Vish Krishnan. Previously, he received his master's degree from Tsinghua University, advised by Prof. Zhiyuan Liu and Prof. Weixing Shen.

Add event to calendar

Apple Google Office 365 Outlook Outlook.com Yahoo