Cohere For AI - Guest Speaker: Niklas Muennighoff, Research Engineer

Date: Jun 11, 2024
Time: 4:00 PM - 5:00 PM
Location: Online
In this talk, we will discuss how to scale language models in data-constrained regimes based on results from "Scaling Data-Constrained Language Models". The current trend of scaling language models suggests that training dataset size may soon be constrained by the amount of text data available on the internet. After a brief background on scaling, we will cover to what extent repeating data can alleviate data constraints. To quantify the effect of repetition, we will discuss a scaling law for compute optimality that accounts for the decreasing value of repeated tokens and excess parameters. We will evaluate this scaling law on a large set of experiments varying the extent of data repetition and compute budget, ranging up to 900 billion training tokens and 9 billion parameter models. We will finish with complementary approaches to lift data constraints, including code augmentation and revising filtering strategies.
Bio: Niklas Muennighoff is a Research Engineer at Contextual AI. His research focuses on making LLMs better and more useful via efforts such as OLMo, BLOOM, and StarCoder. He obtained his Bachelor's from Peking University.






