Cohere Labs - Théo Vincent, Ph.D. student

Date: Jul 18, 2025
Time: 4:00 PM - 5:00 PM
Location: Online
Reinforcement learning is a powerful tool to solve complex sequential decision-making problems. Most advances in reinforcement learning focus on designing a learning algorithm that remains fixed during the training process. In this talk, I will present a new perspective on the training process, considering it as a Markov Decision Process in which actions can be taken to influence the training trajectory of the agent toward a better outcome. I will then explain how this vision materializes, adapting the hyperparameters of the reinforcement learning agent to its learning pace. Then, I will present how the agent's performance can be increased by anticipating its learning trajectory. The idea of optimizing the learning trajectory opens up new possibilities for designing reinforcement learning algorithms that achieve a satisfactory outcome in a single training trajectory.
Théo Vincent is a Ph.D. student at the Technical University of Darmstadt and at DFKI, the German institute for Artificial Intelligence. He is currently working on off-policy Reinforcement Learning methods. He works under the supervision of Jan Peters. Before his Ph.D., Théo graduated from MVA at ENS Paris Saclay. He also worked on Compute Vision problems in a Parisian lab, Saint-Venant lab, and a Swedish start-up, Signality. Théo did an internship in biostatistics at Harvard Medical School.
Add event to calendar