Florian Strub
Florian Strub
About
Detail
Head of RLVR and Post-training Enginering at Cohere
Paris, Île-de-France, France
I am working at Cohere as one of the co-lead of the Post-Training department, where I co-orchestrate the fine-tuning of Large Language Models (LLMs) across the company. In particular, I lead a team of 15 people in charge of RLVR training and developing the engineering of the post-training stack. I was also a key architect of the Command A fine-tuning strategy, coordinating the efforts of over 80 contributors toward a unified goal.
Before joining Cohere, I was at Google DeepMind, where I specialized in training large language models using reinforcement learning and multimodal techniques, and explore game-theory methods for language bootstrapping. My academic journey has taken me to renowned AI laboratories worldwide (Imperial College, Mila, Inria, and KAIST), where I published influential papers (BYOL, FiLM, DeepNash) and book chapters (Oxford Book of Langague Evolution) that have gathered thousands of citations, and co-supervised four PhD theses.
In total, my experience includes over 10 years as a research scientist, 2 years as a project lead, and 2 years as a software developer. This diverse background enables me to navigate seamlessly between low-level engineering, hands-on coding, deep scientific analysis, project management, and high-level company vision.