Florian Strub

Florian Strub

About

Detail

Head of RLVR and Post-training Enginering at Cohere
Paris, Île-de-France, France

Timeline


work
Job
school
Education

Résumé


Jobs verified_user 0% verified
  • Cohere
    Head of RLVR and Post-training enginering at Cohere
    Cohere
    Mar 2025 - Current (1 year 6 months)
  • Cohere
    Co-head of Command A and Command R7B Post-training
    Cohere
    Jul 2024 - Oct 2025 (1 year 4 months)
    My role consist in coordinating the LLM training across Cohere's technical teams in Europe and the US to deliver the next generation of LLMs. aka Command A. I also serve as the primary manager of the RL team in Europe. I co-designed a bottom-up organizational structure to harness the expertise of every individual and foster effective cross-team collaboration,. This decentralized approach, inspired by recent research on team management (Reinventing Organizations), has proven successful in promoting innovation within large R&D groups. The new post-training structure has enabled us to scale the team from 5 to 15 members, attracting over 40 individuals in just a few months. I've harmonized technical processes, evaluation, tooling, and codebas
  • Cohere
    Senior Research Scientist
    Cohere
    Feb 2024 - Oct 2024 (9 months)
    My role was to set-up the foundation of a RL team for LLM at Cohere from both an enginering and research perspective, while developping Cohere presence in France. Codebase: I spearheaded a major refactoring of Cohere's post-training pipeline based on Jax+Ray, enabling the application of new cutting-edge SFT, Offpref, and RL algorithms to our 7B, 35B, and 100B LLM models. - I identified the design bottlenecks, before setting-up comprehensive integrations test. This process guaranted a strict 1:1 behavior, enable a transparent transition to all users (80+) without any production breakage, and it is now the standard training frameowrk for the new generations of LLMs. - I reduced the post-traning codebase from 80%, reduce memory consumption
  • Google Deepmind
    Senior Research Scientist
    Google Deepmind
    Jun 2019 - Apr 2024 (4 years 11 months)
    I coordinate a team of six scientists developing state-of-the-art algorithms for training large-language models with multi-agent RL techniques and vision capabilities. I design self-improving methods to foster network capacities without human intervention. My research journey has been focused on three main directions: 1) Enhancing language agent capabilities: I have pioneered novel RL methods for large language models, surpassing their initial imitation pretraining. With over 10 papers and around 800 citations involves • Scaling up RLHF methods for natural language processing • Devising techniques to counter language drift and analyzing the emergence of communication protocols among conversational agents during co-training. • Training popu
  • Google Deepmind
    Research Intern
    Google Deepmind
    Mar 2018 - Aug 2018 (6 months)
  • Université de Montréal
    Ph.D Visitor in Deep Learning
    Université de Montréal
    Jun 2017 - Dec 2017 (7 months)
  • INRIA
    Ph.D Student in Deep Learning and Reinforcement Learning
    INRIA
    Dec 2015 - Jun 2019 (3 years 7 months)
    I study the deep learning method required to learn consistent multimodal representations and finetune them through interactions with reinforcement learning. This Ph.D. was supervised by Prof. Olivier Pietquin, Dr. Jeremy Marie, with the collaboration of Prof. Aaron Courville. The main contributions of my thesis can be summarized as follows: 1) I co-created one of the largest visually grounded dialogue dataset, namely GuessWhat!?: We set up the infrastructure to collect, clean, and analyze 150k labeled dialogues. This open-source dataset remains one of the most complex dialogue datasets for assessing visual understanding and language generation and has been instrumental in the multimodal learning community. 2) Design of the modulation mec
  • INRIA
    Data scientist, Recommendation System
    INRIA
    Jan 2015 - Dec 2015 (1 year)
    I explored online recommendation systems and collaborative filtering with neural networks. This project aimed to predict future user preferences by leveraging their past ratings and meta-information. It resulted in the publication of one of the pioneering studies employing deep learning techniques in the recommendation system community, outperforming many classic machine learning techniques. The associated open-source codebase garnered over a hundred developers starring it on GitHub, and the scientific publications have accumulated 500 citations.
  • Societe Generale
    Consultant in IT and Finance (Front Office)
    Societe Generale
    Jan 2013 - Feb 2015 (2 years 2 months)
    I implemented a high-frequency trading API to serve low-latency trading automatons. This framework enable connectivity between trading automatons and global markets, executing and verifyig millions of trades daily across financial instruments (Stocks, Futures and options). It involved monthly release in production, which were closely monitored by Quality Assurance (QA) team, the project owners, and the trading desks. Key highlights of my involvement in this project include: • Co-leading, designing and developing the high-frequency API from scratch, aligning with business requirements and real-time constraints. • Providing technical support to IT Quants and Traders, ensuring smooth operations at support levels 2 and 3. • Creating a compreh
  • Imperial College London
    Research Assistant
    Imperial College London
    Sep 2011 - Oct 2012 (1 year 2 months)
    During my 5-months postgraduate project, I enhanced a C++ expert system that aids doctors in delivering precise breast-cancer diagnoses. The input data consisted of pre-processed scientific texts formatted as logic rules. The foundation of my research project was the Argumentation Based Assumption theory pioneered by Francesca Toni, which I enhanced by using genetic algorithms and neural networks. As part of a 4-month part-time independent study, I composed a comprehensive research essay that presented the current state of the art in Bayesian computation for neuronal activity and modeled how multimodal cues may be combined. My study primarily drew from the works of Alexandre Pouget.
  • KAIST
    Research Internship
    KAIST
    May 2011 - Sep 2011 (5 months)
    I developed a web platform aimed at promoting researchers' works and facilitating the testing of cutting-edge algorithms in computer vision. This platform provided users with the ability to upload their pictures and apply state-of-the-art algorithms for analysis. I was also initiated to research vision and robotics while interacting with other scientists daily.
  • Snecma
    Industrial Placement
    Snecma
    Feb 2010 - Mar 2010 (2 months)
    During my time working on assembly lines from 6am to 1pm, I was part of a team responsible for cleaning and cutting plane motor engine pieces. This role provided me with the opportunity to study and observe principles of lean management and human resources, with a focus on optimizing efficiency and productivity while ensuring the well-being and involvement of workers.
  • Planète Sciences
    Summer Camp Leader
    Planète Sciences
    Nov 2006 - Aug 2010 (3 years 10 months)
    Throughout my experience, I have been actively involved in managing various aspects of children's daily lives, including scientific projects, sports activities, and general well-being. Ensuring children's safety has always been a top priority in my work, and I have regularly engaged with parents to address any concerns or provide updates on their child's progress. Additionally, I have played a role in logistics and coordination to ensure smooth operations. I henve supervised and assisted in the management of 40 children aged 10-14 years old or organized scientific events gathering more than 100 children.
Education verified_user 0% verified
  • Imperial College London
    Msc in Advanced Computing, Mathematics and Computer Science
    Imperial College London
    Jan 2011 - Jan 2012 (1 year 1 month)
    Advanced Machine Learning (Genetic Algortihm, Neural Networks, Graph Theory etc.), Bayesian Networks and Probabilistic Inferences, Software Engineering (TDD, Agile development, Jade etc.), Multi-Agent Systems, Computational Finance (Risk Management, Optimization, Portofolio theory), Advanced Neurodynamics.
  • École des Mines de SaintÉtienne
    Double Master in, Engineering and Management
    École des Mines de SaintÉtienne
    Jan 2009 - Jan 2011 (2 years 1 month)
    Software Development, Database, Web Programming, Advanced Algorithms, Probabilistic Mathematic, Electrical Engineering, Project Management, Communication, Economics, Labour law, Psychology
  • L
    Bachelor of Applied Science (B.A.Sc, Mathematics
    Lycee Janson de Sailly
    Jan 2007 - Jan 2009 (2 years 1 month)
    Advanced Mathematics, Advanced Physics, Chemistry.
Awards verified_user 0% verified
  • ErnstYoung
    Second Best Student Project Award in 2011 in Department of Loire
    ErnstYoung
  • E
    Special Rewards
    Ecole Nationale Superieure des Mines de SaintEtienne
    Special Reward in 2013 by the academic committee of Ecole des Mines de Saint-Etienne for outstanding social involvement and excellent academic records
  • F
    Best French Scientific Student Project Award
    French National Student Representative Body CNOUS
    Reward the fulfilment of setting up a new pro-active Robotic Union. The award mainly emphasized about the organization of a scientific and cultural event for two hundred children in May 2010. Children were taught robotic basics and they could build their own robots. Press release: http://www.emse.fr/spip/IMG/pdf/Communique_ENSM-SE_prix_CNOUS_min_bot.pdf
This is a community-created genome.