G

Guillaume Lample

About

Detail

Co-founder & Chief Scientist
Paris, Île-de-France, France

Timeline


work
Job
school
Education
auto_stories
Publication

Résumé


Jobs verified_user 0% verified
  • Mistral AI
    Co-founder & Chief Scientist
    Mistral AI
    May 2023 - Current (3 years 6 months)
  • Facebook
    Research Scientist at Facebook AI Research
    Facebook
    Jan 2020 - Dec 2023 (4 years)
  • Facebook AI
    PhD Student
    Facebook AI
    Sep 2016 - Jan 2020 (3 years 5 months)
  • Facebook
    Research Intern
    Facebook
    May 2015 - Aug 2015 (4 months)
  • J
    Research Intern
    Jane Street Capital
    Apr 2014 - May 2014 (2 months)
Education verified_user 0% verified
  • Carnegie Mellon University
    Master's degree, Intelligence artificielle, Intelligence artificielle
    Carnegie Mellon University
    Jan 2014 - Dec 2016 (3 years)
  • École Polytechnique
    Master's degree, Mathématiques et informatique, Mathématiques et informatique
    École Polytechnique
    Jan 2011 - Dec 2014 (4 years)
  • Pierre and Marie Curie University
    Doctor of Philosophy - PhD, Artificial Intelligence, Artificial Intelligence
    Pierre and Marie Curie University
Publications verified_user 0% verified
  • I
    Deep learning for symbolic mathematics
    ICLR
    Dec 2019
    Neural networks have a reputation for being better at solving statistical or approximate problems than at performing calculations or working with symbolic data. In this paper, we show that they can be surprisingly good at more elaborated tasks in mathematics, such as symbolic integration and solving differential equations. We propose a syntax for representing mathematical problems, and methods for generating large datasets that can be used to train sequence-to-sequence models. We achieve results that outperform commercial Computer Algebra Systems such as Matlab or Mathematica.
  • I
    Word Translation Without Parallel Data
    ICLR
    Jan 2018
    State-of-the-art methods for learning cross-lingual word embeddings have relied on bilingual dictionaries or parallel corpora. Recent studies showed that the need for parallel data supervision can be alleviated with character-level information. While these methods showed encouraging results, they are not on par with their supervised counterparts and are limited to pairs of languages sharing a common alphabet. In this work, we show that we can build a bilingual dictionary between two languages without using any parallel corpora, by aligning monolingual word embedding spaces in an unsupervised way. Without using any character information, our model even outperforms existing supervised methods on cross-lingual tasks for some language pairs. Ou
  • I
    Unsupervised Machine Translation Using Monolingual Corpora Only
    ICLR
    Jan 2018
    Machine translation has recently achieved impressive performance thanks to recent advances in deep learning and the availability of large-scale parallel corpora. There have been numerous attempts to extend these successes to low-resource language pairs, yet requiring tens of thousands of parallel sentences. In this work, we take this research direction to the extreme and investigate whether it is possible to learn to translate even without any parallel data. We propose a model that takes sentences from monolingual corpora in two different languages and maps them into the same latent space. By learning to reconstruct in both languages from this shared feature space, the model effectively learns to translate without using any labeled data. We
  • E
    Phrase-Based & Neural Unsupervised Machine Translation
    EMNLP
    Jan 2018
    Machine translation systems achieve near human-level performance on some languages, yet their effectiveness strongly relies on the availability of large amounts of parallel sentences, which hinders their applicability to the majority of language pairs. This work investigates how to learn to translate when having access to only large monolingual corpora in each language. We propose two model variants, a neural and a phrase-based model. Both versions leverage a careful initialization of the parameters, the denoising effect of language models and automatic generation of parallel data by iterative back-translation. These models are significantly better than methods from the literature, while being simpler and having fewer hyper-parameters. On t
  • N
    Fader Networks: Manipulating Images by Sliding Attributes
    NIPS
    Jan 2017
    This paper introduces a new encoder-decoder architecture that is trained to reconstruct images by disentangling the salient information of the image and the values of attributes directly in the latent space. As a result, after training, our model can generate different realistic versions of an input image by varying the attribute values. By using continuous attribute values, we can choose how much a specific attribute is perceivable in the generated image. This property could allow for applications where users can modify an image using sliding knobs, like faders on a mixing console, to change the facial expression of a portrait, or to update the color of some objects. Compared to the state-of-the-art which mostly relies on training adversar
  • A
    Playing FPS Games with Deep Reinforcement Learning
    AAAI
    Jan 2017
    Advances in deep reinforcement learning have allowed autonomous agents to perform well on Atari games, often outperforming humans, using only raw pixels to make their decisions. However, most of these games take place in 2D environments that are fully observable to the agent. In this paper, we present the first architecture to tackle 3D environments in first-person shooter games, that involve partially observable states. Typically, deep reinforcement learning methods only utilize visual input for training. We present a method to augment these models to exploit game feature information such as the presence of enemies or items, during the training phase. Our model is trained to simultaneously learn these features along with minimizing a Q-lea
  • N
    Neural Architectures for Named Entity Recognition
    NAACL
    Jan 2016
    State-of-the-art named entity recognition systems rely heavily on hand-crafted features and domain-specific knowledge in order to learn effectively from the small, supervised training corpora that are available. In this paper, we introduce two new neural architectures---one based on bidirectional LSTMs and conditional random fields, and the other that constructs and labels segments using a transition-based approach inspired by shift-reduce parsers. Our models rely on two sources of information about words: character-based word representations learned from the supervised corpus and unsupervised word representations learned from unannotated corpora. Our models obtain state-of-the-art performance in NER in four languages without resorting to a
This is a community-created genome.