THE 2010S

Natural Language Processing

Since the 1950s, researchers have been trying to teach computers to understand our language. Natural language processing, more commonly referred to as NLP in specialized circles, has long stumbled over the same obstacles: how to capture the subtleties, double meanings, and changing context of a conversation? Early machines worked with fixed dictionaries and rigid grammatical rules. Then came statistical approaches in the 2000s, which brought their share of improvements without truly solving the fundamental problem.

The year 2017 marks a turning point. On the Mountain View campus, Ashish Vaswani and Jakob Uszkoreit are discussing machine translation in a hallway. A seemingly mundane conversation that will lead to something unexpected. Their team consists of eight Google researchers, including Illia Polosukhin, a science fiction enthusiast. It’s precisely a film, Arrival, that gives him an idea. In this story, aliens communicate with symbols that express entire concepts at once, without following our usual sequential logic. Why couldn’t machines do the same with sentences?

This reflection leads to the concept of self-attention: instead of analyzing words one after another like reading a book, the system grasps the sentence as a whole. All words interact with each other simultaneously. Noam Shazeer, who has been working at Google since 2000 and created the famous “Did you mean?” function, hears about the project in the offices. He joins the venture. The first trials on English-German translations yield encouraging results.

The resulting Transformer architecture is faster and more accurate than previous recurrent neural networks, and it’s not limited to text. The model processes images, computer code, DNA sequences. This versatility surprises its creators.

The following year, another team from Google AI publishes BERT. The acronym conceals a simple idea: reading in both directions. The context of a word depends on what comes before, but also on what follows. BERT performs on eleven different language processing tasks, breaking previous records.

Curiously, all eight researchers behind Transformers leave Google in the following years. Each founds their own company, exploits the technology in their own direction. Aidan Gomez launches Cohere for businesses, Jakob Uszkoreit creates Inceptive and applies Transformers to vaccines, Noam Shazeer develops Character.ai where users design their own conversational agents. Vaswani and Parmar establish Essential.ai. We find here a classic pattern: large organizations struggle to transform their discoveries into commercial products. Researchers leave to create elsewhere.

This architectural foundation has since fueled all modern language models. The ability to process enormous volumes of data, to weave complex links between elements of a sequence, makes Transformers an indispensable tool. Applications multiply: more natural translation, coherent text generation, document analysis, high-performing question-answering systems. Machines are approaching the nuances of human language, better understanding the relationships between ideas.

From word-by-word processing of the early days to contemporary attention systems, each step has brought computers closer to our linguistic complexity. Transformers are not just a technical feat. They serve as a foundation for a cascade of innovations in various fields: medicine, scientific research, education. The scientific community constantly develops new variants.

As these models gain sophistication, they alter our relationship with technology. The boundary between human communication and machine interaction becomes blurrier. Future developments promise new possibilities while raising questions of ethics and responsibility. How far will this capacity of systems to understand and produce natural language go? The answer is being written day by day, in laboratories as in our daily uses.