Listening, Speaking and Writing
30 AI-Speak: Natural Language Processing
Natural language processing has been a topic on which research has worked in length for the past 50 years. This has led to the development of many tools we use every day:
- Word processors
- Automatic grammar and orthography correction
- Automatic completion
- Optical character recognition (OCR)
More recently, chatbots, home assistants and automatic translation tools have been making a huge impact in all areas.
For a long time, research and industry was stalled by the intrinsic complexity of language. At the end of the 20th century, grammars for a language, written by experts, could have up to 50,000 rules. These expert systems were showing that technology could make a difference, but robust solutions were too complex to develop.
On the other hand, speech recognition needed to be able to make use of acoustic data and transform it into text. With the variety of speakers one could find, a hard task indeed!
Researchers understood that if we had a model for the intended language, things would be easier. If we knew the words of the language, how sentences were formed, then it would be easier to find the right sentence from a set of candidates to match a given utterance, or to produce a valid translation from a set of possible sequences of words.
Another crucial aspect has been that of semantics. Most of the work we can do to solve linguistic questions is shallow; the algorithms will produce an answer based on some local syntactic rules. If in the end, the text means nothing, so be it. A similar thing may happen when we read a text by some pupils – we can correct the mistakes without really understanding what the text is about! A real challenge is to associate meaning to text and, when possible, to uttered sentences.
There was a surprising result in 20081. One unique language model could be learnt from a large amount of data and used for a variety of linguistic tasks. In fact, that unique model performed better than models trained for specific tasks.
The model was a deep neural network. Nowhere as deep as the models used today! But enough to convince research and industry that machine learning, and more specifically Deep learning was going to be the answer to many questions in NLP.
Since then, natural language processing has ceased to follow a model-driven approach and has been nearly always based on a data-driven approach.
Traditionally, the main language tasks can be decomposed into 2 families – those involving building models and those involving decoding.
Building models
In order to transcribe, answer questions, generate dialogues or translate, you need to be able to know whether or not “Je parle Français” is indeed a sentence in French. And as with spoken languages, rules of grammar are not always followed accurately, so the answer has to be probabilistic. A sentence can be more or less French. This allows the system to produce different candidate sentences (as the transcripti