one year on
Google researchers propose Transformer network, doing away with recurrence in sequence modeling
Architecture relying solely on attention mechanisms achieves state-of-the-art translation results while training significantly faster than recurrent or convolutional models.
Eight Google researchers today published a paper on arXiv titled ‘Attention Is All You Need,’ proposing a new neural network architecture called the Transformer that does away with recurrence entirely. The architecture relies solely on attention mechanisms, avoiding the sequential processing that has long constrained recurrent and convolutional models.
On the WMT 2014 English-to-German translation task, the Transformer achieved a BLEU score of 28.4, improving over existing best results including ensembles by more than 2 BLEU. On English-to-French, it established a new single-model state-of-the-art score of 41.8 after training for only 3.5 days on eight GPUs — a fraction of the training costs of previous best models. The paper also shows the Transformer generalizes well to English constituency parsing.
The paper says the architecture is more parallelizable and requires significantly less time to train than recurrent or convolutional models.
The record
Submitted the paper to arXiv on June 12, 2017.
Posted a blog entry on August 31, 2017 detailing the Transformer architecture and its advantages over RNNs and CNNs.
One year later — open only if you can handle spoilers
The Transformer will become the foundational architecture for virtually all subsequent large language models, from GPT-2 to GPT-4 and beyond. Its introduction marked a decisive shift away from recurrence in natural language processing.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy