one year on
OpenAI unveils GPT-3, a 175-billion-parameter language model that learns tasks from a few examples
At 175 billion parameters, the model is ten times larger than any previous non-sparse language model and can perform translation, question-answering, cloze tasks, and three-digit arithmetic without fine-tuning.
OpenAI today posted a paper on arXiv describing GPT-3, a language model with 175 billion parameters — ten times more than any previous non-sparse model. The model is trained to predict the next token and, without any gradient updates or fine-tuning, can perform a wide range of tasks from just a handful of examples or natural-language instructions. GPT-3 achieves competitive results on translation, question-answering, cloze tasks, and even three-digit arithmetic, and can generate news articles that human evaluators struggle to distinguish from human-written ones. The paper, titled ‘Language Models are Few-Shot Learners,’ is co-authored by 31 researchers at OpenAI.
The record
One year later — open only if you can handle spoilers
GPT-3 became the foundation of OpenAI's commercial API, launched in 2020, and kicked off a wave of 'foundation model' startups and products. The 'scale is all you need' narrative it inspired was later challenged by the rise of instruction tuning, reinforcement learning from human feedback, and more parameter-efficient models like GPT-3.5. The paper remains one of the most influential AI publications of the decade.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy