OpenAI
GPT-2
Announced Feb 14, 2019
A glimpse of fluent text generation
What could it do?
Continue a supplied passage in a similar style, with some ability to answer questions and summarize through prompting.
What changed?
One language model attempted several tasks without separate training for each task.
Language generation
The release made fluent text generation and the decision to withhold larger model weights part of a public research debate.
Where it fell short
Coherent prose could still repeat itself, change topic or describe impossible events.
February 2019 announcement; model weights followed a staged release. Check sources
- Model type
- Text completion
- Largest model
- 1.5 billion parameters
- Initial access
- Smaller model released first
What this meant in practice
Fluency and reliability are different. A passage can sound natural while contradicting its own opening. Read this milestone as progress in generating language, rather than proof that the system understood every fact it wrote.
Continue a story
Give it an opening paragraph and ask for the next scene. Check whether characters, setting and events stay consistent.
Common question
Was GPT-2 a chatbot?
It was a text completion model. A chat product adds an interface and further behavior around a model.
THE USEFUL CONTEXT
Understanding GPT-2 beyond the headline
Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.
Text completion versus a conversation
The GPT-2 paper studied how language modeling could support different tasks without separate training for each one. This page records that research milestone. A useful way to understand it is to imagine a system continuing a document, rather than an assistant managing a conversation.
Suppose your input starts with a product description. A continuation may imitate the description’s style without respecting a later instruction to use exactly three bullet points. Those are separate things to measure: producing plausible language, keeping the original meaning and obeying a specific request. Treating them as one ability hides the improvements that later models made.
A practical example you can inspect
For an illustrative writing task, begin a fictional story with three fixed facts: the shop closes at six, the main character is called Maya and the parcel is red. Ask for a continuation. Read it once for style, then again only for those facts. A beautiful paragraph that changes the parcel to blue has failed a continuity check.
Try several openings and keep every attempt. Looking only at the best continuation tells you what is possible, but not how often it happens. If you compare models, give each the same opening, output allowance and number of attempts. This is a suggested exercise, not a test we have run or a claim about a success rate.
How to compare this milestone with later models
Start with the job the user wanted done. For creative drafting, variety and coherence may matter most. For extracting an appointment time, exact correctness matters more than style. A result from one task should not stand in for the other.
Our timeline therefore preserves GPT-2 as a historical release even though its modern benchmark fields are empty. Inventing a current score would make the line look smoother while weakening the comparison. If a later evaluation tests an archived model, that should be recorded with its actual evaluation date and setup, separately from the original release date.
FOLLOW THE EVIDENCE
What to watch next
Changes that would make this story worth revisiting:
- A reproducible evaluation of the archived model on a clearly described task.
- Evidence that a claimed improvement holds across many attempts, not just selected examples.
Questions about this milestone
Why keep GPT-2 on a modern AI chart?
It provides historical context for the move from plausible text continuation to more useful assistants. The point records a documented release, even where comparable scores are unavailable.
Can parameter counts tell me how many times better a model is?
No. A parameter count describes model size. To compare usefulness, specify a task and measure the quality of the resulting work.
Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.
OpenAI: GPT-2 announcementBenchmark results
No comparable ECI score is available for this release in our source snapshot. Missing scores are never estimated.