OpenAI
GPT-4
Announced Mar 14, 2023
A step forward in reasoning and images
What could it do?
Work with complex text instructions and, in its research preview, reason about supplied images.
What changed?
OpenAI reported stronger exam and instruction following performance than GPT-3.5.
Image understanding
The release connected stronger language reasoning with visual inputs such as diagrams and screenshots.
Where it fell short
The model could still hallucinate facts, make reasoning mistakes and produce insecure code.
March 2023 announcement. Text access launched before broad image input access. Check sources
- Inputs described
- Text and images
- Output
- Text
- Image access at launch
- Research preview
What this meant in practice
Image understanding creates a different kind of assistance: the user can ask about what they are seeing. Exam performance is a separate measurement. Neither one proves the model can reliably complete an entire professional job.
Explain a diagram
Provide a chart and ask which trend supports a conclusion. Compare the answer with the labels and underlying values; a convincing explanation may still misread the picture.
Common question
Was GPT-4 a qualified lawyer?
No. A reported score on a simulated exam does not establish professional qualification or dependable legal work.
THE USEFUL CONTEXT
What GPT-4 changed in practice
Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.
A benchmark result and a finished job are different
The technical report evaluated GPT-4 on exams and other tests while documenting reliability limitations. Such results help describe a release, but the test conditions are part of the result. The task, examples, scoring and allowed tools all affect what can reasonably be concluded.
Consider a customer support workflow. Answering one written question correctly is only one step. The system also needs the correct policy, the relevant order and a way to identify when information is missing. A better exam score does not establish that these other pieces are present. Our practical interpretation is to evaluate the complete workflow before handing it responsibility.
An example with a clear definition of success
Imagine a fictional store whose policy permits returns within 14 days but excludes personalized items. Supply the policy and ask for draft replies to three customers: an ordinary return on day ten, a personalized item and a message without a purchase date. A useful answer handles the exception and asks for missing information instead of inventing it.
Score policy accuracy separately from tone. Then check whether the response promises an action the system cannot perform. “I have refunded you” is different from “You may be eligible for a refund.” This example is a way to inspect instruction following; it is not a reported GPT-4 evaluation.
What image understanding adds
For a chart or screenshot, break the work into observable steps. First ask for the title, labels and visible values. Then ask for an interpretation. If the extracted values are wrong, a sophisticated explanation does not repair the answer. This approach helps distinguish a reading error from a reasoning error.
Compare a screenshot task with the same information supplied as text. The difference can tell you whether the image itself introduces difficulty. Keep the model version fixed and avoid giving one system extra context. When reading the historical launch, also distinguish capabilities demonstrated by researchers from capabilities broadly available to users that day.
FOLLOW THE EVIDENCE
What to watch next
Changes that would make this story worth revisiting:
- Evaluations that measure complete tasks with the model version and tools disclosed.
- Evidence showing where a newer release improves on the same examples and where it still fails.
Questions about this milestone
Is GPT-4 the same thing as ChatGPT?
GPT-4 is a model. ChatGPT is a product that can combine models with an interface and tools. A product feature should not automatically be attributed to the original model.
Does a stronger model remove the need for review?
The amount of review depends on the consequences of a mistake and the observed error rate on your task. A historical benchmark alone cannot answer that question.
Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.
OpenAI: GPT-4 launch and availability OpenAI: GPT-4 Technical ReportBenchmark results
A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.
Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.
Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026