Early preview
O

OpenAI

GPT-4

Announced Mar 14, 2023

A step forward in reasoning and images

01

What could it do?

Work with complex text instructions and, in its research preview, reason about supplied images.

02

What changed?

OpenAI reported stronger exam and instruction following performance than GPT-3.5.

WHY IT MATTERED

Image understanding

The release connected stronger language reasoning with visual inputs such as diagrams and screenshots.

03

Where it fell short

The model could still hallucinate facts, make reasoning mistakes and produce insecure code.

Which release does this page cover?

March 2023 announcement. Text access launched before broad image input access. Check sources

Inputs described
Text and images
Output
Text
Image access at launch
Research preview

What this meant in practice

Image understanding creates a different kind of assistance: the user can ask about what they are seeing. Exam performance is a separate measurement. Neither one proves the model can reliably complete an entire professional job.

ILLUSTRATIVE TASK · NOT A TEST RESULT

Explain a diagram

Provide a chart and ask which trend supports a conclusion. Compare the answer with the labels and underlying values; a convincing explanation may still misread the picture.

Common question

Was GPT-4 a qualified lawyer?

No. A reported score on a simulated exam does not establish professional qualification or dependable legal work.

THE USEFUL CONTEXT

What GPT-4 changed in practice

Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.

A benchmark result and a finished job are different

The technical report evaluated GPT-4 on exams and other tests while documenting reliability limitations. Such results help describe a release, but the test conditions are part of the result. The task, examples, scoring and allowed tools all affect what can reasonably be concluded.

Consider a customer support workflow. Answering one written question correctly is only one step. The system also needs the correct policy, the relevant order and a way to identify when information is missing. A better exam score does not establish that these other pieces are present. Our practical interpretation is to evaluate the complete workflow before handing it responsibility.

GPT-4 technical report

An example with a clear definition of success

Imagine a fictional store whose policy permits returns within 14 days but excludes personalized items. Supply the policy and ask for draft replies to three customers: an ordinary return on day ten, a personalized item and a message without a purchase date. A useful answer handles the exception and asks for missing information instead of inventing it.

Score policy accuracy separately from tone. Then check whether the response promises an action the system cannot perform. “I have refunded you” is different from “You may be eligible for a refund.” This example is a way to inspect instruction following; it is not a reported GPT-4 evaluation.

What image understanding adds

For a chart or screenshot, break the work into observable steps. First ask for the title, labels and visible values. Then ask for an interpretation. If the extracted values are wrong, a sophisticated explanation does not repair the answer. This approach helps distinguish a reading error from a reasoning error.

Compare a screenshot task with the same information supplied as text. The difference can tell you whether the image itself introduces difficulty. Keep the model version fixed and avoid giving one system extra context. When reading the historical launch, also distinguish capabilities demonstrated by researchers from capabilities broadly available to users that day.

FOLLOW THE EVIDENCE

What to watch next

Changes that would make this story worth revisiting:

  • Evaluations that measure complete tasks with the model version and tools disclosed.
  • Evidence showing where a newer release improves on the same examples and where it still fails.

Questions about this milestone

Is GPT-4 the same thing as ChatGPT?

GPT-4 is a model. ChatGPT is a product that can combine models with an interface and tools. A product feature should not automatically be attributed to the original model.

Does a stronger model remove the need for review?

The amount of review depends on the consequences of a mistake and the observed error rate on your task. A historical benchmark alone cannot answer that question.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.

OpenAI: GPT-4 launch and availability OpenAI: GPT-4 Technical Report

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
125.9index points
Tested variant: GPT-4 (Mar 2023)Source interval: 119.3 to 130.4Variant date in source: Mar 14, 2023

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
1288rating points
Tested: gpt-4-0314Text Arena Overall · October 8, 2026Reported interval: 1283 to 129354,173 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.