Early preview
A

Anthropic

Claude 3.5 Sonnet

Announced Jun 20, 2024

A stronger everyday coding assistant

01

What could it do?

Help with code, writing and visual analysis, including charts and text in images.

02

What changed?

Anthropic reported improved capability and faster responses than Claude 3 Opus.

WHY IT MATTERED

Coding assistance

Artifacts arrived alongside the model, giving generated documents, code and designs a dedicated workspace.

03

Where it fell short

Anthropic’s internal coding evaluation is not interchangeable with SWE bench Verified.

Which release does this page cover?

Original June 2024 release, not the October update. Google Cloud records June 20 availability; Anthropic’s page currently displays June 21. Check sources

Version covered
Original June release
Context at launch
200,000 tokens
Companion feature
Artifacts preview

What this meant in practice

Separate the model from the workspace around it. Artifacts made output easier to inspect and revise, while the model generated that output. Both can make a workflow more useful, but they are different changes.

ILLUSTRATIVE TASK · NOT A TEST RESULT

Draft a booking page

Describe the fields and layout, then inspect the resulting interface. Check keyboard access, validation and actual form submission separately. A convincing screen is only one part of a working site.

Common question

Does this include computer use?

This page covers June 2024. Do not assign features from later releases to this snapshot.

THE USEFUL CONTEXT

Why the original Claude 3.5 Sonnet mattered

Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.

The model and the workspace around it

Anthropic introduced Artifacts alongside Claude 3.5 Sonnet: generated documents, code and designs could be displayed beside the conversation. This is an important distinction for the timeline. The model produces material; the product determines how conveniently a person can inspect and revise it.

Imagine two assistants generating the same page. One returns a long block of code; the other lets you inspect a preview and request changes. The second experience might save time without proving that its underlying model has a higher score on every test. Comparing products means evaluating that experience as well as the model’s output.

Anthropic’s original announcement

A website example with real acceptance criteria

For an illustrative project, request a booking form for a fictional photography business. Specify a name field, email, preferred date and confirmation message. Start by inspecting the page on a narrow screen. Then use only the keyboard, submit empty fields and try a malformed email address. A polished screenshot does not answer any of those questions.

Next, decide what “submitted” means. Does the form only display a message, or does it save a request somewhere? This difference is easy to miss in an impressive demo. Before trusting generated software, define the behavior a person should experience and verify that behavior. These are proposed checks, not results from our own Claude test.

Keep the release version attached to the claim

This entry covers the original June 2024 release. If a comparison uses a later model with a similar name, it needs its own record. Otherwise a historical chart can accidentally make early versions look more capable by giving them features or scores that arrived months later.

For coding results, record the repository, the problem, the tools and the number of attempts. A model that suggests a correct edit in a chat window is being tested differently from an agent that searches files, runs tests and retries. Both can be useful, but the result describes the whole evaluation setup. The benchmark name alone is not enough to establish comparability.

FOLLOW THE EVIDENCE

What to watch next

Changes that would make this story worth revisiting:

  • Releases that change how people inspect, edit or test generated work.
  • Coding evaluations with reproducible tasks and clearly identified model versions.

Questions about this milestone

Does a working preview mean a website is ready to launch?

No. A preview demonstrates the visible interface. Saving data, handling errors, accessibility and behavior across devices require their own checks.

Why does this page discuss Artifacts separately?

A workspace feature can improve a user’s workflow without being a property of the model itself. Separating the two makes the history easier to understand.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models. The announcement page currently displays June 21, 2024.

Anthropic: Introducing Claude 3.5 Sonnet Google Cloud: June 20 launch availability

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
130.0index points
Tested variant: Claude 3.5 SonnetSource interval not suppliedVariant date in source: Jun 20, 2024

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
1343rating points
Tested: claude-3-5-sonnet-20240620Text Arena Overall · October 8, 2026Reported interval: 1340 to 134682,419 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.