DeepSeek
DeepSeek R1
Announced Jan 20, 2025
Reasoning with downloadable weights
What could it do?
Work on reasoning tasks such as mathematics and code using a model trained with reinforcement learning.
What changed?
DeepSeek reported strong reasoning results and released model weights, plus smaller distilled models.
Open reasoning
Downloadable weights gave developers a route to inspect and host a reasoning model themselves.
Where it fell short
A smaller distilled model is a different model. Its performance cannot be assumed to match the full R1.
Original January 20, 2025 release; the paper was first submitted on January 22. Check sources
- Access
- Downloadable weights and hosted service
- Model distinction
- Full R1 versus distilled variants
- Paper date
- January 22, 2025
What this meant in practice
Access to weights changes what developers can build and control. It does not eliminate hardware costs or establish that every deployment will reproduce a published score. Capability and accessibility deserve separate places on a progress timeline.
Compare reasoning approaches
Give two precisely identified models the same logic task and budget. Verify the final answer, record failed attempts and separate the full model from smaller variants.
Common question
Are all models with R1 in their name equivalent?
No. The release includes the main model and distinct distilled variants based on other model families.
THE USEFUL CONTEXT
Understanding DeepSeek R1 and its variants
Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.
What the research actually studied
DeepSeek’s research examines how reinforcement learning can encourage reasoning behavior, including checking and revising an approach. The original release included a main model and distinct smaller models. A family name is therefore not a sufficient description of the system used in a comparison.
A useful analogy is a product range sold under one brand. Shared branding does not make every item interchangeable. If you see an R1 result, look for the exact model name, size or deployment identifier and the evaluation setup. Without those details, the result may not describe the version available to you.
How to judge a reasoning answer
Consider a small scheduling puzzle with a known solution. Give each system the same constraints and ask for a schedule plus a short explanation. Check every constraint against the final schedule. An answer can sound systematic while assigning a person to two places at once.
Then change one constraint and repeat. Keep track of invalid solutions and correct refusals when no solution exists. This is more informative than choosing a single impressive response. A long explanation should not receive credit merely for being long; what matters is whether the result satisfies the task. This exercise is illustrative, and we have not used it to assign a score to R1.
Downloadable weights and practical usefulness
Having access to weights creates choices about deployment. Those choices still need to be evaluated in context. For a proposed local setup, write down the intended hardware, expected workload and the person responsible for keeping it running. For a hosted setup, record the provider and exact version.
A sensible comparison includes operating effort alongside answer quality. If a task only runs occasionally, the cheapest theoretical per request cost may not be the cheapest arrangement overall. Conversely, a team with an established system may value control differently. This is a framework for asking the right questions, not a pricing claim or a recommendation that one deployment is always better.
FOLLOW THE EVIDENCE
What to watch next
Changes that would make this story worth revisiting:
- Evidence about the exact R1 variant and configuration someone can actually run.
- Repeated task evaluations that report incorrect answers, response time and the cost of retries together.
Questions about this milestone
Is a smaller distilled R1 model the full R1?
No. The release distinguishes the main model from smaller distilled models. Keep their scores and deployment requirements separate.
Does more reasoning text prove a better answer?
No. Inspect the final result against explicit criteria. A concise correct result can be more useful than a long explanation with an unnoticed error.
Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models. The announcement was January 20, 2025; the paper appeared January 22.
DeepSeek: R1 release announcement DeepSeek: R1 research paperBenchmark results
A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.
Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.
Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026