by DeepSeek · DeepSeek R1 family · best for exposed-CoT reasoning at a fraction of o-series cost
DeepSeek R1 is the category-defining open-weights reasoning model — the release that broke the "frontier reasoning is expensive" assumption in early 2025 and forced US labs to respond on price. It is reasoning-first: every response includes a full, visible chain-of-thought before the final answer, which is a genuine differentiator for interpretability, distillation, and audit pipelines. The R1-0528 refresh (2025-05-28) added function calling and JSON output and pushed AIME 2025 to 87.5% and GPQA Diamond to 81.0. Built on the V3 671B/37B MoE backbone with RL post-training, open-weights under MIT. The single sentence a buyer needs: when reasoning is the whole job and you want the chain-of-thought exposed, R1 delivers frontier-class math and science at a fraction of o-series cost.
| Benchmark | Score | Source |
|---|---|---|
| Humanity's Last Exam | 17.7% | huggingface.co 2025-05-28T00:00:00.000Z |
| MMLU | 93.4% | huggingface.co 2025-05-28T00:00:00.000Z |
| MMLU-Pro | 85% | huggingface.co 2025-05-28T00:00:00.000Z |
| SimpleQA | 27.8% | huggingface.co 2025-05-28T00:00:00.000Z |
| AIME 2025 | 87.5% | huggingface.co 2025-05-28T00:00:00.000Z |
| TAU-bench | 63.9% | huggingface.co 2025-05-28T00:00:00.000Z |
| LMArena Elo | 1382 | artificialanalysis.ai 2025-05-28T00:00:00.000Z |
| GPQA Diamond | 81% | huggingface.co 2025-05-28T00:00:00.000Z |
| LiveCodeBench | 73.3% | huggingface.co 2025-05-28T00:00:00.000Z |
| Aider Polyglot | 71.6% | huggingface.co 2025-05-28T00:00:00.000Z |
| SWE-bench Verified | 57.6% | huggingface.co 2025-05-28T00:00:00.000Z |
| Artificial Analysis Index | 68 | artificialanalysis.ai 2025-05-28T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“R1 broke the frontier-reasoning price ceiling and gave me an exposed-CoT artifact I can audit — but hybrid models have absorbed most general use.”
R1 was the model that broke the "frontier reasoning is expensive" assumption in early 2025 and forced US labs to respond on price. Strategically, the exposed chain-of-thought is more than a quality story — it lets enterprises build pipelines that consume the reasoning separately from the answer, valuable for audit and distillation. By mid-2026 the hybrid-mode models (V3.1/V3.2/V4) have largely absorbed R1's general-agent use case; R1 remains the right pick when reasoning is the whole job and you want the trace exposed. Sovereignty considerations are identical to the family, and the NIST CAISI safety flag is worth weighing for sensitive deployments.
“R1 is the single most disruptive reasoning release of the cycle — it reset market expectations for what frontier reasoning should cost.”
R1's strategic significance is hard to overstate: it was the open-weights reasoning model that, at roughly 1/30th of contemporaneous o-series pricing, reset the entire market's expectations and triggered a global re-rating of DeepSeek as a lab. Artificial Analysis tied DeepSeek as the #2 lab and undisputed open-weights leader on the back of the 0528 update. Its differentiation — full exposed CoT plus frontier math at commodity price — created a distinct category position no closed o-series model matched. The strategic decay by mid-2026 is that DeepSeek's own hybrid models cannibalize the generalist use case, narrowing R1 to the reasoning-specialist and CoT-artifact niche.
“R1's launch landed at ~1/30th of o-series pricing for comparable AIME/GPQA — the most disruptive pricing event in the reasoning category in years.”
R1's launch pricing ($0.55 in / $2.19 out) landed at roughly 1/30th of the contemporaneous OpenAI o1 rate for comparable AIME/GPQA scores — the single most disruptive pricing event in the reasoning-model category. Cache-hit input at $0.14/M makes repeated-reasoning workloads (tutoring, agent loops) shockingly cheap. The critical line-item caution is reasoning-token volume: R1 burns output tokens at much higher rates than chat models because the chain-of-thought counts, and 0528 nearly doubled per-question reasoning to ~23K tokens — budget on roughly 3-5x the output volume of a comparable non-reasoning workload. Even so, the intelligence-per-dollar on hard reasoning is exceptional.
“The exposed reasoning trace is a developer delight — log it, parse it, display it — and 0528 finally made function calling first-class.”
For a builder, the 0528 refresh fixed the practical complaints: function calling and JSON mode are first-class. The exposed reasoning trace is a genuine delight — you can log it, store it, parse it, or stream it to a "thinking" UI. Open weights and the distilled Qwen3-8B make local development and edge deployment realistic, and the OpenAI-compatible endpoint keeps integration simple. The main friction is latency — reasoning queries take real time, so UX must be designed around streaming the thought process. Tool-call reliability is good but not Claude-grade (Tau-Bench Retail 63.9). No parallel tool calls, no batch API.
“On hard problems R1's visible thinking is genuinely impressive and builds trust — but it's overkill, and slow, for everyday questions.”
For end users on free or low-cost chat, R1's reasoning mode is genuinely impressive on hard problems — it solves AIME problems and GPQA-style science questions that free GPT/Claude tiers struggle with, and the visible thinking process is novel and increases trust in the answer. The downsides are latency (10-30 seconds on hard queries) and that R1 is overkill for everyday questions, where its reasoning-first style is slower and stiffer than a generalist. Best deployed as a "deep think" toggle rather than the default. Content policy follows DeepSeek norms. As a free option via the DeepSeek UI's DeepThink toggle, the value on hard problems is excellent.
“Frontier-class reasoning at commodity price is real — but it's narrow, slow, burns output tokens, and NIST flagged it on safety.”
R1's reasoning scores are well-documented on its own model card and independently tracked by Artificial Analysis, so the headline holds — this is not benchmark theater. The honest caveats are about scope and cost shape. R1 is a specialist: coding (SWE-bench 57.6) and general agent tool use are middling, and creative/instruction work is weak. The chain-of-thought that is its differentiator is also a cost multiplier — ~23K reasoning tokens per hard question means output bills can dwarf a chat model's. And the NIST CAISI evaluation flagged DeepSeek models, R1 included, as more susceptible on certain safety/hijacking evals than US frontier peers — a real consideration for security-sensitive use. The value is genuine; the asterisks are scope, latency, token economics, and safety posture.
The full research notes behind this review — verified against primary sources.
R1-0528 is built on the DeepSeek-V3 base architecture: a 671B-parameter DeepSeekMoE model (685B total counting the Multi-Token-Prediction module), ~37B activated per token, 256 routed experts, 61 layers, using Multi-head Latent Attention (MLA). Artificial Analysis confirms R1-0528 is a post-training update with no change to the V3/R1 architecture — the gains come from a reinforcement-learning post-training pipeline focused on deliberate reasoning. The defining behavior is always-on reasoning with a fully exposed chain-of-thought (reasoning_content), which downstream pipelines can consume as a separate artifact. Open weights are on Hugging Face under MIT; vocab size is 129,280. Text-only.
R1's strongholds justify its reasoning (9.5) and math (9.5) scores: AIME 2025 87.5%, GPQA Diamond 81.0, MMLU-Redux 93.4, and frontier-class competition-math performance. The exposed chain-of-thought is itself a capability — usable for distillation training data, audit trails, and AI tutoring. Coding (7.5) is strong but not specialized — SWE-bench Verified 57.6, LiveCodeBench 73.3, Aider-Polyglot 71.6 — so for pure coding the V3.1+/V4 models fit better. Agentic (7.0) and function-calling (7.0) work post-0528 but Tau-Bench (Retail 63.9 / Airline 53.5) and BFCL multi-turn (37.0) show tool use is competent rather than class-leading. Multilingual (8.0) is strong in English and Chinese. Long-context (7.0) is fine within 128K. Vision, OCR, and real-time data are zero. Creative writing (6.5) and instruction-following (7.5) are not its sharpest modes — R1 is an analyst, not a stylist. Safety calibration (6.5) reflects family norms; NIST CAISI flagged higher susceptibility on some safety evals.
| Benchmark | Score | vs Predecessor | vs Top Competitor | Source |
|---|---|---|---|---|
| AIME 2025 (Pass@1) | 87.5% | +17.5 vs R1-initial (70) | within ~4 pts of o-series | HF card |
| GPQA Diamond (Pass@1) | 81.0 | up significantly | competitive with o3/o4-mini | HF card |
| MMLU-Redux (EM) | 93.4 | n/a | frontier-class | HF card |
| MMLU-Pro (EM) | 85.0 | n/a | within ~2 pts of frontier | HF card |
| LiveCodeBench (Pass@1) | 73.3 | up from 63.5 | strong, not specialized | HF card |
| Aider-Polyglot | 71.6 | n/a | competitive | HF card |
| SWE-bench Verified | 57.6 | n/a | trails frontier coders (~80) | HF card |
| HLE (Pass@1) | 17.7 | n/a | mid-pack reasoning | HF card |
| Tau-Bench (Retail) | 63.9 | n/a | competent tool use | HF card |
| LMArena Elo | 1382 | top-tier at launch | top-10 globally | AA |
| Artificial Analysis Index | 68 | up from 60 (R1-initial) | tied #2 lab at release | AA |
MATH-500 is not in DeepSeek's 0528 table and is left null. The AA Intelligence Index value (68) is on the index version current at R1-0528's release; AA periodically reindexes.
R1 is the slow tier by design. Reasoning-first generation means hard queries take real time — 10-30 seconds is common — and R1-0528 averages ~23K reasoning tokens per AIME question, so both latency and output billing scale with problem difficulty. UX should stream the thought process or show a "thinking" indicator. Latency tier: slow.
| Surface | Cost | Notes |
|---|---|---|
| API input (cache miss) | $0.55 / 1M tok | |
| API input (cache hit) | $0.14 / 1M tok | ~75% discount |
| API output | $2.19 / 1M tok | includes chain-of-thought tokens |
| Direct UI | Free | chat.deepseek.com (R1 / DeepThink toggle) |
| Open weights | $0 | HF download; 8x H200-class node. Distilled Qwen3-8B runs on a single 16GB GPU. |
| Rate limits | Standard (GA) tier |
First-party OpenAI-compatible API at api.deepseek.com (PRC-hosted), with the deepseek-reasoner alias historically pointing at R1. Open weights on Hugging Face under MIT: self-hostable at ~685B/37B MoE (8x H200-class node at FP8, ~400GB+ VRAM), with INT4/GGUF community quants. The official distilled DeepSeek-R1-0528-Qwen3-8B runs reasoning on a single ~16GB GPU for edge/cost-sensitive deployments. Broadly served by neutral inference providers — OpenRouter, DeepInfra, Novita, Fireworks, Together, Hyperbolic, SambaNova. No first-party managed-cloud offering.
Same posture as the family: PRC data storage under Chinese law, trains-on-input by default (de-identified), no documented API opt-out, no SOC2/HIPAA/GDPR/ISO27001 on the first-party service. Content moderation follows PRC norms. Notably, the NIST Center for AI Standards and Innovation (CAISI) evaluation flagged DeepSeek models, including R1, as more susceptible than US frontier models on certain safety/agent-hijacking evals — a relevant data point for security-sensitive buyers. The MIT open weights remain the path to an in-boundary, compliance-controlled deployment.
OpenAI-compatible API with Python/TypeScript SDKs, LangChain / LlamaIndex / Vercel AI SDK integrations, and very broad serving across OpenRouter, DeepInfra, Novita, Fireworks, Together, Hyperbolic, and SambaNova. Used by Perplexity and coding tools (Kilo Code). As the model that reset reasoning-price expectations, R1 has mainstream adoption and one of the most-downloaded open-weights footprints on Hugging Face.
R1 always reasons and exposes the full chain-of-thought as a separate artifact, which is uniquely valuable for audit, distillation, and tutoring. For general agents where you only sometimes need thinking, a hybrid model is usually a better fit.
You can log, parse, store, or display the reasoning independently of the answer — useful for debugging, building distillation datasets, AI-tutoring step-by-step views, and audit trails.
Reasoning tokens count as output, and R1-0528 averages ~23K reasoning tokens per hard question. Budget on roughly 3-5x the output volume of a comparable non-reasoning workload.
The full model needs an 8x H200-class node, but the official distilled DeepSeek-R1-0528-Qwen3-8B runs reasoning on a single ~16GB GPU.
Decent but not specialized — SWE-bench 57.6. Use V3.1+/V4 for pure coding.
The NIST CAISI evaluation flagged DeepSeek models, including R1, as more susceptible on certain safety/hijacking evals than US frontier models. Weigh this for security-sensitive deployments and prefer in-boundary self-host.
stronger on the hardest reasoning, no exposed chain-of-thought, no open weights; materially more expensive (o3) or closer in price but closed (o4-mini).
direct China-origin reasoning peers with open weights and comparable benchmarks; Qwen is more generalist, R1 exposes its CoT more cleanly.
best generalist reasoning with thinking, summary-only CoT; ~10x more expensive and closed.
Primary references used to verify this review.
Last verified 2026-05-27