What are the token costs and operational benchmarks for DeepSeek-R1?

DeepSeek-R1 is priced at $0.55 per million input tokens and $2.19 per million output tokens. It features a 65.536k token context window, an average response latency of 1800ms TTFT, and achieves 49.2% on SWE-bench Verified and 84% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $0.55
Output / 1M $2.19
Context Limit 65.536k
TTFT Latency 1800ms
SWE-bench 49.2%
Throughput 35 t/s

Architectural Overview

Frontier foundation model developed by DeepSeek featuring MoE Reasoning architecture.

Optimal Production Use Cases

Open-weights chain-of-thought reasoning for complex mathematics, algorithmic programming, and scientific research.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

DeepSeek-R1 $5.45

$0.55/1M in · $2.19/1M out

GPT-5.6 Luna $67.50

$0.18/1M in · $0.72/1M out

Projected Monthly Cost Reduction
$62.05 / mo
(91.9% lower cost)

Head-to-Head Comparisons Involving DeepSeek-R1

Versus Comparison

DeepSeek-R1 vs Claude 3.5 Haiku

DeepSeek-R1 is 31.2% cheaper for input tokens ($0.55 vs. $0.80 per 1M tokens) and $2.19 vs. $4.00 for output tokens (1.5x cost difference). In terms of operational performance, Claude 3.5 Haiku delivers faster response latency with 140ms TTFT (1660ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 40.6% for Claude 3.5 Haiku.

Versus Comparison

DeepSeek-R1 vs Claude 3.5 Sonnet

DeepSeek-R1 is 81.7% cheaper for input tokens ($0.55 vs. $3.00 per 1M tokens) and $2.19 vs. $15.00 for output tokens (5.5x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (1480ms faster than DeepSeek-R1). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs Claude 3.7 Sonnet

DeepSeek-R1 is 81.7% cheaper for input tokens ($0.55 vs. $3.00 per 1M tokens) and $2.19 vs. $15.00 for output tokens (5.5x cost difference). In terms of operational performance, Claude 3.7 Sonnet delivers faster response latency with 650ms TTFT (1150ms faster than DeepSeek-R1). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs Claude Opus 5

DeepSeek-R1 is 89.0% cheaper for input tokens ($0.55 vs. $5.00 per 1M tokens) and $2.19 vs. $25.00 for output tokens (9.1x cost difference). In terms of operational performance, Claude Opus 5 delivers faster response latency with 340ms TTFT (1460ms faster than DeepSeek-R1). Claude Opus 5 leads coding benchmarks at 82.4% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs Codestral 25.01

Codestral 25.01 is 45.5% cheaper for input tokens ($0.30 vs. $0.55 per 1M tokens) and $0.90 vs. $2.19 for output tokens (1.8x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (1650ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

DeepSeek-R1 vs Composer 2.5

DeepSeek-R1 is 69.4% cheaper for input tokens ($0.55 vs. $1.80 per 1M tokens) and $2.19 vs. $7.20 for output tokens (3.3x cost difference). In terms of operational performance, Composer 2.5 delivers faster response latency with 180ms TTFT (1620ms faster than DeepSeek-R1). Composer 2.5 leads coding benchmarks at 74.6% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs DeepSeek-V3

DeepSeek-V3 is 74.5% cheaper for input tokens ($0.14 vs. $0.55 per 1M tokens) and $0.28 vs. $2.19 for output tokens (3.9x cost difference). In terms of operational performance, DeepSeek-V3 delivers faster response latency with 340ms TTFT (1460ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 42.0% for DeepSeek-V3.

Versus Comparison

DeepSeek-R1 vs DeepSeek-V4 Flash

DeepSeek-V4 Flash is 74.5% cheaper for input tokens ($0.14 vs. $0.55 per 1M tokens) and $0.56 vs. $2.19 for output tokens (3.9x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (1650ms faster than DeepSeek-R1). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs Fable 5

DeepSeek-R1 is 72.5% cheaper for input tokens ($0.55 vs. $2.00 per 1M tokens) and $2.19 vs. $8.00 for output tokens (3.6x cost difference). In terms of operational performance, Fable 5 delivers faster response latency with 240ms TTFT (1560ms faster than DeepSeek-R1). Fable 5 leads coding benchmarks at 58.0% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs Gemini 2.0 Flash

Gemini 2.0 Flash is 81.8% cheaper for input tokens ($0.10 vs. $0.55 per 1M tokens) and $0.40 vs. $2.19 for output tokens (5.5x cost difference). In terms of operational performance, Gemini 2.0 Flash delivers faster response latency with 380ms TTFT (1420ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 48.0% for Gemini 2.0 Flash.

Versus Comparison

DeepSeek-R1 vs Gemini 3.7 Flash

Gemini 3.7 Flash is 85.5% cheaper for input tokens ($0.08 vs. $0.55 per 1M tokens) and $0.32 vs. $2.19 for output tokens (6.9x cost difference). In terms of operational performance, Gemini 3.7 Flash delivers faster response latency with 75ms TTFT (1725ms faster than DeepSeek-R1). Gemini 3.7 Flash leads coding benchmarks at 68.2% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs GLM 5.3 Flash

GLM 5.3 Flash is 72.7% cheaper for input tokens ($0.15 vs. $0.55 per 1M tokens) and $0.50 vs. $2.19 for output tokens (3.7x cost difference). In terms of operational performance, GLM 5.3 Flash delivers faster response latency with 130ms TTFT (1670ms faster than DeepSeek-R1). GLM 5.3 Flash leads coding benchmarks at 54.2% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs OpenAI GPT-4o

DeepSeek-R1 is 78.0% cheaper for input tokens ($0.55 vs. $2.50 per 1M tokens) and $2.19 vs. $10.00 for output tokens (4.5x cost difference). In terms of operational performance, OpenAI GPT-4o delivers faster response latency with 280ms TTFT (1520ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 38.8% for OpenAI GPT-4o.

Versus Comparison

DeepSeek-R1 vs GPT-5.6 Luna

GPT-5.6 Luna is 67.3% cheaper for input tokens ($0.18 vs. $0.55 per 1M tokens) and $0.72 vs. $2.19 for output tokens (3.1x cost difference). In terms of operational performance, GPT-5.6 Luna delivers faster response latency with 90ms TTFT (1710ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 48.5% for GPT-5.6 Luna.

Versus Comparison

DeepSeek-R1 vs GPT-5.6 Sol

DeepSeek-R1 is 93.1% cheaper for input tokens ($0.55 vs. $8.00 per 1M tokens) and $2.19 vs. $32.00 for output tokens (14.5x cost difference). In terms of operational performance, GPT-5.6 Sol delivers faster response latency with 420ms TTFT (1380ms faster than DeepSeek-R1). GPT-5.6 Sol leads coding benchmarks at 79.5% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs GPT-5.6 Terra

DeepSeek-R1 is 63.3% cheaper for input tokens ($0.55 vs. $1.50 per 1M tokens) and $2.19 vs. $6.00 for output tokens (2.7x cost difference). In terms of operational performance, GPT-5.6 Terra delivers faster response latency with 210ms TTFT (1590ms faster than DeepSeek-R1). GPT-5.6 Terra leads coding benchmarks at 65.4% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs Grok 3

DeepSeek-R1 is 81.7% cheaper for input tokens ($0.55 vs. $3.00 per 1M tokens) and $2.19 vs. $15.00 for output tokens (5.5x cost difference). In terms of operational performance, Grok 3 delivers faster response latency with 850ms TTFT (950ms faster than DeepSeek-R1). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs xAI Grok 4.6

DeepSeek-R1 is 72.5% cheaper for input tokens ($0.55 vs. $2.00 per 1M tokens) and $2.19 vs. $6.00 for output tokens (3.6x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (1520ms faster than DeepSeek-R1). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs Llama 3.3 70B Instruct

Llama 3.3 70B Instruct is 67.3% cheaper for input tokens ($0.18 vs. $0.55 per 1M tokens) and $0.40 vs. $2.19 for output tokens (3.1x cost difference). In terms of operational performance, Llama 3.3 70B Instruct delivers faster response latency with 420ms TTFT (1380ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 38.8% for Llama 3.3 70B Instruct.

Versus Comparison

DeepSeek-R1 vs Mistral Large 2

DeepSeek-R1 is 72.5% cheaper for input tokens ($0.55 vs. $2.00 per 1M tokens) and $2.19 vs. $6.00 for output tokens (3.6x cost difference). In terms of operational performance, Mistral Large 2 delivers faster response latency with 550ms TTFT (1250ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

DeepSeek-R1 vs OpenAI o1

DeepSeek-R1 is 96.3% cheaper for input tokens ($0.55 vs. $15.00 per 1M tokens) and $2.19 vs. $60.00 for output tokens (27.3x cost difference). In terms of operational performance, OpenAI o1 delivers faster response latency with 850ms TTFT (950ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 48.9% for OpenAI o1.

Versus Comparison

DeepSeek-R1 vs o3-mini

DeepSeek-R1 is 50.0% cheaper for input tokens ($0.55 vs. $1.10 per 1M tokens) and $2.19 vs. $4.40 for output tokens (2.0x cost difference). In terms of operational performance, o3-mini delivers faster response latency with 1200ms TTFT (600ms faster than DeepSeek-R1). o3-mini leads coding benchmarks at 49.3% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

DeepSeek-R1 vs Microsoft Phi-4 (14B)

Microsoft Phi-4 (14B) is 78.2% cheaper for input tokens ($0.12 vs. $0.55 per 1M tokens) and $0.36 vs. $2.19 for output tokens (4.6x cost difference). In terms of operational performance, Microsoft Phi-4 (14B) delivers faster response latency with 110ms TTFT (1690ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 42.1% for Microsoft Phi-4 (14B).

Versus Comparison

DeepSeek-R1 vs Qwen 2.5 72B Instruct

Qwen 2.5 72B Instruct is 36.4% cheaper for input tokens ($0.35 vs. $0.55 per 1M tokens) and $0.40 vs. $2.19 for output tokens (1.6x cost difference). In terms of operational performance, Qwen 2.5 72B Instruct delivers faster response latency with 420ms TTFT (1380ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 44.0% for Qwen 2.5 72B Instruct.

Versus Comparison

DeepSeek-R1 vs Qwen 2.5 Max

Qwen 2.5 Max is 49.1% cheaper for input tokens ($0.28 vs. $0.55 per 1M tokens) and $0.84 vs. $2.19 for output tokens (2.0x cost difference). In terms of operational performance, Qwen 2.5 Max delivers faster response latency with 480ms TTFT (1320ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 44.2% for Qwen 2.5 Max.

Versus Comparison

DeepSeek-R1 vs Qwen 3.8 Flash Next

Qwen 3.8 Flash Next is 78.2% cheaper for input tokens ($0.12 vs. $0.55 per 1M tokens) and $0.48 vs. $2.19 for output tokens (4.6x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 120ms TTFT (1680ms faster than DeepSeek-R1). Qwen 3.8 Flash Next leads coding benchmarks at 56.8% SWE-bench vs. 49.2% for DeepSeek-R1.

Frequently Asked Questions & Query Fan-Out

How much does DeepSeek-R1 cost per 1M tokens?

DeepSeek-R1 costs $0.55 per million prompt (input) tokens and $2.19 per million completion (output) tokens.

What is the context window for DeepSeek-R1?

DeepSeek-R1 supports a maximum context window of 65,536 tokens, with a maximum single-generation output of 8,192 tokens.

What are the primary use cases for DeepSeek-R1?

Open-weights chain-of-thought reasoning for complex mathematics, algorithmic programming, and scientific research.