What are the token costs and operational benchmarks for Mistral Large 2?

Mistral Large 2 is priced at $2.00 per million input tokens and $6.00 per million output tokens. It features a 131.072k token context window, an average response latency of 550ms TTFT, and achieves 39% on SWE-bench Verified and 67.5% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $2.00
Output / 1M $6.00
Context Limit 131.072k
TTFT Latency 550ms
SWE-bench 39%
Throughput 90 t/s

Architectural Overview

Frontier foundation model developed by Mistral AI featuring Dense Transformer architecture.

Optimal Production Use Cases

European sovereign enterprise workflows, sophisticated function-calling pipelines, and multi-language code generation.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

Mistral Large 2 $5.45

$2/1M in · $6/1M out

GPT-5.6 Luna $67.50

$0.18/1M in · $0.72/1M out

Projected Monthly Cost Reduction
$62.05 / mo
(91.9% lower cost)

Head-to-Head Comparisons Involving Mistral Large 2

Versus Comparison

Mistral Large 2 vs Claude 3.5 Haiku

Claude 3.5 Haiku is 60.0% cheaper for input tokens ($0.80 vs. $2.00 per 1M tokens) and $4.00 vs. $6.00 for output tokens (2.5x cost difference). In terms of operational performance, Claude 3.5 Haiku delivers faster response latency with 140ms TTFT (410ms faster than Mistral Large 2). Claude 3.5 Haiku leads coding benchmarks at 40.6% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Claude 3.5 Sonnet

Mistral Large 2 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, Claude 3.5 Sonnet delivers faster response latency with 320ms TTFT (230ms faster than Mistral Large 2). Claude 3.5 Sonnet leads coding benchmarks at 63.7% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Claude 3.7 Sonnet

Mistral Large 2 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, Mistral Large 2 delivers faster response latency with 550ms TTFT (100ms faster than Claude 3.7 Sonnet). Claude 3.7 Sonnet leads coding benchmarks at 70.3% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Claude Opus 5

Mistral Large 2 is 60.0% cheaper for input tokens ($2.00 vs. $5.00 per 1M tokens) and $6.00 vs. $25.00 for output tokens (2.5x cost difference). In terms of operational performance, Claude Opus 5 delivers faster response latency with 340ms TTFT (210ms faster than Mistral Large 2). Claude Opus 5 leads coding benchmarks at 82.4% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Codestral 25.01

Codestral 25.01 is 85.0% cheaper for input tokens ($0.30 vs. $2.00 per 1M tokens) and $0.90 vs. $6.00 for output tokens (6.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (400ms faster than Mistral Large 2). Codestral 25.01 leads coding benchmarks at 44.2% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Composer 2.5

Composer 2.5 is 10.0% cheaper for input tokens ($1.80 vs. $2.00 per 1M tokens) and $7.20 vs. $6.00 for output tokens (1.1x cost difference). In terms of operational performance, Composer 2.5 delivers faster response latency with 180ms TTFT (370ms faster than Mistral Large 2). Composer 2.5 leads coding benchmarks at 74.6% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs DeepSeek-R1

DeepSeek-R1 is 72.5% cheaper for input tokens ($0.55 vs. $2.00 per 1M tokens) and $2.19 vs. $6.00 for output tokens (3.6x cost difference). In terms of operational performance, Mistral Large 2 delivers faster response latency with 550ms TTFT (1250ms faster than DeepSeek-R1). DeepSeek-R1 leads coding benchmarks at 49.2% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs DeepSeek-V3

DeepSeek-V3 is 93.0% cheaper for input tokens ($0.14 vs. $2.00 per 1M tokens) and $0.28 vs. $6.00 for output tokens (14.3x cost difference). In terms of operational performance, DeepSeek-V3 delivers faster response latency with 340ms TTFT (210ms faster than Mistral Large 2). DeepSeek-V3 leads coding benchmarks at 42.0% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs DeepSeek-V4 Flash

DeepSeek-V4 Flash is 93.0% cheaper for input tokens ($0.14 vs. $2.00 per 1M tokens) and $0.56 vs. $6.00 for output tokens (14.3x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (400ms faster than Mistral Large 2). DeepSeek-V4 Flash leads coding benchmarks at 62.4% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Fable 5

Both models share identical input pricing at $2.00 per 1M tokens. In terms of operational performance, Fable 5 delivers faster response latency with 240ms TTFT (310ms faster than Mistral Large 2). Fable 5 leads coding benchmarks at 58.0% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Gemini 2.0 Flash

Gemini 2.0 Flash is 95.0% cheaper for input tokens ($0.10 vs. $2.00 per 1M tokens) and $0.40 vs. $6.00 for output tokens (20.0x cost difference). In terms of operational performance, Gemini 2.0 Flash delivers faster response latency with 380ms TTFT (170ms faster than Mistral Large 2). Gemini 2.0 Flash leads coding benchmarks at 48.0% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Gemini 3.7 Flash

Gemini 3.7 Flash is 96.0% cheaper for input tokens ($0.08 vs. $2.00 per 1M tokens) and $0.32 vs. $6.00 for output tokens (25.0x cost difference). In terms of operational performance, Gemini 3.7 Flash delivers faster response latency with 75ms TTFT (475ms faster than Mistral Large 2). Gemini 3.7 Flash leads coding benchmarks at 68.2% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs GLM 5.3 Flash

GLM 5.3 Flash is 92.5% cheaper for input tokens ($0.15 vs. $2.00 per 1M tokens) and $0.50 vs. $6.00 for output tokens (13.3x cost difference). In terms of operational performance, GLM 5.3 Flash delivers faster response latency with 130ms TTFT (420ms faster than Mistral Large 2). GLM 5.3 Flash leads coding benchmarks at 54.2% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs OpenAI GPT-4o

Mistral Large 2 is 20.0% cheaper for input tokens ($2.00 vs. $2.50 per 1M tokens) and $6.00 vs. $10.00 for output tokens (1.2x cost difference). In terms of operational performance, OpenAI GPT-4o delivers faster response latency with 280ms TTFT (270ms faster than Mistral Large 2). Mistral Large 2 leads coding benchmarks at 39.0% SWE-bench vs. 38.8% for OpenAI GPT-4o.

Versus Comparison

Mistral Large 2 vs GPT-5.6 Luna

GPT-5.6 Luna is 91.0% cheaper for input tokens ($0.18 vs. $2.00 per 1M tokens) and $0.72 vs. $6.00 for output tokens (11.1x cost difference). In terms of operational performance, GPT-5.6 Luna delivers faster response latency with 90ms TTFT (460ms faster than Mistral Large 2). GPT-5.6 Luna leads coding benchmarks at 48.5% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs GPT-5.6 Sol

Mistral Large 2 is 75.0% cheaper for input tokens ($2.00 vs. $8.00 per 1M tokens) and $6.00 vs. $32.00 for output tokens (4.0x cost difference). In terms of operational performance, GPT-5.6 Sol delivers faster response latency with 420ms TTFT (130ms faster than Mistral Large 2). GPT-5.6 Sol leads coding benchmarks at 79.5% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs GPT-5.6 Terra

GPT-5.6 Terra is 25.0% cheaper for input tokens ($1.50 vs. $2.00 per 1M tokens) and $6.00 vs. $6.00 for output tokens (1.3x cost difference). In terms of operational performance, GPT-5.6 Terra delivers faster response latency with 210ms TTFT (340ms faster than Mistral Large 2). GPT-5.6 Terra leads coding benchmarks at 65.4% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Grok 3

Mistral Large 2 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, Mistral Large 2 delivers faster response latency with 550ms TTFT (300ms faster than Grok 3). Grok 3 leads coding benchmarks at 58.5% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs xAI Grok 4.6

Both models share identical input pricing at $2.00 per 1M tokens. In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (270ms faster than Mistral Large 2). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Llama 3.3 70B Instruct

Llama 3.3 70B Instruct is 91.0% cheaper for input tokens ($0.18 vs. $2.00 per 1M tokens) and $0.40 vs. $6.00 for output tokens (11.1x cost difference). In terms of operational performance, Llama 3.3 70B Instruct delivers faster response latency with 420ms TTFT (130ms faster than Mistral Large 2). Mistral Large 2 leads coding benchmarks at 39.0% SWE-bench vs. 38.8% for Llama 3.3 70B Instruct.

Versus Comparison

Mistral Large 2 vs OpenAI o1

Mistral Large 2 is 86.7% cheaper for input tokens ($2.00 vs. $15.00 per 1M tokens) and $6.00 vs. $60.00 for output tokens (7.5x cost difference). In terms of operational performance, Mistral Large 2 delivers faster response latency with 550ms TTFT (300ms faster than OpenAI o1). OpenAI o1 leads coding benchmarks at 48.9% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs o3-mini

o3-mini is 45.0% cheaper for input tokens ($1.10 vs. $2.00 per 1M tokens) and $4.40 vs. $6.00 for output tokens (1.8x cost difference). In terms of operational performance, Mistral Large 2 delivers faster response latency with 550ms TTFT (650ms faster than o3-mini). o3-mini leads coding benchmarks at 49.3% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Microsoft Phi-4 (14B)

Microsoft Phi-4 (14B) is 94.0% cheaper for input tokens ($0.12 vs. $2.00 per 1M tokens) and $0.36 vs. $6.00 for output tokens (16.7x cost difference). In terms of operational performance, Microsoft Phi-4 (14B) delivers faster response latency with 110ms TTFT (440ms faster than Mistral Large 2). Microsoft Phi-4 (14B) leads coding benchmarks at 42.1% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Qwen 2.5 72B Instruct

Qwen 2.5 72B Instruct is 82.5% cheaper for input tokens ($0.35 vs. $2.00 per 1M tokens) and $0.40 vs. $6.00 for output tokens (5.7x cost difference). In terms of operational performance, Qwen 2.5 72B Instruct delivers faster response latency with 420ms TTFT (130ms faster than Mistral Large 2). Qwen 2.5 72B Instruct leads coding benchmarks at 44.0% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Qwen 2.5 Max

Qwen 2.5 Max is 86.0% cheaper for input tokens ($0.28 vs. $2.00 per 1M tokens) and $0.84 vs. $6.00 for output tokens (7.1x cost difference). In terms of operational performance, Qwen 2.5 Max delivers faster response latency with 480ms TTFT (70ms faster than Mistral Large 2). Qwen 2.5 Max leads coding benchmarks at 44.2% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

Mistral Large 2 vs Qwen 3.8 Flash Next

Qwen 3.8 Flash Next is 94.0% cheaper for input tokens ($0.12 vs. $2.00 per 1M tokens) and $0.48 vs. $6.00 for output tokens (16.7x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 120ms TTFT (430ms faster than Mistral Large 2). Qwen 3.8 Flash Next leads coding benchmarks at 56.8% SWE-bench vs. 39.0% for Mistral Large 2.

Frequently Asked Questions & Query Fan-Out

How much does Mistral Large 2 cost per 1M tokens?

Mistral Large 2 costs $2.00 per million prompt (input) tokens and $6.00 per million completion (output) tokens.

What is the context window for Mistral Large 2?

Mistral Large 2 supports a maximum context window of 131,072 tokens, with a maximum single-generation output of 8,192 tokens.

What are the primary use cases for Mistral Large 2?

European sovereign enterprise workflows, sophisticated function-calling pipelines, and multi-language code generation.