What are the token costs and operational benchmarks for xAI Grok 4.6?

xAI Grok 4.6 is priced at $2.00 per million input tokens and $6.00 per million output tokens. It features a 512k token context window, an average response latency of 280ms TTFT, and achieves 76.8% on SWE-bench Verified and 89.4% on MMLU-Pro.
Verified daily via automated API latency tests and official documentation.
Input / 1M $2.00
Output / 1M $6.00
Context Limit 512k
TTFT Latency 280ms
SWE-bench 76.8%
Throughput 85 t/s

Architectural Overview

Frontier foundation model developed by xAI featuring Colossus 200k-Cluster Real-Time Frontier architecture.

Optimal Production Use Cases

Frontier real-time web & social intelligence, complex STEM proofs, competitive coding, and deep factual synthesis.

Interactive Monthly Token Economics & ROI Forecaster

Model your expected production workload across prompt (input) and completion (output) tokens.

Live Calculation Engine
10.0M Tokens
100K 50M 250M 500M+
2.5M Tokens
100K 10M 50M 100M+

Estimated Monthly Spend

xAI Grok 4.6 $5.45

$2/1M in · $6/1M out

GPT-5.6 Luna $67.50

$0.18/1M in · $0.72/1M out

Projected Monthly Cost Reduction
$62.05 / mo
(91.9% lower cost)

Head-to-Head Comparisons Involving xAI Grok 4.6

Versus Comparison

xAI Grok 4.6 vs Claude 3.5 Haiku

Claude 3.5 Haiku is 60.0% cheaper for input tokens ($0.80 vs. $2.00 per 1M tokens) and $4.00 vs. $6.00 for output tokens (2.5x cost difference). In terms of operational performance, Claude 3.5 Haiku delivers faster response latency with 140ms TTFT (140ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 40.6% for Claude 3.5 Haiku.

Versus Comparison

xAI Grok 4.6 vs Claude 3.5 Sonnet

xAI Grok 4.6 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (40ms faster than Claude 3.5 Sonnet). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 63.7% for Claude 3.5 Sonnet.

Versus Comparison

xAI Grok 4.6 vs Claude 3.7 Sonnet

xAI Grok 4.6 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (370ms faster than Claude 3.7 Sonnet). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 70.3% for Claude 3.7 Sonnet.

Versus Comparison

xAI Grok 4.6 vs Claude Opus 5

xAI Grok 4.6 is 60.0% cheaper for input tokens ($2.00 vs. $5.00 per 1M tokens) and $6.00 vs. $25.00 for output tokens (2.5x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (60ms faster than Claude Opus 5). Claude Opus 5 leads coding benchmarks at 82.4% SWE-bench vs. 76.8% for xAI Grok 4.6.

Versus Comparison

xAI Grok 4.6 vs Codestral 25.01

Codestral 25.01 is 85.0% cheaper for input tokens ($0.30 vs. $2.00 per 1M tokens) and $0.90 vs. $6.00 for output tokens (6.7x cost difference). In terms of operational performance, Codestral 25.01 delivers faster response latency with 150ms TTFT (130ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 44.2% for Codestral 25.01.

Versus Comparison

xAI Grok 4.6 vs Composer 2.5

Composer 2.5 is 10.0% cheaper for input tokens ($1.80 vs. $2.00 per 1M tokens) and $7.20 vs. $6.00 for output tokens (1.1x cost difference). In terms of operational performance, Composer 2.5 delivers faster response latency with 180ms TTFT (100ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 74.6% for Composer 2.5.

Versus Comparison

xAI Grok 4.6 vs DeepSeek-R1

DeepSeek-R1 is 72.5% cheaper for input tokens ($0.55 vs. $2.00 per 1M tokens) and $2.19 vs. $6.00 for output tokens (3.6x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (1520ms faster than DeepSeek-R1). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 49.2% for DeepSeek-R1.

Versus Comparison

xAI Grok 4.6 vs DeepSeek-V3

DeepSeek-V3 is 93.0% cheaper for input tokens ($0.14 vs. $2.00 per 1M tokens) and $0.28 vs. $6.00 for output tokens (14.3x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (60ms faster than DeepSeek-V3). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 42.0% for DeepSeek-V3.

Versus Comparison

xAI Grok 4.6 vs DeepSeek-V4 Flash

DeepSeek-V4 Flash is 93.0% cheaper for input tokens ($0.14 vs. $2.00 per 1M tokens) and $0.56 vs. $6.00 for output tokens (14.3x cost difference). In terms of operational performance, DeepSeek-V4 Flash delivers faster response latency with 150ms TTFT (130ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 62.4% for DeepSeek-V4 Flash.

Versus Comparison

xAI Grok 4.6 vs Fable 5

Both models share identical input pricing at $2.00 per 1M tokens. In terms of operational performance, Fable 5 delivers faster response latency with 240ms TTFT (40ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 58.0% for Fable 5.

Versus Comparison

xAI Grok 4.6 vs Gemini 2.0 Flash

Gemini 2.0 Flash is 95.0% cheaper for input tokens ($0.10 vs. $2.00 per 1M tokens) and $0.40 vs. $6.00 for output tokens (20.0x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (100ms faster than Gemini 2.0 Flash). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 48.0% for Gemini 2.0 Flash.

Versus Comparison

xAI Grok 4.6 vs Gemini 3.7 Flash

Gemini 3.7 Flash is 96.0% cheaper for input tokens ($0.08 vs. $2.00 per 1M tokens) and $0.32 vs. $6.00 for output tokens (25.0x cost difference). In terms of operational performance, Gemini 3.7 Flash delivers faster response latency with 75ms TTFT (205ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 68.2% for Gemini 3.7 Flash.

Versus Comparison

xAI Grok 4.6 vs GLM 5.3 Flash

GLM 5.3 Flash is 92.5% cheaper for input tokens ($0.15 vs. $2.00 per 1M tokens) and $0.50 vs. $6.00 for output tokens (13.3x cost difference). In terms of operational performance, GLM 5.3 Flash delivers faster response latency with 130ms TTFT (150ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 54.2% for GLM 5.3 Flash.

Versus Comparison

xAI Grok 4.6 vs OpenAI GPT-4o

xAI Grok 4.6 is 20.0% cheaper for input tokens ($2.00 vs. $2.50 per 1M tokens) and $6.00 vs. $10.00 for output tokens (1.2x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (0ms faster than OpenAI GPT-4o). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 38.8% for OpenAI GPT-4o.

Versus Comparison

xAI Grok 4.6 vs GPT-5.6 Luna

GPT-5.6 Luna is 91.0% cheaper for input tokens ($0.18 vs. $2.00 per 1M tokens) and $0.72 vs. $6.00 for output tokens (11.1x cost difference). In terms of operational performance, GPT-5.6 Luna delivers faster response latency with 90ms TTFT (190ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 48.5% for GPT-5.6 Luna.

Versus Comparison

xAI Grok 4.6 vs GPT-5.6 Sol

xAI Grok 4.6 is 75.0% cheaper for input tokens ($2.00 vs. $8.00 per 1M tokens) and $6.00 vs. $32.00 for output tokens (4.0x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (140ms faster than GPT-5.6 Sol). GPT-5.6 Sol leads coding benchmarks at 79.5% SWE-bench vs. 76.8% for xAI Grok 4.6.

Versus Comparison

xAI Grok 4.6 vs GPT-5.6 Terra

GPT-5.6 Terra is 25.0% cheaper for input tokens ($1.50 vs. $2.00 per 1M tokens) and $6.00 vs. $6.00 for output tokens (1.3x cost difference). In terms of operational performance, GPT-5.6 Terra delivers faster response latency with 210ms TTFT (70ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 65.4% for GPT-5.6 Terra.

Versus Comparison

xAI Grok 4.6 vs Grok 3

xAI Grok 4.6 is 33.3% cheaper for input tokens ($2.00 vs. $3.00 per 1M tokens) and $6.00 vs. $15.00 for output tokens (1.5x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (570ms faster than Grok 3). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 58.5% for Grok 3.

Versus Comparison

xAI Grok 4.6 vs Llama 3.3 70B Instruct

Llama 3.3 70B Instruct is 91.0% cheaper for input tokens ($0.18 vs. $2.00 per 1M tokens) and $0.40 vs. $6.00 for output tokens (11.1x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (140ms faster than Llama 3.3 70B Instruct). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 38.8% for Llama 3.3 70B Instruct.

Versus Comparison

xAI Grok 4.6 vs Mistral Large 2

Both models share identical input pricing at $2.00 per 1M tokens. In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (270ms faster than Mistral Large 2). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 39.0% for Mistral Large 2.

Versus Comparison

xAI Grok 4.6 vs OpenAI o1

xAI Grok 4.6 is 86.7% cheaper for input tokens ($2.00 vs. $15.00 per 1M tokens) and $6.00 vs. $60.00 for output tokens (7.5x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (570ms faster than OpenAI o1). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 48.9% for OpenAI o1.

Versus Comparison

xAI Grok 4.6 vs o3-mini

o3-mini is 45.0% cheaper for input tokens ($1.10 vs. $2.00 per 1M tokens) and $4.40 vs. $6.00 for output tokens (1.8x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (920ms faster than o3-mini). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 49.3% for o3-mini.

Versus Comparison

xAI Grok 4.6 vs Microsoft Phi-4 (14B)

Microsoft Phi-4 (14B) is 94.0% cheaper for input tokens ($0.12 vs. $2.00 per 1M tokens) and $0.36 vs. $6.00 for output tokens (16.7x cost difference). In terms of operational performance, Microsoft Phi-4 (14B) delivers faster response latency with 110ms TTFT (170ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 42.1% for Microsoft Phi-4 (14B).

Versus Comparison

xAI Grok 4.6 vs Qwen 2.5 72B Instruct

Qwen 2.5 72B Instruct is 82.5% cheaper for input tokens ($0.35 vs. $2.00 per 1M tokens) and $0.40 vs. $6.00 for output tokens (5.7x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (140ms faster than Qwen 2.5 72B Instruct). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 44.0% for Qwen 2.5 72B Instruct.

Versus Comparison

xAI Grok 4.6 vs Qwen 2.5 Max

Qwen 2.5 Max is 86.0% cheaper for input tokens ($0.28 vs. $2.00 per 1M tokens) and $0.84 vs. $6.00 for output tokens (7.1x cost difference). In terms of operational performance, xAI Grok 4.6 delivers faster response latency with 280ms TTFT (200ms faster than Qwen 2.5 Max). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 44.2% for Qwen 2.5 Max.

Versus Comparison

xAI Grok 4.6 vs Qwen 3.8 Flash Next

Qwen 3.8 Flash Next is 94.0% cheaper for input tokens ($0.12 vs. $2.00 per 1M tokens) and $0.48 vs. $6.00 for output tokens (16.7x cost difference). In terms of operational performance, Qwen 3.8 Flash Next delivers faster response latency with 120ms TTFT (160ms faster than xAI Grok 4.6). xAI Grok 4.6 leads coding benchmarks at 76.8% SWE-bench vs. 56.8% for Qwen 3.8 Flash Next.

Frequently Asked Questions & Query Fan-Out

How much does xAI Grok 4.6 cost per 1M tokens?

xAI Grok 4.6 costs $2.00 per million prompt (input) tokens and $6.00 per million completion (output) tokens.

What is the context window for xAI Grok 4.6?

xAI Grok 4.6 supports a maximum context window of 512,000 tokens, with a maximum single-generation output of 65,536 tokens.

What are the primary use cases for xAI Grok 4.6?

Frontier real-time web & social intelligence, complex STEM proofs, competitive coding, and deep factual synthesis.