What is the latency and uptime reliability of Groq LPU Inference?

Groq LPU Inference maintains an average response latency of 65ms TTFT and an operational uptime SLA of 99.98% across its infrastructure endpoints. Supported deployment regions include ["us-east-1", "us-west-1"] with direct programmatic API access.
Verified daily via automated API latency tests and official documentation.
Average TTFT Latency 65ms P90 Benchmark Window
Historical Uptime SLA 99.98% 30-Day Rolling Telemetry
Regional Deployment ["us-east-1", "us-west-1"] Multi-AZ Availability

Models Available on Groq LPU Inference

GPT-5.6 Luna

99.9 IVR

Frontier foundation model developed by OpenAI featuring Sub-100ms Ultra-Efficient Edge Model architecture.

$0.18/1M in 90ms TTFT

Gemini 3.7 Flash

99.9 IVR

Frontier foundation model developed by Google featuring Native Omni Reasoning Workhorse architecture.

$0.08/1M in 75ms TTFT

Gemini 2.0 Flash

99.9 IVR

Frontier foundation model developed by Google featuring Multimodal Transformer architecture.

$0.10/1M in 380ms TTFT

DeepSeek-V4 Flash

99.9 IVR

Frontier foundation model developed by DeepSeek featuring MLA v3 Ultra-Sparse MoE architecture.

$0.14/1M in 150ms TTFT

DeepSeek-R1

99.9 IVR

Frontier foundation model developed by DeepSeek featuring MoE Reasoning architecture.

$0.55/1M in 1800ms TTFT

DeepSeek-V3

99.9 IVR

Frontier foundation model developed by DeepSeek featuring Multi-head Latent Attention MoE (671B/37B) architecture.

$0.14/1M in 340ms TTFT

Qwen 3.8 Flash Next

99.9 IVR

Frontier foundation model developed by Alibaba Cloud featuring High-Throughput Asian/Global MoE architecture.

$0.12/1M in 120ms TTFT

GLM 5.3 Flash

99.9 IVR

Frontier foundation model developed by Zhipu AI featuring Agent-Centric Fast MoE architecture.

$0.15/1M in 130ms TTFT

Codestral 25.01

99.9 IVR

Frontier foundation model developed by Mistral AI featuring Dense 256k Code Specialist architecture.

$0.30/1M in 150ms TTFT

Microsoft Phi-4 (14B)

99.9 IVR

Frontier foundation model developed by Microsoft featuring Dense Small Language Model architecture.

$0.12/1M in 110ms TTFT

Llama 3.3 70B Instruct

99.9 IVR

Frontier foundation model developed by Meta featuring Dense Transformer architecture.

$0.18/1M in 420ms TTFT

Qwen 2.5 72B Instruct

99.9 IVR

Frontier foundation model developed by Alibaba Cloud featuring Dense Transformer architecture.

$0.35/1M in 420ms TTFT

Qwen 2.5 Max

99.9 IVR

Frontier foundation model developed by Alibaba Cloud featuring MoE architecture.

$0.28/1M in 480ms TTFT

o3-mini

77 IVR

Frontier foundation model developed by OpenAI featuring Reasoning MoE architecture.

$1.10/1M in 1200ms TTFT

xAI Grok 4.6

74.6 IVR

Frontier foundation model developed by xAI featuring Colossus 200k-Cluster Real-Time Frontier architecture.

$2.00/1M in 280ms TTFT

GPT-5.6 Terra

72.5 IVR

Frontier foundation model developed by OpenAI featuring Balanced Production Workhorse MoE architecture.

$1.50/1M in 210ms TTFT

Claude 3.5 Haiku

70.4 IVR

Frontier foundation model developed by Anthropic featuring Lightweight Dense architecture.

$0.80/1M in 140ms TTFT

Composer 2.5

70.4 IVR

Frontier foundation model developed by Anysphere / Cursor featuring Full-Repository Diff & Refactor Specialist architecture.

$1.80/1M in 180ms TTFT

Fable 5

59.8 IVR

Frontier foundation model developed by Fable Studio featuring Generative Simulation & World Agent Engine architecture.

$2.00/1M in 240ms TTFT

Mistral Large 2

53.9 IVR

Frontier foundation model developed by Mistral AI featuring Dense Transformer architecture.

$2.00/1M in 550ms TTFT

Claude 3.7 Sonnet

49.5 IVR

Frontier foundation model developed by Anthropic featuring Hybrid Reasoning architecture.

$3.00/1M in 650ms TTFT

Claude 3.5 Sonnet

46 IVR

Frontier foundation model developed by Anthropic featuring Dense Multimodal Transformer architecture.

$3.00/1M in 320ms TTFT

Grok 3

45.9 IVR

Frontier foundation model developed by xAI featuring Dense Transformer architecture.

$3.00/1M in 850ms TTFT

OpenAI GPT-4o

45.5 IVR

Frontier foundation model developed by OpenAI featuring Omni Multimodal Transformer architecture.

$2.50/1M in 280ms TTFT

Claude Opus 5

42.7 IVR

Frontier foundation model developed by Anthropic featuring Next-Gen Ultra-Dense Hybrid Reasoning architecture.

$5.00/1M in 340ms TTFT

GPT-5.6 Sol

35.3 IVR

Frontier foundation model developed by OpenAI featuring Flagship Autonomous Reasoning Cluster architecture.

$8.00/1M in 420ms TTFT

OpenAI o1

21 IVR

Frontier foundation model developed by OpenAI featuring Large-Scale Reasoning Model architecture.

$15.00/1M in 850ms TTFT