Index126.1
15 live
Cost calculator

What will your run cost?

The two sides of this site, bridged: today’s GPU rental medians against the labs’ token list prices. We supply both price sides live; the throughput is yours to supply, because it varies enormously with model, quantization, and batching — and we don’t print performance numbers we can’t stand behind. The arithmetic is shown under every result.

Your self-hosted cost

$2.875 per million tokens

$7.25/hr × 1,000,000 ÷ (1,000 tok/s × 3600s × 70%)

API modelBlended $/MTokSelf-host vs APIBreak-even tok/s
deepseek-v4-flash deepseek$0.1816.43×16,429
mistral-small-4 mistral$0.2610.95×10,952
gpt-5.6-luna openai$0.456.39×6,389
deepseek-v4-pro deepseek$0.545.29×5,287
mistral-large-3 mistral$0.753.83×3,833
gemini-3.5-flash-lite google$0.853.38×3,382
claude-haiku-4-5 anthropic$2.001.44×1,438
gemini-3.6-flash google$3.000.96×958
mistral-medium-3.5 mistral$3.000.96×958
grok-4.5 xai$3.000.96×958
gemini-3.5-flash google$3.380.85×852
claude-sonnet-5 anthropic$4.000.72×719
gemini-3.1-pro-preview google$4.500.64×639
gpt-5.6-terra openai$4.500.64×639
gpt-5.4 openai$5.630.51×511
claude-opus-4-8 anthropic$10.000.29×288
claude-opus-5 anthropic$10.000.29×288
gpt-5.6-sol openai$11.250.26×256
claude-fable-5 anthropic$20.000.14×144
gpt-5.5-pro openai$67.500.04×43

Below 1.00× self-hosting is cheaper per token than that API at your inputs; the break-even column is the throughput where they match. This compares raw token economics only — it ignores your engineering time, the API’s quality/latency, and that a rented GPU bills whether or not you keep it busy (that’s the utilization slider).