Index127.5
2026-08-12 02:42 UTC · 186 obs today · 16 providers
Cost calculator

What will your run cost?

The two sides of this site, bridged: today’s GPU rental medians against the labs’ token list prices. We supply both price sides live; the throughput is yours to supply, because it varies enormously with model, quantization, and batching — and we don’t print performance numbers we can’t stand behind. The arithmetic is shown under every result.

Your self-hosted cost

$2.694 per million tokens

$6.79/hr × 1,000,000 ÷ (1,000 tok/s × 3600s × 70%)

API modelBlended $/MTokSelf-host vs APIBreak-even tok/s
deepseek-v4-flash deepseek$0.1815.40×15,397
mistral-small-4 mistral$0.2610.26×10,265
gpt-5.6-luna openai$0.455.99×5,988
deepseek-v4-pro deepseek$0.544.96×4,955
mistral-large-3 mistral$0.753.59×3,593
gemini-3.5-flash-lite google$0.853.17×3,170
claude-haiku-4-5 anthropic$2.001.35×1,347
gemini-3.6-flash google$3.000.90×898
mistral-medium-3.5 mistral$3.000.90×898
grok-4.5 xai$3.000.90×898
gemini-3.5-flash google$3.380.80×798
claude-sonnet-5 anthropic$4.000.67×674
gemini-3.1-pro-preview google$4.500.60×599
gpt-5.6-terra openai$4.500.60×599
gpt-5.4 openai$5.630.48×479
claude-opus-4-8 anthropic$10.000.27×269
claude-opus-5 anthropic$10.000.27×269
gpt-5.6-sol openai$11.250.24×240
claude-fable-5 anthropic$20.000.13×135
gpt-5.5-pro openai$67.500.04×40

Below 1.00× self-hosting is cheaper per token than that API at your inputs; the break-even column is the throughput where they match. This compares raw token economics only — it ignores your engineering time, the API’s quality/latency, and that a rented GPU bills whether or not you keep it busy (that’s the utilization slider).