price of compute

Rent the GPU, or pay per token?

The two sides of this site, bridged: today’s GPU rental medians against the labs’ token list prices. We supply both price sides live; the throughput is yours to supply, because it varies enormously with model, quantization, and batching — and we don’t print performance numbers we can’t stand behind. The arithmetic is shown under every result.

Your self-hosted cost

$2.657 per million tokens

$6.70/hr × 1,000,000 ÷ (1,000 tok/s × 3600s × 70%)

API modelBlended $/MTokSelf-host vs APIBreak-even tok/s
deepseek-v4-flash deepseek$0.1815.18×15,181
mistral-small-4 mistral$0.2610.12×10,121
gpt-5.6-luna openai$0.455.90×5,904
deepseek-v4-pro deepseek$0.544.89×4,886
mistral-large-3 mistral$0.753.54×3,542
gemini-3.5-flash-lite google$0.853.13×3,126
claude-haiku-4-5 anthropic$2.001.33×1,328
gemini-3.6-flash google$3.000.89×886
mistral-medium-3.5 mistral$3.000.89×886
grok-4.5 xai$3.000.89×886
gemini-3.5-flash google$3.380.79×787
claude-sonnet-5 anthropic$4.000.66×664
gemini-3.1-pro-preview google$4.500.59×590
gpt-5.6-terra openai$4.500.59×590
gpt-5.4 openai$5.630.47×472
claude-opus-4-8 anthropic$10.000.27×266
claude-opus-5 anthropic$10.000.27×266
gpt-5.6-sol openai$11.250.24×236
claude-fable-5 anthropic$20.000.13×133
gpt-5.5-pro openai$67.500.04×39

Below 1.00× self-hosting is cheaper per token than that API at your inputs; the break-even column is the throughput where they match. This compares raw token economics only — it ignores your engineering time, the API’s quality/latency, and that a rented GPU bills whether or not you keep it busy (that’s the utilization slider).