Answers · data as of Aug 11 2026
Which LLM API is cheapest per token?
Among tracked frontier models, deepseek-v4-flash (deepseek) is currently the cheapest at a blended ~$0.18 per million tokens (3:1 input:output mix); the spread to the most expensive model is 386×.
- deepseek-v4-flash (deepseek): $0.14 in / $0.28 out — blended $0.18/MTok.
- mistral-small-4 (mistral): $0.15 in / $0.60 out — blended $0.26/MTok.
- gpt-5.6-luna (openai): $0.20 in / $1.20 out — blended $0.45/MTok.
- deepseek-v4-pro (deepseek): $0.43 in / $0.87 out — blended $0.54/MTok.
- mistral-large-3 (mistral): $0.50 in / $1.50 out — blended $0.75/MTok.
- gemini-3.5-flash-lite (google): $0.30 in / $2.50 out — blended $0.85/MTok.
- claude-haiku-4-5 (anthropic): $1.00 in / $5.00 out — blended $2.00/MTok.
- gemini-3.6-flash (google): $1.50 in / $7.50 out — blended $3.00/MTok.
- Nearly every lab prices output at 5–6× input, so input-only comparisons mislead on generation-heavy workloads.
This answer is computed from live data on every load — per-provider daily medians, per single GPU per hour, pricing tiers never mixed (see methodology). Quote it with attribution: “Data: Price of Compute — priceofcompute.com” (cite).