Guides · 2026-08-28
LLM API pricing in 2026: what a million tokens really costs
LLM API pricing shifted under everyone’s feet this summer: DeepSeek — the default “cheap tier” of 2025 — roughly tripled its standard rates, Google is running promo pricing with a scheduled increase, and new flagship models landed at aggressive prices. Here is every major model’s current cost per million tokens as of 2026-08-28, and what an actual production workload costs on each.
The full price table
| Model | Provider | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | $10 | $50 | 1000k | Most capable |
| Claude Opus 5 | Anthropic | $5 | $25 | 1000k | |
| Claude Sonnet 5 | Anthropic | $2 | $10 | 1000k | |
| Claude Haiku 4.5 | Anthropic | $1 | $5 | 200k | Fastest |
| GPT-5.6 Sol | OpenAI | $4 | $20 | 400k | Cached input $0.40 |
| GPT-5.6 Terra | OpenAI | $2 | $12 | 400k | |
| GPT-5.6 Luna | OpenAI | $0.2 | $1.2 | 400k | |
| GPT-5.5 | OpenAI | $5 | $30 | 400k | Cached input $0.50 |
| GPT-5.4 mini | OpenAI | $0.75 | $4.5 | 400k | |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1000k | Promo price through 2026; $1.50/$7.50 from 2027 | |
| Gemini 3.1 Pro | $2 | $12 | 1000k | ≤200k prompt; higher above | |
| Gemini 3.5 Flash | $1.5 | $9 | 1000k | ||
| DeepSeek V4 Flash | DeepSeek | $0.44 | $1.32 | 1000k | Off-peak −50%; cache-hit input near-free |
| DeepSeek V4 Pro | DeepSeek | $1.32 | $3.96 | 1000k | Reasoning tier; off-peak −50% |
| Mistral Medium 3.5 | Mistral | $1.5 | $7.5 | 256k | |
| Mistral Large 3 | Mistral | $0.5 | $1.5 | 256k | Open-weight |
| Mistral Small 4 | Mistral | $0.15 | $0.6 | 256k | |
| Grok 4.6 | xAI | $2 | $6 | 500k | ≤200k prompt; higher above |
| Grok 4.3 | xAI | $1.25 | $2.5 | 1000k | ≤200k prompt; higher above |
What changed in 2026
- DeepSeek tripled.V4 Flash went from $0.14/$0.28 to $0.44/$1.32 per million (input/output) at standard rates. Off-peak hours are half price and cache hits stay near-free, but the “10× cheaper than everyone” era is over — budget workloads built on 2025 DeepSeek math need re-costing.
- Promo pricing with expiry dates.Gemini 3.7 Flash lists at $0.75/$3.75 only through 2026 — the posted price doubles to $1.50/$7.50 in January 2027. If you’re budgeting a year ahead, use the 2027 number.
- Cheaper flagships.Claude Sonnet 5 launched at $2/$10 — below the $3/$15 its predecessor charged. GPT-5.6’s tiers (Sol $4/$20, Terra $2/$12, Luna $0.20/$1.20) similarly undercut the 5.4/5.5 generation. Raw intelligence per dollar keeps improving even as the floor rises.
- Caching is the real discount.OpenAI’s cached input is ~10× cheaper, DeepSeek’s near-free, Anthropic’s prompt caching similar. For chat workloads where 90% of each request is repeated context, the cache price matters more than the list price.
A real workload, costed
Per-token prices are abstract, so: a support chatbot handling ~100k conversations a month — roughly 100M input tokens (system prompt, context, history) and 10M output tokens. Uncached monthly cost per model:
| Model | $/month | vs cheapest |
|---|---|---|
| Mistral Small 4 | $21 | cheapest |
| GPT-5.6 Luna | $32 | 1.5× |
| DeepSeek V4 Flash | $57 | 2.7× |
| Mistral Large 3 | $65 | 3.1× |
| Gemini 3.7 Flash | $113 | 5.4× |
| GPT-5.4 mini | $120 | 5.7× |
| Claude Haiku 4.5 | $150 | 7.1× |
| Grok 4.3 | $150 | 7.1× |
| DeepSeek V4 Pro | $172 | 8.2× |
| Mistral Medium 3.5 | $225 | 10.7× |
| Gemini 3.5 Flash | $240 | 11.4× |
| Grok 4.6 | $260 | 12.4× |
| Claude Sonnet 5 | $300 | 14.3× |
| GPT-5.6 Terra | $320 | 15.2× |
| Gemini 3.1 Pro | $320 | 15.2× |
| GPT-5.6 Sol | $600 | 28.6× |
| Claude Opus 5 | $750 | 35.7× |
| GPT-5.5 | $800 | 38.1× |
| Claude Fable 5 | $1,500 | 71.4× |
The spread is the story: the same workload runs from about twenty dollars to fifteen hundred per month depending on the model — a 70× range for the identical token volume. Most production systems land on a two-model split — a budget model (Luna, Haiku, DeepSeek off-peak) for the 80% of easy traffic and a flagship for the hard 20% — plus aggressive prompt caching, which can cut the input bill by an order of magnitude on chat-shaped workloads.
Cost your own workload
Plug your actual token volumes into the LLM API cost calculator — same data, your numbers.
Methodology: prices are from official provider pricing pages (Anthropic, OpenAI, Google, DeepSeek, Mistral, xAI) as of 2026-08-28, date-stamped and refreshed with the rest of our data. Tiered prices use the base tier (≤200k prompt where providers split); cached-input and batch discounts are noted but not included in the workload math. No AI provider pays us anything.