Guides · 2026-08-28

LLM API pricing in 2026: what a million tokens really costs

LLM API pricing shifted under everyone’s feet this summer: DeepSeek — the default “cheap tier” of 2025 — roughly tripled its standard rates, Google is running promo pricing with a scheduled increase, and new flagship models landed at aggressive prices. Here is every major model’s current cost per million tokens as of 2026-08-28, and what an actual production workload costs on each.

The full price table

ModelProviderInput $/MOutput $/MContextNotes
Claude Fable 5Anthropic$10$501000kMost capable
Claude Opus 5Anthropic$5$251000k
Claude Sonnet 5Anthropic$2$101000k
Claude Haiku 4.5Anthropic$1$5200kFastest
GPT-5.6 SolOpenAI$4$20400kCached input $0.40
GPT-5.6 TerraOpenAI$2$12400k
GPT-5.6 LunaOpenAI$0.2$1.2400k
GPT-5.5OpenAI$5$30400kCached input $0.50
GPT-5.4 miniOpenAI$0.75$4.5400k
Gemini 3.7 FlashGoogle$0.75$3.751000kPromo price through 2026; $1.50/$7.50 from 2027
Gemini 3.1 ProGoogle$2$121000k≤200k prompt; higher above
Gemini 3.5 FlashGoogle$1.5$91000k
DeepSeek V4 FlashDeepSeek$0.44$1.321000kOff-peak −50%; cache-hit input near-free
DeepSeek V4 ProDeepSeek$1.32$3.961000kReasoning tier; off-peak −50%
Mistral Medium 3.5Mistral$1.5$7.5256k
Mistral Large 3Mistral$0.5$1.5256kOpen-weight
Mistral Small 4Mistral$0.15$0.6256k
Grok 4.6xAI$2$6500k≤200k prompt; higher above
Grok 4.3xAI$1.25$2.51000k≤200k prompt; higher above

What changed in 2026

  • DeepSeek tripled.V4 Flash went from $0.14/$0.28 to $0.44/$1.32 per million (input/output) at standard rates. Off-peak hours are half price and cache hits stay near-free, but the “10× cheaper than everyone” era is over — budget workloads built on 2025 DeepSeek math need re-costing.
  • Promo pricing with expiry dates.Gemini 3.7 Flash lists at $0.75/$3.75 only through 2026 — the posted price doubles to $1.50/$7.50 in January 2027. If you’re budgeting a year ahead, use the 2027 number.
  • Cheaper flagships.Claude Sonnet 5 launched at $2/$10 — below the $3/$15 its predecessor charged. GPT-5.6’s tiers (Sol $4/$20, Terra $2/$12, Luna $0.20/$1.20) similarly undercut the 5.4/5.5 generation. Raw intelligence per dollar keeps improving even as the floor rises.
  • Caching is the real discount.OpenAI’s cached input is ~10× cheaper, DeepSeek’s near-free, Anthropic’s prompt caching similar. For chat workloads where 90% of each request is repeated context, the cache price matters more than the list price.

A real workload, costed

Per-token prices are abstract, so: a support chatbot handling ~100k conversations a month — roughly 100M input tokens (system prompt, context, history) and 10M output tokens. Uncached monthly cost per model:

Model$/monthvs cheapest
Mistral Small 4$21cheapest
GPT-5.6 Luna$321.5×
DeepSeek V4 Flash$572.7×
Mistral Large 3$653.1×
Gemini 3.7 Flash$1135.4×
GPT-5.4 mini$1205.7×
Claude Haiku 4.5$1507.1×
Grok 4.3$1507.1×
DeepSeek V4 Pro$1728.2×
Mistral Medium 3.5$22510.7×
Gemini 3.5 Flash$24011.4×
Grok 4.6$26012.4×
Claude Sonnet 5$30014.3×
GPT-5.6 Terra$32015.2×
Gemini 3.1 Pro$32015.2×
GPT-5.6 Sol$60028.6×
Claude Opus 5$75035.7×
GPT-5.5$80038.1×
Claude Fable 5$1,50071.4×

The spread is the story: the same workload runs from about twenty dollars to fifteen hundred per month depending on the model — a 70× range for the identical token volume. Most production systems land on a two-model split — a budget model (Luna, Haiku, DeepSeek off-peak) for the 80% of easy traffic and a flagship for the hard 20% — plus aggressive prompt caching, which can cut the input bill by an order of magnitude on chat-shaped workloads.

Cost your own workload

Plug your actual token volumes into the LLM API cost calculator — same data, your numbers.

Methodology: prices are from official provider pricing pages (Anthropic, OpenAI, Google, DeepSeek, Mistral, xAI) as of 2026-08-28, date-stamped and refreshed with the rest of our data. Tiered prices use the base tier (≤200k prompt where providers split); cached-input and batch discounts are noted but not included in the workload math. No AI provider pays us anything.