2026 LLM API Token Pricing, Context Caching & Word-to-Token Benchmark Table
2026 Verified ReferenceIn English prose and source code, `1,000 tokens` equals approximately `750 words` (`1 token ≈ 0.75 words` or `4 characters`). In 2026, enabling Prompt Context Caching reduces repeated system prompt and RAG document input token costs by 50%–90%, while async Batch APIs cut both input and output costs by 50%.
Monthly API Cost ($) = ((Input_Tokens × Input_Rate) + (Cached_Tokens × Cache_Rate) + (Output_Tokens × Output_Rate)) / 1,000,000| 2026 Frontier / Flash Model Tier | Input Cost / 1M Tokens | Cached Input / 1M Tokens | Output Cost / 1M & Context Window |
|---|---|---|---|
| OpenAI GPT-4o (Multimodal Flagship) | $2.50 / 1M | $1.25 / 1M (50% Auto-Cache) | $10.00 / 1M (128k Context Window) |
| OpenAI GPT-4o mini (High-Volume) | $0.15 / 1M | $0.075 / 1M (50% Auto-Cache) | $0.60 / 1M (128k Context Window) |
| Anthropic Claude 3.7 Sonnet | $3.00 / 1M | $0.30 / 1M (90% Cache Read) | $15.00 / 1M (200k Context + Thinking) |
| Google Gemini 2.5 Pro | $1.25 / 1M (≤200k) | $0.31 / 1M (Context Cache) | $10.00 / 1M (Up to 2M Token Context) |
| Google Gemini 2.5 Flash | $0.15 / 1M | $0.0375 / 1M (75% Cache Save) | $0.60 / 1M (1M Token Context Window) |
| DeepSeek V3 / R1 (Open-Weight API) | $0.27 – $0.55 / 1M | $0.07 – $0.14 / 1M (Disk Cache) | $1.10 – $2.19 / 1M (64k–128k Context) |
