Long-context pricing
Several providers raise the per-token price once a single prompt crosses a size threshold (200K or 272K tokens). The table shows normal → long-context prices per 1M tokens, plus the cost of one 500K-token prompt with a 2K-token answer.
| Model | Threshold | Input /1M | Output /1M | One 500K request |
|---|---|---|---|---|
| GPT-6 Luna | 272K | $0.10 → $0.20 | $0.50 → $0.75 | $0.10 |
| GPT-5.6 Luna | 272K | $0.20 → $0.40 | $1.20 → $1.80 | $0.20 |
| Grok Build 0.1 | 200K | $1.00 → $2.00 | $2.00 → $4.00 | $1.01 |
| Grok Code Fast 1 | 200K | $1.00 → $2.00 | $2.00 → $4.00 | $1.01 |
| Gemini 2.5 Pro | 200K | $1.25 → $2.50 | $10.00 → $15.00 | $1.28 |
| Grok 4.20 Non Reasoning | 200K | $1.25 → $2.50 | $2.50 → $5.00 | $1.26 |
| Grok 4.20 Reasoning | 200K | $1.25 → $2.50 | $2.50 → $5.00 | $1.26 |
| Grok 4.3 | 200K | $1.25 → $2.50 | $2.50 → $5.00 | $1.26 |
| Gemini 3.1 Pro Preview | 200K | $2.00 → $4.00 | $12.00 → $18.00 | $2.04 |
| Gemini Pro (latest) | 200K | $2.00 → $4.00 | $12.00 → $18.00 | $2.04 |
| GPT-5.6 Terra | 272K | $2.00 → $4.00 | $12.00 → $18.00 | $2.04 |
| GPT-6 Sol | 272K | $2.00 → $4.00 | $10.00 → $15.00 | $2.03 |
| GPT-6.1 Sol | 272K | $2.00 → $4.00 | $10.00 → $15.00 | $2.03 |
| Grok 4.5 | 200K | $2.00 → $4.00 | $6.00 → $12.00 | $2.02 |
| Grok 4.6 | 200K | $2.00 → $4.00 | $6.00 → $12.00 | $2.02 |
| Grok 4.7 | 200K | $2.00 → $4.00 | $6.00 → $12.00 | $2.02 |
| GPT-5.4 | 272K | $2.50 → $5.00 | $15.00 → $22.50 | $2.54 |
| Claude Sonnet 4.5 | 200K | $3.00 → $6.00 | $15.00 → $22.50 | $3.04 |
| GPT-5.6 Sol | 272K | $4.00 → $8.00 | $20.00 → $30.00 | $4.06 |
| GPT-5.5 | 272K | $5.00 → $10.00 | $30.00 → $45.00 | $5.09 |
| GPT-6 Astra | 272K | $10.00 → $20.00 | $50.00 → $75.00 | $10.15 |
Models not listed either have no published long-context tier in our data or use the same price at every length.
FAQ
Is the whole request billed at the higher rate?
For most providers, yes: once the prompt exceeds the threshold, all input and output tokens of that request use the long-context price, not only the tokens above it.
How do I avoid the surcharge?
Keep prompts under the threshold with retrieval (send only relevant chunks), summarize old conversation turns, or use prompt caching for the repeated part.