5 models across 3 tiers
Current rates for all 5 GPT models, including cache read/write, batch discounts, and long-context tiers.
Models
5
Cheapest input
$0.20/MTok
Cheapest output
$1.20/MTok
Cache discount
90% off
| Model | Tier | Input $/MTok | Output $/MTok | Cache Read | Cache Write | Batch | Context |
|---|---|---|---|---|---|---|---|
| GPT 5.6 Luna | Fast / Budget | $0.20 | $1.20 | $0.02(90% off) | $0.25(25% surcharge) | $0.10(50% off) | 1,050K |
| GPT 5.6 Terra | Mid-tier | $2.00 | $12.00 | $0.20(90% off) | $2.50(25% surcharge) | $1.00(50% off) | 1,050K |
| GPT 5.4 | Mid-tier | $2.50 | $15.00 | $0.25(90% off) | $2.50(no surcharge) | $1.25(50% off) | 1,050K |
| GPT 5.6 Sol | Flagship | $4.00 | $20.00 | $0.40(90% off) | $5.00(25% surcharge) | $2.00(50% off) | 1,050K |
| GPT 5.5 | Flagship | $5.00 | $30.00 | $0.50(90% off) | $5.00(no surcharge) | $2.50(50% off) | 1,050K |
How caching and batch processing affect your GPT API costs.
When a prompt prefix is already cached, the provider charges a reduced rate for those tokens. All GPT models offer a 90% cache read discount, meaning cached input tokens cost just 10% of the standard input rate.
Cache write pricing varies across GPT models. Some charge a surcharge on the first cache write while others charge standard input rates. Check the table above for per-model details.
All GPT models offer a 50% batch discount. Batch requests are processed asynchronously and cost 50% of the standard rate.
Some GPT models apply higher rates once the total prompt exceeds a threshold. Both input and output tokens may be re-priced for the entire request.
| Model | Threshold | Input multiplier | Output multiplier |
|---|---|---|---|
| GPT 5.6 Luna | 272.001K tokens | 2x | 1.5x |
| GPT 5.6 Terra | 272.001K tokens | 2x | 1.5x |
| GPT 5.4 | 272.001K tokens | 2x | 1.5x |
| GPT 5.6 Sol | 272.001K tokens | 2x | 1.5x |
| GPT 5.5 | 272.001K tokens | 2x | 1.5x |
Estimated cost per API call across common workloads, using default token counts for each use case.
GPT 5.6 Sol input decreased
$5.00 to $4.00/MTok
GPT 5.6 Sol output decreased
$30.00 to $20.00/MTok
GPT 5.6 Terra input decreased
$2.50 to $2.00/MTok
GPT 5.6 Terra output decreased
$15.00 to $12.00/MTok
GPT 5.6 Luna input decreased
$1.00 to $0.20/MTok
GPT 5.6 Luna output decreased
$6.00 to $1.20/MTok
7 models from $1.00/MTok
4 models from $0.30/MTok
4 models from $1.00/MTok
See exactly how much you spend on GPT models and where you can save. Two-line wrapper install, no API keys shared.
GPT API pricing starts at $0.20/MTok for input and $1.20/MTok for output. Prices vary by model tier: flagship models cost more but deliver higher quality, while fast/budget models are significantly cheaper for high-volume workloads.
Yes. GPT supports batch processing, which sends requests asynchronously at a reduced rate. The batch discount is 50% off standard pricing across all GPT models.
GPT has a 90% cache read discount. When input tokens are already cached from a previous request, they are charged at 10% of the standard input rate. This matters for applications that send similar prompts repeatedly.
There are currently 5 GPT models available through the API, spanning 3 tiers: Fast / Budget, Mid-tier, Flagship.
GPT API pricing is pay-per-use based on token consumption. There is no permanent free tier, but new accounts typically receive introductory credits. Tokeven helps you track exactly what you spend so nothing is wasted.
Yes. Some GPT models apply a long-context surcharge when the total prompt exceeds a token threshold. Both input and output rates increase for the entire request. Check the long-context pricing table above for exact thresholds and multipliers per model.