4 models across 3 tiers
Current rates for all 4 Grok models, including cache read/write, batch discounts, and long-context tiers.
Models
4
Cheapest input
$1.00/MTok
Cheapest output
$2.00/MTok
Cache discount
75% - 85% off
| Model | Tier | Input $/MTok | Output $/MTok | Cache Read | Cache Write | Batch | Context |
|---|---|---|---|---|---|---|---|
| Grok Build 0.1 | Fast / Budget | $1.00 | $2.00 | $0.20(80% off) | $1.00(no surcharge) | $0.50(50% off) | 256K |
| Grok 4.3 | Mid-tier | $1.25 | $2.50 | $0.20(84% off) | $1.25(no surcharge) | $0.63(50% off) | 1,000K |
| Grok 4.6 | Flagship | $2.00 | $6.00 | $0.50(75% off) | $2.00(no surcharge) | $1.00(50% off) | 500K |
| Grok 4.5 | Flagship | $2.00 | $6.00 | $0.30(85% off) | $2.00(no surcharge) | $1.00(50% off) | 500K |
Grok Build 0.1: xAI's coding-specialised model and the cheapest Grok row on output. Released in early access; xAI publishes release notes at month granularity only.
Grok 4.3: xAI publishes no release date for this model; dated to the March 4.20-generation launch whose pricing and 1M context window it shares. Priced from the live models table, which is what matters for costing traffic.
Grok 4.5: Same list price as Grok 4.6 but a cheaper cached rate ($0.30/MTok vs $0.50), so cache-heavy workloads are cheaper here than on 4.6. xAI publishes release notes at month granularity only; dated to the start of the month.
How caching and batch processing affect your Grok API costs.
When a prompt prefix is already cached, the provider charges a reduced rate for those tokens. Grok cache read discounts range from 75% to 85% off the standard input rate, depending on the model.
Grok does not charge a cache write surcharge. Writing tokens to cache costs the same as standard input.
All Grok models offer a 50% batch discount. Batch requests are processed asynchronously and cost 50% of the standard rate.
Some Grok models apply higher rates once the total prompt exceeds a threshold. Both input and output tokens may be re-priced for the entire request.
| Model | Threshold | Input multiplier | Output multiplier |
|---|---|---|---|
| Grok Build 0.1 | 200K tokens | 2x | 2x |
| Grok 4.3 | 200K tokens | 2x | 2x |
| Grok 4.6 | 200K tokens | 2x | 2x |
| Grok 4.5 | 200K tokens | 2x | 2x |
Estimated cost per API call across common workloads, using default token counts for each use case.
Grok 4.6 input launched
$2.00/MTok at launch
Grok 4.5 input launched
$2.00/MTok at launch
Grok Build 0.1 input launched
$1.00/MTok at launch
Grok 4.3 input launched
$1.25/MTok at launch
7 models from $1.00/MTok
5 models from $0.20/MTok
4 models from $0.30/MTok
See exactly how much you spend on Grok models and where you can save. Two-line wrapper install, no API keys shared.
Grok API pricing starts at $1.00/MTok for input and $2.00/MTok for output. Prices vary by model tier: flagship models cost more but deliver higher quality, while fast/budget models are significantly cheaper for high-volume workloads.
Yes. Grok supports batch processing, which sends requests asynchronously at a reduced rate. The batch discount is 50% off standard pricing across all Grok models.
Grok cache read discounts vary by model, ranging from 75% to 85% off the standard input rate.
There are currently 4 Grok models available through the API, spanning 3 tiers: Fast / Budget, Mid-tier, Flagship.
Grok API pricing is pay-per-use based on token consumption. There is no permanent free tier, but new accounts typically receive introductory credits. Tokeven helps you track exactly what you spend so nothing is wasted.
Yes. Some Grok models apply a long-context surcharge when the total prompt exceeds a token threshold. Both input and output rates increase for the entire request. Check the long-context pricing table above for exact thresholds and multipliers per model.