Tokeven prices every call before your team sends it, across Claude, ChatGPT, Gemini and Grok, and proves each dollar saved. We take 10% of verified savings. You keep 90%.
Works with:
How this is calculated - a worked model from published prices
A team spending $50,000/mo moves 40% of its Claude Opus 4.8 traffic to Claude Sonnet 4.6. At today’s published token prices that traffic costs 60% of what it did (blended 3:1 input/output), so $20,000 becomes $12,000 - $8,000 saved. Tokeven takes 10% of that ($800); you keep the rest. Your own numbers will differ. Measuring them is the whole point.
How it works
Two lines of Python wrap your existing Claude, OpenAI, Google or xAI client. Tokeven captures token counts, model, latency and cost while your calls run exactly as before. Your keys stay local and prompt content stays on your machine.
Captures usage: token counts, model, latency and cost. Prompt content and your provider key never reach Tokeven.
The same 9,000-token system prompt is billed as fresh input on 755 Claude calls a week. Marking it cacheable bills the repeats at the 90%-off cache-read rate instead.
155 of your 370 claude-opus-4-8 calls this week were single-turn lookups. On the identical prompt, claude-sonnet-4-6 prices 40% lower.
3 gemini-3.1-pro workflows re-send full conversation history every turn. Summarising prior turns would cut fresh input tokens by ~35%.
Before each prompt is sent, Tokeven shows the cost on the current model, the cost on alternatives across Claude, GPT, Gemini and Grok, and a confidence score for each. Your developer makes the call in the moment. A weekly digest sums the savings, flags patterns and names the single biggest lever for each engineer.
The savings proof dashboard traces every dollar back to the action that caused it - accepted routing recommendations, cache improvements, prompt optimisations. Per person, per team, fully auditable. Savings are priced against live token prices, so the numbers stay honest as provider pricing moves.
Gainshare pricing
Savings measured against live token prices per model. Published math, open to audit.
Your biggest lever
System prompts repeat across 89% of sessions, all providers
Enable prompt caching where supported
Save $515/wk
gpt-5.5 used for tasks where gemini-3.5-flash matches quality
Route summarisation to Flash
Save $1,030/wk
68% of Opus calls are single-turn Q&A
Switch to Sonnet or Gemini Flash for simple lookups
Save $796/wk
Interactive calculator
Low end of a modelled 15-30% range from routing, caching, and shorter prompts. Modelled, not a customer average - your own numbers are the only ones we bill against.
A price before every call. Accepted recommendations lower the bill. We take 10% of what you actually save.