We cut your AI bill. You pay us from the savings.

Tokeven prices every call before your team sends it, across Claude, ChatGPT, Gemini and Grok, and proves each dollar saved. We take 10% of verified savings. You keep 90%.

2-min installPrompt content never reaches TokevenEvery recommendation is yours to acceptFree tier, forever

Works with:

ClaudeOpenAIGeminiGrokAnthropic SDKPythonREST APIMCPClaudeOpenAIGeminiGrokAnthropic SDKPythonREST APIMCP

The math on a $50k bill

Every figure below comes from published token prices and a worked model. Your own numbers replace it the first week.
16%
lower monthly bill, gross - 14.4% net of our fee
$8,000
saved per month, before our fee
$7,200
you keep - we take 10% of what we save you
2min
to install and start tracking

How this is calculated - a worked model from published prices

A team spending $50,000/mo moves 40% of its Claude Opus 4.8 traffic to Claude Sonnet 4.6. At today’s published token prices that traffic costs 60% of what it did (blended 3:1 input/output), so $20,000 becomes $12,000 - $8,000 saved. Tokeven takes 10% of that ($800); you keep the rest. Your own numbers will differ. Measuring them is the whole point.

How it works

From install to proven savings in minutes

Two lines of code. Before each call your team sees a price and the cheaper alternatives in their terminal. After each call, usage is logged. Every accepted recommendation is traced to the dollars it saved. See the full walkthrough
Step 1

Install the wrapper

Two lines of Python wrap your existing Claude, OpenAI, Google or xAI client. Tokeven captures token counts, model, latency and cost while your calls run exactly as before. Your keys stay local and prompt content stays on your machine.

  • Works with Claude, GPT, Gemini and Grok
  • Runs locally - prompt content stays with you
  • Two lines, then it runs
app.pytracks usage
1# Two lines - wrap your existing Anthropic client
2from anthropic import Anthropic
3from tokeven import TokevenAnthropic
4
5client = Anthropic()
6client = TokevenAnthropic(client, project="my-app")
7
8# That's it - all calls are now tracked
9response = client.messages.create(model="claude-sonnet-4-6", ...)

Captures usage: token counts, model, latency and cost. Prompt content and your provider key never reach Tokeven.

Your Weekly Spend Digest
· tokeven
Mon 9:00 AM
Biggest lever: caching - $104/mo left on the table
3 optimizations found, $183/mo total - 18% of your $1,003/mo run rate
Cache hit rate
34%
Total calls
2,263
Total cost
$234
Avg cost/call
$0.10
Enable cache_control on your repeated system prompt
$104/mo

The same 9,000-token system prompt is billed as fresh input on 755 Claude calls a week. Marking it cacheable bills the repeats at the 90%-off cache-read rate instead.

42% of your Opus calls could use Sonnet
$64/mo

155 of your 370 claude-opus-4-8 calls this week were single-turn lookups. On the identical prompt, claude-sonnet-4-6 prices 40% lower.

Trim re-sent context in 3 workflows
$15/mo

3 gemini-3.1-pro workflows re-send full conversation history every turn. Summarising prior turns would cut fresh input tokens by ~35%.

View full report in dashboard
Step 2

See the price and the alternatives before every call

Before each prompt is sent, Tokeven shows the cost on the current model, the cost on alternatives across Claude, GPT, Gemini and Grok, and a confidence score for each. Your developer makes the call in the moment. A weekly digest sums the savings, flags patterns and names the single biggest lever for each engineer.

  • Price forecast before every call
  • Confidence scores across all four providers
  • Weekly digest with dollar amounts per recommendation
Step 3

Prove every dollar saved

The savings proof dashboard traces every dollar back to the action that caused it - accepted routing recommendations, cache improvements, prompt optimisations. Per person, per team, fully auditable. Savings are priced against live token prices, so the numbers stay honest as provider pricing moves.

  • Only accepted recommendations count
  • Per-person and per-team breakdowns
  • Savings priced against live token prices
Savings Proof, 30 daysExample
Verified Savings
$10,032
Spend Reduction
20%
Weekly Cost Trend
W1W2W3W4W5W6
Cross-Provider Routing
$4,214
Model Downgrades
$3,110
Cache Hits
$2,708

Gainshare pricing

We earn only when you actually save

Free shows every developer the price and the cheaper option. Team proves the savings, per person and per team, and we keep 10% of what we verify. Savings of $0 mean a fee of $0.

Free

$0
per org / month
  • Full cost visibility dashboard
  • Pre-call cost estimates via MCP
  • Model recommendations
  • Local-only privacy (no keys, no prompts)
  • One developer

Team

10%
of verified savings
  • Everything in Free
  • Unlimited tracked calls
  • Org-wide savings dashboard
  • Per-person savings tracking
  • Cache and prompt coaching per engineer
  • Prompt revision suggestions
  • Weekly analysis digest
  • Cross-provider analytics
  • Embeddable widgets
  • CSV, JSON & API export
  • n8n & webhook integrations

Enterprise

Custom
capped or high-touch terms
  • Everything in Team
  • Dedicated account manager
  • Custom integrations
  • SLA & priority support
  • SSO & SAML
  • Custom data retention

Savings measured against live token prices per model. Published math, open to audit.

Your biggest lever

Cached tokens cost 90% less

Repeated system prompts are the largest single saving in most teams and the easiest to miss. Tokeven shows each engineer their cache rate, the dollars it is worth, and a weekly nudge to raise it - alongside per-call cost, model breakdowns and team spend across all four providers.
Cost OverviewExample dashboard
Total Cost (7d)
$9,364
Saved (7d)
$2,341
Total Calls
90,500
Cache Rate
34%
This week’s biggest lever

System prompts repeat across 89% of sessions, all providers

Enable prompt caching where supported

Save $515/wk

gpt-5.5 used for tasks where gemini-3.5-flash matches quality

Route summarisation to Flash

Save $1,030/wk

68% of Opus calls are single-turn Q&A

Switch to Sonnet or Gemini Flash for simple lookups

Save $796/wk

tokeven.com/dashboard/cost
Cost Overview
7 days
30 days
90 days
Total Cost (7d)
$9,364
40 seats
Avg Cost / Call
$0.10
90,500 calls
Cache Hit Rate
34%
↑ from 28%
Daily Cost
MonTueWedThuFriSatSun
Cost by Model
claude-opus-4-8$3,577
gpt-5.6-sol$2,073
claude-sonnet-4-6$1,665
gemini-3.1-pro$773
gpt-5.5$712
gemini-3.5-flash$343
claude-haiku-4-5$221
Most Expensive Calls
claude-opus-4-8
$0.55
gpt-5.6-sol
$0.54
claude-sonnet-4-6
$0.44
gemini-3.1-pro
$0.31

What changes with Tokeven

Cost awareness
Bill arrives at month end
Every call priced before it is sent
Model selection
Developers guess which model fits
Cost and confidence per model before each call, across Claude, GPT, Gemini and Grok
Prompt caching
Repeated system prompts billed at full price
Cacheable patterns flagged, 90% savings on repeats
Team visibility
Spend is one line on the invoice
Per-person cost, calls and savings, fully auditable
Savings proof
"We think we saved money"
Every dollar traced to a recommendation you accepted
Privacy
Prompts sit on vendor servers
Prompt content never reaches Tokeven, metadata only

Interactive calculator

Run your own numbers

Enter your current AI spend and see what better model selection, caching and prompt efficiency would save.

Calculate your savings

$50,000
$5K$500K
25
5500
Assumptions
Modelled savings rate15%
Tokeven fee (gainshare)10% of savings

Low end of a modelled 15-30% range from routing, caching, and shorter prompts. Modelled, not a customer average - your own numbers are the only ones we bill against.

Your projected savings

Annual net savings
$81,000
after Tokeven fee
Monthly savings (gross)
$7,500
Tokeven fee (10%)
$750
You keep monthly$6,750
That's $270 saved per seat per month
Start saving today

Questions buyers always ask

Your team spends less
when the price shows up first

A price before every call. Accepted recommendations lower the bill. We take 10% of what you actually save.

Free tier, foreverCard-free signupTracking from minute one