AI Token Cost Calculator
Price the tokens — monthly input and output at the published per-million rates, with the output share named.
AI Token Cost Calculator
Results recalculate instantly on every keystroke. Nothing you type is transmitted.
What this result does not account for
- Two-rate model — no per-request overheads or tool calls
- Rates move; the presets pin the published 2025–26 lines
In short: 3,000,000 input tokens at $2.50/M bill $7.50 and 7,000,000 output at $10.00/M bill $70.00 — $77.50 a month, of which output is 90.322581%. The model bills 4× for output tokens and this mix sends 2.333333× more of them; that asymmetry is the whole cost story. A year of the mix reads $930.00, and batch would halve it.
Formula
bill = in tokens × in rate ÷ 10⁶ + out tokens × out rate ÷ 10⁶
Token APIs bill per million, in two asymmetrical currencies. The published 2025–26 lines: GPT-4o $2.50 in / $10.00 out, GPT-4o mini $0.15/$0.60, Claude Haiku $1/$5, Claude Sonnet $3/$15, Claude Opus $5/$25 — output costs 4–5× what input does. Two levers reprice the bill without changing models: batch endpoints take 50% off both sides, and cached input reads at about a tenth of fresh.
Worked Example
- Enter monthly input and output tokens from your usage dashboard.
- Pick the rate pair — the presets are the published per-million lines.
- Read the bill, the output share, and what batch and caching would do.
Defaults: $77.50, output 90.322581%. Drive the rates to the Sonnet pair ($3/$15) and the bill reads $114.00 — same traffic, 47% more, because output dominates and Sonnet's output rate carries it.
Strengths & Limits Of This Model
Where this engine is strong
- Output share named — the bill's actual shape
- Batch and caching levers priced, not just mentioned
Where it stops
- Context resend patterns vary; dashboards beat estimates
Practical Use Cases
Budget forecasts
the monthly AI line before it arrives
Model routing
what the cheap tier actually saves
Feature pricing
per-user cost of an AI feature
Methodology & Editorial Standards
Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.
This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.
Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.
AI Token Cost Calculator — 8 Expert FAQs
8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.
What are the current published rates?
The 2025–26 lines: GPT-4o at $2.50/$10.00 per million, GPT-4o mini $0.15/$0.60, Claude Haiku $1/$5, Claude Sonnet $3/$15, Claude Opus $5/$25. Rates move — the fields take any pair, and the presets exist so the defaults are checkable today.
Why is output so much more expensive?
Because generation is autoregressive — every output token is a full forward pass through the model, while input tokens are processed in one parallel sweep. The 4–5× output premium is compute, not marketing.
How many tokens is my text?
Roughly 4 characters a token, so about 750 words to the thousand tokens — a page of prose. Context you resend every turn counts as input every turn, which is why caching exists.
How much do batch and caching actually save?
Batch endpoints take a flat 50% off both input and output for non-realtime work. Prompt caching reads at about a tenth of the fresh-input rate — on stable system prompts and RAG contexts it is the single biggest lever on the input side.
Why does output dominate my bill?
Because it dominates the arithmetic twice: the rate is 4–5× higher AND chatty workloads emit more output tokens than they send input. A 90% output share on the bill is normal for generation-heavy products; the fix is shorter completions, not cheaper prompts.
When does self-hosting beat API pricing?
When utilization is high and steady: the GPU page prices the silicon that replaces the meter. Below roughly full-time utilization, the API's pay-per-token wins because idle GPUs bill the same as busy ones.
Do failed requests bill?
Errored requests generally do not bill, but truncated-and-retried generations bill every attempt — retries are invisible output inflation. Count them in the dashboard number you enter here.
How do I estimate tokens without a dashboard?
Words ÷ 0.75 ≈ tokens (750 words to the thousand). A support ticket is 200–500 tokens in and 300–800 out; a code assistant session is thousands in (context) and hundreds out. Estimate conservatively — context resend is the hidden multiplier.