LLM API price calculator with the cache hit rate priced in
A price table tells you what a million tokens cost. It doesn't tell you what your workload costs once most of your input tokens are served from the provider's prompt cache, and it doesn't tell you where two models trade places. This calculator prices your own token shape per request, per day and per month with the cache hit rate included, then finds the break even between any two models. Every list price was read from two independent records, the OpenRouter models API and the LiteLLM model price table, and converted to dollars per one million tokens in code. Nothing here is typed by hand.
Price your workload
The effective input price is the list input price weighted by the share of input tokens that miss the cache, plus the cached input price weighted by the share that hit. Output tokens are billed at the output price in full. Every result recomputes as you type, in your browser, with no network call.
Break even between two models
Pick two models. The page solves, from the validated list prices and your current inputs, for the output share of tokens at which both cost the same per request, and for the cache hit rate at which they cost the same with your current token shape. Either crossover can fall outside the reachable range, and when it does the page says so instead of printing a figure you can't act on.
At your current inputs Google Gemini 3.8 Flash costs less than xAI Grok 4.3 per request, by a factor of 1.03. Move the output share or the cache hit rate across the crossover above and the order flips.
List prices per one million tokens
Every figure below is the provider's standard list price in dollars per one million tokens, fetched on 2026-09-06. Batch, cached input write, long context and reasoning surcharges are not included. The cached input column is the cached read price only.
| Provider | Model | Input, USD per one million tokens | Cached input, USD per one million tokens | Output, USD per one million tokens | Cached input as a share of input | Output to input ratio |
|---|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra | 10 | 1 | 50 | 10.0% | 5.00x |
| OpenAI | GPT-5.6 Terra | 2 | 0.2 | 12 | 10.0% | 6.00x |
| OpenAI | GPT-5.6 Luna | 0.2 | 0.02 | 1.2 | 10.0% | 6.00x |
| OpenAI | GPT-5.5 Pro | 30 | not in the OpenRouter record | 180 | not in the OpenRouter record | 6.00x |
| OpenAI | GPT-5.5 | 5 | 0.5 | 30 | 10.0% | 6.00x |
| OpenAI | GPT-5.4 Pro | 30 | not in the OpenRouter record | 180 | not in the OpenRouter record | 6.00x |
| OpenAI | GPT-5.4 | 2.5 | 0.25 | 15 | 10.0% | 6.00x |
| OpenAI | GPT-5.4 mini | 0.75 | 0.075 | 4.5 | 10.0% | 6.00x |
| OpenAI | GPT-5.4 nano | 0.2 | 0.02 | 1.25 | 10.0% | 6.25x |
| OpenAI | GPT-5.3 Codex | 1.75 | 0.175 | 14 | 10.0% | 8.00x |
| OpenAI | GPT-5.2 Pro | 21 | not in the OpenRouter record | 168 | not in the OpenRouter record | 8.00x |
| OpenAI | GPT-5.2 | 1.75 | 0.175 | 14 | 10.0% | 8.00x |
| OpenAI | GPT-5.2 Codex | 1.75 | 0.175 | 14 | 10.0% | 8.00x |
| OpenAI | GPT-5.1 Codex Max | 1.25 | 0.125 | 10 | 10.0% | 8.00x |
| OpenAI | GPT-5.1 | 1.25 | 0.125 | 10 | 10.0% | 8.00x |
| OpenAI | GPT-5.1 Codex | 1.25 | held, the two records disagree | 10 | not in the OpenRouter record | 8.00x |
| OpenAI | GPT-5.1 Codex mini | 0.25 | held, the two records disagree | 2 | not in the OpenRouter record | 8.00x |
| OpenAI | GPT-5 Pro | 15 | not in the OpenRouter record | 120 | not in the OpenRouter record | 8.00x |
| OpenAI | GPT-5 | 1.25 | 0.125 | 10 | 10.0% | 8.00x |
| OpenAI | GPT-5 mini | 0.25 | 0.025 | 2 | 10.0% | 8.00x |
| OpenAI | GPT-5 nano | 0.05 | 0.005 | 0.4 | 10.0% | 8.00x |
| OpenAI | o3 Pro | 20 | not in the OpenRouter record | 80 | not in the OpenRouter record | 4.00x |
| OpenAI | o3 | 2 | 0.5 | 8 | 25.0% | 4.00x |
| OpenAI | o4 mini | 1.1 | 0.275 | 4.4 | 25.0% | 4.00x |
| OpenAI | o3 mini | 1.1 | 0.55 | 4.4 | 50.0% | 4.00x |
| OpenAI | o1 Pro | 150 | not in the OpenRouter record | 600 | not in the OpenRouter record | 4.00x |
| OpenAI | o1 | 15 | 7.5 | 60 | 50.0% | 4.00x |
| OpenAI | GPT-4.1 | 2 | 0.5 | 8 | 25.0% | 4.00x |
| OpenAI | GPT-4.1 mini | 0.4 | 0.1 | 1.6 | 25.0% | 4.00x |
| OpenAI | GPT-4.1 nano | 0.1 | 0.025 | 0.4 | 25.0% | 4.00x |
| OpenAI | GPT-4o | 2.5 | 1.25 | 10 | 50.0% | 4.00x |
| OpenAI | GPT-4o mini | 0.15 | 0.075 | 0.6 | 50.0% | 4.00x |
| OpenAI | GPT-4 Turbo | 10 | not in the OpenRouter record | 30 | not in the OpenRouter record | 3.00x |
| OpenAI | GPT-4 | 30 | not in the OpenRouter record | 60 | not in the OpenRouter record | 2.00x |
| OpenAI | GPT-3.5 Turbo | 0.5 | not in the OpenRouter record | 1.5 | not in the OpenRouter record | 3.00x |
| Anthropic | Claude Fable 5.1 | 10 | 0.25 | 50 | 2.5% | 5.00x |
| Anthropic | Claude Fable 5 | 10 | 1 | 50 | 10.0% | 5.00x |
| Anthropic | Claude Opus 5 | 5 | 0.5 | 25 | 10.0% | 5.00x |
| Anthropic | Claude Sonnet 5 | 2 | 0.2 | 10 | 10.0% | 5.00x |
| Anthropic | Claude Opus 4.8 | 5 | 0.5 | 25 | 10.0% | 5.00x |
| Anthropic | Claude Opus 4.7 | 5 | 0.5 | 25 | 10.0% | 5.00x |
| Anthropic | Claude Sonnet 4.6 | 3 | 0.3 | 15 | 10.0% | 5.00x |
| Anthropic | Claude Opus 4.6 | 5 | 0.5 | 25 | 10.0% | 5.00x |
| Anthropic | Claude Opus 4.5 | 5 | 0.5 | 25 | 10.0% | 5.00x |
| Anthropic | Claude Sonnet 4.5 | 3 | 0.3 | 15 | 10.0% | 5.00x |
| Anthropic | Claude Haiku 4.5 | 1 | 0.1 | 5 | 10.0% | 5.00x |
| Gemini 3.8 Flash | 0.75 | 0.075 | 3.75 | 10.0% | 5.00x | |
| Gemini 3.7 Flash | 0.75 | 0.075 | 3.75 | 10.0% | 5.00x | |
| Gemini 3.6 Flash | 0.75 | 0.075 | 3.75 | 10.0% | 5.00x | |
| Gemini 3.5 Flash | 1.5 | 0.15 | 9 | 10.0% | 6.00x | |
| Gemini 3.5 Flash Lite | 0.3 | 0.03 | 2.5 | 10.0% | 8.33x | |
| Gemini 3.1 Pro preview | 2 | 0.2 | 12 | 10.0% | 6.00x | |
| Gemini 3.1 Flash Lite | 0.25 | 0.025 | 1.5 | 10.0% | 6.00x | |
| Gemini 3 Flash preview | 0.5 | 0.05 | 3 | 10.0% | 6.00x | |
| Gemini 2.5 Pro | 1.25 | 0.125 | 10 | 10.0% | 8.00x | |
| Gemini 2.5 Flash | 0.3 | 0.03 | 2.5 | 10.0% | 8.33x | |
| Gemini 2.5 Flash Lite | 0.1 | 0.01 | 0.4 | 10.0% | 4.00x | |
| xAI | Grok 4.6 | 2 | 0.5 | 6 | 25.0% | 3.00x |
| xAI | Grok 4.5 | 2 | 0.3 | 6 | 15.0% | 3.00x |
| xAI | Grok 4.3 | 1.25 | 0.2 | 2.5 | 16.0% | 2.00x |
| xAI | Grok 4.20 | 1.25 | 0.2 | 2.5 | 16.0% | 2.00x |
| xAI | Grok Build 0.1 | 1 | 0.2 | 2 | 20.0% | 2.00x |
| Mistral | Mistral Medium 3.5 | 1.5 | not in the OpenRouter record | 7.5 | not in the OpenRouter record | 5.00x |
| Mistral | Mistral Large 3 | 0.5 | 0.05 | 1.5 | 10.0% | 3.00x |
| Mistral | Mistral Small 4 | 0.15 | 0.015 | 0.6 | 10.0% | 4.00x |
| Mistral | Devstral 2 | 0.4 | held, only one record carries it | 2 | not in the OpenRouter record | 5.00x |
| Mistral | Codestral 2508 | 0.3 | 0.03 | 0.9 | 10.0% | 3.00x |
| Mistral | Ministral 3 14B | 0.2 | 0.02 | 0.2 | 10.0% | 1.00x |
| Mistral | Ministral 3 8B | 0.15 | 0.015 | 0.15 | 10.0% | 1.00x |
| Mistral | Ministral 3 3B | 0.1 | 0.01 | 0.1 | 10.0% | 1.00x |
| Moonshot | Kimi K3 | 3 | 0.3 | 15 | 10.0% | 5.00x |
| Moonshot | Kimi K2.6 | 0.95 | 0.16 | 4 | 16.8% | 4.21x |
| Moonshot | Kimi K2 Thinking | 0.6 | 0.15 | 2.5 | 25.0% | 4.17x |
Google's pricing page lists the Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash rates in this table as valid through December 31, 2026, with higher rates starting January 1, 2027, so those three rows change on that date (Gemini API pricing).
Primary record, the OpenRouter models API, which returns an array of Model objects and states all pricing values in USD per token. In each record the prompt field is the price for input processing and the completion field is the price for output generation, and the live response carries an input_cache_read key for the cached read price. Independent cross check, the LiteLLM model price table, which stores each provider's published prices under input_cost_per_token and output_cost_per_token. The last two columns are computed in your browser from the three prices in the row.
Which published prices disagree
A row appears above only when the OpenRouter record and the LiteLLM record agree within one percent after both are converted to dollars per one million tokens. Two models were left out before the run. GPT-5.6 Sol was left out after a probe on 2026-09-06 showed the two records differ by roughly a factor of two on its price. DeepSeek V4 was left out before the run because the same probe showed the two records disagree on its price, which its hosts set differently, so the two record rule cannot clear it. Ten listed models, GPT-5.5 Pro, GPT-5.4 Pro, GPT-5.2 Pro, GPT-5 Pro, o3 Pro, o1 Pro, GPT-4 Turbo, GPT-4, GPT-3.5 Turbo and Mistral Medium 3.5, have no cached input price in their OpenRouter record, so their cache cell reads not in the OpenRouter record and the calculator treats the cache hit rate as zero for them. Three more cached input prices are held rather than shown: GPT-5.1 Codex and GPT-5.1 Codex mini because the two records disagree on the cache read price, and Devstral 2 because only one record carries one. Their input and output prices passed, so those rows stay, with the cache hit rate treated as zero for them. For Mistral Medium 3.5 the LiteLLM table does carry a cached price, but one record alone does not clear the two record rule, so no figure is shown. No figure is shown for a held value, on purpose. A number the two records can't agree on isn't a number you should plan against.
Every model at your workload
The table below ranks every validated model by monthly cost for the token shape, request volume and cache hit rate you entered above. The order isn't a verdict from this page. It's arithmetic on your inputs, and it changes as they do. Providers price input, cached input and output tokens separately, so a model that leads on a short prompt with a long reply can fall several places on a long cached prompt with a short reply. Move the output tokens or the hit rate and watch the rows reorder.
| Provider | Model | Per request | Per day | Per thirty day month |
|---|---|---|---|---|
| Mistral | Ministral 3 3B | $0.000178 | $0.1780 | $5.340 |
| OpenAI | GPT-5 nano | $0.000264 | $0.2640 | $7.920 |
| Mistral | Ministral 3 8B | $0.000267 | $0.2670 | $8.010 |
| Gemini 2.5 Flash Lite | $0.000328 | $0.3280 | $9.840 | |
| OpenAI | GPT-4.1 nano | $0.000340 | $0.3400 | $10.200 |
| Mistral | Ministral 3 14B | $0.000356 | $0.3560 | $10.680 |
| Mistral | Mistral Small 4 | $0.000492 | $0.4920 | $14.760 |
| OpenAI | GPT-4o mini | $0.000540 | $0.5400 | $16.200 |
| Mistral | Codestral 2508 | $0.000834 | $0.8340 | $25.020 |
| OpenAI | GPT-5.6 Luna | $0.000856 | $0.8560 | $25.680 |
| OpenAI | GPT-5.4 nano | $0.000881 | $0.8810 | $26.430 |
| Gemini 3.1 Flash Lite | $0.001070 | $1.070 | $32.100 | |
| OpenAI | GPT-5 mini | $0.001320 | $1.320 | $39.600 |
| OpenAI | GPT-4.1 mini | $0.001360 | $1.360 | $40.800 |
| Mistral | Mistral Large 3 | $0.001390 | $1.390 | $41.700 |
| OpenAI | GPT-5.1 Codex mini | $0.001500 | $1.500 | $45.000 |
| Gemini 3.5 Flash Lite | $0.001634 | $1.634 | $49.020 | |
| Gemini 2.5 Flash | $0.001634 | $1.634 | $49.020 | |
| OpenAI | GPT-3.5 Turbo | $0.001750 | $1.750 | $52.500 |
| Mistral | Devstral 2 | $0.001800 | $1.800 | $54.000 |
| Moonshot | Kimi K2 Thinking | $0.002090 | $2.090 | $62.700 |
| Gemini 3 Flash preview | $0.002140 | $2.140 | $64.200 | |
| xAI | Grok Build 0.1 | $0.002360 | $2.360 | $70.800 |
| Gemini 3.8 Flash | $0.002835 | $2.835 | $85.050 | |
| Gemini 3.7 Flash | $0.002835 | $2.835 | $85.050 | |
| Gemini 3.6 Flash | $0.002835 | $2.835 | $85.050 | |
| xAI | Grok 4.3 | $0.002910 | $2.910 | $87.300 |
| xAI | Grok 4.20 | $0.002910 | $2.910 | $87.300 |
| OpenAI | GPT-5.4 mini | $0.003210 | $3.210 | $96.300 |
| Moonshot | Kimi K2.6 | $0.003268 | $3.268 | $98.040 |
| OpenAI | o4 mini | $0.003740 | $3.740 | $112.20 |
| Anthropic | Claude Haiku 4.5 | $0.003780 | $3.780 | $113.40 |
| OpenAI | o3 mini | $0.003960 | $3.960 | $118.80 |
| xAI | Grok 4.5 | $0.005640 | $5.640 | $169.20 |
| xAI | Grok 4.6 | $0.005800 | $5.800 | $174.00 |
| Gemini 3.5 Flash | $0.006420 | $6.420 | $192.60 | |
| OpenAI | GPT-5.1 Codex Max | $0.006600 | $6.600 | $198.00 |
| OpenAI | GPT-5.1 | $0.006600 | $6.600 | $198.00 |
| OpenAI | GPT-5 | $0.006600 | $6.600 | $198.00 |
| Gemini 2.5 Pro | $0.006600 | $6.600 | $198.00 | |
| Mistral | Mistral Medium 3.5 | $0.006750 | $6.750 | $202.50 |
| OpenAI | o3 | $0.006800 | $6.800 | $204.00 |
| OpenAI | GPT-4.1 | $0.006800 | $6.800 | $204.00 |
| OpenAI | GPT-5.1 Codex | $0.007500 | $7.500 | $225.00 |
| Anthropic | Claude Sonnet 5 | $0.007560 | $7.560 | $226.80 |
| OpenAI | GPT-5.6 Terra | $0.008560 | $8.560 | $256.80 |
| Gemini 3.1 Pro preview | $0.008560 | $8.560 | $256.80 | |
| OpenAI | GPT-4o | $0.009000 | $9.000 | $270.00 |
| OpenAI | GPT-5.3 Codex | $0.009240 | $9.240 | $277.20 |
| OpenAI | GPT-5.2 | $0.009240 | $9.240 | $277.20 |
| OpenAI | GPT-5.2 Codex | $0.009240 | $9.240 | $277.20 |
| OpenAI | GPT-5.4 | $0.0107 | $10.700 | $321.00 |
| Anthropic | Claude Sonnet 4.6 | $0.0113 | $11.340 | $340.20 |
| Anthropic | Claude Sonnet 4.5 | $0.0113 | $11.340 | $340.20 |
| Moonshot | Kimi K3 | $0.0113 | $11.340 | $340.20 |
| Anthropic | Claude Opus 5 | $0.0189 | $18.900 | $567.00 |
| Anthropic | Claude Opus 4.8 | $0.0189 | $18.900 | $567.00 |
| Anthropic | Claude Opus 4.7 | $0.0189 | $18.900 | $567.00 |
| Anthropic | Claude Opus 4.6 | $0.0189 | $18.900 | $567.00 |
| Anthropic | Claude Opus 4.5 | $0.0189 | $18.900 | $567.00 |
| OpenAI | GPT-5.5 | $0.0214 | $21.400 | $642.00 |
| OpenAI | GPT-4 Turbo | $0.0350 | $35.000 | $1050.00 |
| Anthropic | Claude Fable 5.1 | $0.0372 | $37.200 | $1116.00 |
| OpenAI | GPT-6 Astra | $0.0378 | $37.800 | $1134.00 |
| Anthropic | Claude Fable 5 | $0.0378 | $37.800 | $1134.00 |
| OpenAI | o1 | $0.0540 | $54.000 | $1620.00 |
| OpenAI | o3 Pro | $0.0800 | $80.000 | $2400.00 |
| OpenAI | GPT-5 Pro | $0.0900 | $90.000 | $2700.00 |
| OpenAI | GPT-4 | $0.0900 | $90.000 | $2700.00 |
| OpenAI | GPT-5.2 Pro | $0.1260 | $126.00 | $3780.00 |
| OpenAI | GPT-5.5 Pro | $0.1500 | $150.00 | $4500.00 |
| OpenAI | GPT-5.4 Pro | $0.1500 | $150.00 | $4500.00 |
| OpenAI | o1 Pro | $0.6000 | $600.00 | $18000.00 |
Three worked shapes
Each card states its own token counts and its own assumed cache hit rate, and prices a month of one thousand requests a day on the model you selected in the calculator. Change the model up top and all of them recompute.
Classification call
A short prompt, a label back, nothing shared between calls so nothing to cache.
Input tokens 400, output tokens 5, cache hit rate 0.0%, requests per day 1,000.
Monthly cost on OpenAI GPT-5.4
$32.250
Chat turn with a long shared system prompt
Most of the input is the same instructions and tool definitions every time, which is exactly what prompt caching is for. The assumed hit rate reflects that.
Input tokens 6,000, output tokens 400, cache hit rate 80.0%, requests per day 1,000.
Monthly cost on OpenAI GPT-5.4
$306.00
Long document summary
Every document is different, so the cache doesn't help, and the input side dominates.
Input tokens 40,000, output tokens 800, cache hit rate 0.0%, requests per day 1,000.
Monthly cost on OpenAI GPT-5.4
$3360.00
Common questions
How the price per request is computed
The calculator multiplies your input tokens by an effective input price, adds your output tokens times the output price, and divides by one million. The effective input price is the list input price weighted by the share of input tokens that miss the cache, plus the cached input price weighted by the share that hit. OpenRouter states all pricing values in USD per token, so the page converts to dollars per one million tokens in code before anything is displayed. If you count characters rather than tokens, Anthropic gives a rough estimate of about four characters or three quarters of a word per token in English, and Moonshot estimates roughly three to four English characters per token for typical English text.
Why cached input is cheaper and how each provider bills cache writes
A cached read reuses a prompt prefix the provider already processed, and in every row of the table above the cached read price sits below the base input price. Writes are a different story and vary by provider. Anthropic bills prompt cache writes at 1.25x or 2x the base input price and cache reads at a fraction of it. Google charges a separate per hour storage price for context caching on top of the cached token price. OpenAI discounts reused input tokens by up to ninety percent through prompt caching, and the discount only applies once a prompt prefix reaches a minimum cacheable length, just over a thousand tokens for GPT-5.6 and later and about double that for older models. Mistral says cached input tokens cut input cost by up to ninety percent. None of those write charges or minimums are in the calculator, which is why the cache hit rate is a number you enter and not one the page predicts.
What a list price leaves out
- Batch routes. Anthropic's Batch API gives a half price discount on both input and output tokens, Google's Gemini Batch API is offered at a half price reduction versus standard pricing, OpenAI's Batch API offers a half price discount versus its synchronous APIs, and Mistral's batch processing cuts the price by half.
- Long context. Some Gemini models charge a higher input rate on prompts above a two hundred thousand token threshold, and on xAI models with long context pricing a request whose prompt reaches the threshold is billed at the higher rate for every token in the request. Anthropic's pricing page says a 900k token request is billed at the same per token rate as a 9k token request.
- Reasoning. Google bills Gemini thinking tokens as output tokens, so a reasoning heavy call carries more output tokens than the visible reply.
- Tool use. Anthropic adds tool definition and tool result tokens to the billed input and output counts.
- Tokenizers. Anthropic's pricing page says Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, and that Claude Sonnet 4.6 and earlier models use the previous tokenizer, so the same prompt can bill a different token count across model generations. Type the count your own usage logs report, not a count from another model.
Why the break even moves when the cache hit rate moves
Two models rarely discount their cached input by the same fraction of their list input price. So raising the hit rate lowers the two input costs by different amounts while leaving both output costs alone. The output share at which the two tie has to shift to keep the balance, and it can shift enough to flip which model is cheaper for the same token shape. Run the break even tool at a low hit rate and again at a high one and you'll see the crossover slide.
Why open weight models are missing
The roster is first party API models where one provider publishes one list price that two records can be checked against. An open weight model is served by many hosts at many prices, so the two records rarely agree on one figure. DeepSeek V4 was left out for exactly that reason. Use the custom entry in the model select to cost any host's rate you've been quoted. If you are sizing a self hosted model instead of buying tokens, the LLM memory calculator covers that side.
How often prices are refreshed
Both records are re fetched every time the page is rebuilt, and the fetch date beside the table is the date the record you're reading was pulled. LiteLLM describes its table as a community maintained list, so the cross check is independent of OpenRouter but isn't a provider document. Both are linked under the table so you can open them and read the per token strings yourself.
Whether this page sends any data
No. Every figure is computed in your browser from the prices baked into the page at build time and the numbers you type. No request leaves the page when you change an input, and nothing you enter is stored.
Where every number on this page comes from
Every list price on this page is fetched, not typed. The primary source is the OpenRouter models API record for each model, which publishes the prompt and completion price per token as a decimal string, plus an input_cache_read price for the cached read. The independent cross check is the LiteLLM model price table, a separately maintained JSON file that records each provider's published input and output cost per token under input_cost_per_token and output_cost_per_token. Both values are converted to dollars per one million tokens in code. A model whose two sources disagree by more than one percent is held and never shown. Everything the calculator returns is computed in your browser from the validated per million prices and the token counts you type.
I run scheduled production jobs against pay per token model APIs on a prepaid balance. I've watched a fleet's per cycle budget of about three dollars outrun a balance of one dollar and change, and the jobs stopped mid cycle. That's why this page prices a whole workload per day and per month rather than a single request, and why the cache hit rate is a first class input. I read the price of every model listed here straight from the provider records rather than from a pricing page screenshot.
Estimates use standard published formulas. Real results vary with your data, your settings, and your runtime.
This tool is for planning and teaching. Check a result against your own measurement before you rely on it.
All computation runs client side. No data leaves your browser.
ml0x publishes free machine learning calculators and explainers. Every number on this page is computed in your browser from the inputs you enter. Nothing is sent to a server.