LLM API price calculator with the cache hit rate priced in

A price table tells you what a million tokens cost. It doesn't tell you what your workload costs once most of your input tokens are served from the provider's prompt cache, and it doesn't tell you where two models trade places. This calculator prices your own token shape per request, per day and per month with the cache hit rate included, then finds the break even between any two models. Every list price was read from two independent records, the OpenRouter models API and the LiteLLM model price table, and converted to dollars per one million tokens in code. Nothing here is typed by hand.

Price your workload

Cost per request$0.0107
Cost per day$10.700
Cost per thirty day month$321.00
Share of cost that is output tokens70.1%
Saved per month by the cache, against zero cache$54.000

The effective input price is the list input price weighted by the share of input tokens that miss the cache, plus the cached input price weighted by the share that hit. Output tokens are billed at the output price in full. Every result recomputes as you type, in your browser, with no network call.

Break even between two models

Pick two models. The page solves, from the validated list prices and your current inputs, for the output share of tokens at which both cost the same per request, and for the cache hit rate at which they cost the same with your current token shape. Either crossover can fall outside the reachable range, and when it does the page says so instead of printing a figure you can't act on.

Output share where A and B cost the same, at your cache hit rate21.9%
Cache hit rate where A and B cost the same, at your token shape50.0%

At your current inputs Google Gemini 3.8 Flash costs less than xAI Grok 4.3 per request, by a factor of 1.03. Move the output share or the cache hit rate across the crossover above and the order flips.

List prices per one million tokens

Every figure below is the provider's standard list price in dollars per one million tokens, fetched on 2026-09-06. Batch, cached input write, long context and reasoning surcharges are not included. The cached input column is the cached read price only.

ProviderModelInput, USD per one million tokensCached input, USD per one million tokensOutput, USD per one million tokensCached input as a share of inputOutput to input ratio
OpenAIGPT-6 Astra1015010.0%5.00x
OpenAIGPT-5.6 Terra20.21210.0%6.00x
OpenAIGPT-5.6 Luna0.20.021.210.0%6.00x
OpenAIGPT-5.5 Pro30not in the OpenRouter record180not in the OpenRouter record6.00x
OpenAIGPT-5.550.53010.0%6.00x
OpenAIGPT-5.4 Pro30not in the OpenRouter record180not in the OpenRouter record6.00x
OpenAIGPT-5.42.50.251510.0%6.00x
OpenAIGPT-5.4 mini0.750.0754.510.0%6.00x
OpenAIGPT-5.4 nano0.20.021.2510.0%6.25x
OpenAIGPT-5.3 Codex1.750.1751410.0%8.00x
OpenAIGPT-5.2 Pro21not in the OpenRouter record168not in the OpenRouter record8.00x
OpenAIGPT-5.21.750.1751410.0%8.00x
OpenAIGPT-5.2 Codex1.750.1751410.0%8.00x
OpenAIGPT-5.1 Codex Max1.250.1251010.0%8.00x
OpenAIGPT-5.11.250.1251010.0%8.00x
OpenAIGPT-5.1 Codex1.25held, the two records disagree10not in the OpenRouter record8.00x
OpenAIGPT-5.1 Codex mini0.25held, the two records disagree2not in the OpenRouter record8.00x
OpenAIGPT-5 Pro15not in the OpenRouter record120not in the OpenRouter record8.00x
OpenAIGPT-51.250.1251010.0%8.00x
OpenAIGPT-5 mini0.250.025210.0%8.00x
OpenAIGPT-5 nano0.050.0050.410.0%8.00x
OpenAIo3 Pro20not in the OpenRouter record80not in the OpenRouter record4.00x
OpenAIo320.5825.0%4.00x
OpenAIo4 mini1.10.2754.425.0%4.00x
OpenAIo3 mini1.10.554.450.0%4.00x
OpenAIo1 Pro150not in the OpenRouter record600not in the OpenRouter record4.00x
OpenAIo1157.56050.0%4.00x
OpenAIGPT-4.120.5825.0%4.00x
OpenAIGPT-4.1 mini0.40.11.625.0%4.00x
OpenAIGPT-4.1 nano0.10.0250.425.0%4.00x
OpenAIGPT-4o2.51.251050.0%4.00x
OpenAIGPT-4o mini0.150.0750.650.0%4.00x
OpenAIGPT-4 Turbo10not in the OpenRouter record30not in the OpenRouter record3.00x
OpenAIGPT-430not in the OpenRouter record60not in the OpenRouter record2.00x
OpenAIGPT-3.5 Turbo0.5not in the OpenRouter record1.5not in the OpenRouter record3.00x
AnthropicClaude Fable 5.1100.25502.5%5.00x
AnthropicClaude Fable 51015010.0%5.00x
AnthropicClaude Opus 550.52510.0%5.00x
AnthropicClaude Sonnet 520.21010.0%5.00x
AnthropicClaude Opus 4.850.52510.0%5.00x
AnthropicClaude Opus 4.750.52510.0%5.00x
AnthropicClaude Sonnet 4.630.31510.0%5.00x
AnthropicClaude Opus 4.650.52510.0%5.00x
AnthropicClaude Opus 4.550.52510.0%5.00x
AnthropicClaude Sonnet 4.530.31510.0%5.00x
AnthropicClaude Haiku 4.510.1510.0%5.00x
GoogleGemini 3.8 Flash0.750.0753.7510.0%5.00x
GoogleGemini 3.7 Flash0.750.0753.7510.0%5.00x
GoogleGemini 3.6 Flash0.750.0753.7510.0%5.00x
GoogleGemini 3.5 Flash1.50.15910.0%6.00x
GoogleGemini 3.5 Flash Lite0.30.032.510.0%8.33x
GoogleGemini 3.1 Pro preview20.21210.0%6.00x
GoogleGemini 3.1 Flash Lite0.250.0251.510.0%6.00x
GoogleGemini 3 Flash preview0.50.05310.0%6.00x
GoogleGemini 2.5 Pro1.250.1251010.0%8.00x
GoogleGemini 2.5 Flash0.30.032.510.0%8.33x
GoogleGemini 2.5 Flash Lite0.10.010.410.0%4.00x
xAIGrok 4.620.5625.0%3.00x
xAIGrok 4.520.3615.0%3.00x
xAIGrok 4.31.250.22.516.0%2.00x
xAIGrok 4.201.250.22.516.0%2.00x
xAIGrok Build 0.110.2220.0%2.00x
MistralMistral Medium 3.51.5not in the OpenRouter record7.5not in the OpenRouter record5.00x
MistralMistral Large 30.50.051.510.0%3.00x
MistralMistral Small 40.150.0150.610.0%4.00x
MistralDevstral 20.4held, only one record carries it2not in the OpenRouter record5.00x
MistralCodestral 25080.30.030.910.0%3.00x
MistralMinistral 3 14B0.20.020.210.0%1.00x
MistralMinistral 3 8B0.150.0150.1510.0%1.00x
MistralMinistral 3 3B0.10.010.110.0%1.00x
MoonshotKimi K330.31510.0%5.00x
MoonshotKimi K2.60.950.16416.8%4.21x
MoonshotKimi K2 Thinking0.60.152.525.0%4.17x

Google's pricing page lists the Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash rates in this table as valid through December 31, 2026, with higher rates starting January 1, 2027, so those three rows change on that date (Gemini API pricing).

Primary record, the OpenRouter models API, which returns an array of Model objects and states all pricing values in USD per token. In each record the prompt field is the price for input processing and the completion field is the price for output generation, and the live response carries an input_cache_read key for the cached read price. Independent cross check, the LiteLLM model price table, which stores each provider's published prices under input_cost_per_token and output_cost_per_token. The last two columns are computed in your browser from the three prices in the row.

Which published prices disagree

A row appears above only when the OpenRouter record and the LiteLLM record agree within one percent after both are converted to dollars per one million tokens. Two models were left out before the run. GPT-5.6 Sol was left out after a probe on 2026-09-06 showed the two records differ by roughly a factor of two on its price. DeepSeek V4 was left out before the run because the same probe showed the two records disagree on its price, which its hosts set differently, so the two record rule cannot clear it. Ten listed models, GPT-5.5 Pro, GPT-5.4 Pro, GPT-5.2 Pro, GPT-5 Pro, o3 Pro, o1 Pro, GPT-4 Turbo, GPT-4, GPT-3.5 Turbo and Mistral Medium 3.5, have no cached input price in their OpenRouter record, so their cache cell reads not in the OpenRouter record and the calculator treats the cache hit rate as zero for them. Three more cached input prices are held rather than shown: GPT-5.1 Codex and GPT-5.1 Codex mini because the two records disagree on the cache read price, and Devstral 2 because only one record carries one. Their input and output prices passed, so those rows stay, with the cache hit rate treated as zero for them. For Mistral Medium 3.5 the LiteLLM table does carry a cached price, but one record alone does not clear the two record rule, so no figure is shown. No figure is shown for a held value, on purpose. A number the two records can't agree on isn't a number you should plan against.

Every model at your workload

The table below ranks every validated model by monthly cost for the token shape, request volume and cache hit rate you entered above. The order isn't a verdict from this page. It's arithmetic on your inputs, and it changes as they do. Providers price input, cached input and output tokens separately, so a model that leads on a short prompt with a long reply can fall several places on a long cached prompt with a short reply. Move the output tokens or the hit rate and watch the rows reorder.

ProviderModelPer requestPer dayPer thirty day month
MistralMinistral 3 3B$0.000178$0.1780$5.340
OpenAIGPT-5 nano$0.000264$0.2640$7.920
MistralMinistral 3 8B$0.000267$0.2670$8.010
GoogleGemini 2.5 Flash Lite$0.000328$0.3280$9.840
OpenAIGPT-4.1 nano$0.000340$0.3400$10.200
MistralMinistral 3 14B$0.000356$0.3560$10.680
MistralMistral Small 4$0.000492$0.4920$14.760
OpenAIGPT-4o mini$0.000540$0.5400$16.200
MistralCodestral 2508$0.000834$0.8340$25.020
OpenAIGPT-5.6 Luna$0.000856$0.8560$25.680
OpenAIGPT-5.4 nano$0.000881$0.8810$26.430
GoogleGemini 3.1 Flash Lite$0.001070$1.070$32.100
OpenAIGPT-5 mini$0.001320$1.320$39.600
OpenAIGPT-4.1 mini$0.001360$1.360$40.800
MistralMistral Large 3$0.001390$1.390$41.700
OpenAIGPT-5.1 Codex mini$0.001500$1.500$45.000
GoogleGemini 3.5 Flash Lite$0.001634$1.634$49.020
GoogleGemini 2.5 Flash$0.001634$1.634$49.020
OpenAIGPT-3.5 Turbo$0.001750$1.750$52.500
MistralDevstral 2$0.001800$1.800$54.000
MoonshotKimi K2 Thinking$0.002090$2.090$62.700
GoogleGemini 3 Flash preview$0.002140$2.140$64.200
xAIGrok Build 0.1$0.002360$2.360$70.800
GoogleGemini 3.8 Flash$0.002835$2.835$85.050
GoogleGemini 3.7 Flash$0.002835$2.835$85.050
GoogleGemini 3.6 Flash$0.002835$2.835$85.050
xAIGrok 4.3$0.002910$2.910$87.300
xAIGrok 4.20$0.002910$2.910$87.300
OpenAIGPT-5.4 mini$0.003210$3.210$96.300
MoonshotKimi K2.6$0.003268$3.268$98.040
OpenAIo4 mini$0.003740$3.740$112.20
AnthropicClaude Haiku 4.5$0.003780$3.780$113.40
OpenAIo3 mini$0.003960$3.960$118.80
xAIGrok 4.5$0.005640$5.640$169.20
xAIGrok 4.6$0.005800$5.800$174.00
GoogleGemini 3.5 Flash$0.006420$6.420$192.60
OpenAIGPT-5.1 Codex Max$0.006600$6.600$198.00
OpenAIGPT-5.1$0.006600$6.600$198.00
OpenAIGPT-5$0.006600$6.600$198.00
GoogleGemini 2.5 Pro$0.006600$6.600$198.00
MistralMistral Medium 3.5$0.006750$6.750$202.50
OpenAIo3$0.006800$6.800$204.00
OpenAIGPT-4.1$0.006800$6.800$204.00
OpenAIGPT-5.1 Codex$0.007500$7.500$225.00
AnthropicClaude Sonnet 5$0.007560$7.560$226.80
OpenAIGPT-5.6 Terra$0.008560$8.560$256.80
GoogleGemini 3.1 Pro preview$0.008560$8.560$256.80
OpenAIGPT-4o$0.009000$9.000$270.00
OpenAIGPT-5.3 Codex$0.009240$9.240$277.20
OpenAIGPT-5.2$0.009240$9.240$277.20
OpenAIGPT-5.2 Codex$0.009240$9.240$277.20
OpenAIGPT-5.4$0.0107$10.700$321.00
AnthropicClaude Sonnet 4.6$0.0113$11.340$340.20
AnthropicClaude Sonnet 4.5$0.0113$11.340$340.20
MoonshotKimi K3$0.0113$11.340$340.20
AnthropicClaude Opus 5$0.0189$18.900$567.00
AnthropicClaude Opus 4.8$0.0189$18.900$567.00
AnthropicClaude Opus 4.7$0.0189$18.900$567.00
AnthropicClaude Opus 4.6$0.0189$18.900$567.00
AnthropicClaude Opus 4.5$0.0189$18.900$567.00
OpenAIGPT-5.5$0.0214$21.400$642.00
OpenAIGPT-4 Turbo$0.0350$35.000$1050.00
AnthropicClaude Fable 5.1$0.0372$37.200$1116.00
OpenAIGPT-6 Astra$0.0378$37.800$1134.00
AnthropicClaude Fable 5$0.0378$37.800$1134.00
OpenAIo1$0.0540$54.000$1620.00
OpenAIo3 Pro$0.0800$80.000$2400.00
OpenAIGPT-5 Pro$0.0900$90.000$2700.00
OpenAIGPT-4$0.0900$90.000$2700.00
OpenAIGPT-5.2 Pro$0.1260$126.00$3780.00
OpenAIGPT-5.5 Pro$0.1500$150.00$4500.00
OpenAIGPT-5.4 Pro$0.1500$150.00$4500.00
OpenAIo1 Pro$0.6000$600.00$18000.00

Three worked shapes

Each card states its own token counts and its own assumed cache hit rate, and prices a month of one thousand requests a day on the model you selected in the calculator. Change the model up top and all of them recompute.

Classification call

A short prompt, a label back, nothing shared between calls so nothing to cache.

Input tokens 400, output tokens 5, cache hit rate 0.0%, requests per day 1,000.

Monthly cost on OpenAI GPT-5.4

$32.250

Chat turn with a long shared system prompt

Most of the input is the same instructions and tool definitions every time, which is exactly what prompt caching is for. The assumed hit rate reflects that.

Input tokens 6,000, output tokens 400, cache hit rate 80.0%, requests per day 1,000.

Monthly cost on OpenAI GPT-5.4

$306.00

Long document summary

Every document is different, so the cache doesn't help, and the input side dominates.

Input tokens 40,000, output tokens 800, cache hit rate 0.0%, requests per day 1,000.

Monthly cost on OpenAI GPT-5.4

$3360.00

Common questions

How the price per request is computed

The calculator multiplies your input tokens by an effective input price, adds your output tokens times the output price, and divides by one million. The effective input price is the list input price weighted by the share of input tokens that miss the cache, plus the cached input price weighted by the share that hit. OpenRouter states all pricing values in USD per token, so the page converts to dollars per one million tokens in code before anything is displayed. If you count characters rather than tokens, Anthropic gives a rough estimate of about four characters or three quarters of a word per token in English, and Moonshot estimates roughly three to four English characters per token for typical English text.

Why cached input is cheaper and how each provider bills cache writes

A cached read reuses a prompt prefix the provider already processed, and in every row of the table above the cached read price sits below the base input price. Writes are a different story and vary by provider. Anthropic bills prompt cache writes at 1.25x or 2x the base input price and cache reads at a fraction of it. Google charges a separate per hour storage price for context caching on top of the cached token price. OpenAI discounts reused input tokens by up to ninety percent through prompt caching, and the discount only applies once a prompt prefix reaches a minimum cacheable length, just over a thousand tokens for GPT-5.6 and later and about double that for older models. Mistral says cached input tokens cut input cost by up to ninety percent. None of those write charges or minimums are in the calculator, which is why the cache hit rate is a number you enter and not one the page predicts.

What a list price leaves out

  • Batch routes. Anthropic's Batch API gives a half price discount on both input and output tokens, Google's Gemini Batch API is offered at a half price reduction versus standard pricing, OpenAI's Batch API offers a half price discount versus its synchronous APIs, and Mistral's batch processing cuts the price by half.
  • Long context. Some Gemini models charge a higher input rate on prompts above a two hundred thousand token threshold, and on xAI models with long context pricing a request whose prompt reaches the threshold is billed at the higher rate for every token in the request. Anthropic's pricing page says a 900k token request is billed at the same per token rate as a 9k token request.
  • Reasoning. Google bills Gemini thinking tokens as output tokens, so a reasoning heavy call carries more output tokens than the visible reply.
  • Tool use. Anthropic adds tool definition and tool result tokens to the billed input and output counts.
  • Tokenizers. Anthropic's pricing page says Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, and that Claude Sonnet 4.6 and earlier models use the previous tokenizer, so the same prompt can bill a different token count across model generations. Type the count your own usage logs report, not a count from another model.

Why the break even moves when the cache hit rate moves

Two models rarely discount their cached input by the same fraction of their list input price. So raising the hit rate lowers the two input costs by different amounts while leaving both output costs alone. The output share at which the two tie has to shift to keep the balance, and it can shift enough to flip which model is cheaper for the same token shape. Run the break even tool at a low hit rate and again at a high one and you'll see the crossover slide.

Why open weight models are missing

The roster is first party API models where one provider publishes one list price that two records can be checked against. An open weight model is served by many hosts at many prices, so the two records rarely agree on one figure. DeepSeek V4 was left out for exactly that reason. Use the custom entry in the model select to cost any host's rate you've been quoted. If you are sizing a self hosted model instead of buying tokens, the LLM memory calculator covers that side.

How often prices are refreshed

Both records are re fetched every time the page is rebuilt, and the fetch date beside the table is the date the record you're reading was pulled. LiteLLM describes its table as a community maintained list, so the cross check is independent of OpenRouter but isn't a provider document. Both are linked under the table so you can open them and read the per token strings yourself.

Whether this page sends any data

No. Every figure is computed in your browser from the prices baked into the page at build time and the numbers you type. No request leaves the page when you change an input, and nothing you enter is stored.

Where every number on this page comes from

Every list price on this page is fetched, not typed. The primary source is the OpenRouter models API record for each model, which publishes the prompt and completion price per token as a decimal string, plus an input_cache_read price for the cached read. The independent cross check is the LiteLLM model price table, a separately maintained JSON file that records each provider's published input and output cost per token under input_cost_per_token and output_cost_per_token. Both values are converted to dollars per one million tokens in code. A model whose two sources disagree by more than one percent is held and never shown. Everything the calculator returns is computed in your browser from the validated per million prices and the token counts you type.

I run scheduled production jobs against pay per token model APIs on a prepaid balance. I've watched a fleet's per cycle budget of about three dollars outrun a balance of one dollar and change, and the jobs stopped mid cycle. That's why this page prices a whole workload per day and per month rather than a single request, and why the cache hit rate is a first class input. I read the price of every model listed here straight from the provider records rather than from a pricing page screenshot.

Estimates use standard published formulas. Real results vary with your data, your settings, and your runtime.

This tool is for planning and teaching. Check a result against your own measurement before you rely on it.

All computation runs client side. No data leaves your browser.

ml0x publishes free machine learning calculators and explainers. Every number on this page is computed in your browser from the inputs you enter. Nothing is sent to a server.