What $200 of Claude Max actually buys
We metered one month of agent traffic behind a single Claude Max 20x seat. It came to 33.28 billion input tokens at a 96.4% cache-hit rate. Here's the same workload priced on ten alternatives, and why none of them fits in $200.
Usage measured 2026-07-29 to 2026-08-28 from local Claude Code logs. Prices verified 2026-08-28 against each vendor's own pricing page.
Every number here survived a same-day re-fetch against the vendor's own page; 432 of 446 evidence snippets matched. The complete artifact lives at the full dataset. And the workload is one seat's real traffic, so read it as one data point, not a benchmark.
The measured month
One Claude Max 20x seat absorbed 33.28 billion input tokens and 122 million output tokens in 30 days, measured from local Claude Code session logs.
That's 207,516 deduplicated API messages across 14,973 session files. The parser summed the four billing components per model per day and dropped duplicate request ids, the same dedup rule ccusage applies.
| Billing component | Tokens (30 days) | Share of input |
|---|---|---|
| Cache reads | 32,092M | 96.4% |
| Cache writes | 935M | 2.8% |
| Fresh input | 250M | 0.8% |
| Output | 122M | – |
Agent loops re-read the whole conversation prefix on every call. That's why cache reads are 96% of the bill's raw material, and why the cache-read price ends up deciding everything below.
The same month on ten providers
No pay-as-you-go option reproduces this workload for $200 a month; the floor is $601 on Z.ai's GLM-5.3-Flash, per 2026-08-28 vendor list prices.
Parity is what $200 buys as a share of the measured month. Below 1.0, the option can't replace the plan at the same money.
| Option | Month, billed | Parity at $200 | Reading |
|---|---|---|---|
| Claude Max 20x (this seat) | $200 | 1.00 | delivered the month, 14 throttle events |
| Anthropic API, Opus 5 | $26,187 | 0.008 | like-for-like model, 131× the plan price |
| Anthropic API, Sonnet 5 | $10,475 | 0.019 | cheaper sibling, still 52× |
| OpenAI API, GPT-5.6 Sol | $20,950 | 0.010 | frontier peer, same order as Opus |
| Moonshot API, Kimi K3 | $15,011 | 0.013 | open-weight frontier, $15/M output |
| Z.ai API, GLM-5.3 | $10,539 | 0.019 | the $0.26 cache-read price is the story |
| DeepSeek API, V4 Pro | $2,137 | 0.094 | $0.022 cache reads, off-peak weighted |
| Cloudflare Workers AI, GLM-5.3-Flash | $1,206 | 0.166 | plus $5 Workers Paid, 20 requests/min cap |
| DeepSeek API, V4 Flash | $699 | 0.286 | second-lowest total |
| Z.ai API, GLM-5.3-Flash | $601 | 0.333 | lowest total that runs the whole month |
| Hybrid router, 95% Flash + 5% Opus | $1,973 | 0.101 | LiteLLM budget caps, frontier fallback |
Modelled from the measured token mix and validated list prices. Cache writes bill at each vendor's write rate where one exists, at the input rate where it doesn't. DeepSeek rows carry the observed 23.6% peak-window weighting explained below.
Cache reads decide the bill
Cache-read prices span 37-fold across providers, from DeepSeek V4 Pro's $0.022 to Fable 5's $1.00 per million tokens, and they price 96% of this workload.
Input and output list prices get all the attention. At agent volume they barely matter. Swap the cache-read price and the whole ranking reorders.
| Model | Cache read, $/M | Fresh input, $/M | Output, $/M |
|---|---|---|---|
| DeepSeek V4 Flash | $0.007 | $0.22 | $0.66 |
| GLM-5.3-Flash | $0.015 | $0.075 | $0.25 |
| DeepSeek V4 Pro | $0.022 | $0.66 | $1.98 |
| GPT-5.3-codex | $0.175 | $1.75 | $14.00 |
| GLM-5.3 | $0.26 | $1.40 | $4.40 |
| Kimi K3 | $0.30 | $3.00 | $15.00 |
| GPT-5.6 Sol | $0.40 | $4.00 | $20.00 |
| Claude Opus 5 | $0.50 | $5.00 | $25.00 |
Price your own month
Your mix won't match ours. Put in your own monthly volume and cache-hit rate, and the table reprices on 2026-08-28 list prices. Defaults are the measured month.
| Provider and model | Your month, billed |
|---|---|
| Z.ai GLM-5.3-Flash | $601 |
| DeepSeek V4 Flash, off-peak list | $566 |
| DeepSeek V4 Pro, off-peak list | $1,729 |
| Z.ai GLM-5.3 | $10,539 |
| Moonshot Kimi K2.7-Code | $7,711 |
| Moonshot Kimi K3 | $15,011 |
| OpenAI GPT-5.3-codex | $9,397 |
| OpenAI GPT-5.6 Sol | $20,950 |
| Anthropic Sonnet 5 | $10,475 |
| Anthropic Opus 5 | $26,187 |
Formula per provider: fresh input at the miss rate (or the cache-write rate where the vendor prices writes), cache reads at the hit rate, output at the output rate. DeepSeek rows show off-peak list; weekday peak hours bill 2×.
The subscription break-even
At this cache-heavy mix, Claude Max 20x beats the Anthropic API once monthly volume passes about 255 million tokens, roughly $200 of Opus 5 usage.
Under that line, pay-as-you-go with a spend cap is the better deal and the plan is overpriced for you. Over it, the plan wins. This seat ran 131 times over it.
DeepSeek's peak windows, weighted honestly
DeepSeek bills 2× list during weekday peak windows, 01:00 to 04:00 and 06:00 to 10:00 UTC, per its 2026-08-16 pricing change.
Only 23.6% of this month's tokens landed inside those windows, because agent swarms run nights and weekends too. That weighting puts the effective multiplier at 1.24× off-peak list, which the table above already includes. Shift batch work off-peak and it drops further.
What could flip this verdict
Three live risks, in order of nearness.
The boost lapse
Anthropic's +50% weekly usage boost for Claude Code ends 2026-08-31, per its help center. If delivered capacity drops in week 36, the 131× figure shrinks with it. That's days away, so measure again in September.
Quota opacity
Max limits aren't published in tokens. A rolling 5-hour cap plus an unpublished weekly cap means the multiple can change without a price change. Our logs caught 14 throttle events in the heaviest week.
Thin alternatives
Kimi's top coding tier paused new signups on 2026-07-20. Z.ai's GLM Coding Plan Max covers about 9% of this month in published credits. The escape hatches are narrower than the pricing pages suggest.
Your mix differs
Everything above leans on a 96.4% cache-hit rate. Short sessions and single-shot prompts hit far less cache, and at 60% hit rates the API math moves several multiples in the API's favor.
How we measured this
The workload comes from parsing 14,973 local Claude Code session files covering 2026-07-29 to 2026-08-28, deduplicated by request id, 4.7GB of logs in total. Prices come from each vendor's own pricing page, fetched on 2026-08-28; 432 of 446 evidence snippets re-matched on a second independent fetch the same day. Benchmarks referenced in the ranking are third-party only, from Terminal-Bench, SWE-bench Verified, and Artificial Analysis. A planned live bake-off was skipped because no funded API key was available on the machine, so capability claims rest on those public leaderboards.
One caveat worth repeating. The totals describe one seat's traffic, on one mix, in one month. The method transfers; the totals are yours to re-measure.
Frequently asked
Is Claude Max 20x worth it for agentic coding?
At heavy agent volume, yes by a wide margin. This month's traffic would bill $26,187 at Opus 5 API list prices against a $200 flat fee. The plan beats the API once you pass roughly 255 million tokens a month at a cache-heavy mix.
How many tokens do you get with Claude Max?
Anthropic doesn't say. Limits are a rolling 5-hour session cap plus a weekly cap, neither published in tokens. Our meter shows what one seat absorbed in practice: 33.28 billion input tokens in 30 days, with throttling at the margin.
Is the API cheaper than a Claude subscription?
Below about 255 million tokens a month, yes, and you get hard spend caps with it. Above that line the subscription wins on price and keeps widening. Most casual users sit under the line; most swarm users sit far over it.
What is the cheapest API for high-volume coding agents?
On this measured month, GLM-5.3-Flash at $601 and DeepSeek V4 Flash at $699. Both owe the result to cache-read prices under two cents per million tokens. The frontier names cost 15 to 44 times more for the identical traffic.