Cost study · measured, not modelled

Renting GPUs vs buying tokens

Frontier open-weight models are free to download and expensive to serve. Here is the arithmetic of running one 24/7: real rental prices, the VRAM that actually fits, and the token volume you must reach before owning the compute wins.

Prices verified 2026-08-19. This market re-prices weekly; re-check before committing money.

Every table behind this study, including the ones it summarises away, is published separately: the full dataset and sources. It carries the complete node and card rental grids, model sizes computed from weight files, published throughput measurements, the break-even grid, and the claims an adversarial audit removed.

Break-even volume
4.4B
Output tokens per month needed before an 8x H200 node beats a $4/M API.
Cheapest fitting node
$17,462
Per month, 8x MI300X at $2.99/GPU-hr x 730 hours.
Kimi K3 size
1,561 GB
Native MXFP4. Fits no 8-GPU NVIDIA node in existence.
K3 self-host premium
+34%
$20/M output self-hosted against $15/M list, at full utilization.

What is actually worth renting, and from where

Re-verified on 2026-08-19 against live pricing pages and public catalogue endpoints. Rates on a marketplace move hourly, so medians are quoted rather than the single cheapest listing, which is usually one host in one region and gone by the time you read this.

1. Buy tokens before you rent anything

DeepSeek V4-Flash off-peak runs $0.22 per million input and $0.66 per million output, with cache-hit input at $0.007. Peak covers only seven hours of the day, 01:00 to 04:00 and 06:00 to 10:00 UTC, so off-peak is the default state of the clock and the discount is exactly half. Nothing to provision, no idle burn, no minimum.

DeepSeek pricing

2. If you must hold the weights, use the marketplace

Vast.ai carries the cheapest verified route to a card you control: RTX 3090 at a $0.16 median, RTX 4090 at $0.36, and the 96GB RTX PRO 6000 WS at a $1.20 median against a $0.74 floor. That 96GB card is the only one here that holds a twenty-agent fleet with its context resident. Sort by price, then filter on host reliability, and note that several of the cheapest 4090 listings sit in China.

Vast.ai

3. Pay a little more for a console that behaves

RunPod Community publishes fixed rates: RTX 3090 at $0.22, RTX 4090 at $0.34, RTX PRO 6000 at $1.69. Roughly 35 to 40 percent above the marketplace median for identical silicon, which is a fair price for not running a host lottery. Watch the volume storage, which doubles to $0.20 per GB-month while a pod is stopped.

RunPod pricing

Verified rates worth knowing

ProviderWhatRateStatus
Vast.aiRTX PRO 6000 WS 96GB$0.74 floor, $1.20 medianLive API, today
Vast.aiRTX 5090 32GB$0.29 floor, $0.47 medianLive API, today
RunPodRTX PRO 6000 Community$1.69/hrVerified
KoyebH200 141GB$3.00/hr, scaling linearly to $24.00 for 8xVerified
Hot AisleMI300X 192GB$2.99/GPU-hr VM, $3.39 bare metalPrice rose
SF ComputeH100 clearing market$2.13 average the day of writingMarket, moves daily
MoonshotKimi K3 tokens$0.30 cache hit, $3.00 miss, $15.00 outVerified
LambdaCheapest 24GB card$0.69/hr4.3x the marketplace

Three things this re-check changed. Hot Aisle raised its MI300X rate from $1.99 to $2.99 per GPU-hour for new customers and grandfathered the old price for existing ones, which moves the cheapest fitting node from $11,622 to $17,462 a month. Prime Intellect was dropped: its public pricing block lists H200 four times at the wrong VRAM figure and quotes one identical rate across five different cards, and the live marketplace needs a login, so nothing there could be confirmed. And SF Compute is a clearing market whose seven-day average sat several percent below the day of writing, so it earns a link rather than a hardcoded number.

Two fees that do not appear in any hourly rate. OpenRouter adds 5.5 percent on card credit purchases with a $0.80 minimum, and its advertised cheapest endpoint is frequently deranked, meaning it will not actually serve your traffic. Moonshot quotes K3 excluding tax, which is added at checkout by jurisdiction.

The formula that decides it

A rented GPU bills 730 hours a month whether you send it traffic or not. So break-even does not depend on how fast the hardware is. It depends only on rent and the price of the tokens you would otherwise buy:

break_even_output_tokens = monthly_rent ÷ api_price_per_million

Throughput never appears. It decides something else entirely: whether the required volume is physically reachable on that hardware.

An 8x H200 node at $24.00/hr, the cheapest verifiable rate found, is $17,520 a month. Against a $4.00/M output API that is 4.4 billion output tokens before the node is worth owning. Against a $0.66/M API it is 26.5 billion.

For scale, a developer whose agent generates continuously at 50 tokens per second, eight hours a day, twenty-two days a month, produces 31.7 million output tokens. That is roughly 140x short of the friendlier figure. Even running flat out 24/7 with no pauses, one stream at 50 tok/s yields 131 million a month, still 33x short.

The entry ticket is a node, not a GPU

Every cheap-H100 headline is misleading for this workload, because a frontier open-weight model is a ~1T-parameter mixture of experts. At FP8 that is roughly 1TB of weights before any KV cache.

NodeTotal HBMHolds a ~1T MoE at FP8?
8x H100 80GB640 GBNo. Needs 4-bit, or two nodes.
8x H200 141GB1,128 GBTightly, with thin KV headroom
8x B200 180GB1,440 GBYes
8x MI300X 192GB1,536 GBYes, with real KV headroom
8x MI355X 288GB2,304 GBComfortably

vLLM's own recipe for Kimi K2 at FP8 states the minimum deployment unit is 16 GPUs, not eight. And Kimi K3, at 2.78T parameters and 1,561 GB even in its native MXFP4 form, needs more than 1.5TB before KV cache. It does not fit an 8x B200 node.

The memory-density argument favours AMD hard. Two AWS p5.48xlarge nodes (16x H100) run about $80,358 a month. One 8x MI300X node with 2.4x the HBM runs about $17,462. That is a 4.6x spread before InfiniBand and egress.

Which open-weight models clear a frontier bar

Scores are the Artificial Analysis Intelligence Index, v4.1.1. Where a figure is an estimate rather than a measurement, it is marked, because the difference changes conclusions.

ModelAA IndexActive paramsShipped sizeLicense
Claude Opus 5 closed63n/an/an/a
Kimi K360104B1,561 GBCustom, $20M MaaS gate
Qwen3.8-2.4T-A95B5895B2,496 GBCustom
Claude Opus 4.8 closed57n/an/an/a
GLM-5.253~41B756 GBMIT
DeepSeek-V4-Pro-08135349B893 GBMIT
Qwen3.8-27B5227.8B dense17 GB at Q4Apache-2.0
DeepSeek-V4-Flash-07315213B167 GBMIT
Gemini 3.1 Pro closed48n/an/an/a
Claude Opus 4.6 closed39 estimaten/an/an/a

The ranking inverts when you pay for it. Active parameters drive FLOPs per token, and therefore serving cost. Ranked by index per billion active parameters: GLM-5.2 (1.28) > DeepSeek-V4-Pro (1.09) > Qwen3.8 (0.61) > Kimi K3 (0.57). The two strongest models are the two worst values. If you self-host, DeepSeek-V4-Pro is the pragmatic pick; if you want K3, buy it from the people already running it.

A caution on a comparison currently circulating: Qwen3.8-27B is genuinely remarkable, ranked first of 135 open-weight models in the 4B–40B class against a class median of 9. But claims that it beats Opus 4.6 rest on an Opus figure that Artificial Analysis flags as an estimate, measured in non-reasoning mode. Comparing measured against measured, it is 52 against Opus 5 at 63.

What a real 24/7 workload looks like

Most cost comparisons use a hypothetical user. This one uses a measured month: 1,039,011 assistant messages parsed from the session transcripts of a heavy agentic coding setup running fleets of parallel agents. The shape is the surprise.

Token class30-day totalShareWhat it costs to serve
Cache read61,653,509,01794.1%Near zero. A prefix-cache hit skips prefill entirely.
Cache write2,178,931,3893.3%Full prefill compute, paid once
Fresh input1,306,389,8882.0%Full prefill compute
Output362,119,0580.6%Decode. Memory-bandwidth bound.

Ninety-four percent of the volume is cache-read. That single fact reorders every provider comparison: rank by cache-read price, not by the headline input and output rates. Providers differ by more than 20x on that line while looking near-identical on their pricing pages. Anyone quoting dollars per million output tokens is pricing 0.6% of the bill.

And it cuts both ways. Cache reads are near-free on a self-hosted server, but only if the cache fits in VRAM. A 24GB card leaves roughly 5.3 GB of pooled KV space after weights. Against a 61.7-billion-token working set, contexts cannot stay resident, so every turn re-prefills from scratch. At consumer scale, self-hosting converts the cheapest token class into the most expensive one. At 96GB it starts to work: that card holds about 2.3 million tokens of KV, enough for twenty agents at 110k context each.

Aggregate throughput is not what you feel

Every impressive tokens-per-second figure in the literature is measured at low per-user interactivity. The trade is steep, and it is measured on identical hardware, DeepSeek R1 at FP8 on H200:

Interactivity per usertok/s per GPUAggregate per 8-GPU node
65 tok/s (sluggish)1,0358,280
99 tok/s4883,904
133 tok/s (comfortable)2642,112

Making the model 2.05x faster for the person using it costs 3.92x in aggregate throughput, and therefore 3.92x per token. A widely-quoted 13,860 tok/s peak on an 8x H200 corresponds to a mean time-per-output-token of 203 ms, which is 4.9 tok/s per user and unusable for interactive work.

Two published cost figures do not survive audit. The often-cited $0.21 per million for Kimi K2 counts only the decode GPUs and excludes the four prefill nodes; counting all 128 GPUs gives $0.284, and using today's cheapest on-demand H200 rate gives $0.435, about 2.07x the published number. A $0.33 per million H200 figure implies a GPU cost of $1.23/hr, which is an amortized ownership cost rather than a rental price.

The one figure that does survive is DeepSeek's own published production accounting: 226.75 nodes, 168 billion output tokens in 24 hours, $87,072 per day at a $2.00/GPU-hr internal cost, giving $0.518 per million output tokens at 20 to 22 tok/s per user. That is the floor a self-hoster is chasing, reached with 226 nodes and custom kernels.

Rent or buy

The 2026 GPU market is inflated. The RTX 5090 launched at $1,999 and retails around $4,530, roughly 2.3x list. That inverts the usual advice: buying is weaker than it was, renting stronger.

CardMarketplace rate$/mo at 24/7Buy priceBreakeven
RTX 3090 24GB$0.069–0.138/hr$50–101$1,500117 months
RTX 4090 24GB$0.135–0.336/hr$99–245$3,00015 months
RTX 5090 32GB$0.296–0.402/hr$216–293$4,53019.5 months
RTX PRO 6000 96GB$0.761–1.096/hr$556–800$15,01719.6 months
RX 7900 XTX 24GBn/an/a$710best bandwidth/$

Two findings worth more than the prices. Enterprise clouds are the wrong channel for this workload: one major provider lists no consumer cards at all, and its cheapest part that fits a 27B model costs about 5x the marketplace rate for identical silicon. And a 3090 bought to run 24/7 has a 117-month breakeven against marketplace rates, where European electricity alone nearly matches the rent.

For the buy-it-outright case, a pair of 128GB unified-memory desktop AI boxes lands near $8,000 and runs a 167 GB mixture-of-experts model at a measured 41 tok/s single-stream, 62 to 75 with speculative decoding, and about 350 tok/s aggregate at 32 concurrent streams. Against equivalent-capacity rental that pays back in 3.3 months, and break-even utilization is only about 3.3 hours a day.

Then the API undercuts both. That same model sells for $0.28 per million output tokens. The owned pair, amortized over 24 months and running at its measured peak, works out to $0.437 per million, more expensive at 100% saturation, ignoring your time and all hardware risk. The amortized monthly cost buys 1.44 billion output tokens from the API; the hardware's physical ceiling is 920 million. Buying beats renting a node, and loses to the API. The decision is not financial. It is whether the workload must be local.

Quantization, and where the damage actually lands

Four-bit is not automatically degradation. It depends entirely on the reference precision the vendor shipped.

PrecisionMeasured effectVerdict
FP8Within ~0.4 points of FP16 on MMLU-ProEffectively free
Vendor QAT INT4Trained as the reference checkpoint; ~2x speedup retainedSafe
NVFP4Relative validation-loss error below 1% vs FP8Close to free
Community post-training INT4MMLU-Pro −1.6, HumanEval −8 pointsDamage lands on code

The trap is that post-training 4-bit is the only path that fits some models onto an 8x H100 node, which is exactly the configuration people price when they imagine self-hosting. Across 133 provider endpoints for twelve frontier models, 59 serve FP8, 23 serve FP4, 7 INT4, 3 MXFP4 and 38 do not declare. Roughly a quarter of frontier endpoints are 4-bit, and on aggregator routing, price sorting will silently send you to them. Pin providers explicitly.

What the adversarial audit removed

This study was produced by parallel research agents and then handed to a separate reviewer whose only instruction was to refute them, defaulting to failure on anything it could not verify against a primary source. It killed five claims. They are listed because a cost study that hides its corrections should not be trusted with your money.

Claim as first researchedVerdictCorrection
Cheapest viable node is $9,986/mo at a named vendorFailedThat vendor publishes no public pricing. The rate existed only in SEO aggregators, and two live price trackers list no such figure. Replaced with a verified $1.99/GPU-hr, which a later re-check found had itself risen.
Cheapest fitting node costs $11,622/moCorrected after publicationHot Aisle raised its MI300X rate from $1.99 to $2.99 per GPU-hour for new customers, keeping $1.99 only for existing ones. The node figure is now $17,462. A price verified once is not verified forever.
Cheapest 8x H200 is $28.16/hrCorrectedReal, but not cheapest. A verified $3.00/GPU-hr exists. All break-even figures recomputed.
Opus 4.6 scores 44.9 in adaptive modeFailedNo such figure exists. One variant is published, scoring 39, itself flagged an estimate.
Qwen3.8-27B scores around 30FailedWrong by 22 points. The measured value is 52, first of 135 open-weight models in its size class.
Free EU supercomputing hours cover production useCorrectedFree access is scoped to innovation purposes; commercial production is a separate paid track.

One figure survives with a caveat attached rather than a correction. The claim that self-hosting Kimi K3 costs more than its own API rests on a published throughput number the auditor could not source. Both prices in that calculation are confirmed; the throughput is not. The direction holds regardless, because K3 needs ten to fourteen GPUs, but the exact figure should not be quoted as hard.

How to read this before spending

Rent, do not buy, unless it must be local

At 2.3x list prices, breakeven on a current-generation card is 15 to 20 months against marketplace rental. Data sensitivity is a far stronger argument for owning hardware than economics currently is.

Choose the model before the hardware

Bandwidth-per-dollar and capacity-per-dollar point in opposite directions. A 32GB card is roughly ten times better on bandwidth per dollar; a 128GB box is two and a half times better on capacity. Which matters depends entirely on whether your model fits.

Price the cache, not the output

On a cache-heavy agentic workload the cache-read rate is about 95% of the bill. Providers vary more than 20x on that line and look identical on the front page.

Spot is for training, not serving

Preemption notice ranges from two minutes down to five seconds, and some marketplaces give none at all. A reliable endpoint needs a warm standby, which roughly doubles the effective rate and erases the discount.

Method

Workload figures were produced by parsing the usage field of 1,039,011 assistant messages across 13,587 local session transcript files, covering 2026-07-18 to 2026-08-19. No personal or account-identifying data appears on this page.

Model sizes were computed by summing .safetensors byte counts from the Hugging Face model API, and architecture figures were read from each model's own config.json rather than from a summary. Prices were read from provider pricing pages and public unauthenticated catalogue endpoints on 2026-08-19. All cost arithmetic is shown inline so it can be checked.

Figures described as measured were produced by running code or reading a primary source. Figures marked as estimates were not. Five claims were corrected or removed by adversarial audit after first drafting, and all are disclosed above. GPU and token pricing in this market moves weekly; verify before committing money.

Related