Renting GPUs vs buying tokens
Frontier open-weight models are free to download and expensive to serve. Here is the arithmetic of running one 24/7: real rental prices, the VRAM that actually fits, and the token volume you must reach before owning the compute wins.
Prices verified 2026-08-19. This market re-prices weekly; re-check before committing money.
Every table behind this study, including the ones it summarises away, is published separately: the full dataset and sources. It carries the complete node and card rental grids, model sizes computed from weight files, published throughput measurements, the break-even grid, and the claims an adversarial audit removed.
What is actually worth renting, and from where
Re-verified on 2026-08-19 against live pricing pages and public catalogue endpoints. Rates on a marketplace move hourly, so medians are quoted rather than the single cheapest listing, which is usually one host in one region and gone by the time you read this.
1. Buy tokens before you rent anything
DeepSeek V4-Flash off-peak runs $0.22 per million input and $0.66 per million output, with cache-hit input at $0.007. Peak covers only seven hours of the day, 01:00 to 04:00 and 06:00 to 10:00 UTC, so off-peak is the default state of the clock and the discount is exactly half. Nothing to provision, no idle burn, no minimum.
2. If you must hold the weights, use the marketplace
Vast.ai carries the cheapest verified route to a card you control: RTX 3090 at a $0.16 median, RTX 4090 at $0.36, and the 96GB RTX PRO 6000 WS at a $1.20 median against a $0.74 floor. That 96GB card is the only one here that holds a twenty-agent fleet with its context resident. Sort by price, then filter on host reliability, and note that several of the cheapest 4090 listings sit in China.
3. Pay a little more for a console that behaves
RunPod Community publishes fixed rates: RTX 3090 at $0.22, RTX 4090 at $0.34, RTX PRO 6000 at $1.69. Roughly 35 to 40 percent above the marketplace median for identical silicon, which is a fair price for not running a host lottery. Watch the volume storage, which doubles to $0.20 per GB-month while a pod is stopped.
Verified rates worth knowing
| Provider | What | Rate | Status |
|---|---|---|---|
| Vast.ai | RTX PRO 6000 WS 96GB | $0.74 floor, $1.20 median | Live API, today |
| Vast.ai | RTX 5090 32GB | $0.29 floor, $0.47 median | Live API, today |
| RunPod | RTX PRO 6000 Community | $1.69/hr | Verified |
| Koyeb | H200 141GB | $3.00/hr, scaling linearly to $24.00 for 8x | Verified |
| Hot Aisle | MI300X 192GB | $2.99/GPU-hr VM, $3.39 bare metal | Price rose |
| SF Compute | H100 clearing market | $2.13 average the day of writing | Market, moves daily |
| Moonshot | Kimi K3 tokens | $0.30 cache hit, $3.00 miss, $15.00 out | Verified |
| Lambda | Cheapest 24GB card | $0.69/hr | 4.3x the marketplace |
Three things this re-check changed. Hot Aisle raised its MI300X rate from $1.99 to $2.99 per GPU-hour for new customers and grandfathered the old price for existing ones, which moves the cheapest fitting node from $11,622 to $17,462 a month. Prime Intellect was dropped: its public pricing block lists H200 four times at the wrong VRAM figure and quotes one identical rate across five different cards, and the live marketplace needs a login, so nothing there could be confirmed. And SF Compute is a clearing market whose seven-day average sat several percent below the day of writing, so it earns a link rather than a hardcoded number.
Two fees that do not appear in any hourly rate. OpenRouter adds 5.5 percent on card credit purchases with a $0.80 minimum, and its advertised cheapest endpoint is frequently deranked, meaning it will not actually serve your traffic. Moonshot quotes K3 excluding tax, which is added at checkout by jurisdiction.
The formula that decides it
A rented GPU bills 730 hours a month whether you send it traffic or not. So break-even does not depend on how fast the hardware is. It depends only on rent and the price of the tokens you would otherwise buy:
break_even_output_tokens = monthly_rent ÷ api_price_per_million
Throughput never appears. It decides something else entirely: whether the required volume is physically reachable on that hardware.
An 8x H200 node at $24.00/hr, the cheapest verifiable rate found, is $17,520 a month. Against a $4.00/M output API that is 4.4 billion output tokens before the node is worth owning. Against a $0.66/M API it is 26.5 billion.
For scale, a developer whose agent generates continuously at 50 tokens per second, eight hours a day, twenty-two days a month, produces 31.7 million output tokens. That is roughly 140x short of the friendlier figure. Even running flat out 24/7 with no pauses, one stream at 50 tok/s yields 131 million a month, still 33x short.
The entry ticket is a node, not a GPU
Every cheap-H100 headline is misleading for this workload, because a frontier open-weight model is a ~1T-parameter mixture of experts. At FP8 that is roughly 1TB of weights before any KV cache.
| Node | Total HBM | Holds a ~1T MoE at FP8? |
|---|---|---|
| 8x H100 80GB | 640 GB | No. Needs 4-bit, or two nodes. |
| 8x H200 141GB | 1,128 GB | Tightly, with thin KV headroom |
| 8x B200 180GB | 1,440 GB | Yes |
| 8x MI300X 192GB | 1,536 GB | Yes, with real KV headroom |
| 8x MI355X 288GB | 2,304 GB | Comfortably |
vLLM's own recipe for Kimi K2 at FP8 states the minimum deployment unit is 16 GPUs, not eight. And Kimi K3, at 2.78T parameters and 1,561 GB even in its native MXFP4 form, needs more than 1.5TB before KV cache. It does not fit an 8x B200 node.
The memory-density argument favours AMD hard. Two AWS p5.48xlarge nodes (16x H100) run about $80,358 a month. One 8x MI300X node with 2.4x the HBM runs about $17,462. That is a 4.6x spread before InfiniBand and egress.
Which open-weight models clear a frontier bar
Scores are the Artificial Analysis Intelligence Index, v4.1.1. Where a figure is an estimate rather than a measurement, it is marked, because the difference changes conclusions.
| Model | AA Index | Active params | Shipped size | License |
|---|---|---|---|---|
| Claude Opus 5 closed | 63 | n/a | n/a | n/a |
| Kimi K3 | 60 | 104B | 1,561 GB | Custom, $20M MaaS gate |
| Qwen3.8-2.4T-A95B | 58 | 95B | 2,496 GB | Custom |
| Claude Opus 4.8 closed | 57 | n/a | n/a | n/a |
| GLM-5.2 | 53 | ~41B | 756 GB | MIT |
| DeepSeek-V4-Pro-0813 | 53 | 49B | 893 GB | MIT |
| Qwen3.8-27B | 52 | 27.8B dense | 17 GB at Q4 | Apache-2.0 |
| DeepSeek-V4-Flash-0731 | 52 | 13B | 167 GB | MIT |
| Gemini 3.1 Pro closed | 48 | n/a | n/a | n/a |
| Claude Opus 4.6 closed | 39 estimate | n/a | n/a | n/a |
The ranking inverts when you pay for it. Active parameters drive FLOPs per token, and therefore serving cost. Ranked by index per billion active parameters: GLM-5.2 (1.28) > DeepSeek-V4-Pro (1.09) > Qwen3.8 (0.61) > Kimi K3 (0.57). The two strongest models are the two worst values. If you self-host, DeepSeek-V4-Pro is the pragmatic pick; if you want K3, buy it from the people already running it.
A caution on a comparison currently circulating: Qwen3.8-27B is genuinely remarkable, ranked first of 135 open-weight models in the 4B–40B class against a class median of 9. But claims that it beats Opus 4.6 rest on an Opus figure that Artificial Analysis flags as an estimate, measured in non-reasoning mode. Comparing measured against measured, it is 52 against Opus 5 at 63.
What a real 24/7 workload looks like
Most cost comparisons use a hypothetical user. This one uses a measured month: 1,039,011 assistant messages parsed from the session transcripts of a heavy agentic coding setup running fleets of parallel agents. The shape is the surprise.
| Token class | 30-day total | Share | What it costs to serve |
|---|---|---|---|
| Cache read | 61,653,509,017 | 94.1% | Near zero. A prefix-cache hit skips prefill entirely. |
| Cache write | 2,178,931,389 | 3.3% | Full prefill compute, paid once |
| Fresh input | 1,306,389,888 | 2.0% | Full prefill compute |
| Output | 362,119,058 | 0.6% | Decode. Memory-bandwidth bound. |
Ninety-four percent of the volume is cache-read. That single fact reorders every provider comparison: rank by cache-read price, not by the headline input and output rates. Providers differ by more than 20x on that line while looking near-identical on their pricing pages. Anyone quoting dollars per million output tokens is pricing 0.6% of the bill.
And it cuts both ways. Cache reads are near-free on a self-hosted server, but only if the cache fits in VRAM. A 24GB card leaves roughly 5.3 GB of pooled KV space after weights. Against a 61.7-billion-token working set, contexts cannot stay resident, so every turn re-prefills from scratch. At consumer scale, self-hosting converts the cheapest token class into the most expensive one. At 96GB it starts to work: that card holds about 2.3 million tokens of KV, enough for twenty agents at 110k context each.
Aggregate throughput is not what you feel
Every impressive tokens-per-second figure in the literature is measured at low per-user interactivity. The trade is steep, and it is measured on identical hardware, DeepSeek R1 at FP8 on H200:
| Interactivity per user | tok/s per GPU | Aggregate per 8-GPU node |
|---|---|---|
| 65 tok/s (sluggish) | 1,035 | 8,280 |
| 99 tok/s | 488 | 3,904 |
| 133 tok/s (comfortable) | 264 | 2,112 |
Making the model 2.05x faster for the person using it costs 3.92x in aggregate throughput, and therefore 3.92x per token. A widely-quoted 13,860 tok/s peak on an 8x H200 corresponds to a mean time-per-output-token of 203 ms, which is 4.9 tok/s per user and unusable for interactive work.
Two published cost figures do not survive audit. The often-cited $0.21 per million for Kimi K2 counts only the decode GPUs and excludes the four prefill nodes; counting all 128 GPUs gives $0.284, and using today's cheapest on-demand H200 rate gives $0.435, about 2.07x the published number. A $0.33 per million H200 figure implies a GPU cost of $1.23/hr, which is an amortized ownership cost rather than a rental price.
The one figure that does survive is DeepSeek's own published production accounting: 226.75 nodes, 168 billion output tokens in 24 hours, $87,072 per day at a $2.00/GPU-hr internal cost, giving $0.518 per million output tokens at 20 to 22 tok/s per user. That is the floor a self-hoster is chasing, reached with 226 nodes and custom kernels.
Rent or buy
The 2026 GPU market is inflated. The RTX 5090 launched at $1,999 and retails around $4,530, roughly 2.3x list. That inverts the usual advice: buying is weaker than it was, renting stronger.
| Card | Marketplace rate | $/mo at 24/7 | Buy price | Breakeven |
|---|---|---|---|---|
| RTX 3090 24GB | $0.069–0.138/hr | $50–101 | $1,500 | 117 months |
| RTX 4090 24GB | $0.135–0.336/hr | $99–245 | $3,000 | 15 months |
| RTX 5090 32GB | $0.296–0.402/hr | $216–293 | $4,530 | 19.5 months |
| RTX PRO 6000 96GB | $0.761–1.096/hr | $556–800 | $15,017 | 19.6 months |
| RX 7900 XTX 24GB | n/a | n/a | $710 | best bandwidth/$ |
Two findings worth more than the prices. Enterprise clouds are the wrong channel for this workload: one major provider lists no consumer cards at all, and its cheapest part that fits a 27B model costs about 5x the marketplace rate for identical silicon. And a 3090 bought to run 24/7 has a 117-month breakeven against marketplace rates, where European electricity alone nearly matches the rent.
For the buy-it-outright case, a pair of 128GB unified-memory desktop AI boxes lands near $8,000 and runs a 167 GB mixture-of-experts model at a measured 41 tok/s single-stream, 62 to 75 with speculative decoding, and about 350 tok/s aggregate at 32 concurrent streams. Against equivalent-capacity rental that pays back in 3.3 months, and break-even utilization is only about 3.3 hours a day.
Then the API undercuts both. That same model sells for $0.28 per million output tokens. The owned pair, amortized over 24 months and running at its measured peak, works out to $0.437 per million, more expensive at 100% saturation, ignoring your time and all hardware risk. The amortized monthly cost buys 1.44 billion output tokens from the API; the hardware's physical ceiling is 920 million. Buying beats renting a node, and loses to the API. The decision is not financial. It is whether the workload must be local.
Quantization, and where the damage actually lands
Four-bit is not automatically degradation. It depends entirely on the reference precision the vendor shipped.
| Precision | Measured effect | Verdict |
|---|---|---|
| FP8 | Within ~0.4 points of FP16 on MMLU-Pro | Effectively free |
| Vendor QAT INT4 | Trained as the reference checkpoint; ~2x speedup retained | Safe |
| NVFP4 | Relative validation-loss error below 1% vs FP8 | Close to free |
| Community post-training INT4 | MMLU-Pro −1.6, HumanEval −8 points | Damage lands on code |
The trap is that post-training 4-bit is the only path that fits some models onto an 8x H100 node, which is exactly the configuration people price when they imagine self-hosting. Across 133 provider endpoints for twelve frontier models, 59 serve FP8, 23 serve FP4, 7 INT4, 3 MXFP4 and 38 do not declare. Roughly a quarter of frontier endpoints are 4-bit, and on aggregator routing, price sorting will silently send you to them. Pin providers explicitly.
What the adversarial audit removed
This study was produced by parallel research agents and then handed to a separate reviewer whose only instruction was to refute them, defaulting to failure on anything it could not verify against a primary source. It killed five claims. They are listed because a cost study that hides its corrections should not be trusted with your money.
| Claim as first researched | Verdict | Correction |
|---|---|---|
| Cheapest viable node is $9,986/mo at a named vendor | Failed | That vendor publishes no public pricing. The rate existed only in SEO aggregators, and two live price trackers list no such figure. Replaced with a verified $1.99/GPU-hr, which a later re-check found had itself risen. |
| Cheapest fitting node costs $11,622/mo | Corrected after publication | Hot Aisle raised its MI300X rate from $1.99 to $2.99 per GPU-hour for new customers, keeping $1.99 only for existing ones. The node figure is now $17,462. A price verified once is not verified forever. |
| Cheapest 8x H200 is $28.16/hr | Corrected | Real, but not cheapest. A verified $3.00/GPU-hr exists. All break-even figures recomputed. |
| Opus 4.6 scores 44.9 in adaptive mode | Failed | No such figure exists. One variant is published, scoring 39, itself flagged an estimate. |
| Qwen3.8-27B scores around 30 | Failed | Wrong by 22 points. The measured value is 52, first of 135 open-weight models in its size class. |
| Free EU supercomputing hours cover production use | Corrected | Free access is scoped to innovation purposes; commercial production is a separate paid track. |
One figure survives with a caveat attached rather than a correction. The claim that self-hosting Kimi K3 costs more than its own API rests on a published throughput number the auditor could not source. Both prices in that calculation are confirmed; the throughput is not. The direction holds regardless, because K3 needs ten to fourteen GPUs, but the exact figure should not be quoted as hard.
How to read this before spending
Rent, do not buy, unless it must be local
At 2.3x list prices, breakeven on a current-generation card is 15 to 20 months against marketplace rental. Data sensitivity is a far stronger argument for owning hardware than economics currently is.
Choose the model before the hardware
Bandwidth-per-dollar and capacity-per-dollar point in opposite directions. A 32GB card is roughly ten times better on bandwidth per dollar; a 128GB box is two and a half times better on capacity. Which matters depends entirely on whether your model fits.
Price the cache, not the output
On a cache-heavy agentic workload the cache-read rate is about 95% of the bill. Providers vary more than 20x on that line and look identical on the front page.
Spot is for training, not serving
Preemption notice ranges from two minutes down to five seconds, and some marketplaces give none at all. A reliable endpoint needs a warm standby, which roughly doubles the effective rate and erases the discount.
Method
Workload figures were produced by parsing the usage field of 1,039,011 assistant messages across 13,587 local session transcript files, covering 2026-07-18 to 2026-08-19. No personal or account-identifying data appears on this page.
Model sizes were computed by summing .safetensors byte counts from the Hugging Face model API, and architecture figures were read from each model's own config.json rather than from a summary. Prices were read from provider pricing pages and public unauthenticated catalogue endpoints on 2026-08-19. All cost arithmetic is shown inline so it can be checked.
Figures described as measured were produced by running code or reading a primary source. Figures marked as estimates were not. Five claims were corrected or removed by adversarial audit after first drafting, and all are disclosed above. GPU and token pricing in this market moves weekly; verify before committing money.