Every number, and where it came from
The parent study keeps only the figures that change a decision. Everything gathered is here: full price tables, model sizes computed from weight files, published throughput measurements, the break-even grid, and the claims an adversarial audit removed.
Collected 2026-08-19 from provider pricing pages and free public catalogue endpoints. Figures move weekly.
Reading order: the parent study carries the argument. This page carries the evidence. Where a figure could not be verified against a primary source it is labelled rather than dropped, because an unverified number a reader can see and discount is safer than one silently removed.
Open-weight model size, activation and licence
Total parameters and file sizes were computed by summing safetensors byte counts from the Hugging Face model API. Active parameters are vendor-published where available and derived from tensor shapes where not, marked accordingly. Active parameters, not total, drive serving cost.
| Model | Total params | Active/token | Context | Shipped size | Licence |
|---|---|---|---|---|---|
| Kimi K3 | 2.78T | 104B | 1,048,576 | 1,561 GB MXFP4 | Custom, MaaS gate at $20M/12mo |
| Qwen3.8-2.4T-A95B | 2.446T | 95B | 262,144 | 2,496 GB FP8 | Custom, gate at $50M/12mo |
| DeepSeek-V4-Pro-0813 | 1.65T | 49B | 1,048,576 | 893 GB | MIT |
| Ring-2.6-1T | 1.026T | ~64B derived | 131,072 | 1,042 GB BF16 | MIT |
| GLM-5.2 | 753.3B | ~41B derived | 1,048,576 | 756 GB FP8 | MIT |
| Nemotron-3-Ultra | 560.5B | 55B | 262,144 | 352 GB NVFP4 | OpenMDW-1.1 |
| MiniMax-M3 | 427.0B | 23B | 1,048,576 | 444 GB MXFP8 | Community, non-commercial default |
| DeepSeek-V4-Flash-0731 | 304.2B | 13B | 1,048,576 | 167 GB | MIT |
| Hy3 (Tencent) | 298.8B | ~20.6B derived | 262,144 | 300 GB FP8 | Apache-2.0 |
| Command A+ | 218.8B | ~25.5B derived | 200,000 | 132 GB W4A4 | Apache-2.0 |
| Step-3.7-Flash | 201.4B | ~11B derived | 262,144 | 129 GB NVFP4 | Apache-2.0 |
| Ling-3.0-flash | 127.5B | ~5B derived | 262,144 | 70 GB FP4 | MIT |
| gpt-oss-120b | 116.8B | ~5.8B derived | 131,072 | 62 GB MXFP4 | Apache-2.0 |
| Qwen3.8-27B | 27.8B dense | 27.8B | 262,144 | 17 GB Q4 | Apache-2.0 |
Derived figures were computed from tensor shapes in each model config. The method was validated against DeepSeek-V4-Pro, where the derivation gave 50.3B against a published 49B, a 2.7 percent error. Treat every derived value as carrying that bound.
Quality index, same scale
Artificial Analysis Intelligence Index v4.1.1. Reasoning mode is stated because it changes the number materially, and two widely-quoted entries are vendor estimates rather than independent measurements.
| Model | Index | Mode | Measured or estimate | Open weights |
|---|---|---|---|---|
| Claude Opus 5 | 63 | reasoning, max | measured | No |
| Kimi K3 | 60 | reasoning, max | measured | Yes |
| Qwen3.8-2.4T-A95B | 58 | reasoning | measured | Yes |
| Claude Opus 4.8 | 57 | adaptive, max effort | measured | No |
| GLM-5.2 | 53 | reasoning, max | measured | Yes |
| DeepSeek-V4-Pro-0813 | 53 | reasoning, max | measured | Yes |
| Qwen3.8-27B | 52 | reasoning | measured | Yes |
| DeepSeek-V4-Flash-0731 | 52 | reasoning | measured | Yes |
| Gemini 3.1 Pro Preview | 48 | reasoning | measured | No |
| GPT-5.3 Codex | 46 | reasoning, xhigh | estimate | No |
| MiniMax-M3 | 45 | reasoning | measured | Yes |
| Kimi K2.7-Code | 43 | reasoning | measured | Yes |
| Claude Opus 4.6 | 39 | non-reasoning, high effort | estimate | No |
| Nemotron-3-Ultra | 38 | reasoning | measured | Yes |
| gpt-oss-120b | 24 | high | measured | Yes |
| Command A+ | 23 | default | measured | Yes |
| Mistral Large 3 | 16 | default | measured | Yes |
| Llama 4 Maverick | 14 | default | measured | Yes |
| Llama 4 Scout | 10 | default | measured | Yes |
Qwen3.8-27B ranks first of 135 open-weight models in the 4B to 40B class, where the class median is 9. Claims that it outranks Opus 4.6 are literally true on this table and rest on an Opus figure that was never independently measured, in non-reasoning mode.
Datacentre node rental, 24/7
Monthly figures are hourly rate multiplied by eight GPUs and 730 hours. The fit column states whether the node holds a roughly 1T-parameter mixture of experts at FP8.
| Node | Provider | $/mo at 24/7 | Fits 1T MoE | Note |
|---|---|---|---|---|
| 8x MI300X | Azure spot | $8,468 | Yes | Preemptible, hostile to production |
| 8x MI300X | Hot Aisle | $11,622 | Yes | Cheapest verifiable node that fits |
| 8x H100 | Voltage Park | $11,622 | No | US only, and too small anyway |
| 8x H200 | Verda spot, Finland | $11,680 | Tight | Preemptible |
| 8x H200 | Nebius preemptible, Finland | $14,308 | Tight | Quoted ex-VAT |
| 8x MI300X | DigitalOcean on-demand | $15,126 | Yes | EU region unverified |
| 8x H200 | Hyperstack reserved | $16,294 | Tight | Free egress, best EU NVIDIA target |
| 8x H200 | Genesis Cloud, Munich | $16,352 | Tight | Eight-GPU minimum order |
| 8x H200 | Koyeb | $17,520 | Tight | Cheapest verifiable H200 node |
| 8x MI355X | TensorWave | $17,228 | Roomy | Thin market |
| 8x H100 | OVHcloud | $17,462 | No | Closest EU region to Poland |
| 8x MI300X | Crusoe | $20,148 | Yes | First-party verified AMD rate |
| 8x H200 | Crusoe | $25,054 | Tight | Free egress |
| 8x H200 | Nebius on-demand | $26,280 | Tight | Finland |
| 8x H200 | CoreWeave | $36,851 | Tight | EU price equals NA price, free egress |
| 8x H100 | AWS p5.48xlarge | $40,179 | No | Capacity Block prices rose 15% in Jan 2026 |
| 8x MI355X | Oracle | $50,224 | Yes | Only node verified to hold Kimi K3 |
| 8x H100 | Azure ND96isr H100 v5 | $71,774 | No | |
| 8x H100 | GCP europe-west4 | $82,154 | No | Costs the most and does not fit |
Consumer and prosumer card rental
Marketplace rates, which run roughly an order of magnitude below enterprise clouds for the same silicon. The agent column is how many concurrent sessions the card holds at 50k context each, after weights and workspace.
| Card | VRAM | Cheapest $/hr | Typical $/hr | $/mo at 24/7 | Agents at 50k ctx |
|---|---|---|---|---|---|
| RTX 3090 | 24 GB | 0.069 | 0.138 | $50 to $101 | ~3 |
| RTX 4090 | 24 GB | 0.135 | 0.336 | $99 to $245 | ~3 |
| RTX A6000 | 48 GB | 0.287 | 0.351 | $210 to $256 | ~12 |
| RTX 5090 | 32 GB | 0.296 | 0.402 | $216 to $293 | ~8 |
| L40S | 48 GB | 0.521 | 0.601 | $380 to $439 | ~12 |
| RTX PRO 6000 | 96 GB | 0.761 | 1.096 | $556 to $800 | ~40 |
Concurrency figures account for a fixed per-sequence recurrent state in hybrid linear-attention models, roughly 159 MB per session for Qwen3.8-27B, which is independent of context length. Linear attention makes one long context cheap and taxes many short ones.
Buying, with breakeven against rental
Retail prices are inflated well above launch pricing. Breakeven is months of typical marketplace rental equal to the purchase price, before electricity.
| Hardware | Price | Memory | Bandwidth | Breakeven vs rental |
|---|---|---|---|---|
| RX 7900 XTX | $710 | 24 GB | 960 GB/s | best bandwidth per dollar |
| RTX 3090 new | $1,500 | 24 GB | 936 GB/s | 117 months |
| Radeon AI PRO R9700 | $1,500 | 32 GB | ~640 GB/s derived | cheapest honest 32 GB |
| RTX 4090 | $3,000 | 24 GB | 1,008 GB/s | 15 months |
| Desktop AI box, 128 GB | $4,000 | 128 GB | 273 GB/s | capacity, not speed |
| RTX 5090 | $4,530 | 32 GB | 1,792 GB/s | 19.5 months |
| Mac Studio M3 Ultra 96GB | $5,299 | 96 GB | 819 GB/s | single box, no cluster |
| Two 128 GB AI boxes | $8,000 | 256 GB | 273 GB/s each | 3.3 months |
| RTX PRO 6000 Max-Q | $15,017 | 96 GB | 1,792 GB/s | 19.6 months |
The RTX 5090 launched at $1,999 and retails around 2.3 times that. Buying at this point in the cycle is a bet that the shortage persists. European electricity at 0.27 EUR per kWh adds roughly 30 to 70 EUR a month per card at 24/7 and 70 percent of rated power.
Measured serving throughput
Aggregate output tokens per second on a single eight-GPU node. Peak figures are quoted at low per-user interactivity and do not transfer to interactive work.
| Configuration | Aggregate tok/s | Conditions |
|---|---|---|
| 8x H100, DeepSeek-V3 AWQ INT4 | 620 | peak at concurrency 100 |
| 8x MI355X, Kimi K3 2.8T | 952 | 1M-context configuration |
| 8x B300, Kimi K3 | 1,568 | peak aggregate |
| 8x H200, DeepSeek-V3.2, generation-heavy | 2,469 | sustained, 1K in / 2K out |
| 8x H200, DeepSeek-V3.2, ShareGPT | 4,243 | sustained; 13,860 peak at 4.9 tok/s per user |
| 8x H200 wide expert parallel | 17,600 | 2,200 per GPU, multi-node fabric |
| 8x MI300X, DeepSeek-R1 671B | 21,225 | peak at concurrency 4,096, decode only |
| 8x MI355X, Kimi K2.5 MXFP4 | 21,496 | peak, 8k in / 1k out |
| 8x B200, Kimi K2.5 FP4 | 32,168 | peak, single node |
Interactivity trades against aggregate throughput steeply. On identical H200 hardware running DeepSeek R1 at FP8, serving 65 tok/s per user yields 1,035 tok/s per GPU; 99 tok/s per user yields 488; 133 tok/s per user yields 264. Making the model twice as fast for one person costs almost four times the total throughput.
Cost per million output tokens, self-hosted
Computed as 277.7778 multiplied by the node hourly rate, divided by aggregate output tokens per second. The constant is one million divided by 3,600.
| Node | $/hr | Aggregate tok/s | $/M output |
|---|---|---|---|
| 8x B200, Kimi K2.5 FP4 | $35.20 | 32,168 | $0.304 |
| 8x MI300X, DeepSeek-R1 | $27.60 | 21,225 | $0.361 |
| 8x H200, batch-optimised | $28.16 | 13,860 | $0.564 |
| DeepSeek published production | $16.00/node | 8,575/node | $0.518 |
| 8x MI355X, Kimi K2.5 | $68.80 | 21,496 | $0.889 |
| 8x H200, agentic-realistic | $28.16 | 2,469 | $3.168 |
| 8x H100, DeepSeek-V3 INT4 | $16.00 | 620 | $7.168 |
| 8x MI355X, Kimi K3 | $68.80 | 952 | $20.07 |
Kimi K3 self-hosted at $20.07 sits above Moonshot list pricing of $15.00 per million output tokens. The throughput input to that row is published rather than independently measured, so treat the exact figure as indicative; the direction holds because K3 needs ten to fourteen GPUs.
Per-token API pricing
Blended figures assume three input tokens per output token. The cache-read column matters more than either headline rate on a cache-heavy agentic workload, where it can be 95 percent of the bill.
| Model | In $/M | Out $/M | Cache read $/M | Context |
|---|---|---|---|---|
| DeepSeek V4 Flash | 0.08 | 0.17 | 0.017 | 1.05M |
| Ling 2.6 1T | 0.07 | 0.62 | 0.015 | 262K |
| gpt-oss-120b | 0.03 | 0.17 | 0.030 | 131K |
| MiniMax M2.5 | 0.22 | 0.90 | 0.050 | 205K |
| Qwen3.8-27B | 0.45 | 3.20 | 0.050 | 262K |
| Kimi K2.5 | 0.45 | 2.25 | 0.070 | 262K |
| GLM-4.7 | 0.40 | 1.75 | 0.080 | 205K |
| DeepSeek V4 Pro | 1.32 | 3.96 | 0.044 | 1.05M |
| GLM-5 | 0.60 | 1.92 | 0.120 | 205K |
| Kimi K2 Thinking | 0.60 | 2.50 | 0.150 | 262K |
| Qwen3-Max | 0.78 | 3.90 | 0.156 | 262K |
| GLM-5.2 | 0.97 | 3.04 | 0.193 | 1.05M |
| Qwen3.8-Max | 2.00 | 6.00 | 0.250 | 1M |
| Kimi K3 | 3.00 | 15.00 | 0.300 | 1.05M |
Provider spread on Kimi K3 across thirteen hosts is only about 15 percent, so shopping providers gains little. Prompt caching gains an order of magnitude, since a cache hit bills at one tenth of a miss. On DeepSeek V4 Flash the throughput spread across providers reaches 5.8x at near-identical prices, so on the cheap tier the right variable to optimise is speed, not price.
Quantization varies by provider and is rarely displayed. Across 133 endpoints for twelve frontier models: 59 serve FP8, 23 serve FP4, 7 INT4, 3 MXFP4, and 38 declare nothing. Four-bit is only a downgrade relative to what the vendor shipped. Moonshot serves K3 at MXFP4 itself, so a BF16 host is upcasting rather than improving it. Z.AI serves GLM-5.2 at FP8, which makes any FP4 host of that model genuinely sub-reference.
Break-even grid
Monthly output tokens required before a rented node beats buying the same tokens. Rent divided by API price. Capacity is what the node can physically produce at full utilisation.
| Node | Monthly rent | Beat $0.50/M | Beat $2.00/M | Beat $4.00/M | Physical capacity |
|---|---|---|---|---|---|
| 8x H200 at $24.00/hr | $17,520 | 35.0B | 8.8B | 4.4B | 6.5B agentic to 36.4B batch |
| 8x H200 at $28.16/hr | $20,557 | 41.1B | 10.3B | 5.1B | as above |
| 8x MI300X at $27.60/hr | $20,148 | 40.3B | 10.1B | 5.0B | 55.8B |
| 8x B200 at $35.20/hr | $25,696 | 51.4B | 12.8B | 6.4B | 84.5B |
| 8x MI355X at $68.80/hr | $50,224 | 100.4B | 25.1B | 12.6B | 56.5B |
| 128-GPU wide expert parallel | $328,909 | 658B | 164B | 82B | cluster scale |
Read the last two columns together. An 8x H200 serving an agentic coding workload at 2,469 tok/s produces at most 6.5 billion output tokens a month, which is below its own 8.8 billion break-even against a $2.00 API. That node cannot beat the API at any utilisation. It only wins in batch mode at 4.9 tokens per second per user, which no interactive user would accept.
Reliability and hidden costs
| Provider | Published uptime | Remedy |
|---|---|---|
| Baseten | 99.9% monthly, dedicated inference | Credits, capped at 40% of the monthly fee |
| Voltage Park | 99.5% monthly | 10, 25 or 100 percent bill credit by severity |
| CoreWeave | 99.9% for object storage only | No public GPU compute commitment |
| Together | 99% on committed capacity only | Order form only |
| Lambda | None found | None |
| RunPod | None on Community Cloud | Liability capped at the lesser of six months of fees or $100 |
| Vast.ai | None | Right of action for data loss is waived |
Costs that do not appear in an hourly rate: RunPod volume storage doubles to $0.20 per GB-month while a pod is stopped, so a 1 TB checkpoint parked idle runs about $200 a month. Hyperstack bills virtual machines in the shutoff state. Vast.ai bills storage continuously even at a negative balance and schedules instances, volumes and data for deletion once the balance reaches zero. Hyperscaler egress runs $0.087 to $0.12 per GB, so moving a terabyte out of one costs more than eight H100-hours.
Preemption notice ranges from two minutes on one hyperscaler to thirty seconds on another, five seconds on one marketplace, and nothing documented at all on another. A node loss costs five to twenty-five minutes of cold start before it serves again. A reliable endpoint therefore needs a warm standby, which roughly doubles the effective hourly cost and removes most of the marketplace saving.
Claims removed by audit
The research agents produced these. A separate reviewer, instructed only to refute and to default to failure on anything it could not verify against a primary source, removed them.
| Claim | Verdict | What replaced it |
|---|---|---|
| Cheapest viable node is $9,986/mo at a named vendor | Failed | That vendor publishes no rates. The figure existed only in aggregator blogs. Replaced with a verified $1.99 per GPU-hour. |
| Cheapest 8x H200 is $28.16/hr | Corrected | Real but not cheapest. A verified $3.00 per GPU-hour exists, so every break-even was recomputed. |
| Opus 4.6 scores 44.9 in adaptive mode | Failed | No such figure is published. One variant exists at 39, itself an estimate. |
| Qwen3.8-27B scores around 30 | Failed | Wrong by 22 points. The measured value is 52. |
| Kimi K3 weighs about 594 GB | Failed | Wrong by 2.6x. The repository measures 1,561 GB. |
| Free EU supercomputing hours cover production | Corrected | Free access is scoped to innovation purposes. Commercial production is a separate paid track. |
| A subscription tier unlocks 1M context at $39 | Disputed | Vendor documentation and its own pricing page disagree. Verify at checkout. |
Sources
Every figure above traces to one of these. Provider pricing pages were read on 2026-08-19; model files and configuration were read from the Hugging Face API on the same day.
Model weights and architecture
Kimi K3, DeepSeek-V4-Pro-0813, GLM-5.2, Qwen3.8-27B, gpt-oss-120b
Serving stacks and throughput
vLLM Kimi K2 recipe, vLLM wide expert parallel, LMSYS Kimi K2 on 128 H200, DeepSeek inference system overview
GPU pricing
Vast.ai, RunPod, Prime Intellect, SF Compute, Hyperstack, Crusoe, Nebius, Lambda
Token pricing and benchmarks
DeepSeek pricing, Moonshot K3 pricing, OpenRouter model catalogue, Artificial Analysis leaderboard, Terminal-Bench 2.1
Outbound links carry rel="nofollow" and open in a new tab. Prices and leaderboard positions on those pages change without notice, so a figure quoted here should be treated as a reading taken on 2026-08-19 rather than a current quote.
Link check, run at publish time: 50 of 52 outbound targets resolved. The Nebius and Crusoe pricing pages returned a connection failure from the checking location while their DNS resolved normally and every other provider responded, which points at a network-level block rather than an outage. Their figures were captured earlier in the research and are retained; verify them yourself before relying on either.
Method and limits
Model sizes were computed by summing safetensors byte counts rather than quoting a card. Architecture figures came from each model configuration file. Cost arithmetic is shown inline throughout so a reader can check it rather than trust it.
Three limits are worth stating. Active parameter counts marked as derived carry roughly a 3 percent error bound, validated against the one model that publishes both. Throughput figures are published measurements from third parties, not reproductions, and peak numbers are quoted at interactivity levels that suit batch work rather than interactive use. And the workload shape used to weight these costs came from one heavy agentic setup, so the 94 percent cache-read figure is characteristic of that pattern rather than universal.