Appendix · complete dataset

Every number, and where it came from

The parent study keeps only the figures that change a decision. Everything gathered is here: full price tables, model sizes computed from weight files, published throughput measurements, the break-even grid, and the claims an adversarial audit removed.

Collected 2026-08-19 from provider pricing pages and free public catalogue endpoints. Figures move weekly.

Reading order: the parent study carries the argument. This page carries the evidence. Where a figure could not be verified against a primary source it is labelled rather than dropped, because an unverified number a reader can see and discount is safer than one silently removed.

Open-weight model size, activation and licence

Total parameters and file sizes were computed by summing safetensors byte counts from the Hugging Face model API. Active parameters are vendor-published where available and derived from tensor shapes where not, marked accordingly. Active parameters, not total, drive serving cost.

ModelTotal paramsActive/tokenContextShipped sizeLicence
Kimi K32.78T104B1,048,5761,561 GB MXFP4Custom, MaaS gate at $20M/12mo
Qwen3.8-2.4T-A95B2.446T95B262,1442,496 GB FP8Custom, gate at $50M/12mo
DeepSeek-V4-Pro-08131.65T49B1,048,576893 GBMIT
Ring-2.6-1T1.026T~64B derived131,0721,042 GB BF16MIT
GLM-5.2753.3B~41B derived1,048,576756 GB FP8MIT
Nemotron-3-Ultra560.5B55B262,144352 GB NVFP4OpenMDW-1.1
MiniMax-M3427.0B23B1,048,576444 GB MXFP8Community, non-commercial default
DeepSeek-V4-Flash-0731304.2B13B1,048,576167 GBMIT
Hy3 (Tencent)298.8B~20.6B derived262,144300 GB FP8Apache-2.0
Command A+218.8B~25.5B derived200,000132 GB W4A4Apache-2.0
Step-3.7-Flash201.4B~11B derived262,144129 GB NVFP4Apache-2.0
Ling-3.0-flash127.5B~5B derived262,14470 GB FP4MIT
gpt-oss-120b116.8B~5.8B derived131,07262 GB MXFP4Apache-2.0
Qwen3.8-27B27.8B dense27.8B262,14417 GB Q4Apache-2.0

Derived figures were computed from tensor shapes in each model config. The method was validated against DeepSeek-V4-Pro, where the derivation gave 50.3B against a published 49B, a 2.7 percent error. Treat every derived value as carrying that bound.

Quality index, same scale

Artificial Analysis Intelligence Index v4.1.1. Reasoning mode is stated because it changes the number materially, and two widely-quoted entries are vendor estimates rather than independent measurements.

ModelIndexModeMeasured or estimateOpen weights
Claude Opus 563reasoning, maxmeasuredNo
Kimi K360reasoning, maxmeasuredYes
Qwen3.8-2.4T-A95B58reasoningmeasuredYes
Claude Opus 4.857adaptive, max effortmeasuredNo
GLM-5.253reasoning, maxmeasuredYes
DeepSeek-V4-Pro-081353reasoning, maxmeasuredYes
Qwen3.8-27B52reasoningmeasuredYes
DeepSeek-V4-Flash-073152reasoningmeasuredYes
Gemini 3.1 Pro Preview48reasoningmeasuredNo
GPT-5.3 Codex46reasoning, xhighestimateNo
MiniMax-M345reasoningmeasuredYes
Kimi K2.7-Code43reasoningmeasuredYes
Claude Opus 4.639non-reasoning, high effortestimateNo
Nemotron-3-Ultra38reasoningmeasuredYes
gpt-oss-120b24highmeasuredYes
Command A+23defaultmeasuredYes
Mistral Large 316defaultmeasuredYes
Llama 4 Maverick14defaultmeasuredYes
Llama 4 Scout10defaultmeasuredYes

Qwen3.8-27B ranks first of 135 open-weight models in the 4B to 40B class, where the class median is 9. Claims that it outranks Opus 4.6 are literally true on this table and rest on an Opus figure that was never independently measured, in non-reasoning mode.

Datacentre node rental, 24/7

Monthly figures are hourly rate multiplied by eight GPUs and 730 hours. The fit column states whether the node holds a roughly 1T-parameter mixture of experts at FP8.

NodeProvider$/mo at 24/7Fits 1T MoENote
8x MI300XAzure spot$8,468YesPreemptible, hostile to production
8x MI300XHot Aisle$11,622YesCheapest verifiable node that fits
8x H100Voltage Park$11,622NoUS only, and too small anyway
8x H200Verda spot, Finland$11,680TightPreemptible
8x H200Nebius preemptible, Finland$14,308TightQuoted ex-VAT
8x MI300XDigitalOcean on-demand$15,126YesEU region unverified
8x H200Hyperstack reserved$16,294TightFree egress, best EU NVIDIA target
8x H200Genesis Cloud, Munich$16,352TightEight-GPU minimum order
8x H200Koyeb$17,520TightCheapest verifiable H200 node
8x MI355XTensorWave$17,228RoomyThin market
8x H100OVHcloud$17,462NoClosest EU region to Poland
8x MI300XCrusoe$20,148YesFirst-party verified AMD rate
8x H200Crusoe$25,054TightFree egress
8x H200Nebius on-demand$26,280TightFinland
8x H200CoreWeave$36,851TightEU price equals NA price, free egress
8x H100AWS p5.48xlarge$40,179NoCapacity Block prices rose 15% in Jan 2026
8x MI355XOracle$50,224YesOnly node verified to hold Kimi K3
8x H100Azure ND96isr H100 v5$71,774No
8x H100GCP europe-west4$82,154NoCosts the most and does not fit

Consumer and prosumer card rental

Marketplace rates, which run roughly an order of magnitude below enterprise clouds for the same silicon. The agent column is how many concurrent sessions the card holds at 50k context each, after weights and workspace.

CardVRAMCheapest $/hrTypical $/hr$/mo at 24/7Agents at 50k ctx
RTX 309024 GB0.0690.138$50 to $101~3
RTX 409024 GB0.1350.336$99 to $245~3
RTX A600048 GB0.2870.351$210 to $256~12
RTX 509032 GB0.2960.402$216 to $293~8
L40S48 GB0.5210.601$380 to $439~12
RTX PRO 600096 GB0.7611.096$556 to $800~40

Concurrency figures account for a fixed per-sequence recurrent state in hybrid linear-attention models, roughly 159 MB per session for Qwen3.8-27B, which is independent of context length. Linear attention makes one long context cheap and taxes many short ones.

Buying, with breakeven against rental

Retail prices are inflated well above launch pricing. Breakeven is months of typical marketplace rental equal to the purchase price, before electricity.

HardwarePriceMemoryBandwidthBreakeven vs rental
RX 7900 XTX$71024 GB960 GB/sbest bandwidth per dollar
RTX 3090 new$1,50024 GB936 GB/s117 months
Radeon AI PRO R9700$1,50032 GB~640 GB/s derivedcheapest honest 32 GB
RTX 4090$3,00024 GB1,008 GB/s15 months
Desktop AI box, 128 GB$4,000128 GB273 GB/scapacity, not speed
RTX 5090$4,53032 GB1,792 GB/s19.5 months
Mac Studio M3 Ultra 96GB$5,29996 GB819 GB/ssingle box, no cluster
Two 128 GB AI boxes$8,000256 GB273 GB/s each3.3 months
RTX PRO 6000 Max-Q$15,01796 GB1,792 GB/s19.6 months

The RTX 5090 launched at $1,999 and retails around 2.3 times that. Buying at this point in the cycle is a bet that the shortage persists. European electricity at 0.27 EUR per kWh adds roughly 30 to 70 EUR a month per card at 24/7 and 70 percent of rated power.

Measured serving throughput

Aggregate output tokens per second on a single eight-GPU node. Peak figures are quoted at low per-user interactivity and do not transfer to interactive work.

ConfigurationAggregate tok/sConditions
8x H100, DeepSeek-V3 AWQ INT4620peak at concurrency 100
8x MI355X, Kimi K3 2.8T9521M-context configuration
8x B300, Kimi K31,568peak aggregate
8x H200, DeepSeek-V3.2, generation-heavy2,469sustained, 1K in / 2K out
8x H200, DeepSeek-V3.2, ShareGPT4,243sustained; 13,860 peak at 4.9 tok/s per user
8x H200 wide expert parallel17,6002,200 per GPU, multi-node fabric
8x MI300X, DeepSeek-R1 671B21,225peak at concurrency 4,096, decode only
8x MI355X, Kimi K2.5 MXFP421,496peak, 8k in / 1k out
8x B200, Kimi K2.5 FP432,168peak, single node

Interactivity trades against aggregate throughput steeply. On identical H200 hardware running DeepSeek R1 at FP8, serving 65 tok/s per user yields 1,035 tok/s per GPU; 99 tok/s per user yields 488; 133 tok/s per user yields 264. Making the model twice as fast for one person costs almost four times the total throughput.

Cost per million output tokens, self-hosted

Computed as 277.7778 multiplied by the node hourly rate, divided by aggregate output tokens per second. The constant is one million divided by 3,600.

Node$/hrAggregate tok/s$/M output
8x B200, Kimi K2.5 FP4$35.2032,168$0.304
8x MI300X, DeepSeek-R1$27.6021,225$0.361
8x H200, batch-optimised$28.1613,860$0.564
DeepSeek published production$16.00/node8,575/node$0.518
8x MI355X, Kimi K2.5$68.8021,496$0.889
8x H200, agentic-realistic$28.162,469$3.168
8x H100, DeepSeek-V3 INT4$16.00620$7.168
8x MI355X, Kimi K3$68.80952$20.07

Kimi K3 self-hosted at $20.07 sits above Moonshot list pricing of $15.00 per million output tokens. The throughput input to that row is published rather than independently measured, so treat the exact figure as indicative; the direction holds because K3 needs ten to fourteen GPUs.

Per-token API pricing

Blended figures assume three input tokens per output token. The cache-read column matters more than either headline rate on a cache-heavy agentic workload, where it can be 95 percent of the bill.

ModelIn $/MOut $/MCache read $/MContext
DeepSeek V4 Flash0.080.170.0171.05M
Ling 2.6 1T0.070.620.015262K
gpt-oss-120b0.030.170.030131K
MiniMax M2.50.220.900.050205K
Qwen3.8-27B0.453.200.050262K
Kimi K2.50.452.250.070262K
GLM-4.70.401.750.080205K
DeepSeek V4 Pro1.323.960.0441.05M
GLM-50.601.920.120205K
Kimi K2 Thinking0.602.500.150262K
Qwen3-Max0.783.900.156262K
GLM-5.20.973.040.1931.05M
Qwen3.8-Max2.006.000.2501M
Kimi K33.0015.000.3001.05M

Provider spread on Kimi K3 across thirteen hosts is only about 15 percent, so shopping providers gains little. Prompt caching gains an order of magnitude, since a cache hit bills at one tenth of a miss. On DeepSeek V4 Flash the throughput spread across providers reaches 5.8x at near-identical prices, so on the cheap tier the right variable to optimise is speed, not price.

Quantization varies by provider and is rarely displayed. Across 133 endpoints for twelve frontier models: 59 serve FP8, 23 serve FP4, 7 INT4, 3 MXFP4, and 38 declare nothing. Four-bit is only a downgrade relative to what the vendor shipped. Moonshot serves K3 at MXFP4 itself, so a BF16 host is upcasting rather than improving it. Z.AI serves GLM-5.2 at FP8, which makes any FP4 host of that model genuinely sub-reference.

Break-even grid

Monthly output tokens required before a rented node beats buying the same tokens. Rent divided by API price. Capacity is what the node can physically produce at full utilisation.

NodeMonthly rentBeat $0.50/MBeat $2.00/MBeat $4.00/MPhysical capacity
8x H200 at $24.00/hr$17,52035.0B8.8B4.4B6.5B agentic to 36.4B batch
8x H200 at $28.16/hr$20,55741.1B10.3B5.1Bas above
8x MI300X at $27.60/hr$20,14840.3B10.1B5.0B55.8B
8x B200 at $35.20/hr$25,69651.4B12.8B6.4B84.5B
8x MI355X at $68.80/hr$50,224100.4B25.1B12.6B56.5B
128-GPU wide expert parallel$328,909658B164B82Bcluster scale

Read the last two columns together. An 8x H200 serving an agentic coding workload at 2,469 tok/s produces at most 6.5 billion output tokens a month, which is below its own 8.8 billion break-even against a $2.00 API. That node cannot beat the API at any utilisation. It only wins in batch mode at 4.9 tokens per second per user, which no interactive user would accept.

Reliability and hidden costs

ProviderPublished uptimeRemedy
Baseten99.9% monthly, dedicated inferenceCredits, capped at 40% of the monthly fee
Voltage Park99.5% monthly10, 25 or 100 percent bill credit by severity
CoreWeave99.9% for object storage onlyNo public GPU compute commitment
Together99% on committed capacity onlyOrder form only
LambdaNone foundNone
RunPodNone on Community CloudLiability capped at the lesser of six months of fees or $100
Vast.aiNoneRight of action for data loss is waived

Costs that do not appear in an hourly rate: RunPod volume storage doubles to $0.20 per GB-month while a pod is stopped, so a 1 TB checkpoint parked idle runs about $200 a month. Hyperstack bills virtual machines in the shutoff state. Vast.ai bills storage continuously even at a negative balance and schedules instances, volumes and data for deletion once the balance reaches zero. Hyperscaler egress runs $0.087 to $0.12 per GB, so moving a terabyte out of one costs more than eight H100-hours.

Preemption notice ranges from two minutes on one hyperscaler to thirty seconds on another, five seconds on one marketplace, and nothing documented at all on another. A node loss costs five to twenty-five minutes of cold start before it serves again. A reliable endpoint therefore needs a warm standby, which roughly doubles the effective hourly cost and removes most of the marketplace saving.

Claims removed by audit

The research agents produced these. A separate reviewer, instructed only to refute and to default to failure on anything it could not verify against a primary source, removed them.

ClaimVerdictWhat replaced it
Cheapest viable node is $9,986/mo at a named vendorFailedThat vendor publishes no rates. The figure existed only in aggregator blogs. Replaced with a verified $1.99 per GPU-hour.
Cheapest 8x H200 is $28.16/hrCorrectedReal but not cheapest. A verified $3.00 per GPU-hour exists, so every break-even was recomputed.
Opus 4.6 scores 44.9 in adaptive modeFailedNo such figure is published. One variant exists at 39, itself an estimate.
Qwen3.8-27B scores around 30FailedWrong by 22 points. The measured value is 52.
Kimi K3 weighs about 594 GBFailedWrong by 2.6x. The repository measures 1,561 GB.
Free EU supercomputing hours cover productionCorrectedFree access is scoped to innovation purposes. Commercial production is a separate paid track.
A subscription tier unlocks 1M context at $39DisputedVendor documentation and its own pricing page disagree. Verify at checkout.

Sources

Every figure above traces to one of these. Provider pricing pages were read on 2026-08-19; model files and configuration were read from the Hugging Face API on the same day.

Outbound links carry rel="nofollow" and open in a new tab. Prices and leaderboard positions on those pages change without notice, so a figure quoted here should be treated as a reading taken on 2026-08-19 rather than a current quote.

Link check, run at publish time: 50 of 52 outbound targets resolved. The Nebius and Crusoe pricing pages returned a connection failure from the checking location while their DNS resolved normally and every other provider responded, which points at a network-level block rather than an outage. Their figures were captured earlier in the research and are retained; verify them yourself before relying on either.

Method and limits

Model sizes were computed by summing safetensors byte counts rather than quoting a card. Architecture figures came from each model configuration file. Cost arithmetic is shown inline throughout so a reader can check it rather than trust it.

Three limits are worth stating. Active parameter counts marked as derived carry roughly a 3 percent error bound, validated against the one model that publishes both. Throughput figures are published measurements from third parties, not reproductions, and peak numbers are quoted at interactivity levels that suit batch work rather than interactive use. And the workload shape used to weight these costs came from one heavy agentic setup, so the 94 percent cache-read figure is characteristic of that pattern rather than universal.

Return to the study