ml0x ← Back to the readable version of this audit This page is the raw verified dataset, kept in its audit-artifact form.

AI_SPEND_AUDIT v2, can anything replace $200/mo of subscriptions?

Run 2026-08-28 · 15 candidates · 377 validated citations · workload observed from 30 days of local logs · every figure tagged observed or modelled · no external requests on this page.

What wins, and why, in ten sentences

  1. OpenRouter wins the scorecard at 60/100 because it is the only option that pairs a frontier coding model with every cost control this audit measured: prepaid credits, per-key spend caps, a live usage dashboard, and zero price or limit changes in the last 90 days.
  2. The hybrid router (LiteLLM, 95% DeepSeek Flash with an Opus 5 fallback) is runner-up at 60/100 and becomes the better pick the moment you can enforce a routing policy, because the same $200 then covers 14 times more of the workload.
  3. Neither can replace the subscription at $200, because the measured month is 33.3B tokens and $200 of OpenRouter credits buys 0.7% of it (parity 0.0072).
  4. On raw value the incumbent Claude Max 20x is the real winner: it delivered $26,187 of API-equivalent compute for $200, a 131× multiple.
  5. Max still scores only 56/100 because its quotas are not published in tokens, its pool is shared with the chat product, and the logs show 14 throttle events in the heaviest week.
  6. The entire ranking is decided by cache-read pricing, since 96.4% of measured input tokens were re-read context.
  7. That is why DeepSeek ($0.022/M cache reads) and GLM-5.3-Flash ($0.015/M) are the only APIs under $1,300 a month for this workload, at $2,137 (V4 Pro) and $601 (Flash sibling row).
  8. Kimi Vivace posted the highest raw score, 61, but lost the B3 tie-break on cost control and paused new signups on 2026-07-20.
  9. The one risk that can flip everything is the +50% weekly boost on Max lapsing 2026-08-31; if delivered capacity drops in week 36, the 131× baseline no longer exists.
  10. So the play is to keep Max, add a hard-capped OpenRouter or LiteLLM overflow lane, and re-measure in September before cancelling anything.

1 · Verified sources (377 citations, re-fetched this run)

Every link below returned HTTP 200 during validation on 2026-08-28 and its page body still contained the cited evidence snippet (L1). Unverified sources are in the appendix.

#candidatewhat it provesdomainfetched (UTC)src
1agg-openroutercredit purchase feeopenrouter.ai2026-08-28T07:35link
2agg-openrouterinference markupopenrouter.ai2026-08-28T07:35link
3agg-openrouterbyok feeopenrouter.ai2026-08-28T07:35link
4agg-openrouterbyok allowance basisopenrouter.ai2026-08-28T07:35link
5agg-openrouterprice usd per 1Mopenrouter.ai2026-08-28T07:35link
6agg-openroutercontext tokens/max output tokensopenrouter.ai2026-08-28T07:35link
7agg-openroutercache write 1h noteopenrouter.ai2026-08-28T07:35link
8agg-openrouterprice usd per 1Mopenrouter.ai2026-08-28T07:35link
9agg-openrouterinput cache hitopenrouter.ai2026-08-28T07:35link
10agg-openrouterprice usd per 1Mopenrouter.ai2026-08-28T07:35link
11agg-openrouterinput cache hitopenrouter.ai2026-08-28T07:35link
12agg-openrouterprice disagreement models apiopenrouter.ai2026-08-28T07:35link
13agg-openrouterprice usd per 1Mopenrouter.ai2026-08-28T07:35link
14agg-openrouterinput cache hitopenrouter.ai2026-08-28T07:35link
15agg-openroutermax output tokensopenrouter.ai2026-08-28T07:35link
16agg-openrouterprice disagreement models apiopenrouter.ai2026-08-28T07:35link
17agg-openrouterrpmopenrouter.ai2026-08-28T07:35link
18agg-openrouterpublished in tokens or requestsopenrouter.ai2026-08-28T07:35link
19agg-openrouterspend cap settableopenrouter.ai2026-08-28T07:35link
20agg-openrouterspend cap settableopenrouter.ai2026-08-28T07:35link
21agg-openrouterusage dashboard realtimeopenrouter.ai2026-08-28T07:35link
22agg-openrouterexactoopenrouter.ai2026-08-28T07:35link
23agg-openrouterprovider routingopenrouter.ai2026-08-28T07:35link
24agg-openrouterfallbackopenrouter.ai2026-08-28T07:35link
25agg-openrouterincidentstatus.openrouter.ai2026-08-28T07:35link
26agg-openrouterSWE-bench Verified (mini-SWE-agent, GLM 5 high)www.swebench.com2026-08-28T07:35link
27agg-openrouterSWE-bench Verified (mini-SWE-agent, Kimi K2 5 high)www.swebench.com2026-08-28T07:35link
28agg-openrouterSWE-bench Verified (mini-SWE-agent, DeepSeek V3 2 high)www.swebench.com2026-08-28T07:35link
29agg-openrouterSWE-bench Verified (live-SWE-agent, Claude 4 5 Opus medium)www.swebench.com2026-08-28T07:35link
30agg-openrouterclaude codeopenrouter.ai2026-08-28T07:35link
31agg-openroutercodex cliopenrouter.ai2026-08-28T07:35link
32agg-openrouteropencodeopenrouter.ai2026-08-28T07:35link
33agg-openrouterzdropenrouter.ai2026-08-28T07:35link
34hybrid-routerrouting roledocs.litellm.ai2026-08-28T07:35link
35hybrid-routerrouting roledocs.litellm.ai2026-08-28T07:35link
36hybrid-routerlicenseraw.githubusercontent.com2026-08-28T07:35link
37hybrid-routerself hosted ossgithub.com2026-08-28T07:35link
38hybrid-routerbudgets in free tierdocs.litellm.ai2026-08-28T07:35link
39hybrid-routermax budget and budget durationdocs.litellm.ai2026-08-28T07:35link
40hybrid-routerbudget scopes key team userdocs.litellm.ai2026-08-28T07:35link
41hybrid-routerrpm tpm per keydocs.litellm.ai2026-08-28T07:35link
42hybrid-routerper model budget per key is enterprisedocs.litellm.ai2026-08-28T07:35link
43hybrid-routerspend trackingdocs.litellm.ai2026-08-28T07:35link
44hybrid-routeranthropic v1 messages endpointdocs.litellm.ai2026-08-28T07:35link
45hybrid-routeranthropic endpoint supports routingdocs.litellm.ai2026-08-28T07:35link
46hybrid-routercredit purchase fee pctopenrouter.ai2026-08-28T07:35link
47hybrid-routercrypto fee pctopenrouter.ai2026-08-28T07:35link
48hybrid-routerinference markup pctopenrouter.ai2026-08-28T07:35link
49hybrid-routerper key credit limit provisioningopenrouter.ai2026-08-28T07:35link
50hybrid-routerprovider routing fallbackopenrouter.ai2026-08-28T07:35link
51hybrid-routerusage dashboard realtimedocs.litellm.ai2026-08-28T07:35link
52hybrid-routerenvdocs.litellm.ai2026-08-28T07:35link
53hybrid-routerofficial anthropic sidedocs.anthropic.com2026-08-28T07:35link
54hybrid-routerbase urldevelopers.openai.com2026-08-28T07:35link
55hybrid-routerenvwww.kimi.com2026-08-28T07:35link
56hybrid-routerbase urlwww.kimi.com2026-08-28T07:35link
57hybrid-routerbase urlopencode.ai2026-08-28T07:35link
58hybrid-routeralt hosts statusstatus.openrouter.ai2026-08-28T07:35link
59hybrid-routerdata retention urldocs.litellm.ai2026-08-28T07:35link
60hybrid-routerdata retention auto deletion is enterprisedocs.litellm.ai2026-08-28T07:35link
61sub-kimi-topmodel idwww.kimi.ai2026-08-28T07:35link
62sub-kimi-topmodel mappingwww.kimi.ai2026-08-28T07:35link
63sub-kimi-topk3 accesswww.kimi.com2026-08-28T07:35link
64sub-kimi-topspeedwww.kimi.ai2026-08-28T07:35link
65sub-kimi-topcredit multiplierwww.kimi.ai2026-08-28T07:35link
66sub-kimi-topplan gatewww.kimi.ai2026-08-28T07:35link
67sub-kimi-topfixed fees usd monthwww.kimi.ai2026-08-28T07:35link
68sub-kimi-topannual pricewww.kimi.ai2026-08-28T07:35link
69sub-kimi-topagent creditswww.kimi.ai2026-08-28T07:35link
70sub-kimi-topconcurrencywww.kimi.ai2026-08-28T07:35link
71sub-kimi-topsession windowwww.kimi.ai2026-08-28T07:35link
72sub-kimi-topweekly capwww.kimi.ai2026-08-28T07:35link
73sub-kimi-topspend cap settablewww.kimi.ai2026-08-28T07:35link
74sub-kimi-topusage dashboard realtimewww.kimi.ai2026-08-28T07:35link
75sub-kimi-topusage dashboard realtimewww.kimi.ai2026-08-28T07:35link
76sub-kimi-topshared pool with chatwww.kimi.ai2026-08-28T07:35link
77sub-kimi-topshared pool with chatwww.kimi.ai2026-08-28T07:35link
78sub-kimi-topextra usage capswww.kimi.ai2026-08-28T07:35link
79sub-kimi-topk3 chat 1mwww.kimi.ai2026-08-28T07:35link
80sub-codex-pro20included messages per 5h windowdevelopers.openai.com2026-08-28T07:35link
81sub-codex-pro20credit rate per 1M tokensdevelopers.openai.com2026-08-28T07:35link
82sub-codex-pro20context tokensplatform.openai.com2026-08-28T07:35link
83sub-codex-pro20max output tokensplatform.openai.com2026-08-28T07:35link
84sub-codex-pro20included messages per 5h windowdevelopers.openai.com2026-08-28T07:35link
85sub-codex-pro20credit rate per 1M tokensdevelopers.openai.com2026-08-28T07:35link
86sub-codex-pro20context tokensplatform.openai.com2026-08-28T07:35link
87sub-codex-pro20included messages per 5h windowdevelopers.openai.com2026-08-28T07:35link
88sub-codex-pro20credit rate per 1M tokensdevelopers.openai.com2026-08-28T07:35link
89sub-codex-pro20availabilitydevelopers.openai.com2026-08-28T07:35link
90sub-codex-pro20separate limitdevelopers.openai.com2026-08-28T07:35link
91sub-codex-pro20fixed fees usd monthdevelopers.openai.com2026-08-28T07:35link
92sub-codex-pro20fixed fees usd monthhelp.openai.com2026-08-28T07:35link mirror
93sub-codex-pro20published in tokens or requestsdevelopers.openai.com2026-08-28T07:35link
94sub-codex-pro20published in tokens or requestsdevelopers.openai.com2026-08-28T07:35link
95sub-codex-pro20spend cap settablehelp.openai.com2026-08-28T07:35link
96sub-codex-pro20usage dashboard realtimedevelopers.openai.com2026-08-28T07:35link
97sub-codex-pro20shared pool with chathelp.openai.com2026-08-28T07:35link
98sub-codex-pro20shared pool with chathelp.openai.com2026-08-28T07:35link
99sub-codex-pro20credits overflow existsdevelopers.openai.com2026-08-28T07:35link
100sub-codex-pro20credits validityhelp.openai.com2026-08-28T07:35link
101sub-codex-pro20credit burn ratedevelopers.openai.com2026-08-28T07:35link
102sub-claude-max20context tokensclaude.com2026-08-28T07:35link
103sub-claude-max20context tokens (conflicting primary source)support.claude.com2026-08-28T07:35link
104sub-claude-max20models includedclaude.com2026-08-28T07:35link
105sub-claude-max20fable5 weekly cap sharesupport.claude.com2026-08-28T07:35link
106sub-claude-max20fixed fees usd monthsupport.claude.com2026-08-28T07:35link
107sub-claude-max20session windowsupport.claude.com2026-08-28T07:35link
108sub-claude-max20weekly capsupport.claude.com2026-08-28T07:35link
109sub-claude-max20weekly cap (Opus separate bucket)support.claude.com2026-08-28T07:35link
110sub-claude-max20published in tokens or requestsclaude.com2026-08-28T07:35link
111sub-claude-max20shared pool with chatclaude.com2026-08-28T07:35link
112sub-claude-max20shared pool with chat (Claude Code article)support.claude.com2026-08-28T07:35link
113sub-claude-max20usage multiplier vs proclaude.com2026-08-28T07:35link
114api-openaiprice usd per 1Mplatform.openai.com2026-08-28T07:35link
115api-openaicontext tokensplatform.openai.com2026-08-28T07:35link
116api-openaimax output tokensplatform.openai.com2026-08-28T07:35link
117api-openaipositioningplatform.openai.com2026-08-28T07:35link
118api-openaiprice usd per 1Mplatform.openai.com2026-08-28T07:35link
119api-openaiprice usd per 1Mplatform.openai.com2026-08-28T07:35link
120api-openailong context surchargeplatform.openai.com2026-08-28T07:35link
121api-openaicache writeplatform.openai.com2026-08-28T07:35link
122api-openaipromo lapseplatform.openai.com2026-08-28T07:35link
123api-openaicontext tokensplatform.openai.com2026-08-28T07:35link
124api-openaimax output tokensplatform.openai.com2026-08-28T07:35link
125api-openaiprice usd per 1Mplatform.openai.com2026-08-28T07:35link
126api-openaicontext tokensplatform.openai.com2026-08-28T07:35link
127api-openaimax output tokensplatform.openai.com2026-08-28T07:35link
128api-openaiprice usd per 1Mplatform.openai.com2026-08-28T07:35link
129api-openairpmplatform.openai.com2026-08-28T07:35link
130api-openairpmplatform.openai.com2026-08-28T07:35link
131api-openaitpmplatform.openai.com2026-08-28T07:35link
132api-openaitpmplatform.openai.com2026-08-28T07:35link
133api-openaimonthly usage limit by tierplatform.openai.com2026-08-28T07:35link
134api-openaimonthly usage limit by tierplatform.openai.com2026-08-28T07:35link
135api-openaipublished in tokens or requestsplatform.openai.com2026-08-28T07:35link
136api-openaispend cap settableplatform.openai.com2026-08-28T07:35link
137api-openaispend cap settableplatform.openai.com2026-08-28T07:35link
138api-openaiusage dashboard realtimeplatform.openai.com2026-08-28T07:35link
139self-rented-gpulicensehuggingface.co2026-08-28T07:35link
140self-rented-gpuparams total b/params active brecipes.vllm.ai2026-08-28T07:35link
141self-rented-gpuparams total b (disagreeing source, E4)huggingface.co2026-08-28T07:35link
142self-rented-gpucontext tokensrecipes.vllm.ai2026-08-28T07:35link
143self-rented-gpumin serving configrecipes.vllm.ai2026-08-28T07:35link
144self-rented-gpumin serving config 1Mrecipes.vllm.ai2026-08-28T07:35link
145self-rented-gpupublished throughputrecipes.vllm.ai2026-08-28T07:35link
146self-rented-gpulicensehuggingface.co2026-08-28T07:35link
147self-rented-gpuparams total b/params active bhuggingface.co2026-08-28T07:35link
148self-rented-gpucontext tokenshuggingface.co2026-08-28T07:35link
149self-rented-gpumin serving configrecipes.vllm.ai2026-08-28T07:35link
150self-rented-gpumin serving config amdrecipes.vllm.ai2026-08-28T07:35link
151self-rented-gpulicensehuggingface.co2026-08-28T07:35link
152self-rented-gpuparams total b/params active bhuggingface.co2026-08-28T07:35link
153self-rented-gpucontext tokenshuggingface.co2026-08-28T07:35link
154self-rented-gpumin serving configrecipes.vllm.ai2026-08-28T07:35link
155self-rented-gpulicensehuggingface.co2026-08-28T07:35link
156self-rented-gpuparamshuggingface.co2026-08-28T07:35link
157self-rented-gpucontext tokenshuggingface.co2026-08-28T07:35link
158self-rented-gpuserving recipehuggingface.co2026-08-28T07:35link
159self-rented-gpurunpod H200www.runpod.io2026-08-28T07:35link
160self-rented-gpurunpod H100 SXMwww.runpod.io2026-08-28T07:35link
161self-rented-gpurunpod H100 PCIewww.runpod.io2026-08-28T07:35link
162self-rented-gpurunpod B200www.runpod.io2026-08-28T07:35link
163self-rented-gpurunpod B300www.runpod.io2026-08-28T07:35link
164self-rented-gpurunpod serverless H200www.runpod.io2026-08-28T07:35link
165self-rented-gpurunpod serverless H100www.runpod.io2026-08-28T07:35link
166self-rented-gpurunpod serverless B200www.runpod.io2026-08-28T07:35link
167self-rented-gpulambda H100 8xlambda.ai2026-08-28T07:35link
168self-rented-gpulambda B200 8xlambda.ai2026-08-28T07:35link
169self-rented-gpulambda H100 1xlambda.ai2026-08-28T07:35link
170self-rented-gpuvast H100vast.ai2026-08-28T07:35link
171self-rented-gpuvast H200vast.ai2026-08-28T07:35link
172self-rented-gpuvast rate basisvast.ai2026-08-28T07:35link
173self-rented-gpumodal H200modal.com2026-08-28T07:35link
174self-rented-gpumodal H100modal.com2026-08-28T07:35link
175self-rented-gpumodal B200modal.com2026-08-28T07:35link
176self-rented-gpurunpod serverlessdocs.runpod.io2026-08-28T07:35link
177self-rented-gpurunpod serverless idledocs.runpod.io2026-08-28T07:35link
178self-rented-gpumodalmodal.com2026-08-28T07:35link
179self-rented-gpuconcurrencyrecipes.vllm.ai2026-08-28T07:35link
180self-rented-gpuspend cap settabledocs.runpod.io2026-08-28T07:35link
181self-rented-gpuspend cap settabledocs.runpod.io2026-08-28T07:35link
182self-rented-gpupublished in tokens or requestsdocs.runpod.io2026-08-28T07:35link
183self-rented-gpuArtificial Analysis Intelligence Index v4 1 1, GLM-5 2 (maxartificialanalysis.ai2026-08-28T07:35link
184self-rented-gpuArtificial Analysis Intelligence Index v4 1 1, Kimi K3 (maxartificialanalysis.ai2026-08-28T07:35link
185self-rented-gpuArtificial Analysis Intelligence Index v4 1 1, Kimi K2 7 Coartificialanalysis.ai2026-08-28T07:35link
186self-rented-gpuArtificial Analysis median output speed (API hosts), GLM-5 artificialanalysis.ai2026-08-28T07:35link
187self-rented-gpuSWE-bench (bash-only, mini-SWE-agent), Kimi K2 5 (high)www.swebench.com2026-08-28T07:35link
188self-rented-gpuSWE-bench (bash-only, mini-SWE-agent), GLM 5 (high)www.swebench.com2026-08-28T07:35link
189self-rented-gpuSWE-bench (bash-only, mini-SWE-agent), Kimi K2 Thinkingwww.swebench.com2026-08-28T07:35link
190self-rented-gpuSWE-bench (bash-only, mini-SWE-agent), Qwen3-Coder 480B/A35www.swebench.com2026-08-28T07:35link
191self-rented-gpuTerminal Bench 2 1 (Terminus-2), GLM-5 2 (vendor model cardhuggingface.co2026-08-28T07:35link
192self-rented-gpuTerminal-Bench 2 1, Kimi K3 (vendor model card, Kimi Code hhuggingface.co2026-08-28T07:35link
193agg-coding-planstier price starterwww.atlascloud.ai2026-08-28T07:35link
194agg-coding-planstier quota starterwww.atlascloud.ai2026-08-28T07:35link
195agg-coding-planstier price ultrawww.atlascloud.ai2026-08-28T07:35link
196agg-coding-planstier quota ultrawww.atlascloud.ai2026-08-28T07:35link
197agg-coding-plansquota structurewww.atlascloud.ai2026-08-28T07:35link
198agg-coding-plansweekly resetwww.atlascloud.ai2026-08-28T07:35link
199agg-coding-planspoints formulawww.atlascloud.ai2026-08-28T07:35link
200agg-coding-plansmultiplier glm5www.atlascloud.ai2026-08-28T07:35link
201agg-coding-plansmultiplier kimi k2 6www.atlascloud.ai2026-08-28T07:35link
202agg-coding-planssavings glm5www.atlascloud.ai2026-08-28T07:35link
203agg-coding-plansoveragewww.atlascloud.ai2026-08-28T07:35link
204agg-coding-planspayg pack 99usd 280Mwww.atlascloud.ai2026-08-28T07:35link
205agg-coding-planstier price litenovita.ai2026-08-28T07:35link
206agg-coding-planstier quota lite tokensnovita.ai2026-08-28T07:35link
207agg-coding-planstier rpm litenovita.ai2026-08-28T07:35link
208agg-coding-planstier price pronovita.ai2026-08-28T07:35link
209agg-coding-planstier quota pro tokensnovita.ai2026-08-28T07:35link
210agg-coding-planstier rpm pronovita.ai2026-08-28T07:35link
211agg-coding-planstier price maxnovita.ai2026-08-28T07:35link
212agg-coding-planstier quota max tokensnovita.ai2026-08-28T07:35link
213agg-coding-planstier rpm maxnovita.ai2026-08-28T07:35link
214agg-coding-plansmodels servednovita.ai2026-08-28T07:35link
215agg-coding-plansprice usd per 1Mwww.atlascloud.ai2026-08-28T07:35link
216agg-coding-planscontext tokenswww.atlascloud.ai2026-08-28T07:35link
217agg-coding-planslist price originwww.atlascloud.ai2026-08-28T07:35link
218agg-coding-plansprice usd per 1Mwww.atlascloud.ai2026-08-28T07:35link
219agg-coding-planscontext tokenswww.atlascloud.ai2026-08-28T07:35link
220agg-coding-plansinput and cachenovita.ai2026-08-28T07:35link
221agg-coding-plansoutputnovita.ai2026-08-28T07:35link
222agg-coding-planscontext tokensnovita.ai2026-08-28T07:35link
223agg-coding-plansinputnovita.ai2026-08-28T07:35link
224agg-coding-plansinput cache hitnovita.ai2026-08-28T07:35link
225agg-coding-plansoutputnovita.ai2026-08-28T07:35link
226agg-coding-planscontext tokensnovita.ai2026-08-28T07:35link
227agg-coding-plansinputnovita.ai2026-08-28T07:35link
228agg-coding-plansoutputnovita.ai2026-08-28T07:35link
229agg-coding-plansSWE-bench Verified (mini-SWE-agent, GLM 5 high)www.swebench.com2026-08-28T07:35link
230agg-coding-plansSWE-bench Verified (mini-SWE-agent, Kimi K2 5 high)www.swebench.com2026-08-28T07:35link
231agg-coding-plansSWE-bench Verified (mini-SWE-agent, DeepSeek V3 2 high)www.swebench.com2026-08-28T07:35link
232agg-coding-plansclaude codewww.atlascloud.ai2026-08-28T07:35link
233agg-coding-plansstatus pagestatus.novita.ai2026-08-28T07:35link
234agg-coding-plansdata retentionnovita.ai2026-08-28T07:35link
235sub-zai-codingrouting of older model idsdocs.z.ai2026-08-28T07:35link
236sub-zai-codingcredit multipliers (per 10,000 tokens): GLM-5 3 input 6 9 / docs.z.ai2026-08-28T07:35link
237sub-zai-codingpeak windows/peak multiplierdocs.z.ai2026-08-28T07:35link
238sub-zai-codingoff-peak credit discountdocs.z.ai2026-08-28T07:35link
239sub-zai-codingcredit multipliers: Flash input 2 3 / cached 0 56 / output 8docs.z.ai2026-08-28T07:35link
240sub-zai-codingsession window (5-hour credits table)docs.z.ai2026-08-28T07:35link
241sub-zai-codingsession window reset ruledocs.z.ai2026-08-28T07:35link
242sub-zai-codingweekly cap reset ruledocs.z.ai2026-08-28T07:35link
243sub-zai-codingconcurrency off-peak boostdocs.z.ai2026-08-28T07:35link
244sub-zai-codingspend cap settable false (fixed-fee subscription; hard credidocs.z.ai2026-08-28T07:35link
245sub-zai-codingshared pool with chat false (plan usable only in supported cdocs.z.ai2026-08-28T07:35link
246sub-zai-codingusage visibilitydocs.z.ai2026-08-28T07:35link
247sub-zai-codingLite pricez.ai2026-08-28T07:35link
248sub-zai-codingPro pricez.ai2026-08-28T07:35link
249sub-zai-codingMax pricez.ai2026-08-28T07:35link
250sub-zai-codingbilling-cycle discountsz.ai2026-08-28T07:35link
251sub-zai-codingteam seatsz.ai2026-08-28T07:35link
252sub-zai-codingmonthly list price floordocs.z.ai2026-08-28T07:35link
253sub-zai-codingreferral first-order example (hypothetical)docs.z.ai2026-08-28T07:35link
254api-anthropicprice usd per 1M (input miss, cache write 5m, cache write 1hplatform.claude.com2026-08-28T07:35link
255api-anthropiccontext tokensplatform.claude.com2026-08-28T07:35link
256api-anthropicmax output tokensplatform.claude.com2026-08-28T07:35link
257api-anthropicmodel idplatform.claude.com2026-08-28T07:35link
258api-anthropicprice usd per 1M (input miss, cache write 5m, cache write 1hplatform.claude.com2026-08-28T07:35link
259api-anthropicprice now standardplatform.claude.com2026-08-28T07:35link
260api-anthropicprice usd per 1M (input miss, cache write 5m, cache write 1hplatform.claude.com2026-08-28T07:35link
261api-anthropicrpm/tpm (Claude Opus 5, Start tier; ITPM 2,000,000, OTPM 400platform.claude.com2026-08-28T07:35link
262api-anthropicrpm/tpm higher tiers (Opus 5: Build 5,000/5,000,000/1,000,00platform.claude.com2026-08-28T07:35link
263api-anthropicpublished in tokens or requestsplatform.claude.com2026-08-28T07:35link
264api-anthropicmonthly spend caps by tierplatform.claude.com2026-08-28T07:35link
265api-anthropicspend cap settableplatform.claude.com2026-08-28T07:35link
266api-anthropicspend cap settable (workspace)platform.claude.com2026-08-28T07:35link
267api-anthropicshared pool with chatsupport.claude.com2026-08-28T07:35link
268api-anthropiccache read excluded from itpmplatform.claude.com2026-08-28T07:35link
269api-moonshotprice usd per 1Mplatform.kimi.ai2026-08-28T07:35link
270api-moonshotcontext tokensplatform.kimi.ai2026-08-28T07:35link
271api-moonshotopen weights licensehuggingface.co2026-08-28T07:35link
272api-moonshotparamshuggingface.co2026-08-28T07:35link
273api-moonshotmin serving configrecipes.vllm.ai2026-08-28T07:35link
274api-moonshotmin serving config rocmrecipes.vllm.ai2026-08-28T07:35link
275api-moonshotquantizationhuggingface.co2026-08-28T07:35link
276api-moonshotprice usd per 1Mplatform.kimi.ai2026-08-28T07:35link
277api-moonshotopen weights licensehuggingface.co2026-08-28T07:35link
278api-moonshotparamshuggingface.co2026-08-28T07:35link
279api-moonshotmin serving confighuggingface.co2026-08-28T07:35link
280api-moonshotprice usd per 1Mplatform.kimi.ai2026-08-28T07:35link
281api-moonshotspeedplatform.kimi.ai2026-08-28T07:35link
282api-moonshotprice usd per 1Mplatform.kimi.ai2026-08-28T07:35link
283api-moonshothf repohuggingface.co2026-08-28T07:35link
284api-moonshotprice usd per 1Mplatform.kimi.ai2026-08-28T07:35link
285api-moonshothf repohuggingface.co2026-08-28T07:35link
286api-moonshotrpmplatform.kimi.ai2026-08-28T07:35link
287api-moonshottpmplatform.kimi.ai2026-08-28T07:35link
288api-moonshotconcurrencyplatform.kimi.ai2026-08-28T07:35link
289api-moonshotentry rechargeplatform.kimi.ai2026-08-28T07:35link
290api-moonshotpublished in tokens or requestsplatform.kimi.ai2026-08-28T07:35link
291api-moonshotspend cap settableplatform.kimi.ai2026-08-28T07:35link
292api-moonshotusage dashboard realtimeplatform.kimi.ai2026-08-28T07:35link
293api-moonshotshared pool with chatplatform.kimi.ai2026-08-28T07:35link
294api-moonshotcache hit conditionplatform.kimi.ai2026-08-28T07:35link
295api-deepseekmodel id/model versionapi-docs.deepseek.com2026-08-28T07:35link
296api-deepseekcontext tokens/max output tokensapi-docs.deepseek.com2026-08-28T07:35link
297api-deepseekinput miss (off-peak/peak; cols flash,pro,vision)api-docs.deepseek.com2026-08-28T07:35link
298api-deepseekinput cache hitapi-docs.deepseek.com2026-08-28T07:35link
299api-deepseekoutputapi-docs.deepseek.com2026-08-28T07:35link
300api-deepseekpeak windows utc/peak multiplierapi-docs.deepseek.com2026-08-28T07:35link
301api-deepseekpeak pricing effective dateapi-docs.deepseek.com2026-08-28T07:35link
302api-deepseekopen weights licensehuggingface.co2026-08-28T07:35link
303api-deepseekparams total b/params active b (independent)artificialanalysis.ai2026-08-28T07:35link
304api-deepseekmin serving config recipe url (vLLM/SGLang serve commands onhuggingface.co2026-08-28T07:35link
305api-deepseekopen weights licensehuggingface.co2026-08-28T07:35link
306api-deepseekconcurrency (500 = v4-pro; v4-flash and vision = 2500; accouapi-docs.deepseek.com2026-08-28T07:35link
307api-deepseekconcurrency exceeded behaviorapi-docs.deepseek.com2026-08-28T07:35link
308api-deepseekpublished in tokens or requests (limits published as concurrapi-docs.deepseek.com2026-08-28T07:35link
309api-deepseekbalance model (prepaid: top-up balance + granted balance)api-docs.deepseek.com2026-08-28T07:35link
310cf-workers-aiprice usd per 1M input miss/output/input cache hitdevelopers.cloudflare.com2026-08-28T07:35link
311cf-workers-aineuron unit pricedevelopers.cloudflare.com2026-08-28T07:35link
312cf-workers-aineurons per tokendevelopers.cloudflare.com2026-08-28T07:35link
313cf-workers-aicontext tokensdevelopers.cloudflare.com2026-08-28T07:35link
314cf-workers-aiunit pricing model pagedevelopers.cloudflare.com2026-08-28T07:35link
315cf-workers-aipaid access requireddevelopers.cloudflare.com2026-08-28T07:35link
316cf-workers-aiparams total b/params active bdevelopers.cloudflare.com2026-08-28T07:35link
317cf-workers-aihf repodevelopers.cloudflare.com2026-08-28T07:35link
318cf-workers-aiprice usd per 1M input miss/output/input cache hitdevelopers.cloudflare.com2026-08-28T07:35link
319cf-workers-aineurons per tokendevelopers.cloudflare.com2026-08-28T07:35link
320cf-workers-aiderived price formuladevelopers.cloudflare.com2026-08-28T07:35link
321cf-workers-aicontext tokensdevelopers.cloudflare.com2026-08-28T07:35link
322cf-workers-aicontext tokens native comparisonartificialanalysis.ai2026-08-28T07:35link
323cf-workers-ailaunch datedevelopers.cloudflare.com2026-08-28T07:35link
324cf-workers-aiprice usd per 1M input miss/output/input cache hitdevelopers.cloudflare.com2026-08-28T07:35link
325cf-workers-aineurons per tokendevelopers.cloudflare.com2026-08-28T07:35link
326cf-workers-aiderived price formuladevelopers.cloudflare.com2026-08-28T07:35link
327cf-workers-aicontext tokensdevelopers.cloudflare.com2026-08-28T07:35link
328cf-workers-aiparams total bdevelopers.cloudflare.com2026-08-28T07:35link
329cf-workers-aihf repodevelopers.cloudflare.com2026-08-28T07:35link
330cf-workers-airpmdevelopers.cloudflare.com2026-08-28T07:35link
331cf-workers-airpm default text generationdevelopers.cloudflare.com2026-08-28T07:35link
332cf-workers-airpm note glm53flashdevelopers.cloudflare.com2026-08-28T07:35link
333cf-workers-airpm elevated via prepaiddevelopers.cloudflare.com2026-08-28T07:35link
334cf-workers-aifree daily neuronsdevelopers.cloudflare.com2026-08-28T07:35link
335cf-workers-aidaily reset and hard faildevelopers.cloudflare.com2026-08-28T07:35link
336cf-workers-aifrontier models excluded from freedevelopers.cloudflare.com2026-08-28T07:35link
337cf-workers-aifixed fees usd monthdevelopers.cloudflare.com2026-08-28T07:35link
338cf-workers-aipublished in tokens or requestsdevelopers.cloudflare.com2026-08-28T07:35link
339cf-workers-aispend cap settabledevelopers.cloudflare.com2026-08-28T07:35link
340cf-workers-aiusage dashboard realtimedevelopers.cloudflare.com2026-08-28T07:35link
341api-zaiinput miss/input cache hit/outputdocs.z.ai2026-08-28T07:35link
342api-zaicontext tokens/max output tokensdocs.z.ai2026-08-28T07:35link
343api-zaireasoning always ondocs.z.ai2026-08-28T07:35link
344api-zaiopen weights: none for GLM-5 3 (base is GLM-5 2)docs.z.ai2026-08-28T07:35link
345api-zaiproprietary label (independent)artificialanalysis.ai2026-08-28T07:35link
346api-zaiprices (current 50% promo; list prices struck through: inputdocs.z.ai2026-08-28T07:35link
347api-zaiinput cache hit promo celldocs.z.ai2026-08-28T07:35link
348api-zaicontext tokens/max output (text params = GLM-5 3)docs.z.ai2026-08-28T07:35link
349api-zaiparams total b/params active bhuggingface.co2026-08-28T07:35link
350api-zaiopen weights licensehuggingface.co2026-08-28T07:35link
351api-zaimin serving config recipe urlhuggingface.co2026-08-28T07:35link
352api-zaiconcurrency/rpm/tpm not public; console-gateddocs.z.ai2026-08-28T07:35link
353api-zaiusage dashboard realtime (false: billing shows previous day)docs.z.ai2026-08-28T07:35link
354api-zaicache mechanism beta; hit ~1/5 of pricedocs.z.ai2026-08-28T07:35link
355api-zaibalance model (prepaid recharge)docs.z.ai2026-08-28T07:35link
356self-locallicensehuggingface.co2026-08-28T07:35link
357self-localparamshuggingface.co2026-08-28T07:35link
358self-localquantized footprint gb Q4 K Shuggingface.co2026-08-28T07:35link
359self-localquantized footprint gb UD-Q2 K XLhuggingface.co2026-08-28T07:35link
360self-localquantized footprint gb UD-IQ1 Shuggingface.co2026-08-28T07:35link
361self-localgguf repo sizehuggingface.co2026-08-28T07:35link
362self-locallicensehuggingface.co2026-08-28T07:35link
363self-localparamshuggingface.co2026-08-28T07:35link
364self-localcontext tokenshuggingface.co2026-08-28T07:35link
365self-locallocal runtimeshuggingface.co2026-08-28T07:35link
366self-localquantized footprint gb 4bithuggingface.co2026-08-28T07:35link
367self-localapple silicon pathhuggingface.co2026-08-28T07:35link
368self-localrepresentative config price usdwww.apple.com2026-08-28T07:35link
369self-localrepresentative config namewww.apple.com2026-08-28T07:35link
370self-localbase config price usdwww.apple.com2026-08-28T07:35link
371self-localbase config namewww.apple.com2026-08-28T07:35link
372self-localavailabilitywww.apple.com2026-08-28T07:35link
373self-localArtificial Analysis Intelligence Index, Qwen3 Coder 30B A3Bartificialanalysis.ai2026-08-28T07:35link
374self-localArtificial Analysis Intelligence Index, GLM-4 5-Airartificialanalysis.ai2026-08-28T07:35link
375self-localArtificial Analysis median output speed (API hosts), Qwen3 artificialanalysis.ai2026-08-28T07:35link
376self-localSWE-bench Verified, OpenHands + Qwen3-Coder-30B-A3B-Instrucwww.swebench.com2026-08-28T07:35link
377self-localSWE-bench Verified, EntroPO + R2E + Qwen3-Coder-30B-A3B-Inswww.swebench.com2026-08-28T07:35link

2 · Verdict

Winner: OpenRouter, 60/100. It is the only candidate that pairs a frontier coding model (Opus 5; SWE-bench Verified 79.20 via an independent leaderboard entry) with every cost-control property this audit measured: prepaid credits, per-key spend limits, a real-time activity dashboard, pass-through per-token pricing plus a 5.5% credit fee, and zero recorded price or limit changes in the last 90 days. At the observed workload its modelled cost is $27,628/month, parity 0.0072, meaning $200 buys 0.7% of the volume consumed today. The parity hard-cap (≤60) encodes that: no scored option at $200/month reproduces the observed 33.4B-token month. The incumbent Claude Max 20x (56/100) is the only parity-1.0 row, it demonstrably served that volume for $200, at the cost of the lowest cost-control score in the top half (quotas not published in tokens, pool shared with chat, 14 throttle events in week 32). Runner-up: hybrid-router, 60/100 (LiteLLM proxy, 95/5 token split: DeepSeek V4 Flash default, Opus 5 fallback, hard monthly budget at the proxy). It is the better pick than OpenRouter when a routing policy is enforceable: the same money then covers parity 0.1013 (14× more) and every dollar is capped by LiteLLM max_budget before it is spent. The single risk most likely to change this answer: the Claude Max +50% weekly boost lapses on 2026-08-31. If delivered Max capacity drops in week 36, the parity-1.0 baseline measured here no longer exists, and the hybrid becomes the only path that keeps both capacity and control.

B3 tie-break: raw scores within 3 points of the top (61 / 60 / 60) are ordered by the Cost-control dimension, then Reliability. Kimi Vivace holds the highest raw score (61) but control 65 vs 100/100 for the two winners; its new signups have also been paused since 2026-07-20.

#candidatescorecost @ observedparityconfidence
1OpenRouter
agg-openrouter
60
cap 60: parity 0.0072 < 1.0
$27,628/mo 0.0072 LOW
2Router: low-cost open-weight default + frontier fallback, hard monthly cap
hybrid-router
60
cap 60: parity 0.1013 < 1.0
$1,973/mo 0.1013 LOW
3Kimi membership (top tier)
sub-kimi-top
61
$199/mo n/a LOW
4ChatGPT Pro / Codex
sub-codex-pro20
57
$200/mo n/a LOW
5Claude Max 20x
sub-claude-max20
56
$200/mo 1.0000 LOW
6OpenAI API PAYG
api-openai
55
$20,950/mo 0.0095 LOW
7Self-host open weights on rented GPUs
self-rented-gpu
53
$5,598/mo 0.0357 LOW
8Discount coding plans (resellers)
agg-coding-plans
50
$9,485/mo 0.0211 LOW
9Z.ai GLM Coding Plan
sub-zai-coding
49
$168/mo 0.0914 LOW
10Anthropic API PAYG
api-anthropic
48
$26,187/mo 0.0076 LOW
11Moonshot API
api-moonshot
45
$15,011/mo 0.0133 LOW
12DeepSeek API
api-deepseek
42
$2,137/mo 0.0936 LOW
13Cloudflare Workers AI
cf-workers-ai
39
$1,206/mo 0.1658 LOW
14Z.ai GLM API
api-zai
37
$10,539/mo 0.0190 LOW
15Local hardware
self-local
35
null/mo n/a LOW

3 · What changed since last run

First run, no out/prev/ exists.

4 · Score breakdown (hover a cell for the arithmetic)

candidatecost (25%)control (20%)capability (25%)throughput (10%)tools (10%)reliability (10%)final
agg-openrouter0.0100100.050755060
hybrid-router0.010098.11001002560
sub-kimi-top40.36570100752561
sub-codex-pro2040.05099.250502557
sub-claude-max2040.040100.050355056
api-openai0.05595.9100505055
self-rented-gpu0.04085.7100855053
agg-coding-plans0.05091.950755050
sub-zai-coding49.6582.650505049
api-anthropic0.05100.0100507548
api-moonshot0.00701001007545
api-deepseek0.0070100757542
cf-workers-ai0.0090.430557539
api-zai0.0085.750605037
self-local04051.630852535

Hard caps fired:

Capability uses independent benchmarks only (SWE-bench Verified, Terminal-Bench 2.1, Artificial Analysis Intelligence Index), normalised to best-in-set (83.8 TB / 63.05 AA / 79.2 SWE); fewer than 2 independent entries caps the dimension at 70. The S4 bake-off was skipped (no funded key), so no 50/50 blend was applied.

5 · Cost model, all cells modelled; workload mix observed

Observed 30-day mix (Claude Code logs, 2026-07-29 → 2026-08-28): input-miss 250M, cache-write 935M, cache-read 32.1B, output 122M tokens; cache-hit rate 96.4%. Formula: cost = in_miss·p_in + cache_read·p_hit + cache_write·(p_cw or p_in) + out·p_out, × peak adjustment × (1+fee) + fixed. DeepSeek peak (Mon–Fri 01–04 & 06–10 UTC = 09–12 & 14–18 Kuala Lumpur) carries 2× prices; 23.6% of observed tokens fall inside it. The off-peak-shift column halves that fraction (batch work moved off-peak).

candidatemodel / planobservedLOWMIDHIGHoff-peak shiftparitysub break-even
api-anthropicAnthropic API - Opus 5 $26,187 $192$652 $2,062 - 0.0076 -
api-anthropicAnthropic API - Sonnet 5 $10,475 $77$261 $825 - 0.0191 -
api-anthropicAnthropic API - Fable 5 $52,375 $384$1,305 $4,125 - 0.0038 -
api-openaiOpenAI API - GPT-5.6 Sol $20,950 $154$522 $1,650 - 0.0095 -
api-openaiOpenAI API - GPT-5.3-codex $9,397 $88$307 $984 - 0.0213 -
api-deepseekDeepSeek V4 Pro (peak-weighted) $2,137 $23$76 $234 $1,933 0.0936 -
api-deepseekDeepSeek V4 Flash (peak-weighted) $699 $8$25 $78 $632 0.2861 -
api-zaiZ.ai GLM-5.3 $10,539 $46$156 $492 - 0.0190 -
api-zaiZ.ai GLM-5.3-Flash $601 $3$9 $28 - 0.3329 -
api-moonshotMoonshot Kimi K3 $15,011 $115$392 $1,238 - 0.0133 -
api-moonshotMoonshot K2.7-Code $7,711 $36$123 $390 - 0.0259 -
cf-workers-aiCF Workers AI GLM-5.3-Flash $1,206 $10$22 $60 - 0.1658 -
cf-workers-aiCF Workers AI GLM-5.2 $10,544 $51$161 $498 - 0.0190 -
agg-openrouterOpenRouter Opus 5 (+5.5% fee) $27,628 $203$688 $2,176 - 0.0072 -
agg-openrouterOpenRouter GLM-5.3 (+5.5% fee) $11,119 $49$165 $520 - 0.0180 -
agg-coding-plansAtlas Cloud GLM-5.2 PAYG $9,485 $42$141 $443 - 0.0211 -
hybrid-routerHybrid 80/20 (LiteLLM, no fee) $5,797 $45$151 $475 - 0.0345 -
hybrid-routerHybrid 90/10 (LiteLLM, no fee) $3,248 $26$88 $276 - 0.0616 -
hybrid-routerHybrid 95/5 (LiteLLM, no fee) $1,973 $17$57 $177 - 0.1013 -
sub-claude-max20Claude Max 20x $200 $200$200 $200 - 1.0000 255M
sub-codex-pro20ChatGPT Pro / Codex $200 $200$200 $200 - - 319M
sub-kimi-topKimi membership (top tier) $199 $199$199 $199 - - 443M
sub-zai-codingGLM Coding Plan Lite (list $18.0/mo) $18 $18$18 $18 - 0.0065 -
sub-zai-codingGLM Coding Plan Pro (list $80.0/mo) $80 $80$80 $80 - 0.0392 -
sub-zai-codingGLM Coding Plan Max (list $168.0/mo) $168 $168$168 $168 - 0.0914 -
self-rented-gpuGLM-5.2 8xH200 RunPod, concurrency 1 $24,483 nullnull null - 0.0082 -
self-rented-gpuGLM-5.2 8xH200 RunPod, concurrency 4 $8,295 nullnull null - 0.0241 -
self-rented-gpuGLM-5.2 8xH200 RunPod, concurrency 8 $5,598 nullnull null - 0.0357 -
self-localQwen3-Coder-30B / GLM-4.5-Air local null nullnull null - - -

Sub break-even = token volume at which the same models on PAYG API cost $200/month; the sub is cheaper above it. Observed volume is 33,398M tokens, 131× Claude Max's break-even of 255.1M.

6 · Observed usage & throttle timeline (tool: claude_code; per ISO week, total billed tokens)

W31
4.6B
W32
9.3B
▲▲▲▲▲▲▲▲▲▲▲▲▲▲ 14 throttle
W33
9.0B
W34
5.2B
W35
5.4B
boost lapse 08-31 →

Throttle events observed in logs: 14 (all in week 2026-W32). Codex CLI and Kimi app stores held no token logs on this machine (Codex sqlite threads table: 0 rows; ~/.kimi: config only), their usage is not in these bars; Kimi K3 tokens routed through Claude Code (1,567M input) are.

7 · Candidate cards

OpenRouter, aggregator · score 60 · LOW

observed $27627.79/mo, parity 0.0072, control 100, capability 100.0

Limits as published: rpm: 20
spend-cap:yes · realtime-dash:yes · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • SWE-bench Verified (mini-SWE-agent, GLM 5 high): 72.8
  • SWE-bench Verified (mini-SWE-agent, Kimi K2.5 high): 70.8
  • SWE-bench Verified (mini-SWE-agent, DeepSeek V3.2 high): 70.0
  • SWE-bench Verified (live-SWE-agent, Claude 4.5 Opus medium): 79.2
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

Router: low-cost open-weight default + frontier fallback, hard monthly cap, hybrid · score 60 · LOW

observed $1973.48/mo, parity 0.1013, control 100, capability 98.1

Limits as published: not published
spend-cap:yes · realtime-dash:yes · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • none
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

Kimi membership (top tier), subscription · score 61 · LOW

observed $199.0/mo, parity None, control 65, capability 70

Limits as published: concurrency: 30; session_window: 5-hour window, approx 300-1,200 requests depending on plan; weekly_cap: weekly rate limit + 7-day credit refresh cycle (numeric cap not published)
spend-cap:yes · realtime-dash:yes · shared-with-chat:yes · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Artificial Analysis Intelligence Index v4.1.1 - Kimi K3 (max) (K3 usable on Moderato+ Kimi Code plans): 60
Vendor-claimed (zero weight)
  • Terminal-Bench 2.1 (vendor-run, HF Kimi-K3 model card): 88.3

Bake-off: skipped, no API key in env

ChatGPT Pro / Codex, subscription · score 57 · LOW

observed $200.0/mo, parity None, control 50, capability 99.2

Limits as published: session_window: 5-hour rolling window shared by local messages and cloud chats; per-model included ranges ; weekly_cap: exists but unpublished ('Additional weekly limits may apply')
spend-cap:yes · realtime-dash:yes · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Terminal-Bench 2.1 (Codex agent + GPT-5.5, xhigh): 83.1
  • Terminal-Bench 2.1 (Codex agent + GPT-5.6 Terra, max): 78.4
  • Terminal-Bench 2.1 (Codex agent + GPT-5.6 Luna, max): 75.7
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

Claude Max 20x, subscription · score 56 · LOW

observed $200.0/mo, parity 1.0, control 40, capability 100.0

Limits as published: session_window: 5-hour rolling session; resets every five hours; weekly_cap: weekly usage limit across all models (separate Opus bucket); quantity unpublished; resets
spend-cap:yes · realtime-dash:yes · shared-with-chat:yes · quotas-in-tokens/requests:no

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Terminal-Bench 2.1, Claude Code + Fable 5 (xhigh), accuracy % (rank 1): 83.8
  • Terminal-Bench 2.1, Claude Code + Opus 4.8 (high), accuracy %: 78.9
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

OpenAI API PAYG, api · score 55 · LOW

observed $20949.98/mo, parity 0.0095, control 55, capability 95.9

Limits as published: rpm: 500 (Tier 1) to 15,000 (Tier 5) for gpt-5.3-codex and gpt-5.6 family; tpm: 500,000 (Tier 1) to 40,000,000 (Tier 5) for gpt-5.3-codex and gpt-5.6 family; weekly_cap: none; OpenAI-approved monthly usage limit by tier instead: $100/mo (Free, Tier 1) to $200,
spend-cap:yes · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • SWE-bench Verified (mini-SWE-agent + GPT 5.2 Codex; newest OpenAI-model entry, no gpt-5.3-codex/gpt-5.6 submis: 72.8
  • Terminal-Bench 2.1 (Codex agent + GPT-5.5, xhigh): 83.1
  • Terminal-Bench 2.1 (Codex agent + GPT-5.6 Terra, max): 78.4
  • Artificial Analysis Intelligence Index (GPT-5.6 Sol, max): 60.9
  • Artificial Analysis Agentic Index (GPT-5.6 Sol, max): 57.8
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

Self-host open weights on rented GPUs, self-host · score 53 · LOW

observed $5597.53/mo, parity 0.0357, control 40, capability 85.7

Limits as published: concurrency: operator-set (vLLM --max-num-seqs)
spend-cap:yes · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:no

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Artificial Analysis Intelligence Index v4.1.1, GLM-5.2 (max): 53
  • Artificial Analysis Intelligence Index v4.1.1, Kimi K3 (max): 60
  • Artificial Analysis Intelligence Index v4.1.1, Kimi K2.7 Code: 43
  • Artificial Analysis median output speed (API hosts), GLM-5.2 (max): 69.2
  • SWE-bench (bash-only, mini-SWE-agent), Kimi K2.5 (high): 70.8
  • SWE-bench (bash-only, mini-SWE-agent), GLM 5 (high): 72.8
  • SWE-bench (bash-only, mini-SWE-agent), Kimi K2 Thinking: 63.4
  • SWE-bench (bash-only, mini-SWE-agent), Qwen3-Coder 480B/A35B Instruct: 55.4
  • Terminal-Bench 2.1 (tbench.ai, Claude Code agent), GLM-5.1: 58.65
Vendor-claimed (zero weight)
  • Terminal Bench 2.1 (Terminus-2), GLM-5.2 (vendor model card): 81.0
  • Terminal-Bench 2.1, Kimi K3 (vendor model card, Kimi Code harness): 88.3

Bake-off: skipped, no API key in env

Discount coding plans (resellers), aggregator · score 50 · LOW

observed $9485.17/mo, parity 0.0211, control 50, capability 91.9

Limits as published: rpm: 45; weekly_cap: 16500000
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • SWE-bench Verified (mini-SWE-agent, GLM 5 high): 72.8
  • SWE-bench Verified (mini-SWE-agent, Kimi K2.5 high): 70.8
  • SWE-bench Verified (mini-SWE-agent, DeepSeek V3.2 high): 70.0
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

Z.ai GLM Coding Plan, subscription · score 49 · LOW

observed $168.0/mo, parity 0.0914, control 5, capability 82.6

Limits as published: session_window: 5-hour rolling credit window: Lite 2,000 / Pro 12,000 / Max 28,000 credits; resets 5 hours; weekly_cap: Lite 10,000 / Pro 60,000 / Max 140,000 credits per week (Team: Standard 66,000 / Premium 1
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Artificial Analysis Intelligence Index (GLM-5.3 max) 60, rank 9/187: 60
  • Terminal-Bench 2.1 official leaderboard (GLM-5.1 via Claude Code; no GLM-5.3 entry): 58.7
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

Anthropic API PAYG, api · score 48 · LOW

observed $26187.48/mo, parity 0.0076, control 5, capability 100.0

Limits as published: rpm: 1000; tpm: 2000000; weekly_cap: none weekly; monthly spend caps by tier: Start $500, Build $1,000, Scale $200,000, Custom
spend-cap:yes · realtime-dash:yes · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Terminal-Bench 2.1 (tbench.ai), Claude Code + Fable 5 (xhigh), accuracy % (rank 1 of 17): 83.8
  • Terminal-Bench 2.1 (tbench.ai), Claude Code + Opus 4.8 (high), accuracy %: 78.9
  • Terminal-Bench 2.1 (tbench.ai), Claude Code + Sonnet 5 (high), accuracy %: 74.6
  • Artificial Analysis Intelligence Index v4.1.1, Claude Opus 5 (max), highest-scoring model on the index: 63.05
  • Artificial Analysis Intelligence Index v4.1.1, Claude Fable 5 (with fallback): 62.07
  • SWE-bench Verified (swebench.com), best Claude entry: live-SWE-agent + Claude 4.5 Opus medium (20251101), % r: 79.2
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

Moonshot API, api · score 45 · LOW

observed $15011.39/mo, parity 0.0133, control 0, capability 70

Limits as published: rpm: 10000; tpm: 5000000; concurrency: 1000
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Artificial Analysis Intelligence Index v4.1.1 - Kimi K3 (max), rank #1/110: 60
  • Artificial Analysis output speed (tokens/s) - Kimi K3 (max): 35.5
Vendor-claimed (zero weight)
  • Terminal-Bench 2.1 - Kimi K3 (vendor-run on H20 GPUs, HF model card): 88.3

Bake-off: skipped, no API key in env

DeepSeek API, api · score 42 · LOW

observed $2136.84/mo, parity 0.0936, control 0, capability 70

Limits as published: concurrency: 500
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Artificial Analysis Intelligence Index (DeepSeek V4 Pro 0813, max) 53, rank 5/110: 53
  • Artificial Analysis output speed tok/s (DeepSeek V4 Pro 0813, max): 66.3
Vendor-claimed (zero weight)
  • Terminal Bench 2.1 (DeepSeek-V4-Pro-0813, self-run via DeepSeek Harness): 87.9
  • DeepSWE (DeepSeek-V4-Pro-0813, self-run): 62.7
  • Terminal Bench 2.1 (deepseek-v4-flash, self-run): 82.7

Bake-off: skipped, account balance negative; completions would 402

Cloudflare Workers AI, serverless · score 39 · LOW

observed $1206.44/mo, parity 0.1658, control 0, capability 90.4

Limits as published: rpm: 20
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Artificial Analysis Intelligence Index v4.1.1 (GLM-5.3-Flash): 57
  • Artificial Analysis Intelligence Index rank (GLM-5.3-Flash): 3
  • Artificial Analysis output speed tokens/sec (GLM-5.3-Flash): 49.8
  • Artificial Analysis Intelligence Index v4.1.1 (GLM-5.2 max): 53
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, CLOUDFLARE_API_TOKEN lacks Workers AI permission; CLOUDFLARE_ACCOUNT_ID also absent from env

Z.ai GLM API, api · score 37 · LOW

observed $10539.08/mo, parity 0.019, control 0, capability 85.7

Limits as published: not published
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:no

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Artificial Analysis Intelligence Index (GLM-5.3 max) 60, rank 9/187: 60
  • SWE-bench Verified bash-only (GLM 5 high, mini-SWE-agent), older sibling, no GLM-5.3 entry: 72.8
  • Terminal-Bench 2.1 official leaderboard (GLM-5.1 max via Claude Code), older sibling, no GLM-5.3 entry: 58.7
Vendor-claimed (zero weight)
  • Terminal Bench 2.1 (GLM-5.2, as reported by DeepSeek on its V4 model card; not Z.ai and not an E3 evaluator): 81.0

Bake-off: skipped, no API key in env

Local hardware, self-host · score 35 · LOW

observed $None/mo, parity None, control 40, capability 51.6

Limits as published: concurrency: operator-set (local server)
spend-cap:yes · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:no

Integrations:

Changes, last 90 days:

Throttling incidents, 180 days:

Benchmarks, independent
  • Artificial Analysis Intelligence Index, Qwen3 Coder 30B A3B: 14
  • Artificial Analysis Intelligence Index, GLM-4.5-Air: 17
  • Artificial Analysis median output speed (API hosts), Qwen3 Coder 30B A3B: 91.5
  • SWE-bench Verified, OpenHands + Qwen3-Coder-30B-A3B-Instruct: 51.6
  • SWE-bench Verified, EntroPO + R2E + Qwen3-Coder-30B-A3B-Instruct: 60.4
Vendor-claimed (zero weight)
  • none

Bake-off: skipped, no API key in env

8 · Migration plan (winner: OpenRouter, with the hybrid as its enforcement layer)

  1. Day 1, spend cap first. Create an OpenRouter account; buy $50 of credits, no auto-top-up. Create one provisioning key per tool with a per-key credit limit: claude_code=$20, opencode=$15, experiments=$15. Total exposure: $50, hard.
  2. Stand up LiteLLM (MIT, self-hosted) as the router: default route deepseek/deepseek-v4-flash (off-peak aware), fallback anthropic/claude-opus-5 via OpenRouter; set proxy max_budget=200, budget_duration=30d.
  3. Point tools at the proxy: ANTHROPIC_BASE_URL=http://localhost:4000 for Claude Code (LiteLLM /v1/messages passthrough); OpenAI-compatible base URL for opencode.
  4. Days 1–7, parallel run. Keep Claude Max 20x active. Route 3 real projects through the proxy. Compare per task: pass rate, wall time, and router-metered USD vs the sub's flat $200. The S4 fixture in out/S4/fixture/ (hash 2d78…8cb0) is the controlled comparison set: 3 tasks, 2 attempts each.
  5. Decision gate, day 7: if proxy pass-rate ≥ sub pass-rate on the fixture AND projected 30-day proxy spend ≤ $200 for the workload actually routed, downgrade one sub tier ($200→$100), never all at once. Observed data says full replacement will NOT fit $200 (parity 0.0072–0.10); the realistic outcome is sub + capped overflow lane.
  6. Watch 2026-08-31 (Max boost lapse) and week-36 throttle counts in /usage; if delivered capacity drops, shift the routing split rather than buying a second sub.
  7. Rollback: unset ANTHROPIC_BASE_URL (tools revert to the sub), leave remaining OpenRouter credits parked, they do not expire on a schedule the audit found, and the per-key caps keep them inert.

9 · Assumptions & gaps (how missing data moved scores)

10 · Appendix — unverified / blocked sources (14)

USER_LINK resolved via mirror (x.com blocks direct fetch; retrieved through Twitter's syndication CDN, mirror:true): @zebassembly, 2026-08-27 — “i was so excited that local models like qwen 3.8 27B exist, and then I did the math on how much it costs on my power bill.... looks like i'll be using glm 5.3 flash on Cloudflare instead” Author carries a Cloudflare business label on the platform; the audit scored Cloudflare Workers AI on its own artifacts (B1): final 39, parity 0.1658, rpm 20.

urlstatusunmatched snippet
https://platform.claude.com/docs/en/api/rate-limits200You can monitor your rate limit usage on the Usage page of the Claude
https://www.tbench.ai/leaderboard/terminal-bench/2.1200"label":"GLM-5.1"},"reasoning_effort":"max"},"metrics":{"accuracy":58.
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813200"safetensors":{"parameters":{"BF16":2954820352,"I64":2327040,"F32":905
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731200"total":304180418494,"sharded":true
https://huggingface.co/moonshotai/Kimi-K2.7-Code200"createdAt":"2026-06-11T07:51:47.000Z"
https://developers.openai.com/codex/pricing200On ChatGPT plans, local messages and cloud chats share a five-hour win
https://developers.openai.com/codex/pricing200If you want to see your remaining limits during an active Codex CLI se
https://developers.openai.com/codex/pricing200All users may also run extra local chats using an API key, with usage
https://developers.cloudflare.com/workers-ai/platform/pricing/200USD per 1M tokens = (neurons per M tokens) x ($0.011 / 1000 neurons).
https://docs.runpod.io/pods/pricing200GPUs are dedicated to your Pod and cannot be displaced by other users.
https://support.claude.com/en/articles/9797557-usage-limit-best-practices200navigate to Settings > Usage to view progress bars showing how much of
https://support.claude.com/en/articles/12429409-manage-usage-credits-for-paid-claude-plans200Usage dashboard: View real-time consumption in Settings > Usage.
https://docs.z.ai/devpack/overview200All plans support **GLM-5.3**, GLM-5-Flash.
https://docs.z.ai/devpack/usage-policy200The platform dynamically adjusts these limits based on resource availa

Run fb2d5acf-a818-4a52-9c17-b4da13b7f15b · inputs_hash 6ac7f23a16c9727d… · S0–S7 per AI_SPEND_AUDIT v2 · every figure tagged observed/modelled · generated 2026-08-28.