Run 2026-08-28 · 15 candidates · 377 validated citations · workload observed from 30 days of local logs · every figure tagged observed or modelled · no external requests on this page.
Every link below returned HTTP 200 during validation on 2026-08-28 and its page body still contained the cited evidence snippet (L1). Unverified sources are in the appendix.
| # | candidate | what it proves | domain | fetched (UTC) | src |
|---|---|---|---|---|---|
| 1 | agg-openrouter | credit purchase fee | openrouter.ai | 2026-08-28T07:35 | link |
| 2 | agg-openrouter | inference markup | openrouter.ai | 2026-08-28T07:35 | link |
| 3 | agg-openrouter | byok fee | openrouter.ai | 2026-08-28T07:35 | link |
| 4 | agg-openrouter | byok allowance basis | openrouter.ai | 2026-08-28T07:35 | link |
| 5 | agg-openrouter | price usd per 1M | openrouter.ai | 2026-08-28T07:35 | link |
| 6 | agg-openrouter | context tokens/max output tokens | openrouter.ai | 2026-08-28T07:35 | link |
| 7 | agg-openrouter | cache write 1h note | openrouter.ai | 2026-08-28T07:35 | link |
| 8 | agg-openrouter | price usd per 1M | openrouter.ai | 2026-08-28T07:35 | link |
| 9 | agg-openrouter | input cache hit | openrouter.ai | 2026-08-28T07:35 | link |
| 10 | agg-openrouter | price usd per 1M | openrouter.ai | 2026-08-28T07:35 | link |
| 11 | agg-openrouter | input cache hit | openrouter.ai | 2026-08-28T07:35 | link |
| 12 | agg-openrouter | price disagreement models api | openrouter.ai | 2026-08-28T07:35 | link |
| 13 | agg-openrouter | price usd per 1M | openrouter.ai | 2026-08-28T07:35 | link |
| 14 | agg-openrouter | input cache hit | openrouter.ai | 2026-08-28T07:35 | link |
| 15 | agg-openrouter | max output tokens | openrouter.ai | 2026-08-28T07:35 | link |
| 16 | agg-openrouter | price disagreement models api | openrouter.ai | 2026-08-28T07:35 | link |
| 17 | agg-openrouter | rpm | openrouter.ai | 2026-08-28T07:35 | link |
| 18 | agg-openrouter | published in tokens or requests | openrouter.ai | 2026-08-28T07:35 | link |
| 19 | agg-openrouter | spend cap settable | openrouter.ai | 2026-08-28T07:35 | link |
| 20 | agg-openrouter | spend cap settable | openrouter.ai | 2026-08-28T07:35 | link |
| 21 | agg-openrouter | usage dashboard realtime | openrouter.ai | 2026-08-28T07:35 | link |
| 22 | agg-openrouter | exacto | openrouter.ai | 2026-08-28T07:35 | link |
| 23 | agg-openrouter | provider routing | openrouter.ai | 2026-08-28T07:35 | link |
| 24 | agg-openrouter | fallback | openrouter.ai | 2026-08-28T07:35 | link |
| 25 | agg-openrouter | incident | status.openrouter.ai | 2026-08-28T07:35 | link |
| 26 | agg-openrouter | SWE-bench Verified (mini-SWE-agent, GLM 5 high) | www.swebench.com | 2026-08-28T07:35 | link |
| 27 | agg-openrouter | SWE-bench Verified (mini-SWE-agent, Kimi K2 5 high) | www.swebench.com | 2026-08-28T07:35 | link |
| 28 | agg-openrouter | SWE-bench Verified (mini-SWE-agent, DeepSeek V3 2 high) | www.swebench.com | 2026-08-28T07:35 | link |
| 29 | agg-openrouter | SWE-bench Verified (live-SWE-agent, Claude 4 5 Opus medium) | www.swebench.com | 2026-08-28T07:35 | link |
| 30 | agg-openrouter | claude code | openrouter.ai | 2026-08-28T07:35 | link |
| 31 | agg-openrouter | codex cli | openrouter.ai | 2026-08-28T07:35 | link |
| 32 | agg-openrouter | opencode | openrouter.ai | 2026-08-28T07:35 | link |
| 33 | agg-openrouter | zdr | openrouter.ai | 2026-08-28T07:35 | link |
| 34 | hybrid-router | routing role | docs.litellm.ai | 2026-08-28T07:35 | link |
| 35 | hybrid-router | routing role | docs.litellm.ai | 2026-08-28T07:35 | link |
| 36 | hybrid-router | license | raw.githubusercontent.com | 2026-08-28T07:35 | link |
| 37 | hybrid-router | self hosted oss | github.com | 2026-08-28T07:35 | link |
| 38 | hybrid-router | budgets in free tier | docs.litellm.ai | 2026-08-28T07:35 | link |
| 39 | hybrid-router | max budget and budget duration | docs.litellm.ai | 2026-08-28T07:35 | link |
| 40 | hybrid-router | budget scopes key team user | docs.litellm.ai | 2026-08-28T07:35 | link |
| 41 | hybrid-router | rpm tpm per key | docs.litellm.ai | 2026-08-28T07:35 | link |
| 42 | hybrid-router | per model budget per key is enterprise | docs.litellm.ai | 2026-08-28T07:35 | link |
| 43 | hybrid-router | spend tracking | docs.litellm.ai | 2026-08-28T07:35 | link |
| 44 | hybrid-router | anthropic v1 messages endpoint | docs.litellm.ai | 2026-08-28T07:35 | link |
| 45 | hybrid-router | anthropic endpoint supports routing | docs.litellm.ai | 2026-08-28T07:35 | link |
| 46 | hybrid-router | credit purchase fee pct | openrouter.ai | 2026-08-28T07:35 | link |
| 47 | hybrid-router | crypto fee pct | openrouter.ai | 2026-08-28T07:35 | link |
| 48 | hybrid-router | inference markup pct | openrouter.ai | 2026-08-28T07:35 | link |
| 49 | hybrid-router | per key credit limit provisioning | openrouter.ai | 2026-08-28T07:35 | link |
| 50 | hybrid-router | provider routing fallback | openrouter.ai | 2026-08-28T07:35 | link |
| 51 | hybrid-router | usage dashboard realtime | docs.litellm.ai | 2026-08-28T07:35 | link |
| 52 | hybrid-router | env | docs.litellm.ai | 2026-08-28T07:35 | link |
| 53 | hybrid-router | official anthropic side | docs.anthropic.com | 2026-08-28T07:35 | link |
| 54 | hybrid-router | base url | developers.openai.com | 2026-08-28T07:35 | link |
| 55 | hybrid-router | env | www.kimi.com | 2026-08-28T07:35 | link |
| 56 | hybrid-router | base url | www.kimi.com | 2026-08-28T07:35 | link |
| 57 | hybrid-router | base url | opencode.ai | 2026-08-28T07:35 | link |
| 58 | hybrid-router | alt hosts status | status.openrouter.ai | 2026-08-28T07:35 | link |
| 59 | hybrid-router | data retention url | docs.litellm.ai | 2026-08-28T07:35 | link |
| 60 | hybrid-router | data retention auto deletion is enterprise | docs.litellm.ai | 2026-08-28T07:35 | link |
| 61 | sub-kimi-top | model id | www.kimi.ai | 2026-08-28T07:35 | link |
| 62 | sub-kimi-top | model mapping | www.kimi.ai | 2026-08-28T07:35 | link |
| 63 | sub-kimi-top | k3 access | www.kimi.com | 2026-08-28T07:35 | link |
| 64 | sub-kimi-top | speed | www.kimi.ai | 2026-08-28T07:35 | link |
| 65 | sub-kimi-top | credit multiplier | www.kimi.ai | 2026-08-28T07:35 | link |
| 66 | sub-kimi-top | plan gate | www.kimi.ai | 2026-08-28T07:35 | link |
| 67 | sub-kimi-top | fixed fees usd month | www.kimi.ai | 2026-08-28T07:35 | link |
| 68 | sub-kimi-top | annual price | www.kimi.ai | 2026-08-28T07:35 | link |
| 69 | sub-kimi-top | agent credits | www.kimi.ai | 2026-08-28T07:35 | link |
| 70 | sub-kimi-top | concurrency | www.kimi.ai | 2026-08-28T07:35 | link |
| 71 | sub-kimi-top | session window | www.kimi.ai | 2026-08-28T07:35 | link |
| 72 | sub-kimi-top | weekly cap | www.kimi.ai | 2026-08-28T07:35 | link |
| 73 | sub-kimi-top | spend cap settable | www.kimi.ai | 2026-08-28T07:35 | link |
| 74 | sub-kimi-top | usage dashboard realtime | www.kimi.ai | 2026-08-28T07:35 | link |
| 75 | sub-kimi-top | usage dashboard realtime | www.kimi.ai | 2026-08-28T07:35 | link |
| 76 | sub-kimi-top | shared pool with chat | www.kimi.ai | 2026-08-28T07:35 | link |
| 77 | sub-kimi-top | shared pool with chat | www.kimi.ai | 2026-08-28T07:35 | link |
| 78 | sub-kimi-top | extra usage caps | www.kimi.ai | 2026-08-28T07:35 | link |
| 79 | sub-kimi-top | k3 chat 1m | www.kimi.ai | 2026-08-28T07:35 | link |
| 80 | sub-codex-pro20 | included messages per 5h window | developers.openai.com | 2026-08-28T07:35 | link |
| 81 | sub-codex-pro20 | credit rate per 1M tokens | developers.openai.com | 2026-08-28T07:35 | link |
| 82 | sub-codex-pro20 | context tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 83 | sub-codex-pro20 | max output tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 84 | sub-codex-pro20 | included messages per 5h window | developers.openai.com | 2026-08-28T07:35 | link |
| 85 | sub-codex-pro20 | credit rate per 1M tokens | developers.openai.com | 2026-08-28T07:35 | link |
| 86 | sub-codex-pro20 | context tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 87 | sub-codex-pro20 | included messages per 5h window | developers.openai.com | 2026-08-28T07:35 | link |
| 88 | sub-codex-pro20 | credit rate per 1M tokens | developers.openai.com | 2026-08-28T07:35 | link |
| 89 | sub-codex-pro20 | availability | developers.openai.com | 2026-08-28T07:35 | link |
| 90 | sub-codex-pro20 | separate limit | developers.openai.com | 2026-08-28T07:35 | link |
| 91 | sub-codex-pro20 | fixed fees usd month | developers.openai.com | 2026-08-28T07:35 | link |
| 92 | sub-codex-pro20 | fixed fees usd month | help.openai.com | 2026-08-28T07:35 | link mirror |
| 93 | sub-codex-pro20 | published in tokens or requests | developers.openai.com | 2026-08-28T07:35 | link |
| 94 | sub-codex-pro20 | published in tokens or requests | developers.openai.com | 2026-08-28T07:35 | link |
| 95 | sub-codex-pro20 | spend cap settable | help.openai.com | 2026-08-28T07:35 | link |
| 96 | sub-codex-pro20 | usage dashboard realtime | developers.openai.com | 2026-08-28T07:35 | link |
| 97 | sub-codex-pro20 | shared pool with chat | help.openai.com | 2026-08-28T07:35 | link |
| 98 | sub-codex-pro20 | shared pool with chat | help.openai.com | 2026-08-28T07:35 | link |
| 99 | sub-codex-pro20 | credits overflow exists | developers.openai.com | 2026-08-28T07:35 | link |
| 100 | sub-codex-pro20 | credits validity | help.openai.com | 2026-08-28T07:35 | link |
| 101 | sub-codex-pro20 | credit burn rate | developers.openai.com | 2026-08-28T07:35 | link |
| 102 | sub-claude-max20 | context tokens | claude.com | 2026-08-28T07:35 | link |
| 103 | sub-claude-max20 | context tokens (conflicting primary source) | support.claude.com | 2026-08-28T07:35 | link |
| 104 | sub-claude-max20 | models included | claude.com | 2026-08-28T07:35 | link |
| 105 | sub-claude-max20 | fable5 weekly cap share | support.claude.com | 2026-08-28T07:35 | link |
| 106 | sub-claude-max20 | fixed fees usd month | support.claude.com | 2026-08-28T07:35 | link |
| 107 | sub-claude-max20 | session window | support.claude.com | 2026-08-28T07:35 | link |
| 108 | sub-claude-max20 | weekly cap | support.claude.com | 2026-08-28T07:35 | link |
| 109 | sub-claude-max20 | weekly cap (Opus separate bucket) | support.claude.com | 2026-08-28T07:35 | link |
| 110 | sub-claude-max20 | published in tokens or requests | claude.com | 2026-08-28T07:35 | link |
| 111 | sub-claude-max20 | shared pool with chat | claude.com | 2026-08-28T07:35 | link |
| 112 | sub-claude-max20 | shared pool with chat (Claude Code article) | support.claude.com | 2026-08-28T07:35 | link |
| 113 | sub-claude-max20 | usage multiplier vs pro | claude.com | 2026-08-28T07:35 | link |
| 114 | api-openai | price usd per 1M | platform.openai.com | 2026-08-28T07:35 | link |
| 115 | api-openai | context tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 116 | api-openai | max output tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 117 | api-openai | positioning | platform.openai.com | 2026-08-28T07:35 | link |
| 118 | api-openai | price usd per 1M | platform.openai.com | 2026-08-28T07:35 | link |
| 119 | api-openai | price usd per 1M | platform.openai.com | 2026-08-28T07:35 | link |
| 120 | api-openai | long context surcharge | platform.openai.com | 2026-08-28T07:35 | link |
| 121 | api-openai | cache write | platform.openai.com | 2026-08-28T07:35 | link |
| 122 | api-openai | promo lapse | platform.openai.com | 2026-08-28T07:35 | link |
| 123 | api-openai | context tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 124 | api-openai | max output tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 125 | api-openai | price usd per 1M | platform.openai.com | 2026-08-28T07:35 | link |
| 126 | api-openai | context tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 127 | api-openai | max output tokens | platform.openai.com | 2026-08-28T07:35 | link |
| 128 | api-openai | price usd per 1M | platform.openai.com | 2026-08-28T07:35 | link |
| 129 | api-openai | rpm | platform.openai.com | 2026-08-28T07:35 | link |
| 130 | api-openai | rpm | platform.openai.com | 2026-08-28T07:35 | link |
| 131 | api-openai | tpm | platform.openai.com | 2026-08-28T07:35 | link |
| 132 | api-openai | tpm | platform.openai.com | 2026-08-28T07:35 | link |
| 133 | api-openai | monthly usage limit by tier | platform.openai.com | 2026-08-28T07:35 | link |
| 134 | api-openai | monthly usage limit by tier | platform.openai.com | 2026-08-28T07:35 | link |
| 135 | api-openai | published in tokens or requests | platform.openai.com | 2026-08-28T07:35 | link |
| 136 | api-openai | spend cap settable | platform.openai.com | 2026-08-28T07:35 | link |
| 137 | api-openai | spend cap settable | platform.openai.com | 2026-08-28T07:35 | link |
| 138 | api-openai | usage dashboard realtime | platform.openai.com | 2026-08-28T07:35 | link |
| 139 | self-rented-gpu | license | huggingface.co | 2026-08-28T07:35 | link |
| 140 | self-rented-gpu | params total b/params active b | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 141 | self-rented-gpu | params total b (disagreeing source, E4) | huggingface.co | 2026-08-28T07:35 | link |
| 142 | self-rented-gpu | context tokens | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 143 | self-rented-gpu | min serving config | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 144 | self-rented-gpu | min serving config 1M | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 145 | self-rented-gpu | published throughput | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 146 | self-rented-gpu | license | huggingface.co | 2026-08-28T07:35 | link |
| 147 | self-rented-gpu | params total b/params active b | huggingface.co | 2026-08-28T07:35 | link |
| 148 | self-rented-gpu | context tokens | huggingface.co | 2026-08-28T07:35 | link |
| 149 | self-rented-gpu | min serving config | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 150 | self-rented-gpu | min serving config amd | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 151 | self-rented-gpu | license | huggingface.co | 2026-08-28T07:35 | link |
| 152 | self-rented-gpu | params total b/params active b | huggingface.co | 2026-08-28T07:35 | link |
| 153 | self-rented-gpu | context tokens | huggingface.co | 2026-08-28T07:35 | link |
| 154 | self-rented-gpu | min serving config | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 155 | self-rented-gpu | license | huggingface.co | 2026-08-28T07:35 | link |
| 156 | self-rented-gpu | params | huggingface.co | 2026-08-28T07:35 | link |
| 157 | self-rented-gpu | context tokens | huggingface.co | 2026-08-28T07:35 | link |
| 158 | self-rented-gpu | serving recipe | huggingface.co | 2026-08-28T07:35 | link |
| 159 | self-rented-gpu | runpod H200 | www.runpod.io | 2026-08-28T07:35 | link |
| 160 | self-rented-gpu | runpod H100 SXM | www.runpod.io | 2026-08-28T07:35 | link |
| 161 | self-rented-gpu | runpod H100 PCIe | www.runpod.io | 2026-08-28T07:35 | link |
| 162 | self-rented-gpu | runpod B200 | www.runpod.io | 2026-08-28T07:35 | link |
| 163 | self-rented-gpu | runpod B300 | www.runpod.io | 2026-08-28T07:35 | link |
| 164 | self-rented-gpu | runpod serverless H200 | www.runpod.io | 2026-08-28T07:35 | link |
| 165 | self-rented-gpu | runpod serverless H100 | www.runpod.io | 2026-08-28T07:35 | link |
| 166 | self-rented-gpu | runpod serverless B200 | www.runpod.io | 2026-08-28T07:35 | link |
| 167 | self-rented-gpu | lambda H100 8x | lambda.ai | 2026-08-28T07:35 | link |
| 168 | self-rented-gpu | lambda B200 8x | lambda.ai | 2026-08-28T07:35 | link |
| 169 | self-rented-gpu | lambda H100 1x | lambda.ai | 2026-08-28T07:35 | link |
| 170 | self-rented-gpu | vast H100 | vast.ai | 2026-08-28T07:35 | link |
| 171 | self-rented-gpu | vast H200 | vast.ai | 2026-08-28T07:35 | link |
| 172 | self-rented-gpu | vast rate basis | vast.ai | 2026-08-28T07:35 | link |
| 173 | self-rented-gpu | modal H200 | modal.com | 2026-08-28T07:35 | link |
| 174 | self-rented-gpu | modal H100 | modal.com | 2026-08-28T07:35 | link |
| 175 | self-rented-gpu | modal B200 | modal.com | 2026-08-28T07:35 | link |
| 176 | self-rented-gpu | runpod serverless | docs.runpod.io | 2026-08-28T07:35 | link |
| 177 | self-rented-gpu | runpod serverless idle | docs.runpod.io | 2026-08-28T07:35 | link |
| 178 | self-rented-gpu | modal | modal.com | 2026-08-28T07:35 | link |
| 179 | self-rented-gpu | concurrency | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 180 | self-rented-gpu | spend cap settable | docs.runpod.io | 2026-08-28T07:35 | link |
| 181 | self-rented-gpu | spend cap settable | docs.runpod.io | 2026-08-28T07:35 | link |
| 182 | self-rented-gpu | published in tokens or requests | docs.runpod.io | 2026-08-28T07:35 | link |
| 183 | self-rented-gpu | Artificial Analysis Intelligence Index v4 1 1, GLM-5 2 (max | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 184 | self-rented-gpu | Artificial Analysis Intelligence Index v4 1 1, Kimi K3 (max | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 185 | self-rented-gpu | Artificial Analysis Intelligence Index v4 1 1, Kimi K2 7 Co | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 186 | self-rented-gpu | Artificial Analysis median output speed (API hosts), GLM-5 | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 187 | self-rented-gpu | SWE-bench (bash-only, mini-SWE-agent), Kimi K2 5 (high) | www.swebench.com | 2026-08-28T07:35 | link |
| 188 | self-rented-gpu | SWE-bench (bash-only, mini-SWE-agent), GLM 5 (high) | www.swebench.com | 2026-08-28T07:35 | link |
| 189 | self-rented-gpu | SWE-bench (bash-only, mini-SWE-agent), Kimi K2 Thinking | www.swebench.com | 2026-08-28T07:35 | link |
| 190 | self-rented-gpu | SWE-bench (bash-only, mini-SWE-agent), Qwen3-Coder 480B/A35 | www.swebench.com | 2026-08-28T07:35 | link |
| 191 | self-rented-gpu | Terminal Bench 2 1 (Terminus-2), GLM-5 2 (vendor model card | huggingface.co | 2026-08-28T07:35 | link |
| 192 | self-rented-gpu | Terminal-Bench 2 1, Kimi K3 (vendor model card, Kimi Code h | huggingface.co | 2026-08-28T07:35 | link |
| 193 | agg-coding-plans | tier price starter | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 194 | agg-coding-plans | tier quota starter | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 195 | agg-coding-plans | tier price ultra | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 196 | agg-coding-plans | tier quota ultra | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 197 | agg-coding-plans | quota structure | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 198 | agg-coding-plans | weekly reset | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 199 | agg-coding-plans | points formula | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 200 | agg-coding-plans | multiplier glm5 | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 201 | agg-coding-plans | multiplier kimi k2 6 | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 202 | agg-coding-plans | savings glm5 | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 203 | agg-coding-plans | overage | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 204 | agg-coding-plans | payg pack 99usd 280M | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 205 | agg-coding-plans | tier price lite | novita.ai | 2026-08-28T07:35 | link |
| 206 | agg-coding-plans | tier quota lite tokens | novita.ai | 2026-08-28T07:35 | link |
| 207 | agg-coding-plans | tier rpm lite | novita.ai | 2026-08-28T07:35 | link |
| 208 | agg-coding-plans | tier price pro | novita.ai | 2026-08-28T07:35 | link |
| 209 | agg-coding-plans | tier quota pro tokens | novita.ai | 2026-08-28T07:35 | link |
| 210 | agg-coding-plans | tier rpm pro | novita.ai | 2026-08-28T07:35 | link |
| 211 | agg-coding-plans | tier price max | novita.ai | 2026-08-28T07:35 | link |
| 212 | agg-coding-plans | tier quota max tokens | novita.ai | 2026-08-28T07:35 | link |
| 213 | agg-coding-plans | tier rpm max | novita.ai | 2026-08-28T07:35 | link |
| 214 | agg-coding-plans | models served | novita.ai | 2026-08-28T07:35 | link |
| 215 | agg-coding-plans | price usd per 1M | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 216 | agg-coding-plans | context tokens | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 217 | agg-coding-plans | list price origin | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 218 | agg-coding-plans | price usd per 1M | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 219 | agg-coding-plans | context tokens | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 220 | agg-coding-plans | input and cache | novita.ai | 2026-08-28T07:35 | link |
| 221 | agg-coding-plans | output | novita.ai | 2026-08-28T07:35 | link |
| 222 | agg-coding-plans | context tokens | novita.ai | 2026-08-28T07:35 | link |
| 223 | agg-coding-plans | input | novita.ai | 2026-08-28T07:35 | link |
| 224 | agg-coding-plans | input cache hit | novita.ai | 2026-08-28T07:35 | link |
| 225 | agg-coding-plans | output | novita.ai | 2026-08-28T07:35 | link |
| 226 | agg-coding-plans | context tokens | novita.ai | 2026-08-28T07:35 | link |
| 227 | agg-coding-plans | input | novita.ai | 2026-08-28T07:35 | link |
| 228 | agg-coding-plans | output | novita.ai | 2026-08-28T07:35 | link |
| 229 | agg-coding-plans | SWE-bench Verified (mini-SWE-agent, GLM 5 high) | www.swebench.com | 2026-08-28T07:35 | link |
| 230 | agg-coding-plans | SWE-bench Verified (mini-SWE-agent, Kimi K2 5 high) | www.swebench.com | 2026-08-28T07:35 | link |
| 231 | agg-coding-plans | SWE-bench Verified (mini-SWE-agent, DeepSeek V3 2 high) | www.swebench.com | 2026-08-28T07:35 | link |
| 232 | agg-coding-plans | claude code | www.atlascloud.ai | 2026-08-28T07:35 | link |
| 233 | agg-coding-plans | status page | status.novita.ai | 2026-08-28T07:35 | link |
| 234 | agg-coding-plans | data retention | novita.ai | 2026-08-28T07:35 | link |
| 235 | sub-zai-coding | routing of older model ids | docs.z.ai | 2026-08-28T07:35 | link |
| 236 | sub-zai-coding | credit multipliers (per 10,000 tokens): GLM-5 3 input 6 9 / | docs.z.ai | 2026-08-28T07:35 | link |
| 237 | sub-zai-coding | peak windows/peak multiplier | docs.z.ai | 2026-08-28T07:35 | link |
| 238 | sub-zai-coding | off-peak credit discount | docs.z.ai | 2026-08-28T07:35 | link |
| 239 | sub-zai-coding | credit multipliers: Flash input 2 3 / cached 0 56 / output 8 | docs.z.ai | 2026-08-28T07:35 | link |
| 240 | sub-zai-coding | session window (5-hour credits table) | docs.z.ai | 2026-08-28T07:35 | link |
| 241 | sub-zai-coding | session window reset rule | docs.z.ai | 2026-08-28T07:35 | link |
| 242 | sub-zai-coding | weekly cap reset rule | docs.z.ai | 2026-08-28T07:35 | link |
| 243 | sub-zai-coding | concurrency off-peak boost | docs.z.ai | 2026-08-28T07:35 | link |
| 244 | sub-zai-coding | spend cap settable false (fixed-fee subscription; hard credi | docs.z.ai | 2026-08-28T07:35 | link |
| 245 | sub-zai-coding | shared pool with chat false (plan usable only in supported c | docs.z.ai | 2026-08-28T07:35 | link |
| 246 | sub-zai-coding | usage visibility | docs.z.ai | 2026-08-28T07:35 | link |
| 247 | sub-zai-coding | Lite price | z.ai | 2026-08-28T07:35 | link |
| 248 | sub-zai-coding | Pro price | z.ai | 2026-08-28T07:35 | link |
| 249 | sub-zai-coding | Max price | z.ai | 2026-08-28T07:35 | link |
| 250 | sub-zai-coding | billing-cycle discounts | z.ai | 2026-08-28T07:35 | link |
| 251 | sub-zai-coding | team seats | z.ai | 2026-08-28T07:35 | link |
| 252 | sub-zai-coding | monthly list price floor | docs.z.ai | 2026-08-28T07:35 | link |
| 253 | sub-zai-coding | referral first-order example (hypothetical) | docs.z.ai | 2026-08-28T07:35 | link |
| 254 | api-anthropic | price usd per 1M (input miss, cache write 5m, cache write 1h | platform.claude.com | 2026-08-28T07:35 | link |
| 255 | api-anthropic | context tokens | platform.claude.com | 2026-08-28T07:35 | link |
| 256 | api-anthropic | max output tokens | platform.claude.com | 2026-08-28T07:35 | link |
| 257 | api-anthropic | model id | platform.claude.com | 2026-08-28T07:35 | link |
| 258 | api-anthropic | price usd per 1M (input miss, cache write 5m, cache write 1h | platform.claude.com | 2026-08-28T07:35 | link |
| 259 | api-anthropic | price now standard | platform.claude.com | 2026-08-28T07:35 | link |
| 260 | api-anthropic | price usd per 1M (input miss, cache write 5m, cache write 1h | platform.claude.com | 2026-08-28T07:35 | link |
| 261 | api-anthropic | rpm/tpm (Claude Opus 5, Start tier; ITPM 2,000,000, OTPM 400 | platform.claude.com | 2026-08-28T07:35 | link |
| 262 | api-anthropic | rpm/tpm higher tiers (Opus 5: Build 5,000/5,000,000/1,000,00 | platform.claude.com | 2026-08-28T07:35 | link |
| 263 | api-anthropic | published in tokens or requests | platform.claude.com | 2026-08-28T07:35 | link |
| 264 | api-anthropic | monthly spend caps by tier | platform.claude.com | 2026-08-28T07:35 | link |
| 265 | api-anthropic | spend cap settable | platform.claude.com | 2026-08-28T07:35 | link |
| 266 | api-anthropic | spend cap settable (workspace) | platform.claude.com | 2026-08-28T07:35 | link |
| 267 | api-anthropic | shared pool with chat | support.claude.com | 2026-08-28T07:35 | link |
| 268 | api-anthropic | cache read excluded from itpm | platform.claude.com | 2026-08-28T07:35 | link |
| 269 | api-moonshot | price usd per 1M | platform.kimi.ai | 2026-08-28T07:35 | link |
| 270 | api-moonshot | context tokens | platform.kimi.ai | 2026-08-28T07:35 | link |
| 271 | api-moonshot | open weights license | huggingface.co | 2026-08-28T07:35 | link |
| 272 | api-moonshot | params | huggingface.co | 2026-08-28T07:35 | link |
| 273 | api-moonshot | min serving config | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 274 | api-moonshot | min serving config rocm | recipes.vllm.ai | 2026-08-28T07:35 | link |
| 275 | api-moonshot | quantization | huggingface.co | 2026-08-28T07:35 | link |
| 276 | api-moonshot | price usd per 1M | platform.kimi.ai | 2026-08-28T07:35 | link |
| 277 | api-moonshot | open weights license | huggingface.co | 2026-08-28T07:35 | link |
| 278 | api-moonshot | params | huggingface.co | 2026-08-28T07:35 | link |
| 279 | api-moonshot | min serving config | huggingface.co | 2026-08-28T07:35 | link |
| 280 | api-moonshot | price usd per 1M | platform.kimi.ai | 2026-08-28T07:35 | link |
| 281 | api-moonshot | speed | platform.kimi.ai | 2026-08-28T07:35 | link |
| 282 | api-moonshot | price usd per 1M | platform.kimi.ai | 2026-08-28T07:35 | link |
| 283 | api-moonshot | hf repo | huggingface.co | 2026-08-28T07:35 | link |
| 284 | api-moonshot | price usd per 1M | platform.kimi.ai | 2026-08-28T07:35 | link |
| 285 | api-moonshot | hf repo | huggingface.co | 2026-08-28T07:35 | link |
| 286 | api-moonshot | rpm | platform.kimi.ai | 2026-08-28T07:35 | link |
| 287 | api-moonshot | tpm | platform.kimi.ai | 2026-08-28T07:35 | link |
| 288 | api-moonshot | concurrency | platform.kimi.ai | 2026-08-28T07:35 | link |
| 289 | api-moonshot | entry recharge | platform.kimi.ai | 2026-08-28T07:35 | link |
| 290 | api-moonshot | published in tokens or requests | platform.kimi.ai | 2026-08-28T07:35 | link |
| 291 | api-moonshot | spend cap settable | platform.kimi.ai | 2026-08-28T07:35 | link |
| 292 | api-moonshot | usage dashboard realtime | platform.kimi.ai | 2026-08-28T07:35 | link |
| 293 | api-moonshot | shared pool with chat | platform.kimi.ai | 2026-08-28T07:35 | link |
| 294 | api-moonshot | cache hit condition | platform.kimi.ai | 2026-08-28T07:35 | link |
| 295 | api-deepseek | model id/model version | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 296 | api-deepseek | context tokens/max output tokens | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 297 | api-deepseek | input miss (off-peak/peak; cols flash,pro,vision) | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 298 | api-deepseek | input cache hit | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 299 | api-deepseek | output | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 300 | api-deepseek | peak windows utc/peak multiplier | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 301 | api-deepseek | peak pricing effective date | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 302 | api-deepseek | open weights license | huggingface.co | 2026-08-28T07:35 | link |
| 303 | api-deepseek | params total b/params active b (independent) | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 304 | api-deepseek | min serving config recipe url (vLLM/SGLang serve commands on | huggingface.co | 2026-08-28T07:35 | link |
| 305 | api-deepseek | open weights license | huggingface.co | 2026-08-28T07:35 | link |
| 306 | api-deepseek | concurrency (500 = v4-pro; v4-flash and vision = 2500; accou | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 307 | api-deepseek | concurrency exceeded behavior | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 308 | api-deepseek | published in tokens or requests (limits published as concurr | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 309 | api-deepseek | balance model (prepaid: top-up balance + granted balance) | api-docs.deepseek.com | 2026-08-28T07:35 | link |
| 310 | cf-workers-ai | price usd per 1M input miss/output/input cache hit | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 311 | cf-workers-ai | neuron unit price | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 312 | cf-workers-ai | neurons per token | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 313 | cf-workers-ai | context tokens | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 314 | cf-workers-ai | unit pricing model page | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 315 | cf-workers-ai | paid access required | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 316 | cf-workers-ai | params total b/params active b | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 317 | cf-workers-ai | hf repo | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 318 | cf-workers-ai | price usd per 1M input miss/output/input cache hit | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 319 | cf-workers-ai | neurons per token | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 320 | cf-workers-ai | derived price formula | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 321 | cf-workers-ai | context tokens | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 322 | cf-workers-ai | context tokens native comparison | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 323 | cf-workers-ai | launch date | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 324 | cf-workers-ai | price usd per 1M input miss/output/input cache hit | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 325 | cf-workers-ai | neurons per token | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 326 | cf-workers-ai | derived price formula | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 327 | cf-workers-ai | context tokens | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 328 | cf-workers-ai | params total b | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 329 | cf-workers-ai | hf repo | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 330 | cf-workers-ai | rpm | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 331 | cf-workers-ai | rpm default text generation | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 332 | cf-workers-ai | rpm note glm53flash | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 333 | cf-workers-ai | rpm elevated via prepaid | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 334 | cf-workers-ai | free daily neurons | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 335 | cf-workers-ai | daily reset and hard fail | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 336 | cf-workers-ai | frontier models excluded from free | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 337 | cf-workers-ai | fixed fees usd month | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 338 | cf-workers-ai | published in tokens or requests | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 339 | cf-workers-ai | spend cap settable | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 340 | cf-workers-ai | usage dashboard realtime | developers.cloudflare.com | 2026-08-28T07:35 | link |
| 341 | api-zai | input miss/input cache hit/output | docs.z.ai | 2026-08-28T07:35 | link |
| 342 | api-zai | context tokens/max output tokens | docs.z.ai | 2026-08-28T07:35 | link |
| 343 | api-zai | reasoning always on | docs.z.ai | 2026-08-28T07:35 | link |
| 344 | api-zai | open weights: none for GLM-5 3 (base is GLM-5 2) | docs.z.ai | 2026-08-28T07:35 | link |
| 345 | api-zai | proprietary label (independent) | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 346 | api-zai | prices (current 50% promo; list prices struck through: input | docs.z.ai | 2026-08-28T07:35 | link |
| 347 | api-zai | input cache hit promo cell | docs.z.ai | 2026-08-28T07:35 | link |
| 348 | api-zai | context tokens/max output (text params = GLM-5 3) | docs.z.ai | 2026-08-28T07:35 | link |
| 349 | api-zai | params total b/params active b | huggingface.co | 2026-08-28T07:35 | link |
| 350 | api-zai | open weights license | huggingface.co | 2026-08-28T07:35 | link |
| 351 | api-zai | min serving config recipe url | huggingface.co | 2026-08-28T07:35 | link |
| 352 | api-zai | concurrency/rpm/tpm not public; console-gated | docs.z.ai | 2026-08-28T07:35 | link |
| 353 | api-zai | usage dashboard realtime (false: billing shows previous day) | docs.z.ai | 2026-08-28T07:35 | link |
| 354 | api-zai | cache mechanism beta; hit ~1/5 of price | docs.z.ai | 2026-08-28T07:35 | link |
| 355 | api-zai | balance model (prepaid recharge) | docs.z.ai | 2026-08-28T07:35 | link |
| 356 | self-local | license | huggingface.co | 2026-08-28T07:35 | link |
| 357 | self-local | params | huggingface.co | 2026-08-28T07:35 | link |
| 358 | self-local | quantized footprint gb Q4 K S | huggingface.co | 2026-08-28T07:35 | link |
| 359 | self-local | quantized footprint gb UD-Q2 K XL | huggingface.co | 2026-08-28T07:35 | link |
| 360 | self-local | quantized footprint gb UD-IQ1 S | huggingface.co | 2026-08-28T07:35 | link |
| 361 | self-local | gguf repo size | huggingface.co | 2026-08-28T07:35 | link |
| 362 | self-local | license | huggingface.co | 2026-08-28T07:35 | link |
| 363 | self-local | params | huggingface.co | 2026-08-28T07:35 | link |
| 364 | self-local | context tokens | huggingface.co | 2026-08-28T07:35 | link |
| 365 | self-local | local runtimes | huggingface.co | 2026-08-28T07:35 | link |
| 366 | self-local | quantized footprint gb 4bit | huggingface.co | 2026-08-28T07:35 | link |
| 367 | self-local | apple silicon path | huggingface.co | 2026-08-28T07:35 | link |
| 368 | self-local | representative config price usd | www.apple.com | 2026-08-28T07:35 | link |
| 369 | self-local | representative config name | www.apple.com | 2026-08-28T07:35 | link |
| 370 | self-local | base config price usd | www.apple.com | 2026-08-28T07:35 | link |
| 371 | self-local | base config name | www.apple.com | 2026-08-28T07:35 | link |
| 372 | self-local | availability | www.apple.com | 2026-08-28T07:35 | link |
| 373 | self-local | Artificial Analysis Intelligence Index, Qwen3 Coder 30B A3B | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 374 | self-local | Artificial Analysis Intelligence Index, GLM-4 5-Air | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 375 | self-local | Artificial Analysis median output speed (API hosts), Qwen3 | artificialanalysis.ai | 2026-08-28T07:35 | link |
| 376 | self-local | SWE-bench Verified, OpenHands + Qwen3-Coder-30B-A3B-Instruc | www.swebench.com | 2026-08-28T07:35 | link |
| 377 | self-local | SWE-bench Verified, EntroPO + R2E + Qwen3-Coder-30B-A3B-Ins | www.swebench.com | 2026-08-28T07:35 | link |
Winner: OpenRouter, 60/100.
It is the only candidate that pairs a frontier coding model (Opus 5; SWE-bench Verified 79.20 via an independent leaderboard entry) with every cost-control property this audit measured: prepaid credits, per-key spend limits, a real-time activity dashboard, pass-through per-token pricing plus a 5.5% credit fee, and zero recorded price or limit changes in the last 90 days.
At the observed workload its modelled cost is $27,628/month, parity 0.0072, meaning $200 buys 0.7% of the volume consumed today. The parity hard-cap (≤60) encodes that: no scored option at $200/month reproduces the observed 33.4B-token month.
The incumbent Claude Max 20x (56/100) is the only parity-1.0 row, it demonstrably served that volume for $200, at the cost of the lowest cost-control score in the top half (quotas not published in tokens, pool shared with chat, 14 throttle events in week 32).
Runner-up: hybrid-router, 60/100 (LiteLLM proxy, 95/5 token split: DeepSeek V4 Flash default, Opus 5 fallback, hard monthly budget at the proxy). It is the better pick than OpenRouter when a routing policy is enforceable: the same money then covers parity 0.1013 (14× more) and every dollar is capped by LiteLLM max_budget before it is spent.
The single risk most likely to change this answer: the Claude Max +50% weekly boost lapses on 2026-08-31. If delivered Max capacity drops in week 36, the parity-1.0 baseline measured here no longer exists, and the hybrid becomes the only path that keeps both capacity and control.
† B3 tie-break: raw scores within 3 points of the top (61 / 60 / 60) are ordered by the Cost-control dimension, then Reliability. Kimi Vivace holds the highest raw score (61) but control 65 vs 100/100 for the two winners; its new signups have also been paused since 2026-07-20.
| # | candidate | score | cost @ observed | parity | confidence |
|---|---|---|---|---|---|
| 1 | OpenRouter† agg-openrouter |
60 cap 60: parity 0.0072 < 1.0 | $27,628/mo | 0.0072 | LOW |
| 2 | Router: low-cost open-weight default + frontier fallback, hard monthly cap† hybrid-router |
60 cap 60: parity 0.1013 < 1.0 | $1,973/mo | 0.1013 | LOW |
| 3 | Kimi membership (top tier)† sub-kimi-top |
61 | $199/mo | n/a | LOW |
| 4 | ChatGPT Pro / Codex sub-codex-pro20 |
57 | $200/mo | n/a | LOW |
| 5 | Claude Max 20x sub-claude-max20 |
56 | $200/mo | 1.0000 | LOW |
| 6 | OpenAI API PAYG api-openai |
55 | $20,950/mo | 0.0095 | LOW |
| 7 | Self-host open weights on rented GPUs self-rented-gpu |
53 | $5,598/mo | 0.0357 | LOW |
| 8 | Discount coding plans (resellers) agg-coding-plans |
50 | $9,485/mo | 0.0211 | LOW |
| 9 | Z.ai GLM Coding Plan sub-zai-coding |
49 | $168/mo | 0.0914 | LOW |
| 10 | Anthropic API PAYG api-anthropic |
48 | $26,187/mo | 0.0076 | LOW |
| 11 | Moonshot API api-moonshot |
45 | $15,011/mo | 0.0133 | LOW |
| 12 | DeepSeek API api-deepseek |
42 | $2,137/mo | 0.0936 | LOW |
| 13 | Cloudflare Workers AI cf-workers-ai |
39 | $1,206/mo | 0.1658 | LOW |
| 14 | Z.ai GLM API api-zai |
37 | $10,539/mo | 0.0190 | LOW |
| 15 | Local hardware self-local |
35 | null/mo | n/a | LOW |
First run, no out/prev/ exists.
| candidate | cost (25%) | control (20%) | capability (25%) | throughput (10%) | tools (10%) | reliability (10%) | final |
|---|---|---|---|---|---|---|---|
| agg-openrouter | 0.0 | 100 | 100.0 | 50 | 75 | 50 | 60 |
| hybrid-router | 0.0 | 100 | 98.1 | 100 | 100 | 25 | 60 |
| sub-kimi-top | 40.3 | 65 | 70 | 100 | 75 | 25 | 61 |
| sub-codex-pro20 | 40.0 | 50 | 99.2 | 50 | 50 | 25 | 57 |
| sub-claude-max20 | 40.0 | 40 | 100.0 | 50 | 35 | 50 | 56 |
| api-openai | 0.0 | 55 | 95.9 | 100 | 50 | 50 | 55 |
| self-rented-gpu | 0.0 | 40 | 85.7 | 100 | 85 | 50 | 53 |
| agg-coding-plans | 0.0 | 50 | 91.9 | 50 | 75 | 50 | 50 |
| sub-zai-coding | 49.6 | 5 | 82.6 | 50 | 50 | 50 | 49 |
| api-anthropic | 0.0 | 5 | 100.0 | 100 | 50 | 75 | 48 |
| api-moonshot | 0.0 | 0 | 70 | 100 | 100 | 75 | 45 |
| api-deepseek | 0.0 | 0 | 70 | 100 | 75 | 75 | 42 |
| cf-workers-ai | 0.0 | 0 | 90.4 | 30 | 55 | 75 | 39 |
| api-zai | 0.0 | 0 | 85.7 | 50 | 60 | 50 | 37 |
| self-local | 0 | 40 | 51.6 | 30 | 85 | 25 | 35 |
Hard caps fired:
Capability uses independent benchmarks only (SWE-bench Verified, Terminal-Bench 2.1, Artificial Analysis Intelligence Index), normalised to best-in-set (83.8 TB / 63.05 AA / 79.2 SWE); fewer than 2 independent entries caps the dimension at 70. The S4 bake-off was skipped (no funded key), so no 50/50 blend was applied.
Observed 30-day mix (Claude Code logs, 2026-07-29 → 2026-08-28): input-miss 250M, cache-write 935M, cache-read 32.1B, output 122M tokens; cache-hit rate 96.4%. Formula: cost = in_miss·p_in + cache_read·p_hit + cache_write·(p_cw or p_in) + out·p_out, × peak adjustment × (1+fee) + fixed. DeepSeek peak (Mon–Fri 01–04 & 06–10 UTC = 09–12 & 14–18 Kuala Lumpur) carries 2× prices; 23.6% of observed tokens fall inside it. The off-peak-shift column halves that fraction (batch work moved off-peak).
| candidate | model / plan | observed | LOW | MID | HIGH | off-peak shift | parity | sub break-even |
|---|---|---|---|---|---|---|---|---|
| api-anthropic | Anthropic API - Opus 5 | $26,187 | $192 | $652 | $2,062 | - | 0.0076 | - |
| api-anthropic | Anthropic API - Sonnet 5 | $10,475 | $77 | $261 | $825 | - | 0.0191 | - |
| api-anthropic | Anthropic API - Fable 5 | $52,375 | $384 | $1,305 | $4,125 | - | 0.0038 | - |
| api-openai | OpenAI API - GPT-5.6 Sol | $20,950 | $154 | $522 | $1,650 | - | 0.0095 | - |
| api-openai | OpenAI API - GPT-5.3-codex | $9,397 | $88 | $307 | $984 | - | 0.0213 | - |
| api-deepseek | DeepSeek V4 Pro (peak-weighted) | $2,137 | $23 | $76 | $234 | $1,933 | 0.0936 | - |
| api-deepseek | DeepSeek V4 Flash (peak-weighted) | $699 | $8 | $25 | $78 | $632 | 0.2861 | - |
| api-zai | Z.ai GLM-5.3 | $10,539 | $46 | $156 | $492 | - | 0.0190 | - |
| api-zai | Z.ai GLM-5.3-Flash | $601 | $3 | $9 | $28 | - | 0.3329 | - |
| api-moonshot | Moonshot Kimi K3 | $15,011 | $115 | $392 | $1,238 | - | 0.0133 | - |
| api-moonshot | Moonshot K2.7-Code | $7,711 | $36 | $123 | $390 | - | 0.0259 | - |
| cf-workers-ai | CF Workers AI GLM-5.3-Flash | $1,206 | $10 | $22 | $60 | - | 0.1658 | - |
| cf-workers-ai | CF Workers AI GLM-5.2 | $10,544 | $51 | $161 | $498 | - | 0.0190 | - |
| agg-openrouter | OpenRouter Opus 5 (+5.5% fee) | $27,628 | $203 | $688 | $2,176 | - | 0.0072 | - |
| agg-openrouter | OpenRouter GLM-5.3 (+5.5% fee) | $11,119 | $49 | $165 | $520 | - | 0.0180 | - |
| agg-coding-plans | Atlas Cloud GLM-5.2 PAYG | $9,485 | $42 | $141 | $443 | - | 0.0211 | - |
| hybrid-router | Hybrid 80/20 (LiteLLM, no fee) | $5,797 | $45 | $151 | $475 | - | 0.0345 | - |
| hybrid-router | Hybrid 90/10 (LiteLLM, no fee) | $3,248 | $26 | $88 | $276 | - | 0.0616 | - |
| hybrid-router | Hybrid 95/5 (LiteLLM, no fee) | $1,973 | $17 | $57 | $177 | - | 0.1013 | - |
| sub-claude-max20 | Claude Max 20x | $200 | $200 | $200 | $200 | - | 1.0000 | 255M |
| sub-codex-pro20 | ChatGPT Pro / Codex | $200 | $200 | $200 | $200 | - | - | 319M |
| sub-kimi-top | Kimi membership (top tier) | $199 | $199 | $199 | $199 | - | - | 443M |
| sub-zai-coding | GLM Coding Plan Lite (list $18.0/mo) | $18 | $18 | $18 | $18 | - | 0.0065 | - |
| sub-zai-coding | GLM Coding Plan Pro (list $80.0/mo) | $80 | $80 | $80 | $80 | - | 0.0392 | - |
| sub-zai-coding | GLM Coding Plan Max (list $168.0/mo) | $168 | $168 | $168 | $168 | - | 0.0914 | - |
| self-rented-gpu | GLM-5.2 8xH200 RunPod, concurrency 1 | $24,483 | null | null | null | - | 0.0082 | - |
| self-rented-gpu | GLM-5.2 8xH200 RunPod, concurrency 4 | $8,295 | null | null | null | - | 0.0241 | - |
| self-rented-gpu | GLM-5.2 8xH200 RunPod, concurrency 8 | $5,598 | null | null | null | - | 0.0357 | - |
| self-local | Qwen3-Coder-30B / GLM-4.5-Air local | null | null | null | null | - | - | - |
Sub break-even = token volume at which the same models on PAYG API cost $200/month; the sub is cheaper above it. Observed volume is 33,398M tokens, 131× Claude Max's break-even of 255.1M.
Throttle events observed in logs: 14 (all in week 2026-W32). Codex CLI and Kimi app stores held no token logs on this machine (Codex sqlite threads table: 0 rows; ~/.kimi: config only), their usage is not in these bars; Kimi K3 tokens routed through Claude Code (1,567M input) are.
observed $27627.79/mo, parity 0.0072, control 100, capability 100.0
Limits as published: rpm: 20
spend-cap:yes · realtime-dash:yes · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
ANTHROPIC_BASE_URL=https://openrouter.ai/api; ANTHROPIC_AUTH_TOKEN=$OPENROUTER_API_KEY; ANTHROPIC_API_KEY="" (must be explicitly empty) https://openrouter.ai/api anthropic/claude-opus-5OPENROUTER_API_KEY via [model_providers.openrouter] in config.toml https://openrouter.ai/api/v1 openai/gpt-latest (any OpenRouter model id)OPENROUTER_API_KEY (built-in provider via /connect) Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $1973.48/mo, parity 0.1013, control 100, capability 98.1
Limits as published: not published
spend-cap:yes · realtime-dash:yes · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
ANTHROPIC_BASE_URL http://0.0.0.0:4000 (LiteLLM proxy; serves Anthropic /v1/messages) as configured in LiteLLM config.yaml model_listenv_key = "OPENAI_API_KEY" (config.toml model_providers) http://<litellm-proxy>:4000/v1 (OpenAI-compatible) via model_providers.<id>.base_url model_provider = "proxy" + model = <litellm model_name>ANTHROPIC_API_KEY, ANTHROPIC_BASE_URL LiteLLM proxy base via [providers.anthropic] type="anthropic" base_url, or OpenAI-compatible provider, in config.toml declared per-model in config.toml [models] over the providernone required; opencode.json provider config (npm @ai-sdk/openai-compatible) http://<litellm-proxy>:4000/v1 via provider.<id>.options.baseURL in opencode.json declared under provider.<id>.models in opencode.jsonChanges, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $199.0/mo, parity None, control 65, capability 70
Limits as published: concurrency: 30; session_window: 5-hour window, approx 300-1,200 requests depending on plan; weekly_cap: weekly rate limit + 7-day credit refresh cycle (numeric cap not published)
spend-cap:yes · realtime-dash:yes · shared-with-chat:yes · quotas-in-tokens/requests:yes
Integrations:
ANTHROPIC_BASE_URL, ANTHROPIC_API_KEY https://api.kimi.com/coding/v1 kimi-for-coding/login device authorization in Kimi Code CLI, or API key from Kimi Console kimi-for-codingAPI key (OpenAI-compatible provider config) https://api.kimi.com/coding/v1 kimi-for-codingChanges, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $200.0/mo, parity None, control 50, capability 99.2
Limits as published: session_window: 5-hour rolling window shared by local messages and cloud chats; per-model included ranges ; weekly_cap: exists but unpublished ('Additional weekly limits may apply')
spend-cap:yes · realtime-dash:yes · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
gpt-5.6 (Sol/Terra/Luna) via Sign in with ChatGPT ChatGPT Plus/Pro OAuth via /connect -> OpenAIChanges, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $200.0/mo, parity 1.0, control 40, capability 100.0
Limits as published: session_window: 5-hour rolling session; resets every five hours; weekly_cap: weekly usage limit across all models (separate Opus bucket); quantity unpublished; resets
spend-cap:yes · realtime-dash:yes · shared-with-chat:yes · quotas-in-tokens/requests:no
Integrations:
(none, /login with claude.ai subscription credentials; a set ANTHROPIC_API_KEY overrides the subscription) claude-opus-5 Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $20949.98/mo, parity 0.0095, control 55, capability 95.9
Limits as published: rpm: 500 (Tier 1) to 15,000 (Tier 5) for gpt-5.3-codex and gpt-5.6 family; tpm: 500,000 (Tier 1) to 40,000,000 (Tier 5) for gpt-5.3-codex and gpt-5.6 family; weekly_cap: none; OpenAI-approved monthly usage limit by tier instead: $100/mo (Free, Tier 1) to $200,
spend-cap:yes · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
OPENAI_API_KEY https://api.openai.com gpt-5.3-codex / gpt-5.6 family (model availability follows the API models available to your key) provider 'openai' via /connect -> Manually enter API KeyChanges, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $5597.53/mo, parity 0.0357, control 40, capability 85.7
Limits as published: concurrency: operator-set (vLLM --max-num-seqs)
spend-cap:yes · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:no
Integrations:
ANTHROPIC_BASE_URL http://<host>:8000 (vLLM serves Anthropic /v1/messages) <served-model-name e.g. glm-5.2-fp8>model_providers.<id>.base_url in config.toml http://<host>:8000/v1 <served-model-name>providers.<name>.base_url in ~/.kimi/config.toml, type=openai_legacy or anthropic http://<host>:8000/v1 <served-model-name>provider.<id>.options.baseURL with npm @ai-sdk/openai-compatible http://<host>:8000/v1 <served-model-name>Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $9485.17/mo, parity 0.0211, control 50, capability 91.9
Limits as published: rpm: 45; weekly_cap: 16500000
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $168.0/mo, parity 0.0914, control 5, capability 82.6
Limits as published: session_window: 5-hour rolling credit window: Lite 2,000 / Pro 12,000 / Max 28,000 credits; resets 5 hours; weekly_cap: Lite 10,000 / Pro 60,000 / Max 140,000 credits per week (Team: Standard 66,000 / Premium 1
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
ANTHROPIC_BASE_URL=https://api.z.ai/api/anthropic, ANTHROPIC_AUTH_TOKEN=<Z.ai API key>; default maps Opus/Sonnet/Haiku to GLM-5.3-Flash; or npx @z_ai/coding-helper https://api.z.ai/api/anthropic glm-5.3API key via opencode auth login built-in provider: opencode auth login -> 'Z.AI Coding Plan' glm-5.3Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $26187.48/mo, parity 0.0076, control 5, capability 100.0
Limits as published: rpm: 1000; tpm: 2000000; weekly_cap: none weekly; monthly spend caps by tier: Start $500, Build $1,000, Scale $200,000, Custom
spend-cap:yes · realtime-dash:yes · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
ANTHROPIC_API_KEY https://api.anthropic.com claude-opus-5 claude-opus-5 (select via /models after /connect > Anthropic > Manually enter API Key)Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $15011.39/mo, parity 0.0133, control 0, capability 70
Limits as published: rpm: 10000; tpm: 5000000; concurrency: 1000
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL, CLAUDE_CODE_AUTO_COMPACT_WINDOW=1048576 https://api.moonshot.ai/anthropic kimi-k3[1m]via CC Switch provider config (API key) https://api.moonshot.ai/v1 kimi-k3Kimi account or 'a callable API key' opencode auth login -> built-in 'Moonshot AI' provider kimi-k3Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $2136.84/mo, parity 0.0936, control 0, capability 70
Limits as published: concurrency: 500
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL=deepseek-v4-pro[1m], ANTHROPIC_DEFAULT_HAIKU_MODEL=deepseek-v4-flash, CLAUDE_CODE_SUBAGENT_MODEL=deepseek-v4-flash, CLAUDE_CODE_EFFORT_LEVEL=max, CLAUDE_CODE_AUTO_COMPACT_WINDOW=786432 https://api.deepseek.com/anthropic deepseek-v4-pro[1m]~/.codex/config.toml written by official setup script https://api.deepseek.com (Responses API; one-click script cdn.deepseek.com/api-docs/codex-deepseek-setup-en.sh) deepseek-v4-proAPI key via /connect flow built-in provider: /connect -> deepseek (OpenCode >= v1.14.24) DeepSeek-V4-ProChanges, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, account balance negative; completions would 402
observed $1206.44/mo, parity 0.1658, control 0, capability 90.4
Limits as published: rpm: 20
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:yes
Integrations:
(needs an OpenAI->Anthropic router/proxy; Cloudflare documents OpenAI-compatible /v1/chat/completions only, no Anthropic-format endpoint) https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1 @cf/zai-org/glm-5.3-flashcustom model_provider with OpenAI-compatible base_url; not documented by Cloudflare or OpenAI for each other https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1 @cf/zai-org/glm-5.3-flashOpenAI-compatible endpoint serves @cf/moonshotai/kimi-k2.6; no vendor doc pairs it with kimi_code https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1 @cf/moonshotai/kimi-k2.6CLOUDFLARE_ACCOUNT_ID + CLOUDFLARE_API_KEY (or /connect) selected via /modelsChanges, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, CLOUDFLARE_API_TOKEN lacks Workers AI permission; CLOUDFLARE_ACCOUNT_ID also absent from env
observed $10539.08/mo, parity 0.019, control 0, capability 85.7
Limits as published: not published
spend-cap:no · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:no
Integrations:
ANTHROPIC_BASE_URL=https://api.z.ai/api/anthropic, ANTHROPIC_AUTH_TOKEN=<Z.ai API key>; defaults map Opus/Sonnet/Haiku to GLM-5.3-Flash; switch to GLM-5.3 per docs https://api.z.ai/api/anthropic glm-5.3 https://api.z.ai/api/v1 (OpenAI Response Protocol documented; no Codex-specific official guide found) glm-5.3API key via opencode auth login built-in provider: opencode auth login -> Z.AI (or Z.AI Coding Plan) glm-5.3Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
observed $None/mo, parity None, control 40, capability 51.6
Limits as published: concurrency: operator-set (local server)
spend-cap:yes · realtime-dash:no · shared-with-chat:no · quotas-in-tokens/requests:no
Integrations:
ANTHROPIC_BASE_URL http://localhost:<port> (server must expose Anthropic /v1/messages, e.g. vLLM) <served-model-name>oss_provider / model_providers.<id>.base_url local via --oss (lmstudio | ollama) or model_providers.<id>.base_url <local model>providers.<name>.base_url in ~/.kimi/config.toml http://localhost:<port>/v1 (type=openai_legacy) <local model>provider.<id>.options.baseURL with npm @ai-sdk/openai-compatible http://localhost:11434/v1 (Ollama example in official docs) <local model>Changes, last 90 days:
Throttling incidents, 180 days:
Bake-off: skipped, no API key in env
claude_code=$20, opencode=$15, experiments=$15. Total exposure: $50, hard.deepseek/deepseek-v4-flash (off-peak aware), fallback anthropic/claude-opus-5 via OpenRouter; set proxy max_budget=200, budget_duration=30d.ANTHROPIC_BASE_URL=http://localhost:4000 for Claude Code (LiteLLM /v1/messages passthrough); OpenAI-compatible base URL for opencode.out/S4/fixture/ (hash 2d78…8cb0) is the controlled comparison set: 3 tasks, 2 attempts each./usage; if delivered capacity drops, shift the routing split rather than buying a second sub.ANTHROPIC_BASE_URL (tools revert to the sub), leave remaining OpenRouter credits parked, they do not expire on a schedule the audit found, and the per-key caps keep them inert.is_available:false); Cloudflare token lacks Workers AI scope; no other keys in env. Effect: capability came from public benchmarks alone; candidates with <2 independent entries (api-deepseek, api-moonshot, sub-kimi-top) were capped at 70 and may be understated.USER_LINK resolved via mirror (x.com blocks direct fetch; retrieved through Twitter's syndication CDN, mirror:true): @zebassembly, 2026-08-27 — “i was so excited that local models like qwen 3.8 27B exist, and then I did the math on how much it costs on my power bill....
looks like i'll be using glm 5.3 flash on Cloudflare instead” Author carries a Cloudflare business label on the platform; the audit scored Cloudflare Workers AI on its own artifacts (B1): final 39, parity 0.1658, rpm 20.
| url | status | unmatched snippet |
|---|---|---|
| https://platform.claude.com/docs/en/api/rate-limits | 200 | You can monitor your rate limit usage on the Usage page of the Claude |
| https://www.tbench.ai/leaderboard/terminal-bench/2.1 | 200 | "label":"GLM-5.1"},"reasoning_effort":"max"},"metrics":{"accuracy":58. |
| https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813 | 200 | "safetensors":{"parameters":{"BF16":2954820352,"I64":2327040,"F32":905 |
| https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 | 200 | "total":304180418494,"sharded":true |
| https://huggingface.co/moonshotai/Kimi-K2.7-Code | 200 | "createdAt":"2026-06-11T07:51:47.000Z" |
| https://developers.openai.com/codex/pricing | 200 | On ChatGPT plans, local messages and cloud chats share a five-hour win |
| https://developers.openai.com/codex/pricing | 200 | If you want to see your remaining limits during an active Codex CLI se |
| https://developers.openai.com/codex/pricing | 200 | All users may also run extra local chats using an API key, with usage |
| https://developers.cloudflare.com/workers-ai/platform/pricing/ | 200 | USD per 1M tokens = (neurons per M tokens) x ($0.011 / 1000 neurons). |
| https://docs.runpod.io/pods/pricing | 200 | GPUs are dedicated to your Pod and cannot be displaced by other users. |
| https://support.claude.com/en/articles/9797557-usage-limit-best-practices | 200 | navigate to Settings > Usage to view progress bars showing how much of |
| https://support.claude.com/en/articles/12429409-manage-usage-credits-for-paid-claude-plans | 200 | Usage dashboard: View real-time consumption in Settings > Usage. |
| https://docs.z.ai/devpack/overview | 200 | All plans support **GLM-5.3**, GLM-5-Flash. |
| https://docs.z.ai/devpack/usage-policy | 200 | The platform dynamically adjusts these limits based on resource availa |
Run fb2d5acf-a818-4a52-9c17-b4da13b7f15b · inputs_hash 6ac7f23a16c9727d… · S0–S7 per AI_SPEND_AUDIT v2 · every figure tagged observed/modelled · generated 2026-08-28.