OpenRouter Pay-as-You-Go vs LLM Subscriptions: Measured Costs and a Workload Split
Who this is forDevelopers who run heavy agentic coding harnesses such as Claude Code or opencode and are choosing between flat-rate subscriptions and pay-as-you-go routing through OpenRouter.
Heavy agentic coding sessions can consume billions of tokens in a single month, and the price you pay depends on whether you buy a flat-rate subscription or pay per token through a gateway such as OpenRouter. This article reports what I measured on one heavy Claude Code setup, why a $10 OpenRouter top-up ran out in a single week, and how subscription and API prices stood on September 5, 2026. The short answer is that this is not an either-or decision. Subscriptions absorb the cache-heavy volume that dominates agentic work, while pay-as-you-go is only safe for unattended runs inside a harness with spend limits. You will get the measured figures, the price tables, and a workload split you can apply to your own setup.
Key diagram

Key data
About 95% of the tokens in my agentic coding work were cache reads. That single fact drives most of the analysis. An agent such as Claude Code repeatedly resends a long working context, and most of that context is served from the provider’s cache rather than processed fresh. The cost model I used prices a cache write at 1.25 times the base input price and a cache read at 0.1 times the base input price, based on the standard 5-minute cache TTL. Under that model, a volume that looks enormous in tokens can still be modest in dollars, but only if you pay for it in a way that matches the discount. A flat subscription does that by design. A metered API bills every token at the rate the provider lists, so the same workflow can become expensive quickly.
1. Measured usage: what the subscription absorbed (Claude Max 20x, $200 per month)
| Month | Total tokens | Cache-read share | API-equivalent cost | Versus $200 |
|---|---|---|---|---|
| April 2026 (confirmed in report) | 2.05 billion | — | $5,199 | 26× |
| July 2026 | 14.6 billion | 95% | $17,297 | 86× |
| August 2026 (20 active days) | 11.16 billion | 95% (10.6 billion) | $10,298 | 51× |
- Calculation: Claude Code JSONL usage aggregated from
claude-monthly-review/data/<month>/by_model.csv, multiplied by Anthropic’s published API prices (Opus 5 $5 input / $25 output, Fable 5 $10 / $50, Sonnet 5 $2 / $10, per 1M tokens), then adjusted by the cache multipliers (cache write 1.25×, cache read 0.1×, standard 5-minute TTL). The August model split was Fable 5 $4,889, Opus 5 $4,141, and Sonnet 5 $1,254. - The same August volume, routed through other paths, with the 10.6 billion cache-read tokens as the variable:
| Path | Without cache discount | With 90% cache-read discount |
|---|---|---|
| DeepSeek v4-flash (lowest OpenRouter price, $0.08 / $0.17) | $897 | $134 |
| GLM-5.3 ($1.40 / $4.40) | $15,777 | $2,418 |
| Sonnet 5 direct API ($2 / $10, cache read 0.2) | $3,895 | — |
The comparison shows why the cheapest listed model is not the cheapest workflow. DeepSeek v4-flash looks inexpensive at $0.08 per million input tokens, but once the cache-read volume is counted at face value it still reaches $897 for the month. Only with a 90% cache-read discount does it fall to $134. Routing the same workload through a metered path changes the bill far more than the choice of model does.
2. OpenRouter in practice: how $10 disappeared in one week
| Item | Value | Source |
|---|---|---|
| Top-up / usage | $10.00 / $10.08 (balance −$0.08) | /api/v1/credits, September 5, 2026 |
| Weekly usage | $9.96 (all within the last 7 days) | /api/v1/auth/key |
| Purpose | age-of-steam LLM arena fifth-round screening + openrouter-lab benchmarks | Internal report |
| Normal cost per judgment | GLM 5.3 about $0.5, DeepSeek v4-flash $0.06 | Arena report, fifth round |
| Causes of depletion | ① Empty-content responses were still billed (reasoning consumed 3,800 to 4,096 tokens, then stopped at the length limit). ② Reasoning had no upper bound. ③ With a low balance, max_tokens reservations for concurrent requests returned 402, which wiped out the pilot. | Arena report, fifth round (fix commits 25d06d3, a668115) |
| Resume estimate | flash 24 judgments ≈ $1.5, GLM 6 judgments ≈ $5.4 | Same |
The incident was not caused by unit prices. The arena pilot was a batch job, and its harness had three defects that interacted badly with a small balance. The first was billing for empty responses: reasoning tokens were consumed up to the length limit and the request still charged for them. The second was that reasoning had no upper bound at all, so a single request could spend far more than the expected per-judgment cost. The third was reservation behavior: when several requests ran concurrently against a balance that was already low, OpenRouter reserved max_tokens for each one and rejected requests with 402 errors, which stopped the pilot entirely.
OpenRouter’s fee structure, according to its official FAQ as checked on September 5, 2026, works as follows. Card top-ups carry a 5.5% fee with an $0.80 minimum. Crypto top-ups carry a 5% fee. Model prices pass through from providers with no markup. BYOK usage above $25,000 per month carries a 5% fee, and unused credits can expire after one year. For a $10 top-up, the $0.80 minimum fee works out to 8% of the amount, and the lower the balance, the more often reservation-based 402 errors occur.
3. Subscription price table (checked September 5, 2026, USD per month)
| Service and plan | Monthly price | Coding agent and limits | Source tier |
|---|---|---|---|
| Claude Pro / Max 5x / Max 20x | $20 (annual $17) / $100 / $200 | Claude Code included; web, desktop, and code share one pool. 5x and 20x multiply the Pro 5-hour window. | Primary: claude.com/pricing |
| ChatGPT Go / Plus / Pro | KRW 13,000 / KRW 29,000 / “5× or 20×” (USD $8 / $20 / $100 or $200) | Codex included on all paid plans (Free is limited). GPT-5.6 family; Pro includes GPT-5.6 Pro. | Primary: chatgpt.com/pricing (Korean locale, KRW and structure) + secondary, three sources agree on USD |
| Google AI Plus / AI Pro / AI Ultra | $4.99 / $19.99 / $99.99 or $199.99 | Jules and Antigravity included. Gemini CLI: AI Pro 1,500 requests/day, Ultra 2,000 requests/day. The free and Google One paths were replaced by Antigravity CLI per a June 18 notice. | Primary: gemini.google/subscriptions, geminicli.com docs |
| Z.ai GLM Coding Plan Lite / Pro / Max | $18 / $80 / $168 (annual $12.6 / $56 / $117.6) | GLM-5.3 and 5.3-Flash. Lite 2,000 credits per 5 hours and 10,000 per week; Pro 6×; Max 14×. Works with Claude Code, OpenCode, Cline, and 20+ tools. | Primary: z.ai/subscribe, docs.z.ai |
| MiniMax Token Plan Plus / Max / Ultra | $22 / $55 / $132 | M3, M2.7, image, and voice shared. Works with Claude Code, Cursor, Hermes. | Primary: platform.minimax.io docs |
| Kimi Code (Moonshot) Moderato / Allegretto / Allegro / Vivace | $19 / $39 / $99 / $199 (annual $15 / $31 / $79 / $159) | K3 requires Moderato or higher. The current button shows a Chinese-language label meaning “subscription reservation”: new sign-ups were halted on July 19 due to GPU shortages, and batch resumption is in progress. | Primary: kimi.com/code + secondary (halt reporting) |
| Alibaba Model Studio Coding Plan Pro | $50 (Lite at $10 closed to new sign-ups on March 20, 2026) | qwen3.7-plus, kimi-k2.5, glm-5, MiniMax-M2.5. 6,000 requests per 5 hours, 45,000 per week, 90,000 per month. Works with Claude Code, Qwen Code, Cursor. Terms explicitly prohibit automation scripts and backend use. | Primary: alibabacloud.com |
| DeepSeek | No subscription (API only) | v4-flash: cached input $0.014 / uncached input $0.44 / output $1.32 (peak; off-peak is half), v4-pro $1.32 / $3.96. Provides an Anthropic-compatible endpoint. | Primary: api-docs.deepseek.com |
| StepFun Step Plan | $6.99 to $99 | Claude Code compatible | Secondary: cc-compatible-models |
| Xiaomi MiMo Token Plan | About $6 to $100 | No 5-hour window limit | Secondary |
| Ollama Cloud Pro / Max | $20 / $100 | Local execution is unlimited and free | Secondary |
The table makes two things visible. Entry prices for coding plans cluster around $18 to $22, and most of the low-cost tiers come with conditions that matter more than the headline price. Alibaba’s terms prohibit automation and backend use outright, and Kimi has paused new sign-ups. Price alone does not tell you whether a plan fits a given workload.
4. Pay-as-you-go API price snapshot (per 1M tokens, September 5, 2026)
| Model | Input / output | Notes |
|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 / $0.40 | Lowest current price; free tier available |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | |
| Gemini 3.8 Flash | $0.75 / $3.75 (through 12/31; then $1.50 / $7.50) | The “Flash means cheap” rule breaks down in 3.x |
| Gemini 3.1 Pro | $2 / $12 (≤200K) | No free tier |
| DeepSeek v4-flash (OpenRouter listed price) | $0.08 / $0.17 | Third-party hosts are cheaper than DeepSeek direct |
| GLM-5.3-Flash / GLM-5.3 (OpenRouter) | $0.07 / $0.25 · $1.40 / $4.40 | |
| MiniMax M3 (OpenRouter) | $0.30 / $1.20 | |
| Kimi K3 | $3 / $15 | Expensive among Chinese providers |
| Claude Sonnet 5 / Opus 5 / Fable 5 | $2 / $10 · $5 / $25 · $10 / $50 | Cache read 0.1× |
Insights
-
A subscription is a product that absorbs cache reads at a flat price. Of the 11.16 billion tokens used in August, 10.6 billion were cache reads. Even at the 0.1× discounted rate, that volume costs $2,632 on Opus 5 and $2,961 on Fable 5. This is why users who run a harness such as Claude Code or opencode for several hours a day see subscription returns of 50× to 86×. Moving the same workflow to pay-as-you-go costs several hundred dollars a month even on the cheapest models.
-
OpenRouter felt expensive because of usage patterns and incidents, not unit prices. Three factors combined. Batch experiments gain almost nothing from caching, so the discount that makes a subscription efficient does not help them. Harness defects (billing empty responses, unbounded reasoning, and 402 reservations) amplified the charges. A small $10 top-up carries an 8% fee and is vulnerable to 402 errors when the balance runs low. Pay-as-you-go is only safe in a harness with
--max-spendcaps, reasoning limits, and retries for empty responses. The openrouter-lab operating rules already require these controls. -
Subscription tokens cannot be used for automation, so the answer is workload splitting, not a single choice. Anthropic prohibits using subscription OAuth in third-party harnesses, according to a July 22, 2026 research ruling. Alibaba’s coding plan explicitly bans automation scripts and backends. Interactive and coding work where a human watches the screen belongs on a subscription. Unattended batch, backend, and benchmark work belongs on pay-as-you-go. Low-cost parallel experiments belong on a coding plan at $18 to $22.
-
Chinese-provider coding plans offer Claude Code compatibility at an entry price of $18 to $22, but the jurisdictional risk remains. Routing them through US hosting, for example by pinning an OpenRouter provider, turns them into metered usage and removes the subscription advantage. Use them only for experiments and parallel runs that involve no sensitive data. Kimi now requires reservations for new sign-ups, and Alibaba closed its Lite tier, which reflects a broader shrinking of low-cost tiers.
-
Gemini has lost its position as a free or low-cost coding CLI. The free login path was replaced by Antigravity CLI on June 18, and 3.x Flash prices are 3 to 7 times the 2.5 generation. The 2.5 Flash-Lite price ($0.10 / $0.40) and the API free tier (250 requests per day) remain useful only for small experiments.
Decisions for this setup
| Item | Decision | Basis |
|---|---|---|
| Claude Max 20x, $200 | Keep | August equivalent $10,298. Downgrading to Max 5x ($100 saving) will be decided after the opencode parallel-run results (previously deferred). |
| ChatGPT Plus, KRW 29,000 | Keep | Codex account is already on Plus. No additional tier needed. |
| Gemini AI Pro, $19.99 | Do not add | The coding CLI path moved to Antigravity and has no place in the current workflow. |
| GLM Lite $18 / MiniMax Plus $22 | Conditional | Only if GLM-5.3 replaces deepseek-v3.2 on a flat rate during the opencode parallel run. Cancel once the Max 5x decision is made. |
| OpenRouter | Keep, with reduced role | Small pool dedicated to unattended batch and benchmarks. Top up at $25 or more to reduce fees and 402 errors. Harness spend caps are mandatory. Arena resume cost $1.5 to $5.4. |
| DeepSeek direct API | Alternative for batch experiments | Anthropic-compatible endpoint, $0.014 cache-hit price, half price off-peak. Excluded for sensitive data due to jurisdictional risk. |
Claims that need deeper verification: ChatGPT Pro USD pricing (the official page shows only KRW and “From” prices), and the actual 5-hour limit of Claude Max 5x (not publicly disclosed; it can only be determined through parallel runs).
Bottom line
For this heavy user, the evidence does not support moving interactive coding to pay-as-you-go. The subscription absorbed usage that would have cost $10,298 in August at API rates, and the OpenRouter overspend came from harness defects and a small top-up, not from unit prices. The subscription should stay for work a person watches. OpenRouter should be kept as a small, capped pool for unattended batch and benchmark runs, with top-ups of $25 or more. Low-cost coding plans are worth testing for non-sensitive parallel runs, but their terms and jurisdictional risks mean they should not carry sensitive data.
Sources
Retrieved September 5, 2026. Primary sources are provider official pages, my own account APIs, and internal measurements. Secondary sources are community summaries.
Primary: official
- Anthropic Claude pricing: https://claude.com/pricing (Pro $20, annual $17; Max 5x/20x “From $100”; Claude Code included)
- Anthropic API prices:
claude-apiskill cache (June 24, 2026, based on official pricing): Opus 5 $5 / $25, Fable 5 $10 / $50, Sonnet 5 $2 / $10, Haiku 4.5 $1 / $5 - OpenAI ChatGPT pricing (Korean locale): https://chatgpt.com/pricing/ (Go KRW 13,000, Plus KRW 29,000, Pro “5× or 20×”, Codex on all paid plans; collected via Playwright)
- Google AI plans: https://gemini.google/subscriptions/ (Plus $4.99, Pro $19.99, Ultra $99.99 / $199.99)
- Gemini CLI quotas: https://geminicli.com/docs/resources/quota-and-pricing/ (AI Pro 1,500/day, Ultra 2,000/day, API free 250/day, June 18 Antigravity CLI replacement notice)
- Gemini API prices: https://ai.google.dev/gemini-api/docs/pricing
- Z.ai GLM Coding Plan: https://z.ai/subscribe (Playwright) · https://docs.z.ai/devpack/overview
- MiniMax Token Plan: https://platform.minimax.io/docs/coding-plan/intro
- Kimi Code Plan: https://www.kimi.com/code/ (Playwright; Chinese-language page)
- Alibaba Model Studio Coding Plan: https://www.alibabacloud.com/help/en/model-studio/coding-plan
- DeepSeek prices: https://api-docs.deepseek.com/quick_start/pricing (insane-search)
- OpenRouter fees: https://openrouter.ai/docs/faq · Listed model prices: https://openrouter.ai/api/v1/models
- OpenRouter account data:
/api/v1/credits,/api/v1/auth/key(September 5, 2026)
Primary: internal measurements
~/Projects/claude-monthly-review/data/2026-07,2026-08by_model.csv: token aggregation by model~/Projects/claude-monthly-review/docs/2026-04-claude-max-report.md: April 2026 confirmed figure of $5,199 / 26×~/Projects/age-of-steam-web/docs/reports/llm-arena-eval-2026-08-31.md, fifth round: three credit-depletion incidents and cost per judgment~/Projects/openrouter-lab/CLAUDE.md,docs/handoff/2026-09-01-session-handoff.md: spend-cap rules and the deferred Max 5x downgraderesearches/2026-07-22-hermes-llm-backend-selection.md: ruling that subscription tokens are banned in third-party harnesses; jurisdictional risk of Chinese providers
Secondary: community (cross-checks)
- https://github.com/Alorse/cc-compatible-models: StepFun, MiMo, and Ollama plans; provider API prices
- https://www.cloudzero.com/blog/openai-codex-pricing/ · https://www.morphllm.com/codex-pricing · https://codingplan.org/en: ChatGPT Go $8 / Plus $20 / Pro $100 or $200 (three sources agree)
- https://www.pymnts.com/news/artificial-intelligence/2026/moonshot-halts-new-kimi-k3-subscriptions-demand-overwhelms-compute/: Kimi new-subscription halt (July 19, 2026)
- https://www.tembo.io/blog/gemini-cli-pricing: explanation of the end of the free Gemini CLI path
Frequently asked questions
- Should a heavy agentic coding user switch from a flat-rate subscription to OpenRouter pay-as-you-go?
- Not for interactive coding. In August 2026, a $200 Claude Max 20x plan corresponded to $10,298 of API-equivalent usage because about 95% of tokens were cache reads. Keep pay-as-you-go for unattended batch and benchmark runs.
- Why did a $10 OpenRouter balance run out in one week?
- Three harness defects amplified spending: empty responses were billed after reasoning hit the length limit, reasoning had no cap, and concurrent max_tokens reservations returned 402 errors at low balance. The $0.80 minimum card fee also took 8% of a $10 top-up.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›