Insights

OpenRouter Pay-as-You-Go vs LLM Subscriptions: Measured Costs and a Workload Split

11 min read#llm-pricing#subscriptions#openrouter#coding-agents#cost-optimization

Who this is forDevelopers who run heavy agentic coding harnesses such as Claude Code or opencode and are choosing between flat-rate subscriptions and pay-as-you-go routing through OpenRouter.

Heavy agentic coding sessions can consume billions of tokens in a single month, and the price you pay depends on whether you buy a flat-rate subscription or pay per token through a gateway such as OpenRouter. This article reports what I measured on one heavy Claude Code setup, why a $10 OpenRouter top-up ran out in a single week, and how subscription and API prices stood on September 5, 2026. The short answer is that this is not an either-or decision. Subscriptions absorb the cache-heavy volume that dominates agentic work, while pay-as-you-go is only safe for unattended runs inside a harness with spend limits. You will get the measured figures, the price tables, and a workload split you can apply to your own setup.

Key diagram

Subscriptions absorb cache reads: measured cost structure and workload split

Key data

About 95% of the tokens in my agentic coding work were cache reads. That single fact drives most of the analysis. An agent such as Claude Code repeatedly resends a long working context, and most of that context is served from the provider’s cache rather than processed fresh. The cost model I used prices a cache write at 1.25 times the base input price and a cache read at 0.1 times the base input price, based on the standard 5-minute cache TTL. Under that model, a volume that looks enormous in tokens can still be modest in dollars, but only if you pay for it in a way that matches the discount. A flat subscription does that by design. A metered API bills every token at the rate the provider lists, so the same workflow can become expensive quickly.

1. Measured usage: what the subscription absorbed (Claude Max 20x, $200 per month)

Month Total tokens Cache-read share API-equivalent cost Versus $200
April 2026 (confirmed in report) 2.05 billion — $5,199 26×
July 2026 14.6 billion 95% $17,297 86×
August 2026 (20 active days) 11.16 billion 95% (10.6 billion) $10,298 51×
  • Calculation: Claude Code JSONL usage aggregated from claude-monthly-review/data/<month>/by_model.csv, multiplied by Anthropic’s published API prices (Opus 5 $5 input / $25 output, Fable 5 $10 / $50, Sonnet 5 $2 / $10, per 1M tokens), then adjusted by the cache multipliers (cache write 1.25×, cache read 0.1×, standard 5-minute TTL). The August model split was Fable 5 $4,889, Opus 5 $4,141, and Sonnet 5 $1,254.
  • The same August volume, routed through other paths, with the 10.6 billion cache-read tokens as the variable:
Path Without cache discount With 90% cache-read discount
DeepSeek v4-flash (lowest OpenRouter price, $0.08 / $0.17) $897 $134
GLM-5.3 ($1.40 / $4.40) $15,777 $2,418
Sonnet 5 direct API ($2 / $10, cache read 0.2) $3,895 —

The comparison shows why the cheapest listed model is not the cheapest workflow. DeepSeek v4-flash looks inexpensive at $0.08 per million input tokens, but once the cache-read volume is counted at face value it still reaches $897 for the month. Only with a 90% cache-read discount does it fall to $134. Routing the same workload through a metered path changes the bill far more than the choice of model does.

2. OpenRouter in practice: how $10 disappeared in one week

Item Value Source
Top-up / usage $10.00 / $10.08 (balance −$0.08) /api/v1/credits, September 5, 2026
Weekly usage $9.96 (all within the last 7 days) /api/v1/auth/key
Purpose age-of-steam LLM arena fifth-round screening + openrouter-lab benchmarks Internal report
Normal cost per judgment GLM 5.3 about $0.5, DeepSeek v4-flash $0.06 Arena report, fifth round
Causes of depletion ① Empty-content responses were still billed (reasoning consumed 3,800 to 4,096 tokens, then stopped at the length limit). ② Reasoning had no upper bound. ③ With a low balance, max_tokens reservations for concurrent requests returned 402, which wiped out the pilot. Arena report, fifth round (fix commits 25d06d3, a668115)
Resume estimate flash 24 judgments ≈ $1.5, GLM 6 judgments ≈ $5.4 Same

The incident was not caused by unit prices. The arena pilot was a batch job, and its harness had three defects that interacted badly with a small balance. The first was billing for empty responses: reasoning tokens were consumed up to the length limit and the request still charged for them. The second was that reasoning had no upper bound at all, so a single request could spend far more than the expected per-judgment cost. The third was reservation behavior: when several requests ran concurrently against a balance that was already low, OpenRouter reserved max_tokens for each one and rejected requests with 402 errors, which stopped the pilot entirely.

OpenRouter’s fee structure, according to its official FAQ as checked on September 5, 2026, works as follows. Card top-ups carry a 5.5% fee with an $0.80 minimum. Crypto top-ups carry a 5% fee. Model prices pass through from providers with no markup. BYOK usage above $25,000 per month carries a 5% fee, and unused credits can expire after one year. For a $10 top-up, the $0.80 minimum fee works out to 8% of the amount, and the lower the balance, the more often reservation-based 402 errors occur.

3. Subscription price table (checked September 5, 2026, USD per month)

Service and plan Monthly price Coding agent and limits Source tier
Claude Pro / Max 5x / Max 20x $20 (annual $17) / $100 / $200 Claude Code included; web, desktop, and code share one pool. 5x and 20x multiply the Pro 5-hour window. Primary: claude.com/pricing
ChatGPT Go / Plus / Pro KRW 13,000 / KRW 29,000 / “5× or 20×” (USD $8 / $20 / $100 or $200) Codex included on all paid plans (Free is limited). GPT-5.6 family; Pro includes GPT-5.6 Pro. Primary: chatgpt.com/pricing (Korean locale, KRW and structure) + secondary, three sources agree on USD
Google AI Plus / AI Pro / AI Ultra $4.99 / $19.99 / $99.99 or $199.99 Jules and Antigravity included. Gemini CLI: AI Pro 1,500 requests/day, Ultra 2,000 requests/day. The free and Google One paths were replaced by Antigravity CLI per a June 18 notice. Primary: gemini.google/subscriptions, geminicli.com docs
Z.ai GLM Coding Plan Lite / Pro / Max $18 / $80 / $168 (annual $12.6 / $56 / $117.6) GLM-5.3 and 5.3-Flash. Lite 2,000 credits per 5 hours and 10,000 per week; Pro 6×; Max 14×. Works with Claude Code, OpenCode, Cline, and 20+ tools. Primary: z.ai/subscribe, docs.z.ai
MiniMax Token Plan Plus / Max / Ultra $22 / $55 / $132 M3, M2.7, image, and voice shared. Works with Claude Code, Cursor, Hermes. Primary: platform.minimax.io docs
Kimi Code (Moonshot) Moderato / Allegretto / Allegro / Vivace $19 / $39 / $99 / $199 (annual $15 / $31 / $79 / $159) K3 requires Moderato or higher. The current button shows a Chinese-language label meaning “subscription reservation”: new sign-ups were halted on July 19 due to GPU shortages, and batch resumption is in progress. Primary: kimi.com/code + secondary (halt reporting)
Alibaba Model Studio Coding Plan Pro $50 (Lite at $10 closed to new sign-ups on March 20, 2026) qwen3.7-plus, kimi-k2.5, glm-5, MiniMax-M2.5. 6,000 requests per 5 hours, 45,000 per week, 90,000 per month. Works with Claude Code, Qwen Code, Cursor. Terms explicitly prohibit automation scripts and backend use. Primary: alibabacloud.com
DeepSeek No subscription (API only) v4-flash: cached input $0.014 / uncached input $0.44 / output $1.32 (peak; off-peak is half), v4-pro $1.32 / $3.96. Provides an Anthropic-compatible endpoint. Primary: api-docs.deepseek.com
StepFun Step Plan $6.99 to $99 Claude Code compatible Secondary: cc-compatible-models
Xiaomi MiMo Token Plan About $6 to $100 No 5-hour window limit Secondary
Ollama Cloud Pro / Max $20 / $100 Local execution is unlimited and free Secondary

The table makes two things visible. Entry prices for coding plans cluster around $18 to $22, and most of the low-cost tiers come with conditions that matter more than the headline price. Alibaba’s terms prohibit automation and backend use outright, and Kimi has paused new sign-ups. Price alone does not tell you whether a plan fits a given workload.

4. Pay-as-you-go API price snapshot (per 1M tokens, September 5, 2026)

Model Input / output Notes
Gemini 2.5 Flash-Lite $0.10 / $0.40 Lowest current price; free tier available
Gemini 3.5 Flash-Lite $0.30 / $2.50
Gemini 3.8 Flash $0.75 / $3.75 (through 12/31; then $1.50 / $7.50) The “Flash means cheap” rule breaks down in 3.x
Gemini 3.1 Pro $2 / $12 (≤200K) No free tier
DeepSeek v4-flash (OpenRouter listed price) $0.08 / $0.17 Third-party hosts are cheaper than DeepSeek direct
GLM-5.3-Flash / GLM-5.3 (OpenRouter) $0.07 / $0.25 · $1.40 / $4.40
MiniMax M3 (OpenRouter) $0.30 / $1.20
Kimi K3 $3 / $15 Expensive among Chinese providers
Claude Sonnet 5 / Opus 5 / Fable 5 $2 / $10 · $5 / $25 · $10 / $50 Cache read 0.1×

Insights

  1. A subscription is a product that absorbs cache reads at a flat price. Of the 11.16 billion tokens used in August, 10.6 billion were cache reads. Even at the 0.1× discounted rate, that volume costs $2,632 on Opus 5 and $2,961 on Fable 5. This is why users who run a harness such as Claude Code or opencode for several hours a day see subscription returns of 50× to 86×. Moving the same workflow to pay-as-you-go costs several hundred dollars a month even on the cheapest models.

  2. OpenRouter felt expensive because of usage patterns and incidents, not unit prices. Three factors combined. Batch experiments gain almost nothing from caching, so the discount that makes a subscription efficient does not help them. Harness defects (billing empty responses, unbounded reasoning, and 402 reservations) amplified the charges. A small $10 top-up carries an 8% fee and is vulnerable to 402 errors when the balance runs low. Pay-as-you-go is only safe in a harness with --max-spend caps, reasoning limits, and retries for empty responses. The openrouter-lab operating rules already require these controls.

  3. Subscription tokens cannot be used for automation, so the answer is workload splitting, not a single choice. Anthropic prohibits using subscription OAuth in third-party harnesses, according to a July 22, 2026 research ruling. Alibaba’s coding plan explicitly bans automation scripts and backends. Interactive and coding work where a human watches the screen belongs on a subscription. Unattended batch, backend, and benchmark work belongs on pay-as-you-go. Low-cost parallel experiments belong on a coding plan at $18 to $22.

  4. Chinese-provider coding plans offer Claude Code compatibility at an entry price of $18 to $22, but the jurisdictional risk remains. Routing them through US hosting, for example by pinning an OpenRouter provider, turns them into metered usage and removes the subscription advantage. Use them only for experiments and parallel runs that involve no sensitive data. Kimi now requires reservations for new sign-ups, and Alibaba closed its Lite tier, which reflects a broader shrinking of low-cost tiers.

  5. Gemini has lost its position as a free or low-cost coding CLI. The free login path was replaced by Antigravity CLI on June 18, and 3.x Flash prices are 3 to 7 times the 2.5 generation. The 2.5 Flash-Lite price ($0.10 / $0.40) and the API free tier (250 requests per day) remain useful only for small experiments.

Decisions for this setup

Item Decision Basis
Claude Max 20x, $200 Keep August equivalent $10,298. Downgrading to Max 5x ($100 saving) will be decided after the opencode parallel-run results (previously deferred).
ChatGPT Plus, KRW 29,000 Keep Codex account is already on Plus. No additional tier needed.
Gemini AI Pro, $19.99 Do not add The coding CLI path moved to Antigravity and has no place in the current workflow.
GLM Lite $18 / MiniMax Plus $22 Conditional Only if GLM-5.3 replaces deepseek-v3.2 on a flat rate during the opencode parallel run. Cancel once the Max 5x decision is made.
OpenRouter Keep, with reduced role Small pool dedicated to unattended batch and benchmarks. Top up at $25 or more to reduce fees and 402 errors. Harness spend caps are mandatory. Arena resume cost $1.5 to $5.4.
DeepSeek direct API Alternative for batch experiments Anthropic-compatible endpoint, $0.014 cache-hit price, half price off-peak. Excluded for sensitive data due to jurisdictional risk.

Claims that need deeper verification: ChatGPT Pro USD pricing (the official page shows only KRW and “From” prices), and the actual 5-hour limit of Claude Max 5x (not publicly disclosed; it can only be determined through parallel runs).

Bottom line

For this heavy user, the evidence does not support moving interactive coding to pay-as-you-go. The subscription absorbed usage that would have cost $10,298 in August at API rates, and the OpenRouter overspend came from harness defects and a small top-up, not from unit prices. The subscription should stay for work a person watches. OpenRouter should be kept as a small, capped pool for unattended batch and benchmark runs, with top-ups of $25 or more. Low-cost coding plans are worth testing for non-sensitive parallel runs, but their terms and jurisdictional risks mean they should not carry sensitive data.

Sources

Retrieved September 5, 2026. Primary sources are provider official pages, my own account APIs, and internal measurements. Secondary sources are community summaries.

Primary: official

Primary: internal measurements

  • ~/Projects/claude-monthly-review/data/2026-07, 2026-08 by_model.csv: token aggregation by model
  • ~/Projects/claude-monthly-review/docs/2026-04-claude-max-report.md: April 2026 confirmed figure of $5,199 / 26×
  • ~/Projects/age-of-steam-web/docs/reports/llm-arena-eval-2026-08-31.md, fifth round: three credit-depletion incidents and cost per judgment
  • ~/Projects/openrouter-lab/CLAUDE.md, docs/handoff/2026-09-01-session-handoff.md: spend-cap rules and the deferred Max 5x downgrade
  • researches/2026-07-22-hermes-llm-backend-selection.md: ruling that subscription tokens are banned in third-party harnesses; jurisdictional risk of Chinese providers

Secondary: community (cross-checks)

Frequently asked questions

Should a heavy agentic coding user switch from a flat-rate subscription to OpenRouter pay-as-you-go?
Not for interactive coding. In August 2026, a $200 Claude Max 20x plan corresponded to $10,298 of API-equivalent usage because about 95% of tokens were cache reads. Keep pay-as-you-go for unattended batch and benchmark runs.
Why did a $10 OpenRouter balance run out in one week?
Three harness defects amplified spending: empty responses were billed after reasoning hit the length limit, reasoning had no cap, and concurrent max_tokens reservations returned 402 errors at low balance. The $0.80 minimum card fee also took 8% of a $10 top-up.