Insights

Claude API Cost Breakdown: How Cache Hit Ratio Changes What You Pay

6 min read#claude-api#token-pricing#prompt-caching#cost-optimization#cache-hit-ratio

Who this is forDevelopers and non-developer power users who use the Claude API or Claude Code and want to understand why a short question can cost several dollars.

If you use the Claude API, a single request can cost far more than its length suggests. A 15-token question cost me $2.58 in one call, which is hard to explain from the prompt alone. The answer is in the billing breakdown. Each token type has a very different price, and the share of tokens served from cache determines most of the total. This article takes apart one real Claude Opus 4.6 call, shows the per-token-type prices, explains cache hit ratio, and simulates how the cost changes as that ratio moves. By the end, you should understand where the money goes and which habits reduce it.

Measured Cost Breakdown (Single API Call)

The table below splits one API call into its four token types. The per-token prices are estimates derived from the billed amounts, so they should be read as approximations rather than published rates.

Item Tokens Cost Price per token (estimated)
Input 15 $0.000 $15/1M tokens
Output 2,241 $0.168 ~$75/1M tokens
Cache Read 219,064 $0.329 ~$1.50/1M tokens
Cache Write 111,033 $2.082 ~$18.75/1M tokens
Total 332,353 $2.58 —

The call used 332,353 tokens in total, but almost none of them were the new question. Input was only 15 tokens. Most of the volume was cached content: 219,064 tokens read from cache and 111,033 tokens written to it. That split is the key to the whole cost picture. The question itself was negligible, and the bill was driven by the material the model had to work with.

Cost Share vs. Token Share

Comparing where the tokens are with where the money goes shows why the cost looks surprising.

Item Share of cost Share of tokens
Cache Write 80.7% 33.4%
Cache Read 12.8% 65.9%
Output 6.5% 0.7%
Input 0.0% 0.0%

Two patterns stand out:

  • About 80% of the cost is Cache Write. This is the cost of putting large data into the cache for the first time.
  • About 66% of the tokens are Cache Read. These are cached tokens reused in the call. Because their unit price is low, they make up only about 13% of the cost.

Cache Write carries the highest per-token price among the cached token types, so a large first load dominates the bill even when most of the volume is reads. Output is also priced well above cached reads, but it contributed only 0.7% of the tokens, which keeps its share of the cost at 6.5%. The practical reading is that volume alone does not determine cost. The price per token type matters just as much, and the cheapest tokens in this call were the ones that appeared most often.

Cache Hit Ratio

Definition: the share of data requests served directly from cache (hits) out of all data requests.

$$\text{Cache hit ratio} = \frac{\text{Cache Read tokens}}{\text{Cache Read tokens} + \text{Cache Write tokens}} \times 100$$

Ratio for this call:

$$\frac{219,064}{219,064 + 111,033} \times 100 = 66.4%$$

The formula counts only cache reads and cache writes, not Input or Output tokens. That is deliberate. The ratio measures how well the cache is being reused for the bulky material, which is where the cost sits. A high ratio means most of the cached content was pulled from cache rather than loaded again. A low ratio means the system keeps paying the more expensive Cache Write price to load content that it could have reused.

Cost Simulation by Cache Hit Ratio (Based on 332,000 Tokens)

The simulation holds the total volume at roughly 332,000 tokens and changes only the cache hit ratio. The Cache Read and Cache Write costs in each row follow from that ratio, and the totals are estimates.

Cache hit ratio Cache Read cost Cache Write cost Total cost (estimated)
0% (no cache) $0 ~$6.19 ~$6.37
50% $0.25 ~$3.09 ~$3.52
66.4% (measured) $0.33 ~$2.08 $2.58
90% $0.45 ~$0.62 ~$1.25
  • Moving from a 0% hit ratio to the measured 66.4% saves about 60% of the cost.
  • Moving to a 90% hit ratio saves about 80% of the cost.

The 0% row is the most instructive. With no cache, the same volume would cost roughly $6.37, about 2.5 times the measured total. Caching is therefore not a minor optimization in this case. It is the difference between a call that costs a few dollars and one that costs several times that. The 90% row shows the other side: when most of the cached content is reused, the Cache Write share shrinks and the total falls to around $1.25.

Insights

1. Cache Write is the main cost driver

The question was only 15 tokens, yet it cost $2.58. The cause was the 111,033 Cache Write tokens. Cost concentrates where large documents or code are loaded for the first time. The first call is the most expensive, and later repeated calls become cheap because they read from cache. This means the first request in a workflow often looks disproportionately expensive, and that impression is accurate: it is paying for the setup that makes later requests cheap.

2. Cache hit ratio is the key efficiency metric

The more times you ask questions about the same context, the higher the hit ratio and the lower the cost. The pattern is expensive at the start of a session and gets cheaper as the session continues. Put simply, the first question about a document is expensive, and the rest are cheap. If you want a single number to watch when you judge whether a workflow is efficient, the hit ratio is a better guide than the length of any individual prompt.

3. An analogy for non-developers

  • Cache Write is the cost of copying a thick book and placing it on the AI’s desk.
  • Cache Read is the cost of the AI flipping through a book already on its desk, which is much cheaper.
  • A high cache hit ratio means you place the book once and ask many questions about it, which is efficient.

The analogy also explains why the prices differ. Copying the book takes more effort than looking through a copy that is already open, and the pricing reflects that difference. It also explains why a short question can be expensive. The question is small, but the book it refers to may be very large.

4. Practical optimization tips

  • When analyzing the same document, ask all your questions within one session.
  • If you end sessions often, you pay the Cache Write cost again each time.
  • Claude Code works the same way: CLAUDE.md and project files are cached automatically.

These tips follow directly from the cost structure. Splitting one document across many short sessions multiplies the expensive step, while keeping related questions together lets the cached material do its job. The same logic applies to Claude Code, where project context such as CLAUDE.md and project files is cached automatically, so the way you structure sessions affects how often that material has to be written to the cache again.

Bottom line

For this call, the cost came almost entirely from loading a large body of content into the cache, not from the short question. Cache Write accounted for $2.082 of $2.58, while Cache Read tokens, which were about two-thirds of the volume, cost only $0.329. The simulation shows that a higher cache hit ratio lowers the total: from an estimated $6.37 with no cache, to $2.58 at the measured 66.4%, to about $1.25 at 90%. If you plan to ask several questions about the same material, keep them in one session so the expensive first load is paid once.

Sources

Frequently asked questions

Why did a 15-token question cost $2.58?
Most of the cost came from 111,033 Cache Write tokens, which accounted for $2.082 of the $2.58 total. Cache Write is the cost of loading a large document into the cache for the first time.
How is cache hit ratio calculated, and how much can it save?
Cache hit ratio is Cache Read tokens divided by the sum of Cache Read and Cache Write tokens, times 100. In this case it was 66.4%. In the simulation, raising the ratio from 0% to 66.4% cut the cost by about 60%, and raising it to 90% cut it by about 80%.