Insights

Fable 5 vs Opus 4.8: Real Session Cost and Agent Behavior Compared

6 min read#claude#llm-cost#agentic-coding#model-comparison#fable-5

Who this is forDevelopers and automation builders who run long autonomous Claude sessions and need to estimate real cost per session rather than per-token price.

If you run Claude for long coding or automation sessions, the per-token price is the wrong number to plan around. Fable 5 is priced at exactly twice the token rate of Opus 4.8, yet the measured sessions cost between 1.0 and 1.4 times as much. This article walks through the token, cost, and tool-call data behind that gap, along with what the sessions show about how each model behaves when it works autonomously. You will see where the cost difference comes from, which kinds of tasks benefit most, and where the evidence stops.

Summary

Fable 5 is priced at $10/$50 per 1M tokens, while Opus 4.8 is priced at $5/$25 per 1M tokens. That makes the Fable rate exactly 2x the Opus rate. Even so, the measured session cost was only 1.0 to 1.4 times the Opus cost. The reason is that Fable produced 40 to 60 percent fewer output tokens on the same workload. Since the conversation context grew more slowly, fewer tokens were also read from cache on each turn. On autonomy, the Fable sessions completed building, deploying, and verifying their output with only four to five user interventions.

Test sessions

The comparison uses two Fable sessions and a comparison group of five recent Opus sessions. The Opus sessions ran between June 4 and June 9, 2026, and covered three Discord automation runs, a security check, and social media collection. The Fable sessions were each about 80 minutes long, which put them in the same time range as the longer Opus sessions.

Category Task Duration Result
Fable ① Interactive Seoul apartment actual-transaction site (public API → build → GitHub Pages deploy) 80 min Live deployment completed (ggplab.github.io/seoul-apt-trends)
Fable ② Research on an air-gapped n8n delivery playbook → docx → published to Google Docs 82 min Published + formatting verified + saved to memory
Opus comparison group Five most recent sessions (three Discord automation runs, a security check, social media collection), June 4–9, 2026 33–83 min -

The Fable ② session published its output to Google Docs and then opened the uploaded result to check formatting. The Fable ① session deployed a live site. Both tasks were long, multi-step, and tool-heavy, which is why they make a useful test of autonomous behavior.

Token usage patterns

The token table below compares the two Fable sessions with the two Opus sessions that were closest in duration and tool-call volume: the 83-minute market report and the 70-minute security check.

Metric Fable ① Fable ② Opus market report (83 min) Opus security check (70 min)
Output tokens 171K 223K 552K 500K
Cache reads 32.4M 23.6M 70.9M 25.6M
Assistant messages 270 169 342 174
Output tokens per message 633 1,322 1,614 2,873
Tool calls 148 78 131 65
Tool errors 14 5 4 3

When sessions with similar duration and tool-call counts are compared, Fable’s output tokens come to roughly one-third to one-half of Opus’s. This matters for cost beyond the output line itself. Each assistant turn carries the accumulated conversation forward. Less output means less context added for the next turn, so the cache-read volume shrinks as well. The cache-read figures in the table show this pattern: the Fable ① session read 32.4M tokens from cache, compared with 70.9M for the Opus market report.

The per-message figures make the same point from a different angle. The Fable ① session averaged 633 output tokens per message, while the Opus market report averaged 1,614 and the Opus security check averaged 2,873.

Cost

The cost figures below are API-equivalent estimates, not invoices. They assume that cache writes cost 1.25 times the standard rate.

Session Cost Cost per minute
Fable ① (80 min) $51.2 $0.64
Fable ② (82 min) $42.7 $0.52
Opus market report (83 min) $61.9 $0.75
Opus security check (70 min) $33.7 $0.48
Opus Modusign (47 min) $23.7 $0.50
Opus social media collection (33 min) $9.4 $0.29

Modusign is a South Korean e-signature service. Its session is included here because it was one of the Opus comparison runs.

Several patterns stand out in this table:

  • Cache reads, which are independent of the model, make up 55 to 63 percent of total cost. This makes them the largest single cost item in every session.
  • In the first like-for-like pair, Fable ① came in at $51.2 and the Opus market report came in at $61.9. Fable was therefore 17 percent cheaper. Both sessions ran about 80 minutes and used more than 130 tool calls.
  • In the second pair, Fable ② came in at $42.7 and the Opus security check came in at $33.7. Fable was therefore 27 percent more expensive.
  • Taken together, session cost was close to parity, within ±30 percent. The doubled unit price did not translate directly into doubled session cost.

The per-minute column shows the same spread. Fable sessions ranged from $0.52 to $0.64 per minute, while the Opus sessions ranged from $0.29 to $0.75 per minute. The per-minute number depends heavily on what the session does, so it is more useful for comparing similar tasks than for judging the model itself.

Stability

Tool-call error rates were higher for Fable overall: 8.4 percent (19 of 226 calls) for Fable compared with 4.2 percent (14 of 334 calls) for Opus. Most of the Fable ① errors came from Playwright browser automation, which is an area that is unstable by nature. The Fable ② session had no browser work, and its error rate was 6.4 percent, which is close to the Opus average.

This means the higher overall error rate for Fable is largely explained by the browser-heavy task rather than by the model’s general reliability in tool use. The sample is too small to separate the two effects cleanly, but the pattern is consistent across the two Fable sessions.

Behavior insights

1. Completed with few interventions. The Fable ① session needed only four real user messages: the initial request, a label fix, a request for recent data, and a wrap-up. Those messages covered data collection, deployment, and screenshot verification. For an 80-minute build-and-deploy task, four interventions is markedly fewer than in the Opus sessions.

2. Verified its own work. After deployment, the Fable ① session checked the live site itself, taking 14 Playwright screenshots. The Fable ② session opened its Google Docs upload to check formatting, and it saved a pitfall it found to memory. A verify-each-step behavior rule worked in both sessions without any separate enforcement.

3. Less narration, same pace. Output tokens per message were less than half of the Opus figures. That means much less relay commentary and exposed reasoning between tool calls. Tool-call density per hour stayed similar or rose: the Fable ① session made 111 calls per hour, compared with 95 per hour for the Opus sessions. The pattern is less talk with the same amount of work.

4. Where the advantage fits. Long autonomous work, such as building and deploying an app or researching and publishing a document, is where Fable’s behavior looks most useful. In these tasks, the cost does not rise by the full unit price, and the number of interventions falls. Short, single-shot tasks may offer less advantage over Opus, because there is less accumulated context for Fable’s lower output to save on.

Limitations

  • The sample is two Fable sessions against five Opus sessions, and the tasks differ. This is not a same-task A/B test, so a quality advantage cannot be concluded. The supported claim is that Fable completed its tasks with fewer user interventions.
  • The cost figures are API-equivalent. Under a subscription plan, they indicate how quickly usage limits are consumed, not the amount actually billed.
  • Cache writes are assumed to use the 5-minute TTL, which carries the 1.25 times multiplier. If some writes use the 1-hour TTL, which costs 2 times, both models’ costs would rise slightly.

Bottom line

The evidence supports a narrow conclusion. Fable 5 is priced at twice Opus 4.8’s token rate, but measured session cost stayed close to Opus, at 1.0 to 1.4 times. The main reason is that Fable produced fewer output tokens, which reduced the context carried forward and the cache reads that followed. In the long build-and-publish sessions measured here, Fable also finished with few user interventions and verified its own output. The data does not establish a quality difference, and the advantage is likely smaller for short, single-shot tasks.

Sources

  • Local session transcripts (~/.claude/projects/): Fable ① alice-samsung-2606/7315e399, Fable ② tech-research-hub/0eeb0838, Opus comparison group task-orchestrator/{061cf967, 9ef35d4a, c81dea60, af3700b0}, ggplab-ops/74e67e18
  • Model pricing: Anthropic official documentation (model catalog from the claude-api skill, cached as of May 26, 2026) https://platform.claude.com/docs/en/pricing

Frequently asked questions

Why doesn't Fable 5's doubled token price double the session cost?
Fable 5 wrote 40 to 60 percent fewer output tokens for the same workload. Because context accumulated more slowly, it also read less from cache. Measured session costs came out at 1.0 to 1.4 times Opus 4.8.
Does this comparison prove Fable 5 produces better results?
No. The sample was two Fable sessions against five Opus sessions on different tasks, not a controlled A/B test. The data supports only the claim that Fable completed long work with fewer user interventions, not a quality advantage.