Insights

Claude Code Token Optimization: Why a 2-Hour Session Exhausted the Max x5 Budget

5 min read#claude-code#token-optimization#cost-analysis#model-selection#productivity

Who this is forDevelopers and content creators who use Claude Code for large multi-file edits or long PDF processing and want to control token costs.

If you run Claude Code on long, multi-file jobs, you may hit your usage limit much sooner than you expect. On April 6, 2026, I ran two unrelated jobs in the same session: revising a large set of generative AI lecture materials and extracting a back-of-book index from a PDF of more than 300 pages. Within two hours, the max x5 token budget was fully used up. This article walks through what consumed the tokens, how I measured the usage, and the strategy I derived to prevent it. You will get a method for estimating tokens per task, a split-and-delegate pipeline for PDF work, and practical rules for when to switch models or start a new session.

Summary

Running generative AI lecture-material production (wrtn_contents) and PDF indexing (n8n-book-index) at the same time exhausted the max x5 token budget within two hours. Analyzing the cause produced a framework for quantifying tokens per task and a set of improvement strategies.

Key Data

The Problem (April 6, 2026 Session)

Two tasks ran concurrently, and together they used up the max x5 token budget within two hours.

Task A: wrtn_contents Lecture Materials (The Main Token Consumer)

Time Task Files Changed Notes
10:31 Fact-checking the generative AI lecture plan, splitting it into six parts, and linking scripts 70 Large-scale file creation and modification
11:47 Fixing page layout drift in the lecture plan 73 Recognized token shortage; requested delegation to Codex
11:55 Delegated the remaining work to Codex 73 Saves Claude tokens
  • The lecture structure was: fact-check the existing lecture materials, split them into six parts, and map each part to its script.
  • The problem was that the script pages had drifted out of alignment. The whole layout had to be rearranged, which meant revising 70 to 73 files over and over.
  • Main inefficiency: every revision resent the full file context cumulatively.

Task B: n8n-book-index PDF Indexing

Item Value
Task Extracting the back-of-book index from the PDF of the Korean-language n8n book published by Hanbit Media (a Korean tech publisher)
PDF text size Approximately 259KB
Book length 300+ pages

Token Consumption Analysis (April 7, 2026 Session, Measured)

Model Cache Reads Cache Writes Notes
Claude Opus 4,379,381 tokens 228,823 tokens Main source of consumption
Claude Sonnet 538,872 tokens 93,650 tokens Secondary work
Total 5,264,903 tokens — Estimated cost $12.54

Causes of Token Inefficiency

  1. Repeated edits to many files (lecture materials): Modifying 70 to 73 files several times meant the full context of each file was resent cumulatively on every change. Cache-read tokens grew sharply; Opus cache reads alone reached 4,379,381 tokens.
  2. Inefficient Korean tokenization: The note reports that Korean uses two to three times as many tokens as English for the same text.
  3. Two tasks in one session: Lecture materials and PDF indexing drew from the same token budget at the same time.
  4. Model selection: Mechanical work such as rearranging page layouts ran on Opus, the most expensive option for that kind of task.

Quantifying Work Units

Criterion Value
1K tokens (Opus) About 5% of the total budget
1 million tokens (Opus) About 15% of the total budget
1 million tokens ≈ About 750 A4 pages of text
Recommended work unit Set work budgets in units of 1 million tokens

The two per-unit figures do not scale linearly with each other, so read them as rough observations from this session rather than a conversion rate. The practical value is the recommended unit. If you set a work budget in 1-million-token increments before starting, you can judge whether a task fits within the budget you have left.

Insights

1. Model Selection Is the Core Lever

  • Opus: Work that requires judgment, such as analysis, decisions, and architecture design
  • Sonnet: Execution-oriented work, such as text extraction, code generation, and documentation
  • You can switch models mid-session in Claude Code with the /model command.

The principle is to match the model to the kind of thinking a task needs. Rearranging page layouts requires no architectural judgment, so running it on Opus spends the most expensive tokens on the least demanding work.

2. PDF Processing: Split and Delegate

For token-intensive work like PDF indexing, the best pipeline I found has three steps:

  1. Split the PDF into chapter or section units and extract the text locally with Python.
  2. Delegate term extraction for each section to Codex, since it does not face the same token constraints.
  3. Have Claude handle only the merging, de-duplication, and sorting of the results.

This keeps the expensive model out of the bulk-reading step and reserves it for the part of the job that benefits from judgment.

3. Automate Per-Session Cost Tracking

  • Add a cost tab to the Obsidian session log.
  • Remove the dependency on n8n and switch to a local hook approach.
  • Automatically save token usage at the end of every session.

Measuring usage per session turns cost from a vague feeling into a number you can compare across tasks.

4. Delegate Bulk File Edits to Codex First

Lecture materials are a case where dozens of files are revised repeatedly. Claude Opus is weak at this pattern because every turn includes the entire set of target files in context, so token use grows rapidly. The fix is a division of labor: Claude defines the rules and patterns, and Codex applies them in bulk. This switch was effective in the 11:55 session.

5. Practical Guidelines

  • Before processing large text (200KB or more), design a splitting strategy before you start.
  • If you expect to modify 20 or more files, delegate to Codex or switch to Sonnet.
  • Separate heterogeneous work, such as lecture materials and PDF indexing, into different sessions.
  • Avoid single sessions longer than two hours, because accumulated context cost rises sharply.
  • Before starting, ask: “Is this a tens-of-thousands-token job?” Make the estimate a habit.

Bottom Line

The evidence in this note points to one conclusion: the token budget is consumed mainly by how much context is resent on each turn, not by how much work you think you are doing. Repeated edits across dozens of files, and mixing unrelated tasks in one long session, multiply that cost. The measurements support four fixes: keep sessions focused, match the model to the task type, delegate bulk edits once the rules are defined, and set a token budget before you start.

Sources

  • Measured data: 2026-04-07 Obsidian session records (2026-04-07-Projects-0f4f16e4.md, 2026-04-07-wrtn_contents-ad68db76.md)
  • Lecture-material sessions: Three wrtn_contents sessions from April 6, 2026 (10:31, 11:47, 11:55)
  • Project folders: ~/Projects/n8n-book-index/, ~/Projects/wrtn_contents/
  • Token usage: Output of the /cost command within the session

Frequently asked questions

Why did the token budget run out so quickly?
The main causes were repeated edits across 70 to 73 files, where the full file context was resent on every change, and running a heavy lecture-material task and PDF indexing in the same session.
What should I do instead?
Keep unrelated tasks in separate sessions, switch to Sonnet or delegate when a change touches 20 or more files, have Claude define the rules while Codex applies them in bulk, and budget work in 1-million-token units.