Claude Code Token Optimization: Why a 2-Hour Session Exhausted the Max x5 Budget
Who this is forDevelopers and content creators who use Claude Code for large multi-file edits or long PDF processing and want to control token costs.
If you run Claude Code on long, multi-file jobs, you may hit your usage limit much sooner than you expect. On April 6, 2026, I ran two unrelated jobs in the same session: revising a large set of generative AI lecture materials and extracting a back-of-book index from a PDF of more than 300 pages. Within two hours, the max x5 token budget was fully used up. This article walks through what consumed the tokens, how I measured the usage, and the strategy I derived to prevent it. You will get a method for estimating tokens per task, a split-and-delegate pipeline for PDF work, and practical rules for when to switch models or start a new session.
Summary
Running generative AI lecture-material production (wrtn_contents) and PDF indexing (n8n-book-index) at the same time exhausted the max x5 token budget within two hours. Analyzing the cause produced a framework for quantifying tokens per task and a set of improvement strategies.
Key Data
The Problem (April 6, 2026 Session)
Two tasks ran concurrently, and together they used up the max x5 token budget within two hours.
Task A: wrtn_contents Lecture Materials (The Main Token Consumer)
| Time | Task | Files Changed | Notes |
|---|---|---|---|
| 10:31 | Fact-checking the generative AI lecture plan, splitting it into six parts, and linking scripts | 70 | Large-scale file creation and modification |
| 11:47 | Fixing page layout drift in the lecture plan | 73 | Recognized token shortage; requested delegation to Codex |
| 11:55 | Delegated the remaining work to Codex | 73 | Saves Claude tokens |
- The lecture structure was: fact-check the existing lecture materials, split them into six parts, and map each part to its script.
- The problem was that the script pages had drifted out of alignment. The whole layout had to be rearranged, which meant revising 70 to 73 files over and over.
- Main inefficiency: every revision resent the full file context cumulatively.
Task B: n8n-book-index PDF Indexing
| Item | Value |
|---|---|
| Task | Extracting the back-of-book index from the PDF of the Korean-language n8n book published by Hanbit Media (a Korean tech publisher) |
| PDF text size | Approximately 259KB |
| Book length | 300+ pages |
Token Consumption Analysis (April 7, 2026 Session, Measured)
| Model | Cache Reads | Cache Writes | Notes |
|---|---|---|---|
| Claude Opus | 4,379,381 tokens | 228,823 tokens | Main source of consumption |
| Claude Sonnet | 538,872 tokens | 93,650 tokens | Secondary work |
| Total | 5,264,903 tokens | — | Estimated cost $12.54 |
Causes of Token Inefficiency
- Repeated edits to many files (lecture materials): Modifying 70 to 73 files several times meant the full context of each file was resent cumulatively on every change. Cache-read tokens grew sharply; Opus cache reads alone reached 4,379,381 tokens.
- Inefficient Korean tokenization: The note reports that Korean uses two to three times as many tokens as English for the same text.
- Two tasks in one session: Lecture materials and PDF indexing drew from the same token budget at the same time.
- Model selection: Mechanical work such as rearranging page layouts ran on Opus, the most expensive option for that kind of task.
Quantifying Work Units
| Criterion | Value |
|---|---|
| 1K tokens (Opus) | About 5% of the total budget |
| 1 million tokens (Opus) | About 15% of the total budget |
| 1 million tokens ≈ | About 750 A4 pages of text |
| Recommended work unit | Set work budgets in units of 1 million tokens |
The two per-unit figures do not scale linearly with each other, so read them as rough observations from this session rather than a conversion rate. The practical value is the recommended unit. If you set a work budget in 1-million-token increments before starting, you can judge whether a task fits within the budget you have left.
Insights
1. Model Selection Is the Core Lever
- Opus: Work that requires judgment, such as analysis, decisions, and architecture design
- Sonnet: Execution-oriented work, such as text extraction, code generation, and documentation
- You can switch models mid-session in Claude Code with the
/modelcommand.
The principle is to match the model to the kind of thinking a task needs. Rearranging page layouts requires no architectural judgment, so running it on Opus spends the most expensive tokens on the least demanding work.
2. PDF Processing: Split and Delegate
For token-intensive work like PDF indexing, the best pipeline I found has three steps:
- Split the PDF into chapter or section units and extract the text locally with Python.
- Delegate term extraction for each section to Codex, since it does not face the same token constraints.
- Have Claude handle only the merging, de-duplication, and sorting of the results.
This keeps the expensive model out of the bulk-reading step and reserves it for the part of the job that benefits from judgment.
3. Automate Per-Session Cost Tracking
- Add a cost tab to the Obsidian session log.
- Remove the dependency on n8n and switch to a local hook approach.
- Automatically save token usage at the end of every session.
Measuring usage per session turns cost from a vague feeling into a number you can compare across tasks.
4. Delegate Bulk File Edits to Codex First
Lecture materials are a case where dozens of files are revised repeatedly. Claude Opus is weak at this pattern because every turn includes the entire set of target files in context, so token use grows rapidly. The fix is a division of labor: Claude defines the rules and patterns, and Codex applies them in bulk. This switch was effective in the 11:55 session.
5. Practical Guidelines
- Before processing large text (200KB or more), design a splitting strategy before you start.
- If you expect to modify 20 or more files, delegate to Codex or switch to Sonnet.
- Separate heterogeneous work, such as lecture materials and PDF indexing, into different sessions.
- Avoid single sessions longer than two hours, because accumulated context cost rises sharply.
- Before starting, ask: “Is this a tens-of-thousands-token job?” Make the estimate a habit.
Bottom Line
The evidence in this note points to one conclusion: the token budget is consumed mainly by how much context is resent on each turn, not by how much work you think you are doing. Repeated edits across dozens of files, and mixing unrelated tasks in one long session, multiply that cost. The measurements support four fixes: keep sessions focused, match the model to the task type, delegate bulk edits once the rules are defined, and set a token budget before you start.
Sources
- Measured data: 2026-04-07 Obsidian session records (
2026-04-07-Projects-0f4f16e4.md,2026-04-07-wrtn_contents-ad68db76.md) - Lecture-material sessions: Three wrtn_contents sessions from April 6, 2026 (10:31, 11:47, 11:55)
- Project folders:
~/Projects/n8n-book-index/,~/Projects/wrtn_contents/ - Token usage: Output of the
/costcommand within the session
Frequently asked questions
- Why did the token budget run out so quickly?
- The main causes were repeated edits across 70 to 73 files, where the full file context was resent on every change, and running a heavy lecture-material task and PDF indexing in the same session.
- What should I do instead?
- Keep unrelated tasks in separate sessions, switch to Sonnet or delegate when a change touches 20 or more files, have Claude define the rules while Codex applies them in bulk, and budget work in 1-million-token units.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›