Karpathy CLAUDE.md vs Anthropic's Opus 4.7 Guide: What to Add to Your Claude Code Setup
Who this is forDevelopers and solo builders who run Claude Code daily and want to tune their CLAUDE.md files and agent rules for Opus 4.7.
If you use Claude Code every day, your CLAUDE.md file and agent rules shape how the model behaves on each task. Two sources offer guidance that overlaps but does not match. One is the four principles in Andrej Karpathy’s CLAUDE.md, as adapted by Forrest Chang. The other is Anthropic’s official prompting guide for Opus 4.7. This article maps the first against the second, shows where the principles fall short, and lists six concrete changes you can make to a working setup. You will get a clear picture of which gaps matter and where each change belongs.
Summary
The four Karpathy/Chang principles are a subset of Anthropic’s Opus 4.7 guide. Two of them, “state assumptions” and “verify step by step,” are worth writing down as explicit rules, because the principles state them more clearly than Anthropic’s sections do.
Running Opus 4.7 well also requires operational rules that the four principles do not cover. These include effort-level operation, emphasis tone, avoiding frontend defaults, and tracking state across multiple context windows. These belong to the API and harness layer, so they need their own standard operating procedures (SOPs).
Core Diagram

Core Data
1. Mapping the Karpathy Principles to the Anthropic Guide
| Karpathy principle | Corresponding section in the Anthropic guide | Relationship |
|---|---|---|
| #1 Think Before Coding (state assumptions, branch on interpretations, stop when uncertain) | “Be clear and direct” plus the hypothesis tracking in “Investigate before answering” and “Long-horizon reasoning” | Partial match. Karpathy is clearer on stating assumptions and presenting branches. |
| #2 Simplicity First (no features beyond the request, 200 lines down to 50) | “Overeagerness / Avoid overengineering” section (lines 731–747) | Nearly identical. Anthropic’s sample prompt is more specific, with four axes: Scope, Documentation, Defensive coding, and Abstractions. |
| #3 Surgical Changes (no improving adjacent code, keep existing style) | “More literal instruction following,” “Reduce file creation,” and the literal behavior changes in 4.7 | Naturally satisfied. Opus 4.7 interprets instructions narrowly by default. |
| #4 Goal-Driven Execution (verifiable success criteria, step-by-step verification) | Success criteria in “Research and information gathering” and the tests.json approach in “Multi-context window” | Partial match. Karpathy is clearer on step-by-step verification. |
The table shows that Karpathy’s principles are not a replacement for Anthropic’s guide. Principles #2 and #3 already have close counterparts in Anthropic’s text. Principles #1 and #4 are where Karpathy’s wording adds something, because the Anthropic sections imply assumption tracking and verification without stating them as firm rules.
2. Items the Anthropic 4.7 Guide Adds: Status in a User System
The second table lists what the Anthropic guide adds beyond the four principles, and whether a typical user system already reflects each item. The status markers are: ✅ reflected, ⚠️ partly reflected, ❌ not reflected.
| Area | Key guidance | Status in user system |
|---|---|---|
| effort-level operation | low/medium/high/xhigh/max. Coding and agentic work should use xhigh. |
⚠️ Mentioned in memory, but no SOP in CLAUDE.md or AGENT_GUIDE |
| literal interpretation | “Fix this section” applies only to that section. State the scope explicitly (“every section”). | ⚠️ No explicit rule (memory reference only) |
| response length tuning | If you want concise output, say so. “Skip non-essential context.” | ❌ No explicit rule |
| emphasis tone (MUST/CRITICAL) | Triggers over-triggering. “Use this tool when…” is usually enough. | ❌ Many Korean must-type phrasings (the words for “must” and “required”) across existing CLAUDE.md files |
| parallel tool calling | Run calls with no dependencies in parallel. Use the <use_parallel_tool_calls> pattern. |
✅ Included in Claude Code’s default behavior. No separate SOP needed. |
| subagent spawn control | 4.7 spawns subagents less by default. State when you want them. | ❌ No explicit rule |
| code review harness recall | A single “high-severity only” instruction lowers recall. Separate coverage from post-processing filtering. | ❌ Not applied to /review or /security-review |
| frontend default avoidance | Defaults to cream, serif, and terracotta. Specify a concrete palette. | ⚠️ A design system exists, but no rule requires stating a design input |
| code and file creation restraint | “Reduce file creation.” Clean up temporary files explicitly. | ⚠️ Scattered in AGENT_GUIDE (for example, Codex delegation rules) |
| anti-overengineering | Four axes: Scope, Documentation, Defensive coding, Abstractions | ❌ No explicit rule (pair with Karpathy #2) |
| hallucination prevention | “Never speculate about code you have not opened. MUST read the file before answering.” | ⚠️ Partial. Data integrity rules exist, but no explicit “open before claim” rule |
| multi-context window state | tests.json, progress.txt, and git as the long-horizon workflow | ❌ No explicit rule |
| balancing autonomy and safety | Consider reversibility. Confirm destructive actions. | ✅ Partly reflected (AGENT_GUIDE section 12 on shell context preservation and data integrity) |
| prefilled response deprecation | Not supported from 4.6 onward. Replaced by Structured Outputs. | ✅ No effect on the user workflow |
| progress update automation | 4.7 handles this automatically. Remove scaffolding such as “After every 3 tool calls, summarize.” | ⚠️ Possibly still present in some skills and rules. Needs a grep check. |
Most of the ❌ rows share one cause: the user system has rules for what to do but not for how strongly or how specifically to phrase them. The emphasis-tone, subagent, and state-tracking rows all depend on wording and structure, not on new capabilities.
3. Where the Two Sources Disagree
| Issue | Karpathy CLAUDE.md | Anthropic Opus 4.7 guide |
|---|---|---|
| Pausing at the model call | Makes “stop when uncertain” the first principle | Provides two contrasting samples, <do_not_act_before_instructions> and <default_to_action>. Treats the choice as the user’s explicit decision. |
| Emphasis words | Frequently uses “MUST” and “NEVER” | Warns since 4.5 about over-triggering and recommends gentler instructions |
| Tool calls | No explicit guidance | 4.7 makes fewer tool calls. Raise the effort level or state the need explicitly. |
| Design and frontend | Not covered | Recommends specifying concrete palettes, fonts, and motion to avoid “AI slop” |
The disagreement on emphasis words matters most in practice. Karpathy’s style is forceful, and that works for a model that under-applies rules. Opus 4.7 tends in the opposite direction, so the same forceful wording can cause over-application.
Insights
Below are six concrete changes to apply to your own Claude Code system. Each one states what to change, where to put it, and why.
1. A New File for Assumption, Interpretation, and Verification Behavior
(a) What: Create a new behavior rules file that condenses Karpathy’s principles #1 and #4 into a short set of rules. The core has three parts.
# Behavior — assumptions, interpretation, verification
## State assumptions
If a request is ambiguous, before starting work, echo in 1–3 lines
(a) which assumptions you are proceeding under and
(b) what other reasonable interpretations exist.
## Present interpretation branches
If two or more clear interpretations exist, do not pick one silently.
Show them as "Option A / Option B" and wait for a "go" signal.
## Verify step by step
For multi-step work, define in advance how each step will be verified,
and actually run that verification before moving to the next step.
After a deployment or an external call, confirm with a real check
(curl, fetch, or execution).
(b) Where: Create ~/.claude/rules/behavior.md. Add only the import path to ~/CLAUDE.md or ~/Projects/shared/AGENT_GUIDE.md.
(c) Why: Your memory already contains a rule called feedback_scope_echo_before_delegate, which echoes scope before delegating. The same pattern is needed for ordinary work that is not delegated. The “verify after deployment” rule in deployment.md applies only to deployments, so the duty to verify is missing for code, documents, and research in general. Karpathy’s principles #1 and #4 fill exactly this gap.
2. Formalizing an Effort-Level Operating SOP
(a) What: Formalize an effort-level matrix by task type as an SOP table.
| Task type | Recommended effort | Reason |
|---|---|---|
| Simple lookup, short answers | low/medium | 4.7 adjusts automatically to task complexity |
| Content writing, research, general analysis | high | Minimum level for intelligence-sensitive work |
| Coding, agentic loops, multi-step work | xhigh | Default for coding and agentic work in the official guide |
| Critical decisions, complex debugging | max | Use sparingly, with awareness of overthinking risk |
(b) Where: Add an “effort-level operation” subsection inside section 12 (operating rules) of ~/Projects/shared/AGENT_GUIDE.md. The memory file reference_opus_4_7_prompting_guide.md is a reference, so it is not suitable as an SOP. AGENT_GUIDE serves as the single source of truth (SSOT).
(c) Why: A memory reference tells you that the guidance exists, but it does not settle which effort level to use at the start of each task, so you have to decide again every time. With an SOP table, both you and Claude can answer immediately. As a solo operator, you are sensitive to token costs, and running a simple lookup at xhigh wastes money.
3. Auditing and Downgrading Emphasis Words
(a) What: Search existing CLAUDE.md files, AGENT_GUIDE.md, memory, and skills for emphasis patterns, then review each match. Include the Korean must-type and never-type phrasings in your search, since the English patterns below will not catch them:
# Audit command (add the Korean equivalents of must/required/never to the pattern list)
rg -n "MUST|CRITICAL|NEVER|ALWAYS|required|mandatory" \
~/CLAUDE.md \
~/CLAUDE.local.md \
~/Projects/shared/AGENT_GUIDE.md \
~/.claude/rules/ \
~/.claude/skills/
Classify each match as either (i) truly critical (data loss or security) or (ii) at risk of over-triggering because an ordinary guideline was emphasized. Downgrade the items in (ii) to a neutral tone. For example:
- “Must follow the structure in templates/research-template.md” → “Follow the templates/research-template.md structure”
- “Never fabricate, estimate, or use placeholder data” → keep as is (data integrity is genuinely critical)
(b) Where: Edit existing files in place. Do not do everything at once. Split the work across one or two sessions per file.
(c) Why: The Anthropic guide warns consistently across 4.5, 4.6, and 4.7 (line 492). The more emphasis words there are, the more Opus 4.7 over-triggers. It reads an ordinary guideline as “this is critical, so apply it to every decision.” Your system already records concerns about context accumulation and token efficiency over roughly two-hour sessions, and reducing emphasis words is a cost reduction of the same kind.
4. Review Work: Separating Coverage from a Post-Processing Filter
(a) What: For review tasks (code review, content review, document review), replace a single conservative instruction with a two-step pattern.
# Step 1 — coverage mode (collect every issue)
"Report every issue you find with confidence (low/med/high) and severity
(info/warn/error/critical) tags. Do not pre-filter."
# Step 2 — post-processing filter (separate step)
"From the previous list, surface only items where severity ≥ warn AND
confidence ≥ med. Group by file."
(b) Where: Apply the same two-step pattern to the system prompts of review-type skills in ~/.claude/skills/, starting with code-review, simplify, review-video, and review-inquiry. Also add one line to the AGENT_GUIDE.md operating rules: “For review tasks, separate coverage from filtering.”
(c) Why: Lines 148–163 of the Anthropic guide explicitly warn about this. A single conservative instruction such as “only report high-severity issues” lowers recall in 4.7. Among your workflows, /security-review and simplify follow exactly this pattern, so checking them is worthwhile.
5. An SOP for Frontend and Design Work
(a) What: When starting design, frontend, slide, or diagram work, require the following inputs explicitly.
# Design brief input checklist
- [ ] Palette: tokens.js from ggplab-design-system, or 3–5 explicit hex values
- [ ] Fonts: specify body, heading, and monospace (for example, Pretendard, Georgia, SF Mono)
- [ ] Identity tone: one of editorial / fintech / dashboard / portfolio
- [ ] One or two reference images or URLs (optional)
If only vague instructions such as "clean," "minimal," or "stylish" are given
without the information above, Claude falls back to the 4.7 default
(cream, serif, terracotta). Before starting, ask the user for the four items
above, or apply ggplab-design-system as the default.
(b) Where: Add a “frontend and design work” section at the end of the existing ~/.claude/rules/content-creation.md. Alternatively, separate it into a new ~/.claude/rules/design.md.
(c) Why: You produce many design outputs, including the ggplab.xyz site, the content-designer-challenge site, slide decks, diagrams, and Notion pages. Lines 97–141 of the Anthropic guide state that 4.7 tends to drift toward cream or off-white (around #F4F1EA), serif fonts (Georgia or Fraunces), and a terracotta accent. That works for editorial design but not for dashboards or fintech. Negative instructions such as “don’t use cream” just move the output to another fixed palette. Specifying a concrete palette is the only real fix. Without an SOP that feeds ggplab-design-system in as an explicit input, the output falls back to the default every time.
6. Long-Horizon Work Across Multiple Context Windows
(a) What: For work that spans multiple context windows, such as long refactors, book writing, or continuous analysis, standardize the following three files and git.
# Standard working directory structure
<project>/
├── tests.json # verification items — structured (id/name/status)
├── progress.txt # progress notes — free format, appended per session
├── init.sh # environment setup (start server, lint, test)
└── .git/ # checkpoints — commit at each step
# First prompt pattern when starting a new window
"Check pwd. Read and write only in this directory. First read progress.txt,
tests.json, and git log. Run the integration test once before starting new work."
(b) Where: Add a “multi-context-window long-horizon work” subsection inside section 12 (operating rules) of ~/Projects/shared/AGENT_GUIDE.md. Long-horizon projects such as claude-book-hanbit (book writing) and n8n-playbook (guidebook) benefit directly.
(c) Why: Lines 590–666 of the Anthropic guide explicitly recommend this pattern. You already run work across multiple context windows, including claude-book-hanbit, n8n-playbook, and monthly-data-note. Your memory includes rules about a “two-hour session limit” and separating sessions, but it does not formalize a handoff pattern between sessions. With three elements, tests.json, progress.txt, and git, a new session can recover its context immediately, which reduces cost and prevents omissions.
Sources
- Anthropic official prompting guide (SSOT): https://docs.claude.com/en/docs/build-with-claude/prompt-engineering/claude-4-best-practices
- Full text of the Anthropic guide (local cache, 904 lines):
/Users/limjung/.claude/projects/-Users-limjung-Projects/1558ddcb-d4f7-4801-9eb4-9b569a75a24d/tool-results/toolu_01REarmep6QzghdT1deECNib.txt - Karpathy/Chang CLAUDE.md (109k+ stars): https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md
- Opus 4.7 release notes: https://www.anthropic.com/news/claude-opus-4-7
- User memory: summary of 4.7 behavior changes:
/Users/limjung/.claude/projects/-Users-limjung-Projects/memory/reference_opus_4_7_prompting_guide.md - User memory: echo-before-delegate rule:
/Users/limjung/.claude/projects/-Users-limjung-Projects/memory/feedback_scope_echo_before_delegate.md
Bottom line
The four Karpathy/Chang principles cover part of what Anthropic’s Opus 4.7 guide recommends, and the gaps are concrete. Assumption echoing and step-by-step verification should be written as explicit rules for all work, not only for deployments. Effort levels, emphasis words, design inputs, and multi-window handoffs each need their own SOP at the harness level. Start with the behavior rules file and the emphasis-word audit, since they are the smallest changes with the widest effect.
Frequently asked questions
- Do Karpathy's four CLAUDE.md principles cover everything in Anthropic's Opus 4.7 guide?
- No. The four principles are a subset of Anthropic's guide. Assumption echoing and step-by-step verification deserve their own rules. Effort-level operation, emphasis tone, frontend defaults, and multi-window state tracking need separate SOPs.
- What is the most important change to make for Opus 4.7?
- The highest-value steps are adding a behavior rule for assumptions and verification, defining an effort-level SOP, and downgrading overused emphasis words such as MUST and CRITICAL, since Opus 4.7 tends to over-trigger on them.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›