Insights

Claude Code Harness Audit: Finding Silent Failures in Hooks, npx, and YAML

6 min read#claude-code#prompt-engineering#hooks#troubleshooting#ops#ssot

Who this is forDevelopers and ops engineers who maintain Claude Code configuration, hooks, skills, and shared prompt documentation and suspect some of it has quietly stopped working.

Over three months, a Claude Code harness grew into a large set of files: global CLAUDE.md, rules, skills, agents, and settings; a shared SSOT (single source of truth) document; per-project CLAUDE.md files; and a Mac mini integration. We audited the whole harness against Anthropic’s official prompting guides for Opus 4.8 and Fable 5. The main finding was that almost every failure happened without an error. Non-standard approaches, including a non-standard environment variable hook, npx auto-install, and idiomatic YAML parsing, looked healthy. In practice they either did not run or ran something other than what we intended. This article covers what we found, how we tested it, and the patterns to watch for, so you can check your own setup before a quiet failure costs you time.

Failure pattern map

Silent failure vs. immediately visible failure: positioning of six harness drift patterns

The diagram places the drift patterns on two axes: whether a failure is silent, and whether it surfaces immediately. The patterns in the lower-left area are the ones that deserve attention first, because nothing in your normal workflow will tell you they are broken.

Audit scope and results

We reviewed seven units: hooks, skill diet, frontmatter parsing, the SSOT registry, credentials and IP addresses, verification of subagent audit results, and the workflow catalog. The table below summarizes the measured results.

Item Result
Inspection units 7 (hooks, skill diet, frontmatter parsing, SSOT registry, credentials/IP, subagent audit verification, workflow catalog)
Commits 10 (3 in shared, 7 across projects)
~/Projects/shared/workflows/ verification 98/98 passed
Secret exposure 1 found and removed immediately
Skill renames 3 (aligned with the naming.md convention)
Fable 5 skill diet 377 lines to 214 lines; 375 lines to 132 lines (removed progress-narration scaffolding, duplicate branches, and micro-step enumerations)
frontmatter strict YAML parse failures 5 of 48 skills (unquoted values containing colons or brackets)
Reasoning-extraction risk instructions (“move your reasoning into the response” style) 0 found in a full audit

The numbers show the scale of the work, but the more useful information is in the failure types behind them.

The common thread

When the harness depended on convenience instead of standard interfaces (stdin JSON, local node_modules/.bin binaries, and strict YAML), failures did not show up as error logs. They showed up as nothing happening. The audit was less about finding bugs to fix and more about finding automation that had quietly died. That is the central lesson of this work. If you expect a broken hook to print an error, you will miss most of the problems in a harness like this one.

1. PostToolUse hooks fail silently

A hook that depended on a non-standard environment variable ($TOOL_INPUT_FILE) ran without error and did nothing. The standard approach is JSON delivered on stdin, with the file path parsed by jq from .tool_input.file_path.

You can only discover that a hook is broken by testing it with deliberately faulty input. If you test only with normal input, a hook that never runs will pass every time, and you cannot tell it apart from a working hook. A passing test therefore proves less than it seems to. Feed the hook an input you designed to be wrong, and confirm that it reacts the way you intended. Only then can you say the hook works.

2. The npx auto-install squatting trap

If typescript is not installed in the current directory, npx tsc automatically installs and runs a similarly named package, [email protected], which is a deprecated squatting package. The command looks like it ran your type checker, but it ran a different program.

In hooks and scripts, always point to the project-local binary directly, using a relative or absolute path such as:

node_modules/.bin/tsc

Using npx <cmd> directly in a hook can create the false impression that the type check passed, while a different binary is actually running. This is the same pattern as the hook problem: the step appears to succeed, and nothing in the output tells you otherwise.

3. Skill diet for Fable 5

Anthropic’s official guide states that over-prescriptive skills aimed at older models can actually reduce quality. We applied that guidance. We removed progress-narration scaffolding (instructions like “now output ~”), duplicate branch descriptions, and micro-step enumerations. That cut one skill from 377 to 214 lines and another from 375 to 132 lines. Shorter skills gave the model room to decide how to proceed, instead of following a script that no longer matched the task.

During this pass we also checked for instructions that ask the model to move its reasoning into the response. Such instructions can trigger a reasoning-extraction refusal, so we audited every skill for them. The full audit found none. Because new skills can introduce this pattern, we keep it on the checklist and recheck it whenever a skill is written. A clean audit today does not guarantee a clean file next month.

4. Lenient frontmatter parsing

Claude Code itself tolerates unquoted YAML values that contain colons or brackets. Strict YAML tools, such as yaml.safe_load in automation scripts, fail on the same values. The belief that a file is valid because it works in Claude Code is where this problem starts. Claude Code runs the file, so the file looks correct, but any tool that parses it strictly will break.

If automation will read your skill inventory, run a strict-parse check across all files in advance. In this audit, that check found 5 failures among 48 skills. None of them had caused a visible problem, because Claude Code had been handling them leniently the whole time.

5. SSOT drift in hand-maintained registries

Hand-maintained inventory documents, such as UTILITIES.md and project registries, are not updated when assets are deleted or renamed. Stale references build up over time, and no one notices until someone follows one.

The fix is not to update the registry more diligently. It is to separate roles. Keep the hand-maintained registry minimal and limited to core projects. Generate the complete list with a script (generate_project_index.py produces PROJECT_INDEX.md), and treat the generated file as one that people do not edit. Each document then has one owner: a person or a script, never both.

6. Hardcoded DHCP IP addresses

The Mac mini’s LAN IP address appeared in three different documents with three different values: .104, .125, and the measured .226. The root cause is that DHCP assigns addresses that change. Updating the number to the latest value is a temporary patch, and the next address change will bring the same problem back.

The correct fix was a rule: reference only the Tailscale alias, and never hardcode a numeric IP address anywhere. Any value that can change should pass through an alias, so that a change happens in one place instead of in every document that mentions it.

7. A subagent false positive

A subagent audit flagged defaultMode: auto as non-standard and matcher: Agent as unmatched. When we checked the official documentation, both were valid settings. An audit produced by a subagent should not be trusted on its own. Verify each finding against a primary source, such as the official documentation, before you act on it. Auditors can produce false positives too, and acting on a false positive means changing a working configuration.

Reusable principles

  • Test automation with deliberately broken input. Reliability is shown by how it responds to input you broke on purpose, not only by whether it passes with normal input.
  • Separate apparent success from actual success for convenience tools such as npx and lenient parsers.
  • Do not mix human-maintained documents with machine-generated ones.
  • Do not hardcode values that can change, such as IP addresses. Always route them through an alias.

Sources

Bottom line

The most serious risks in a Claude Code harness are the ones that produce no error. Hooks that read the wrong input, npx commands that run the wrong binary, lenient parsers that hide invalid YAML, and registries that keep stale references can all look healthy while doing nothing useful. Test automation with deliberately broken input, use standard interfaces and project-local binaries, validate frontmatter with a strict parser, keep generated and hand-written documents separate, and reference IP addresses only through aliases. The audit showed that these checks find the failures that would otherwise stay hidden.

Frequently asked questions

Why do Claude Code harness failures often go unnoticed?
In this audit, nearly every failure happened silently. Non-standard env var hooks, npx auto-install, and lenient YAML parsing looked fine but either never ran or executed the wrong thing. Testing with only normal input hid the problem.
How should I test a Claude Code hook so I know it actually works?
Test it with deliberately broken input, not only normal input. A hook that never runs will still pass tests that use valid input. Read the file path from stdin JSON with jq, such as .tool_input.file_path, and confirm the hook reacts to bad input as intended.