Claude Code Harness Audit: Finding Silent Failures in Hooks, npx, and YAML
A three-month drift audit of a Claude Code harness. Learn which failures stay silent and how to test hooks, avoid npx squatting, and validate YAML.
Hands-on notes on AI agents and work automation: guides, concepts and what actually worked.
A three-month drift audit of a Claude Code harness. Learn which failures stay silent and how to test hooks, avoid npx squatting, and validate YAML.
Autonomous agent loops need a machine verifier and explicit context. Learn the two-axis test, the evidence, and when humans should write rubrics first.
Compare a markdown-based LLM wiki pipeline with Google's Open Knowledge Format v0.1. Learn where they match, where they differ, and when to align your schema.
Anthropic's analysis of about 400,000 Claude Code sessions shows domain knowledge amplifies agentic coding results more than coding training does.
Why MCP and SDKs are complementary layers, not rivals, and how Google and n8n each handled protocol support versus building a framework.
A four-gate delivery playbook for automation projects in closed networks. Catch tool limits and security constraints in week one, not week three.
Fable 5 costs twice as much per token as Opus 4.8, but measured sessions cost 1.0 to 1.4 times as much. See the token, cost, and intervention data behind it.
Why Claude Code hooks guarantee deterministic control over AI agents, traced from git hooks and Design by Contract to neuro-symbolic guardrails.
Why MCP moved to Linux Foundation governance, how it complements Google's A2A, and what the adoption and security evidence supports.
Teardown means destroying one-off resources, not keeping them for later. Here is which term fits a service you shut down temporarily and plan to restart.