Claude Code Prompt Hooks Let Through What They Were Told to Block
Version 2.1.294 fixed instruction-style hooks. The lasting lesson is about where a safety rule lives, not which build you run.
A safety rule written in plain English has to be read before it can be enforced. That sounds obvious until you notice how many agent guardrails are now written exactly that way, as a sentence handed to a model with the expectation that the model will say no at the right moment.
Late on October 7, Claude Code shipped a fix that shows what happens when the reading goes wrong. The first line of the 2.1.294 changelog reads: "Fixed prompt and agent hooks written as instructions (such as "Block commands that...") allowing what they should block." npm records the publish at 23:42 ET that night.
Read that sentence twice. A hook that someone wrote to block commands was allowing them. Not crashing, not timing out, not throwing an error you would see in a log. Allowing.
Why this one matters more than a normal bug fix
Most changelog fixes are about a feature misbehaving. This one is about a guardrail misbehaving in the permissive direction, which is the worst direction a guardrail can fail. If a formatter hook breaks, your code looks ugly. If a "block commands that touch production" hook breaks open, the agent does the thing you built the hook to stop, and nothing tells you.
The changelog does not say which versions carried the bug or how long it lived. It does not say whether any specific phrasing was safe. And when I checked the npm registry around 07:25 ET on October 8, the latest tag still pointed at 2.1.293 and the stable tag at 2.1.285. So if your team pins stable, or updates on its own schedule, the fix may not be on your machines yet.
My position is simple. Upgrade, yes. But the upgrade is the small lesson. The big one is that a guard a model interprets inherits the model's failure modes, and any action you cannot undo deserves a guard that does not need interpreting.
How Claude Code hooks actually decide
Claude Code supports five hook handler types, per its hooks reference: command, http, mcp_tool, prompt and agent. Two of them involve a model making the call.
A prompt hook sends your prompt text, with the hook's input JSON substituted for $ARGUMENTS, to a Claude model for a single-turn evaluation. By default it uses the model Claude Code runs for background work, and you can set a different one with the model field. Its default timeout is 30 seconds.
The hooks guide spells out the contract: "The model's only job is to return its decision as JSON." That JSON is "ok": true to let the action proceed or "ok": false with a reason. On PreToolUse, ok: false denies the tool call, and by default the turn ends with the reason shown as a warning.
An agent hook goes further. It spawns a subagent that can Read, Grep and Glob, with a 60-second default timeout and up to 50 tool-use turns. The docs flag agent hooks as experimental and say, in a warning box, "For production workflows, prefer command hooks."
Now look at the shape of the bug again. The decision schema asks the model a yes-or-no question: is this OK? A hook written as an instruction, "Block commands that write outside the repo," asks the model to do something else, to act as an enforcer. The model has to translate a command into a verdict, and the verdict into the polarity of a boolean called ok. That translation is where a hook written as an instruction can come out backward. The changelog does not explain the root cause, so treat that as my reading of the mechanism, not Anthropic's. What the changelog does confirm is the symptom.
The deterministic path that was there all along
A command hook is a shell command. It receives the event JSON on stdin and answers through its exit code and stdout. No model reads it.
The blocking rules are blunt, which is what you want from a guard. Per the hooks reference, exit code 2 is a blocking error on events that can block, and on PreToolUse it blocks the tool call with stderr as the reason. It blocks even if the same hook prints a JSON permissionDecision of "allow". Exit 2 wins.
The matcher is just as literal. A matcher made only of letters, digits, underscores, hyphens, spaces, commas and pipes is an exact match, so Bash matches the Bash tool and nothing else. Any other character turns it into an unanchored JavaScript regular expression, which is where people get surprised: Edit.* matches both Edit and NotebookEdit, so the docs recommend ^Edit$ when you mean exactly one tool. For finer filtering, a handler's if field takes permission-rule syntax such as Bash(git *).
None of that is glamorous. All of it does the same thing on the thousandth run as on the first.
There is one fail-open edge the docs state plainly: a timed-out command, http or mcp_tool hook does not block a PreToolUse tool call, and the call continues through the normal permission flow. The docs do not say what a timed-out prompt or agent hook does, so do not assume it fails closed. A slow guard is an absent guard. Keep guards fast, and for prompt hooks remember the 30-second default.
Put this into practice
Here is the lowest-friction path, in the order I would do it.
1. Check your build. Run claude --version. If you are below 2.1.294 and you use prompt or agent hooks for anything that blocks, update, or move those guards to command hooks before you rely on them again.
2. Inventory your model-judged guards. Search your settings files for "type": "prompt" and "type": "agent". For each one, ask one question: if this hook silently returned ok: true every time, what could happen? If the answer involves deleted data, pushed code, sent messages or spent money, it is in the wrong tier.
3. Rewrite the destructive ones as command hooks. The pattern is small. Match the tool exactly, read the command from stdin, and exit 2 on anything that matches your deny list.
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/deny-destructive.sh" }
]
}
]
}
}
The script reads the JSON, pulls out the command, checks it against patterns like rm -rf, git push --force or your production hostnames, and exits 2 with a one-line reason on stderr when one matches. It is twenty lines of shell or Python, and it never needs to understand your intent.
4. If you keep a prompt hook, phrase it as the question the schema asks. The fix should cover instruction-style prompts, but there is no reason to make the model translate. Instead of "Block commands that write outside the repo," write "Does this command write outside the repository? If it does, respond with ok false and the path as the reason. Otherwise respond with ok true." Spell out the polarity.
5. Test the guard against the thing it guards. Feed it a command it must block and confirm the block appears. Then feed it one it must allow. Do this again after every Claude Code upgrade. A guard you have never seen fire is a guess.
Where prompt hooks still earn their place
None of this means prompt hooks are a mistake. They exist for judgment, and some questions only judgment can answer. "Did Claude finish every task the user asked for?" is not a regex. The hooks guide's own example uses a Stop hook for exactly that, and on Stop a wrong ok: false just makes Claude keep working, while a wrong ok: true ends a turn early. Both are annoying. Neither is destructive.
That is the dividing line I would draw. Use a model to decide things where a wrong answer costs time. Use code to decide things where a wrong answer costs something you cannot get back. If you want both, stack them: a command hook enforces the hard deny list, and a prompt hook on the same event adds a softer review on top.
The honest limits of this advice
Command hooks are not a security boundary against a determined adversary. A deny list of command patterns can be evaded by an agent that writes the same operation a different way: a script file, an alias, a language runtime that shells out. Pattern matching catches the obvious and the accidental. It does not catch a model that has been prompt-injected into working around you. For that you need sandboxing and least-privilege credentials underneath the hooks, not more hooks.
Exact matchers have their own trap. A guard on Bash does nothing about an MCP tool that runs shell commands under another name, and a regex matcher that is broader or narrower than you meant fails without a sound. Read the matcher rules once, carefully.
And the evidence here is thin in one important way. The changelog gives a one-line fix with no version range, no example of a failing phrasing and no root cause. Everything I have said about why instruction-style prompts went wrong is inference from the decision schema. It is a reasonable inference, but it is not a postmortem, and if Anthropic publishes one it should replace this section.
Where the rule should live
Agents are getting faster at doing things, and the rules around them are increasingly written in the same medium the agents run on: language. That is convenient, and for a lot of guidance it is the right call. A skill, a CLAUDE.md, a review prompt all work because a model reads them.
A brake is different. A brake should work when the driver misreads the road.
Open your settings file today and find the hooks that stand between your agent and something you cannot undo. For each one, decide whether you want it read or run. Then make it so.
Sources: Claude Code changelog, npm registry entry for 2.1.294, Claude Code hooks reference, Claude Code hooks guide.
Medium metadata
- Title: Claude Code Prompt Hooks Let Through What They Were Told to Block
- Subtitle: Version 2.1.294 fixed instruction-style hooks. The lasting lesson is about where a safety rule lives, not which build you run.
- Tags: Claude Code, AI Agents, AI Safety, Developer Tools, Software Engineering
- Canonical URL: import from the fervorai.dev post
- Estimated read time: 8 minutes