OpenAPPA Takes the Security Call Away From the Model
How a deterministic data-flow policy sits between Claude Code and its tools, what its benchmark proves, and what it leaves to you
Most agent safety features today ask a model to judge another model. An agent proposes a tool call, a classifier or a second prompt looks at it, and something decides whether it seems safe. OpenAPPA, an MIT-licensed project from Archestra that sat at #14 on Trendshift's daily board this morning, refuses that whole arrangement. It asks one question before every tool call, "is this data allowed to go to this destination?", and answers it without consulting a model at all.
That sounds like an old idea from operating system security, and it is. Which is exactly why it is interesting to see it arrive in an agent stack.
The problem it is aimed at
The agent leak that security researchers keep publishing rarely starts with a model "deciding" to be malicious. It starts with a chain. The agent reads something it should not trust (a support ticket, an email, a forum post), that content carries instructions, and a few steps later the agent sends private data somewhere public. Every individual step looks reasonable. The violation only shows up when you track where the data came from and where it is going.
Model-based guards look at one step at a time and make a judgment call. OpenAPPA tracks the whole trajectory and applies a rule. My view is that the rule-based approach is the right foundation for anything that touches customer data, with model judgment layered on top rather than the other way around. The interesting question is what it costs you.
How the policy works
OpenAPPA policies are TOML files, and the policy reference shows the full shape. Every piece of data carries a security label with two independent parts: audience (who may receive it) and trust (how reliable it is). Tools declare what they do to that label.
A source tool restricts the label with a delta:
[[policy.tool]]
name = "get_ticket_from_crm"
delta = { trust = "suspicious", audience = ["internal"] }
A sink tool states what it needs with requires:
[[policy.tool]]
name = "send_email"
parameters = { type = "object", properties = { recipient = { type = "string" }, body = { type = "string" } }, required = ["recipient", "body"] }
requires = { audience = { contains = ["$recipient"] } }
delta = {}
Here the email tool reads its own recipient argument and checks that the current audience includes that person. Once the agent has read an internal CRM ticket, the trajectory's audience is internal, so an email to an outside address fails the check before the tool ever runs.
The key property is that labels only move one way. In the docs' words, "Reading restricted data limits where the trajectory can send data later," and "Reading a later result marked trusted does not undo the earlier drop." A tool's delta.trust "can lower the trajectory's trust, but cannot raise it." An injected instruction cannot talk its way back up.
Decisions come only from the event log, with no network or file calls, which is what makes them deterministic and replayable.
When a flow is denied, there are three outcomes:
- Block. Some requirements have no remedy at all.
- Ask. An "authority" can approve an exception, but only for requirements listed in its
permitssection. The built-inhitlauthority "asks a person to review the exact call and the requirements to be approved." - Sanitize. A sanitizer transforms restricted data so it can cross a boundary, and "the integration keeps the original hidden from the agent and delivers the transformed result."
That middle option matters. It turns a hard block into a narrow, logged approval for one specific call, which is a much better experience than a blanket "allow this tool forever" prompt.
How it plugs into Claude Code
The Claude Code integration uses Claude Code's own lifecycle hooks (SessionStart, UserPromptSubmit, PreToolUse, PostToolUse and subagent events) and covers built-in tools like Bash, Read, Edit and Write plus external MCP tools. You launch protected sessions with clappa instead of claude.
Setup is a guided step. The /appa-guide skill inspects your tools, asks "focused questions when an account identity, data sensitivity, or boundary needs clarification," and writes the policy. When a call breaks policy, OpenAPPA blocks it and returns the conflict with remedies, such as a sanitizer or operator approval. The project also lists an Archestra proxy layer for Claude Desktop, Cursor, Copilot CLI, n8n and others, and an embeddable runtime for custom agents.
What the benchmark says, and what it does not
The README's headline: on Bench-Corp and OWASP Top 10 evaluations across 1,320 runs, OpenAPPA reached 89% task completion with 0% successful attacks, against 90% and 10% for Claude Auto mode and 41% and 31% for Microsoft FIDES.
The evaluation page fills in the method. Bench-Corp is "20 multi-step workplace tasks" across HR, finance, support, vendors, email, forums and task tracking, testing sensitive-data sharing, prompt injection, fake approvals and tenant isolation. Each model ran every scenario five times with standard prompts and five times with adversarial ones. Across GPT-5.6 Luna, DeepSeek V4 Flash and Gemini 3.7 Flash, OpenAPPA scored between 88% and 90% utility with 0% attack success, while FIDES variants landed between 37% and 44.5% utility and 28% to 34.5% attack success. One design choice deserves credit: "The benchmark checks what the agent actually changed or sent. It does not use an LLM judge."
Now the honest reading. The authors built the benchmark, wrote the policies for it, and ran the comparison. Twenty scenarios is small. A 0% attack rate means the policies covered the attacks in this suite. It says nothing about the attack you did not think to label for.
Put this into practice
The lowest-friction way in is the Claude Code path, on a throwaway project first.
curl -fsSL https://openappa.com/install.sh | sh &&
~/.local/bin/appa plugin install claude-code
clappa
/appa-guide
Read the install script before you pipe it to a shell. That advice applies doubly to a security tool.
Then do three things in order:
Label your sources before your sinks. List every tool that reads private or untrusted data (CRM, email, tickets, shared docs) and give it a delta. Most of the protection comes from honest labels on the read side.
Mark the exits. For every tool that sends data out (email, Slack, GitHub comments, HTTP calls), add a requires that names who may receive it.
Replay before you trust. OpenAPPA ships appa describe --config appa.toml --check and appa replay --config appa.toml policy-tests/. Write a test trajectory that reads an internal ticket and tries to post it publicly, and confirm it fails. A policy you have not replayed against a bad trajectory is a policy you are guessing about.
Honest limitations
It is a preview. The README calls OpenAPPA "a preview and an RFC" and warns that "configuration and interfaces may change without backward compatibility guarantees." Pin the version and expect to rewrite policies.
Your labels are the security. The engine is deterministic, and that cuts both ways. If you forget to label a tool that reads customer data, OpenAPPA has no way to know it should care. A model-based guard might catch something you never anticipated. A rule-based one will not.
The bypass is one command away. Protection depends on launching with clappa. The docs say plainly that resuming with plain claude, or a project with disableAllHooks: true, gives you an unprotected session. On a team, that is a training and configuration problem, not a technical guarantee.
The numbers are self-reported. Every benchmark figure comes from the project's own suite. Treat 0% as "covered what they tested," and test your own flows.
The real choice
Agent security keeps getting framed as a smarter model watching a less careful one. OpenAPPA offers the other bargain: a dumb, predictable check that you configure, test and read. It will never surprise you with good judgment, and it will never be talked out of a rule.
If your agents touch data you would be embarrassed to leak, that predictability is worth the labeling work. The question to ask this week is simple. Could you write down, right now, which of your agent's tools read private data and which ones send it out? If you cannot, that list is the first security feature you need, whether or not you install anything.
Sources: OpenAPPA on GitHub · OpenAPPA policy reference · OpenAPPA evaluation · OpenAPPA Claude Code integration
Medium metadata
- Title: OpenAPPA Takes the Security Call Away From the Model
- Subtitle: How a deterministic data-flow policy sits between Claude Code and its tools, what its benchmark proves, and what it leaves to you
- Tags: AI Agents, AI Security, Claude Code, Prompt Injection, Open Source
- Canonical: import from the fervorai.dev URL