BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage News New Archestra's OpenAPPA Saturates Two Major Security Benchmarks with a 0% Attack Success Rate

New Archestra's OpenAPPA Saturates Two Major Security Benchmarks with a 0% Attack Success Rate

Listen to this article -  0:00

Archestra released OpenAPPA, an open-source security engine designed to stop data exfiltration caused by prompt injection or model hallucination. OpenAPPA runs outside the agent’s prompt and execution loop. Its configuration details concepts such as data sources, audiences, trust levels, and authorities, along with their associated deterministic security enforcement rules. The team reports zero successful attacks when running security benchmarks Bench-Corp (20 multi-step enterprise workflows) and AgentThreatBench, vs. 10% for Claude Code’s auto mode and 31% for Microsoft FIDES.

OpenAPPA’s documentation explains why a stochastic approach to automated policy enforcement fails:

The industry’s answer to approval fatigue is a second model that judges each tool call: Claude Code’s auto mode, Codex’s auto-review, and other auto-modes.

By design, they cannot track data flow across tool calls. Because classifiers are prompt-injectable themselves, harnesses hide tool outputs from them, so the judge never sees the data at all.

Because of their probabilistic design, even the best top out at 99.3%: at millions of calls, 0.7% is a lot of breaches.

[…] Rule sets end up either so tight they break the agent or so intricate nobody can audit what they permit.

On the one hand, agents have proven skilled at working around simple but common approaches like allowlists or denylists of tools: a denied rm -rf may be replaced by an equivalent Python script. On the other hand, extending denylists or overly restricting policies to protect against eager agents results in decreased utility (e.g., while the agent does not leak data, it does not perform the task successfully because of the restrictions). OpenAPPA’s GitHub repository reminds developers:

Agent security has two axes: an agent that permits unauthorized flows is unsafe, and an agent that refuses valid work is useless.

Archestra seeks to resolve the tension between strict enforcement and operational utility with what it calls an Agentic Permissions Policy Algebra (APPA), described in a paper by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, and Matvey Kukuy.

OpenAPPA implements this approach with a pluggable engine that is executed outside the agent’s loop, thus defeating any attempts by the underlying language model to inspect, negotiate with, or bypass policy rules. The security policies take the form of a single appa.toml configuration file that details data sources, audiences, trust levels, and authorities.

OpenAPPA jointly labels and monitors both audience (the authorized set of consumers) and trust (the degree of data verification). Labels compose monotonically using lattice algebra; labels can only become more restrictive; reading restricted records narrows the audience; reading unvetted external web pages lowers trust.

Each tool contract defines three primary operational attributes: requires (the audience membership and trust levels necessary to run the tool), delta (the restrictions applied when the tool returns data), and effects (an audit trail of successful actions).

In the following example, reading a ticket via get_ticket_from_crm restricts the trajectory’s audience to internal. The subsequent process_internal_data call requires the audience to be within internal.

[[policy.tool]]
name = "get_ticket_from_crm"
delta = { audience = ["internal"] }

[[policy.tool]]
name = "publish_update"
requires = { audience = { contains = ["public"] } }

[[policy.tool]]
name = "process_internal_data"
requires = { audience = { within = ["internal"] } }

In the following example, the read_web_page tool’s result is marked as suspicious. Upon reception of a result from the read_web_page tool, OpenAPPA blocks apply_db_migration because it requires trusted data.

[policy]
version = 2
trust_chain = ["suspicious", "trusted"]

[[policy.tool]]
name = "read_web_page"
delta = { trust = "suspicious" }

[[policy.tool]]
name = "apply_db_migration"
requires = { trust = "trusted" }

OpenAPPA additionally has explicit recovery semantics. When an agent attempts an illegal action, the engine halts dispatch and provides structured pathways to proceed. Sanitizers may edit payloads, e.g., stripping personally identifiable information, to safely expand the permitted audience. Authorities route requests to human operators or internal verification APIs for single-action approval. Disposable Child Branches enable on-demand confinement: when an agent must ingest untrusted data, the engine isolates the read in a transient subagent branch, returning only schema-attested, sanitized outputs to the parent.

The Bench-Corp and OWASP AgentThreatBench benchmarks show that OpenAPPA maintains high utility while under strict security constraints, reporting a 0% attack success rate and an 89% task completion rate. Claude Code’s native auto mode yielded a 10% attack success rate with a 90% completion rate. Microsoft FIDES permitted 31% of attacks to succeed and completed only 41% of tasks. The team of researchers reports in the paper that ablation experiments seem to validate the value of recovery strategies: task completion fell to 35.0% when remedy plans were completely disabled.

Bench-Corp and AgentThreatBench test explicit policy breaches: sensitive data sharing, prompt injection, approval and ordering, and tenant isolation. Bench-Corp is a highly specialized corporate-assistant benchmark designed to evaluate how security policies hold up across complex, multi-step enterprise workflows. AgentThreatBench operationalizes the OWASP Top 10 for Agentic Applications (2026) into executable tasks. It was recently merged into the official UK AI Safety Institute’s inspect_evals repository. AgentThreatBench uses a dual-metric scoring system, scoring both utility and security. OpenAPPA achieves a perfect security score on the benchmark with zero successful attacks.

OpenAPPA is currently a preview. The formal algebra and recovery guarantees are published on arXiv (APPA: Recoverable Information-Flow Control for Real-World LLM Agents). The reader is encouraged to read the paper and the documentation, both of which contain plenty of additional technical details, accompanying illustrations, and empirical results.

About the Author

Rate this Article

Adoption
Style

BT