OpenAPPA Security Engine Aims to Block LLM Data Leaks
Archestra has launched OpenAPPA, an open-source security engine designed to prevent data exfiltration in AI agents through deterministic enforcement.

Archestra has introduced OpenAPPA, an open-source security engine aimed at mitigating data exfiltration risks caused by model hallucination or prompt injection in AI applications. Operating outside the main prompt and execution loop of an agent, the engine relies on deterministic enforcement rules rather than probabilistic second models.
The project stems from research detailed in a paper by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, and Matvey Kukuy. According to the creators, stochastic approaches like auto-review modes often fail because secondary classifiers can be prompt-injectable and cannot adequately track data flow across multiple tool calls. Furthermore, simple allowlists or denylists frequently lead to operational rigidity or complex policy management.
Deterministic Policy Architecture
OpenAPPA uses a single configuration file named appa.toml to define data sources, trust levels, audiences, and authorities. The engine tracks both the audience, representing authorized consumers, and trust, denoting the verification degree of data. These labels compose monotonically using lattice algebra, meaning security constraints only become more restrictive during execution. For instance, reading external unvetted web pages lowers trust, while accessing restricted records narrows the audience scope.
Every tool contract specifies three operational attributes: requires (mandating audience membership and trust levels for execution), delta (applying restrictions upon data return), and effects (generating an audit trail). When an agent triggers an unauthorized action, OpenAPPA halts dispatch and invokes structured recovery semantics. These include sanitizers to strip sensitive information, authorities to route requests to human operators or internal verification APIs, and Disposable Child Branches to isolate untrusted data ingestion within transient subagent branches.
Benchmark Performance
In evaluations across Bench-Corp, which features 20 multi-step enterprise workflows, and AgentThreatBench, which operationalizes the OWASP Top 10 for Agentic Applications and is merged into the UK AI Safety Institute’s inspect_evals repository, OpenAPPA reported a 0% attack success rate. The engine achieved an 89% task completion rate on these benchmarks, which test explicit policy breaches such as tenant isolation, prompt injection, sensitive data sharing, and ordering anomalies.
Comparatively, Claude Code’s auto mode yielded a 10% attack success rate with a 90% completion rate, while Microsoft FIDES recorded a 31% attack success rate and a 41% completion rate. Ablation experiments noted by the researchers showed that task completion dropped to 35% when remedy plans were completely disabled.
What it means for developers
For engineers building production-grade autonomous agents, balancing security with operational utility remains a central challenge. Overly restrictive policies can render an agent useless, while loose policies invite exploits. OpenAPPA's pluggable architecture runs completely outside the language model loop, preventing underlying models from bypassing rules. As developers work to secure complex multi-step workflows against prompt injection and data exfiltration, exploring such deterministic enforcement engines offers a promising path forward. To test these capabilities across models, developers can try top AI models cheaply through one API at https://apixoai.online.
The OpenAPPA repository is currently available in preview, accompanied by formal documentation and a paper on arXiv titled APPA: Recoverable Information-Flow Control for Real-World LLM Agents.
Source: New Archestra's OpenAPPA Saturates Two Major Security Benchmarks with a 0% Attack Success Rate — InfoQ AI & ML. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

