Designing LLM Applications That Stay Safe When The Model Is Fooled
If prompt injection cannot be fully prevented, the system around the model has to contain it. This article covers excessive agency, the lethal trifecta and the design patterns that limit damage.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
The earlier articles in this group reach an uncomfortable conclusion. As of October 2026 no one has shown a reliable way to prevent prompt injection inside the model, jailbreaks keep finding new routes, and poisoned or tampered models can hide behaviour that testing misses. The UK National Cyber Security Centre calls LLMs “inherently confusable” and says the realistic goal is to reduce the likelihood or impact of attacks.1
The project leads of the OWASP Top 10 for LLM Applications open the 2026 edition on the same theme. Their advice is to stop treating the model as the line of defence and to put the protection in the surrounding system instead.2 This article sets out what that means in practice for applications that use an LLM as a component. Systems where models plan and act autonomously raise further issues, which we cover in Agentic AI Security.
Excessive Agency: The Multiplier
A fooled model that can only write text in a chat window does limited harm. A fooled model that can send email, query a database or run code does far more. OWASP calls this Excessive Agency and ranks it third in 2026, up from sixth in 2025.3 Its 2026 entry identifies three root causes:
- Excessive functionality. The model can reach tools or functions the task does not need, such as a document reader that can also delete documents, or a test tool left connected after development.
- Excessive permissions. Tools connect to downstream systems with more rights than needed, such as a generic administrator account instead of the current user’s identity.
- Excessive autonomy. High-impact actions happen without independent verification or a person’s approval.
OWASP is explicit that filtering inputs and outputs is not the root control here. The fix is to give the model less to work with.3
The Lethal Trifecta
Excessive Agency asks how much a fooled model can do. A second test asks when a fooled model can leak data. In June 2025 the developer and writer Simon Willison named three capabilities that are dangerous together: access to private data, exposure to untrusted content, and a way to communicate externally.4 With all three, anyone who can put text in front of the model may be able to make it gather private data and send it out. OWASP’s 2026 prompt injection entry uses the same idea as a check before deployment.5
The test turns a vague worry into a design question. If a feature needs private data and must read untrusted content, remove or tightly restrict its ability to send anything out, including less obvious channels such as a rendered image or link whose address can carry data. The Agentic AI Security group explains the trifecta in detail, alongside a closely related rule of thumb, so this article focuses on what to build once you have applied it.
Patterns That Contain Injection
Researchers have proposed architectures that assume the model can be fooled and try to stop that from mattering.
Separate trusted planning from untrusted reading. NIST’s 2025 taxonomy suggests designing systems on the assumption that injection is possible, for example by using several models with different permissions, or by letting models touch untrusted data only through well-defined interfaces.6 The CaMeL design, published by Edoardo Debenedetti and colleagues in March 2025, is a concrete version.7 A privileged model reads the user’s request, but never the untrusted data, and turns it into a short program. A second, quarantined model reads emails, web pages and tool results, but it has no tools of its own and can only hand back structured values. Every value carries labels recording where it came from and who may see it, and a policy check runs before each tool call. Planted text can change a value, such as the wording of a summary, but it should not be able to add a step to the plan or send data where the policy forbids. The protection is only as good as the policies the developer writes, and the authors acknowledge that some indirect leaks, known as side channels, remain possible.
On the AgentDojo benchmark, the authors report that CaMeL completed 77 percent of tasks with provable security under their threat model, against 84 percent for an undefended system.7 The cost is some lost capability and more engineering.
Mark untrusted content. Keegan Hines and colleagues proposed spotlighting, which rewrites outside text before the model sees it, for example by wrapping it in markers or encoding it, so the model can keep telling which parts came from an outside source. In their tests on GPT-family models it cut attack success from over 50 percent to under 2 percent with little effect on task performance.8 That is a large improvement, but it is a probabilistic defence, not a guarantee.
Defence In Depth For LLM Applications
Putting these ideas together gives a layered design. Each layer assumes the one above it can fail.
- Input HandlingLabel and separate untrusted content, apply classifiers, strip hidden text where practical.
- Model And PromptClear instructions and safety training. Helpful, but assume they can be overridden.
- Output ValidationCheck output against expected formats, encode it for its destination, block unexpected links and requests.
- Least Privilege For Tools And DataOnly the tools, functions and data the task needs, under the current user's identity.
- Break The TrifectaNever combine private data, untrusted content and outbound communication without strong separation.
- Human Approval For High-Impact ActionsPayments, deletions, external messages and permission changes need a person to confirm.
- Logging, Monitoring And ResponseRecord prompts, retrieved content and tool calls so incidents can be detected and investigated.
None of this is exotic. It is least privilege, separation of duties and monitoring applied to a component that takes instructions from whatever text it reads. The hard part is discipline: resisting the pull to connect one more tool or data source because the model could use it.
Footnotes
-
UK National Cyber Security Centre, D. Chismon, “Prompt injection is not SQL injection (it may be worse)”, 8 December 2025. ncsc.gov.uk ↩
-
S. Wilson and R. Lambros, “Letter from the Project Leads”, OWASP Top 10 for LLM Applications 2026, August 2026. github.com ↩
-
OWASP GenAI Security Project, “LLM03:2026 Excessive Agency”, OWASP Top 10 for LLM Applications 2026, August 2026. github.com ↩ ↩2
-
S. Willison, “The lethal trifecta for AI agents: private data, untrusted content, and external communication”, 16 June 2025. simonwillison.net ↩
-
OWASP GenAI Security Project, “LLM01:2026 Prompt Injection”, OWASP Top 10 for LLM Applications 2026, August 2026. github.com ↩
-
NIST, AI 100-2 E2025, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations”, March 2025, section 3.4. csrc.nist.gov ↩
-
E. Debenedetti et al., “Defeating Prompt Injections by Design”, arXiv 2503.18813, March 2025. arxiv.org ↩ ↩2
-
K. Hines et al., “Defending Against Indirect Prompt Injection Attacks With Spotlighting”, arXiv 2403.14720, March 2024. arxiv.org ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.