The Lethal Trifecta And The Rule Of Two: A Simple Test For Risky Agents
Two short rules of thumb explain which agents an attacker can turn against you. Both reach the same answer, which is to limit the capabilities an agent can combine.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
Nobody yet knows how to make a language model reliably ignore instructions hidden in the content it reads. Until that changes, the practical question for anyone deploying an agent is a different one: if the agent is tricked, how much harm can it do?
Two short frameworks answer that question in a way that business and technical readers can both apply. The first is the “lethal trifecta”, described by the developer and writer Simon Willison in June 2025. The second is Meta’s “Agents Rule of Two”, published in October 2025. This article explains both, shows how they relate, and gives you a quick test to run on any agent design.
The Lethal Trifecta
Willison’s post, published on 16 June 2025, names three capabilities that become dangerous together.1
- Access to private data. The agent can read things that should stay inside the organisation, such as email, files or customer records.
- Exposure to untrusted content. Text or images that an outsider could have written reach the model, for example a web page, an inbound email or a shared document.
- The ability to communicate externally. The agent can send something out, by email, by calling a web address, or by rendering a link or image that fetches a remote URL.
If an agent has all three, an attacker does not need to break in. They only need to place instructions where the agent will read them, asking it to collect private data and send it somewhere. The model may follow those instructions because it cannot reliably tell them apart from legitimate ones. Willison’s advice to people combining tools is blunt: “avoid that lethal trifecta combination entirely”.1
The Agents Rule Of Two
Meta turned the same idea into a design rule in a post titled “Agents Rule of Two: A Practical Approach to AI Agent Security”, dated 31 October 2025.2 It lists three properties:
- [A] the agent processes untrustworthy inputs;
- [B] the agent can access sensitive systems or private data;
- [C] the agent can change state or communicate externally.
Within a single session, an agent should have no more than two of these. Where a task genuinely needs all three, Meta’s answer is that the agent must not be left to act alone: someone, or some dependable check, has to confirm what it does before it happens.2 A fresh session with a clean context counts as a new start, which is why splitting work into separate sessions is one way to stay within the rule.
Meta credits two influences: a similarly named rule used by the Chromium browser project, and Willison’s trifecta. The main difference is the third property. Willison focuses on sending data out. Meta widens it to any change of state, so an agent that can delete files, approve a payment or push code counts in the same way as one that can send email.
Applying The Test
The value of both rules is that they move the conversation away from asking whether the model is clever enough to resist an attack. They ask a design question instead, which you can answer from an architecture diagram.
Can any input to this session come from outside your control?
- Yes:
Can the session access sensitive systems or private data?
- Yes:
Can the session change state or communicate externally without a person approving each action?
- Yes:
All three properties. Do not run this autonomously. Split the task across sessions, remove one capability, or require human approval of each consequential action.
- No:
Two properties, with supervision on the third. Make sure approvals show the real action and its target, not a summary written by the agent.
- Yes:
- No:
Two or fewer properties. Keep it that way. Check that new tools do not quietly add access to private data.
- Yes:
- No:
Can the session access sensitive systems or data and also change state or send data out?
- Yes:
Two properties. Acceptable only if every input really is trusted. Recheck whenever a new data source is added.
- No:
Low risk on this test. Still apply least privilege and logging.
- Yes:
In practice, teams tend to remove one leg of the triangle. A research agent that browses the open web can be denied access to internal documents. An email assistant that reads private mail can be restricted to drafting replies that a person sends. A coding agent that reads untrusted issues can be prevented from making network calls outside an approved list.
Where The Rules Fall Short
Neither framework claims to be complete. An agent that reads untrusted input and can delete records in a production system can do serious damage without stealing anything. Willison’s trifecta is about data theft, so it does not set out to cover that case. Meta’s rule does, because write access to a sensitive system counts as its second property and the deletion itself counts as its third.
They also assume you know which inputs are untrusted. In real systems that boundary blurs. A document in an internal drive may have been uploaded by a supplier. A memory store may hold text that arrived from the web weeks ago. Treat anything that a person outside the organisation could have influenced as untrusted.
Finally, both rules stop at a single agent. When agents pass work to each other, one agent’s output becomes another’s untrusted input, and the combined system can satisfy all three properties even when each part satisfies two. Later articles in this group cover that problem.
The broader principle is shared with Google’s guidance that an agent’s powers should be carefully limited.4 Three independent sources reached the same conclusion from different directions, which is a good sign that it will last longer than any particular product feature.
Footnotes
-
Simon Willison, “The lethal trifecta for AI agents: private data, untrusted content, and external communication”, 16 June 2025. simonwillison.net ↩ ↩2
-
Meta AI, “Agents Rule of Two: A Practical Approach to AI Agent Security”, 31 October 2025. ai.meta.com ↩ ↩2
-
I. Ravia, Aim Labs, “Breaking down ‘EchoLeak’, the First Zero-Click AI Vulnerability Enabling Data Exfiltration from Microsoft 365 Copilot”, 2025, now published by Cato Networks. catonetworks.com ↩
-
S. Díaz, C. Kern and K. Olive, Google, “Google’s Approach for Secure AI Agents: An Introduction”, May 2025. research.google ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.