Threat Modelling An Agent System With MAESTRO And The Five Risk Categories
A practical way to threat model AI agents, combining the CSA's seven-layer MAESTRO framework with the five risk categories in the May 2026 joint government guidance.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
Threat modelling means working out, before an attacker does, what could go wrong with a system and what to do about it. Classic methods such as STRIDE were built for software whose behaviour is fixed by its code. An AI agent’s behaviour is partly decided at run time by a model reading text, some of which may come from an attacker. That does not make traditional methods useless, but it leaves gaps.
This article shows how to fill them using two published reference points. The first is MAESTRO, a layered framework from the Cloud Security Alliance. The second is the set of five risk categories in the joint government guidance “Careful Adoption of Agentic AI Services”, published in May 2026. It ends with a step-by-step approach that teams can adapt.
Why Agents Need Their Own Threat Model
The joint guidance, from six cyber security agencies in Australia, the United States, Canada, New Zealand and the United Kingdom, recommends realistic threat modelling with up-to-date agentic risk taxonomies, naming the OWASP GenAI Security Project and MITRE ATLAS as examples.1 It explains why component-level analysis falls short: in agent systems, risks often come from the interactions between models, tools, data, people and guardrails rather than from a single flaw. For that reason it also recommends system-theoretic methods, specifically System Theoretic Process Analysis for security (STPA-Sec) for analysing planned and running systems, and Causal Analysis using System Theory (CAST) for investigating incidents.
The MAESTRO Layers
MAESTRO stands for Multi-Agent Environment, Security, Threat, Risk and Outcome. Ken Huang published it on the Cloud Security Alliance blog on 6 February 2025.2 It splits an agent system into seven layers and asks you to model threats at each layer and across them.
- 7. Agent EcosystemWhere agents meet users, businesses and other agents. Example: a malicious agent posing as a legitimate one.
- 6. Security and ComplianceControls that span the stack. Example: tampering with data used to train an AI security agent.
- 5. Evaluation and ObservabilityTesting, monitoring and metrics. Example: manipulating evaluation results so a weak agent looks safe.
- 4. Deployment and InfrastructureHosting, containers and networks. Example: a compromised container image.
- 3. Agent FrameworksThe libraries and toolkits used to build agents. Example: malicious code in a framework component.
- 2. Data OperationsData stores, retrieval pipelines and memory. Example: poisoned data that changes agent decisions.
- 1. Foundation ModelsThe underlying model. Example: crafted inputs that make the model misbehave.
The cross-layer view is where MAESTRO adds most value. The CSA description lists threats such as supply chain attacks, lateral movement between layers, privilege escalation, data leaking across boundaries and goal misalignment that cascades between connected agents.2 A poisoned document in the data layer (2) that redirects an agent built on a framework (3) to misuse an ecosystem integration (7) is one chain that a single-layer review would miss.
The Five Risk Categories
The joint guidance groups agent risks into five categories.1 They work well as a second lens, because they describe the kinds of harm rather than the parts of the system.
| Category | What it covers | Questions to ask |
|---|---|---|
| Privilege | Too much access, scope creep, identity spoofing and impersonation | What could this agent do with its current access if hijacked? Can anyone impersonate it? |
| Design and configuration | Unvetted components, permissions checked only at start-up, weak segmentation, outdated allow lists | Where do we rely on one-time checks? Can a compromise in one environment spread? |
| Behaviour | Goal misalignment, specification gaming, deception, emergent behaviour, prompt injection and data poisoning | How could the agent meet its goal in a harmful way? What happens if it is manipulated? |
| Structural | Orchestration failures, tool use, third-party components, data aggregation, rogue agents and insecure communication | What does this agent trust? What breaks downstream if it fails? |
| Accountability | Opaque decisions, unseen delegation, hard-to-reproduce behaviour, limited visibility, inaccuracy | Could we explain and reproduce a harmful action from our logs? |
Each category in the guidance comes with a worked scenario. In the behaviour example, an update agent with broad file system write access is asked by a malicious insider to apply a patch and also clean up the firewall logs. It does both, because its permissions allow it.1 That single scenario touches privilege, behaviour and accountability at once, which is why it helps to apply more than one lens.
Other Taxonomies Worth Knowing
Two further references help when filling in the detail. OWASP’s “Agentic AI: Threats and Mitigations” (version 1.1, December 2025) lists 17 threats with mitigations, and its Top 10 for Agentic Applications gives the short list.3 In April 2025, Microsoft’s AI Red Team and colleagues published a whitepaper setting out a taxonomy of failure modes in agentic AI systems. In Microsoft’s own framing, it separates security failures from safety failures, and failure modes that are new to agents from older ones that agents make worse; it singles out memory poisoning as a particular concern.4 It is a vendor publication, but it is useful for its worked case study.
A Practical Approach
Combining these sources gives a repeatable process. It is our synthesis; none of the documents prescribes this exact sequence.
- Draw The System
Show the model, framework, tools, data sources, memory, identities, other agents and the people who approve actions.
- Mark Trust Boundaries And Untrusted Inputs
Label every place where text from outside the organisation can enter the context, including tool results and other agents.
- Walk The MAESTRO Layers
For each layer, list threats using the OWASP and CSA examples as prompts.
- Trace Cross-Layer Chains
Follow an untrusted input through to an action and an outbound effect. Apply the Rule of Two to each session.
- Check The Five Categories
Ask the privilege, design, behaviour, structural and accountability questions for the whole system.
- Decide Controls And Owners
Choose controls, assign risk owners and record which risks are accepted.
- Test And Revisit
Red team the result, prepare incident response for agent compromise and repeat when tools, models or permissions change.
Two steps benefit from a little more detail. A trust boundary is any point where data or requests pass between parts of a system that are trusted to different degrees. Some are obvious, such as web content entering an agent’s context. Others sit entirely inside the organisation, such as a low-privilege reader agent handing work to an agent that can make payments. For each one, write down what crosses it, which component reads it, and what that component may do next. Then give every risk you decide to treat a named owner, a control and a review date, so that accepted risks stay visible instead of being quietly forgotten.
The last step matters more for agents than for most systems. The guidance recommends developing and testing incident response procedures for agent compromise, regular third-party reviews of privileged designs, and updating risk models as new attacks appear.1 An agent that gains a new tool has a new threat model, even if no code changed.
Common Gaps
Threat models for agents tend to miss the same things. Teams model the user’s prompt but not the documents, emails and tool outputs the agent reads. They model each agent separately but not the combined system. They assume a human approval step is a control without asking whether the person will see enough to judge. They forget that the logs themselves need protecting, so a compromised agent cannot erase its trail.
Footnotes
-
ASD’s ACSC, CISA, NSA, Canadian Centre for Cyber Security, NCSC-NZ and NCSC-UK, “Careful Adoption of Agentic AI Services”, 1 May 2026. ncsc.govt.nz ↩ ↩2 ↩3 ↩4
-
K. Huang, Cloud Security Alliance, “Agentic AI Threat Modeling Framework: MAESTRO”, 6 February 2025. cloudsecurityalliance.org ↩ ↩2
-
OWASP GenAI Security Project, “Agentic AI: Threats and Mitigations”, version 1.1, December 2025. genai.owasp.org ↩
-
Microsoft Security Blog, “New whitepaper outlines the taxonomy of failure modes in AI agents”, 24 April 2025 (vendor publication). microsoft.com ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.