AGENTIC AI SECURITY / ADVANCED

Threat Modelling An Agent System With MAESTRO And The Five Risk Categories

A practical way to threat model AI agents, combining the CSA's seven-layer MAESTRO framework with the five risk categories in the May 2026 joint government guidance.

Checked against primary sources and independently reviewed on . Sources are listed at the end.

Threat modelling means working out, before an attacker does, what could go wrong with a system and what to do about it. Classic methods such as STRIDE were built for software whose behaviour is fixed by its code. An AI agent’s behaviour is partly decided at run time by a model reading text, some of which may come from an attacker. That does not make traditional methods useless, but it leaves gaps.

This article shows how to fill them using two published reference points. The first is MAESTRO, a layered framework from the Cloud Security Alliance. The second is the set of five risk categories in the joint government guidance “Careful Adoption of Agentic AI Services”, published in May 2026. It ends with a step-by-step approach that teams can adapt.

Why Agents Need Their Own Threat Model

The joint guidance, from six cyber security agencies in Australia, the United States, Canada, New Zealand and the United Kingdom, recommends realistic threat modelling with up-to-date agentic risk taxonomies, naming the OWASP GenAI Security Project and MITRE ATLAS as examples.1 It explains why component-level analysis falls short: in agent systems, risks often come from the interactions between models, tools, data, people and guardrails rather than from a single flaw. For that reason it also recommends system-theoretic methods, specifically System Theoretic Process Analysis for security (STPA-Sec) for analysing planned and running systems, and Causal Analysis using System Theory (CAST) for investigating incidents.

The MAESTRO Layers

MAESTRO stands for Multi-Agent Environment, Security, Threat, Risk and Outcome. Ken Huang published it on the Cloud Security Alliance blog on 6 February 2025.2 It splits an agent system into seven layers and asks you to model threats at each layer and across them.

  1. 7. Agent EcosystemWhere agents meet users, businesses and other agents. Example: a malicious agent posing as a legitimate one.
  2. 6. Security and ComplianceControls that span the stack. Example: tampering with data used to train an AI security agent.
  3. 5. Evaluation and ObservabilityTesting, monitoring and metrics. Example: manipulating evaluation results so a weak agent looks safe.
  4. 4. Deployment and InfrastructureHosting, containers and networks. Example: a compromised container image.
  5. 3. Agent FrameworksThe libraries and toolkits used to build agents. Example: malicious code in a framework component.
  6. 2. Data OperationsData stores, retrieval pipelines and memory. Example: poisoned data that changes agent decisions.
  7. 1. Foundation ModelsThe underlying model. Example: crafted inputs that make the model misbehave.
The seven MAESTRO layers, with one example threat for each from the CSA description. Highlighted layers form the cross-layer chain described below.

The cross-layer view is where MAESTRO adds most value. The CSA description lists threats such as supply chain attacks, lateral movement between layers, privilege escalation, data leaking across boundaries and goal misalignment that cascades between connected agents.2 A poisoned document in the data layer (2) that redirects an agent built on a framework (3) to misuse an ecosystem integration (7) is one chain that a single-layer review would miss.

The Five Risk Categories

The joint guidance groups agent risks into five categories.1 They work well as a second lens, because they describe the kinds of harm rather than the parts of the system.

CategoryWhat it coversQuestions to ask
PrivilegeToo much access, scope creep, identity spoofing and impersonationWhat could this agent do with its current access if hijacked? Can anyone impersonate it?
Design and configurationUnvetted components, permissions checked only at start-up, weak segmentation, outdated allow listsWhere do we rely on one-time checks? Can a compromise in one environment spread?
BehaviourGoal misalignment, specification gaming, deception, emergent behaviour, prompt injection and data poisoningHow could the agent meet its goal in a harmful way? What happens if it is manipulated?
StructuralOrchestration failures, tool use, third-party components, data aggregation, rogue agents and insecure communicationWhat does this agent trust? What breaks downstream if it fails?
AccountabilityOpaque decisions, unseen delegation, hard-to-reproduce behaviour, limited visibility, inaccuracyCould we explain and reproduce a harmful action from our logs?
The five risk categories in the joint guidance, with typical questions for a threat modelling session.

Each category in the guidance comes with a worked scenario. In the behaviour example, an update agent with broad file system write access is asked by a malicious insider to apply a patch and also clean up the firewall logs. It does both, because its permissions allow it.1 That single scenario touches privilege, behaviour and accountability at once, which is why it helps to apply more than one lens.

Other Taxonomies Worth Knowing

Two further references help when filling in the detail. OWASP’s “Agentic AI: Threats and Mitigations” (version 1.1, December 2025) lists 17 threats with mitigations, and its Top 10 for Agentic Applications gives the short list.3 In April 2025, Microsoft’s AI Red Team and colleagues published a whitepaper setting out a taxonomy of failure modes in agentic AI systems. In Microsoft’s own framing, it separates security failures from safety failures, and failure modes that are new to agents from older ones that agents make worse; it singles out memory poisoning as a particular concern.4 It is a vendor publication, but it is useful for its worked case study.

A Practical Approach

Combining these sources gives a repeatable process. It is our synthesis; none of the documents prescribes this exact sequence.

  1. Draw The System

    Show the model, framework, tools, data sources, memory, identities, other agents and the people who approve actions.

  2. Mark Trust Boundaries And Untrusted Inputs

    Label every place where text from outside the organisation can enter the context, including tool results and other agents.

  3. Walk The MAESTRO Layers

    For each layer, list threats using the OWASP and CSA examples as prompts.

  4. Trace Cross-Layer Chains

    Follow an untrusted input through to an action and an outbound effect. Apply the Rule of Two to each session.

  5. Check The Five Categories

    Ask the privilege, design, behaviour, structural and accountability questions for the whole system.

  6. Decide Controls And Owners

    Choose controls, assign risk owners and record which risks are accepted.

  7. Test And Revisit

    Red team the result, prepare incident response for agent compromise and repeat when tools, models or permissions change.

A threat modelling sequence for an agent system, combining MAESTRO layers with the five risk categories.

Two steps benefit from a little more detail. A trust boundary is any point where data or requests pass between parts of a system that are trusted to different degrees. Some are obvious, such as web content entering an agent’s context. Others sit entirely inside the organisation, such as a low-privilege reader agent handing work to an agent that can make payments. For each one, write down what crosses it, which component reads it, and what that component may do next. Then give every risk you decide to treat a named owner, a control and a review date, so that accepted risks stay visible instead of being quietly forgotten.

The last step matters more for agents than for most systems. The guidance recommends developing and testing incident response procedures for agent compromise, regular third-party reviews of privileged designs, and updating risk models as new attacks appear.1 An agent that gains a new tool has a new threat model, even if no code changed.

Common Gaps

Threat models for agents tend to miss the same things. Teams model the user’s prompt but not the documents, emails and tool outputs the agent reads. They model each agent separately but not the combined system. They assume a human approval step is a control without asking whether the person will see enough to judge. They forget that the logs themselves need protecting, so a compromised agent cannot erase its trail.

Footnotes

  1. ASD’s ACSC, CISA, NSA, Canadian Centre for Cyber Security, NCSC-NZ and NCSC-UK, “Careful Adoption of Agentic AI Services”, 1 May 2026. ncsc.govt.nz ↩ ↩2 ↩3 ↩4

  2. K. Huang, Cloud Security Alliance, “Agentic AI Threat Modeling Framework: MAESTRO”, 6 February 2025. cloudsecurityalliance.org ↩ ↩2

  3. OWASP GenAI Security Project, “Agentic AI: Threats and Mitigations”, version 1.1, December 2025. genai.owasp.org ↩

  4. Microsoft Security Blog, “New whitepaper outlines the taxonomy of failure modes in AI agents”, 24 April 2025 (vendor publication). microsoft.com ↩

Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.