Testing Agents Before Attackers Do: Benchmarks And Large-Scale Red Teaming
How AgentDojo, InjecAgent and a large public red-teaming competition measure agent hijacking, what they found, and how to use their results without being misled.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
Agent products often claim some resistance to prompt injection. The only way to know how much is to attack it under controlled conditions and count. That is harder for agents than for chat models, because success is not a bad sentence in a reply. It is an action: an email sent, a file read, a payment made.
This article explains how the main public agent security benchmarks work, what a large public red-teaming competition found, and how to read the numbers. It is aimed at teams who commission tests, read vendor claims or build their own evaluations. The general practice of red teaming language models is covered in the AI Security group.
How Agent Benchmarks Work
Agent benchmarks give an agent a realistic task in a simulated environment with working tools, then hide an attacker’s instruction somewhere the agent will read it. The test records two things: whether the agent completed the user’s task, and whether it also carried out the attacker’s.
- Set Up An Environment
Simulated email, calendar, files, banking, travel booking or chat, with tools the agent can call.
- Give A User Task
For example, summarise today’s messages or pay an outstanding bill.
- Plant An Injection Task
An attacker instruction is placed in data the agent will read, such as an email body or a review on a web page.
- Run The Agent
The agent plans and calls tools as it would in real use.
- Score Both Outcomes
Utility: was the user task done? Attack success: did the agent perform the attacker’s action?
- Repeat And Vary
Try different attacks, defences, phrasings and multiple attempts, then report results per task as well as in aggregate.
Measuring both outcomes matters. A defence that blocks every attack by refusing to do anything useful scores well on security and badly on utility. Good evaluations report the trade-off.
AgentDojo And InjecAgent
Two academic benchmarks underpin much of the public work.
AgentDojo, from researchers including Edoardo Debenedetti and Florian Tramèr, describes itself as a dynamic environment rather than a fixed test set. It contains 97 realistic user tasks and 629 security test cases across settings such as email, e-banking and travel booking.1 Two results stood out to its authors. Leading models struggled with many of the tasks even when nobody was attacking them, and the attacks then available succeeded against some of the security goals they measured but not every one. NIST later used and extended it.
InjecAgent, by Qiusi Zhan, Zhixiang Liang, Zifan Ying and Daniel Kang, appeared in the Findings of ACL 2024. It has 1,054 test cases built from 17 user tools and 62 attacker tools, and covers two harms: direct harm to the user and theft of private data. In its evaluation of 30 agents, a GPT-4 agent using the ReAct prompting style, in which the model alternates between reasoning steps and tool calls, followed the injected instruction 24 percent of the time, and the rate nearly doubled when the attacker’s text was reinforced with an extra “hacking prompt”.2
| Benchmark | Size | What it measures | Headline finding |
|---|---|---|---|
| AgentDojo (2024) | 97 user tasks, 629 security test cases | Utility and attack success together, in an extensible environment | Known attacks only partly succeeded, and models often failed the user task even with no attack present |
| InjecAgent (2024) | 1,054 test cases, 17 user tools, 62 attacker tools | Indirect injection leading to user harm or data theft | ReAct-prompted GPT-4 was vulnerable 24 percent of the time, nearly doubling with a reinforced prompt |
What NIST Learned From AgentDojo
In January 2025, NIST published lessons from testing an agent built on Claude 3.5 Sonnet (the October 2024 update) in AgentDojo’s workspace, travel, Slack and banking environments.3 Four findings stand out.
First, shared frameworks need constant improvement; NIST added new scenarios to AgentDojo. Second, defences tuned to known attacks can look strong until new attacks appear: on held-out workspace tasks, the strongest existing attack succeeded 11 percent of the time, while the strongest new attack, developed with red teamers from the UK AI Security Institute, succeeded 81 percent of the time. Third, aggregate figures hide detail. Across five injection tasks, the average success rate was 57 percent, but individual tasks varied widely. Fourth, a single attempt understates risk: allowing 25 attempts raised average success from 57 to 80 percent.
A Large Public Red-Teaming Competition
Gray Swan hosted a large public red-teaming competition aimed at agents. NIST’s Center for AI Standards and Innovation (CAISI) then analysed the data with Gray Swan, the UK AI Security Institute and several frontier AI developers, and summarised the findings in a blog post on 23 March 2026.4 More than 400 participants made over 250,000 attack attempts against 13 frontier models, in scenarios covering tool-using agents, coding agents and computer-use agents.
The findings are sobering for anyone hoping a model upgrade will solve the problem.4
- At least one successful attack was found against every model tested.
- Resistance to attack varied a lot between models and did not track general capability in any simple way.
- Some attack patterns worked across many scenarios and models, which CAISI suggests may reflect shared weaknesses in how models follow instructions.
- Attacks that beat the more resistant models tended to work on weaker ones, but not the other way round.
CAISI has also reported, in earlier work cited in the same post, that the DeepSeek models it evaluated were easier to hijack than leading US models.4
Using Evaluations In Practice
The joint government guidance from May 2026 treats testing as a core control. It recommends red teaming in sandboxed environments before production, probing for unexpected abilities, multi-agent red teaming and chaos testing, regular assessment of whether an agent can get around guardrails, monitors, approvals and input filters, and continuous evaluation as agents change.5 It also notes the limits of current methods: results can be sensitive to small changes in wording, vary by scenario and only partly reflect real deployments.
For a team commissioning or running tests, a few practices help:
- Test the agent as deployed, with its real tools, permissions and data sources, not the bare model.
- Report utility and attack success together, per task, with the number of attempts stated.
- Include adaptive attacks written for your system, not only published ones.
- Re-test whenever the model, tools, prompts or permissions change.
- Treat results as a measure of how often defences fail, then size the blast radius on the assumption that they will.
That last point ties the series together. The evidence from benchmarks and from the competition CAISI analysed shows that, as of October 2026, every frontier model tested fell to at least one attack in the scenarios tried. Testing tells you how often. Design choices such as the Rule of Two, least privilege and approval gates decide how much damage each success can do.
Footnotes
-
E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer and F. Tramèr, “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents”, arXiv 2406.13352, June 2024. arxiv.org ↩
-
Q. Zhan, Z. Liang, Z. Ying and D. Kang, “InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents”, Findings of ACL 2024, arXiv 2403.02691. arxiv.org ↩
-
NIST, “Technical Blog: Strengthening AI Agent Hijacking Evaluations”, 17 January 2025. nist.gov ↩
-
NIST CAISI, “Insights into AI Agent Security from a Large-Scale Red-Teaming Competition”, 23 March 2026. nist.gov ↩ ↩2 ↩3
-
ASD’s ACSC, CISA, NSA, Canadian Centre for Cyber Security, NCSC-NZ and NCSC-UK, “Careful Adoption of Agentic AI Services”, 1 May 2026. ncsc.govt.nz ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.