Poisoned Data, Leaked Data And Stolen Models: Attacks Across The AI Lifecycle
Some attacks on AI target the model itself rather than the prompt. This article maps data poisoning, training data extraction and model theft to the lifecycle stage each one exploits.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
Prompt injection attacks the conversation. A second family of attacks goes after the model itself: the data it learned from, the behaviour baked into its weights, and the commercial value of the model as an asset. These attacks were studied long before today’s chat assistants, and in several cases they are cheap to carry out.
NIST’s taxonomy of adversarial machine learning gives a neutral way to organise them. It sorts attacks by the stage of the machine learning lifecycle they target and by what the attacker wants to break: availability, integrity or privacy.1 This article follows that framing and pins each attack to the stage where it happens, so you can see where controls belong.
Where Each Attack Strikes
Poisoning: Changing What The Model Learns
Data poisoning means slipping manipulated examples into the data a model learns from, so that it learns the wrong lesson. A backdoor is a targeted version: the model behaves normally until it sees a specific trigger, then does what the attacker planted. OWASP lists Data and Model Poisoning as LLM05 in its 2026 edition. Its entry also covers poisoned retrieval data and automated retraining loops, and notes that fixing poisoning can mean revalidating data, retraining or replacing the model, which is far harder than patching code.2
Two findings show why this matters for models trained on internet data. In 2023 Nicholas Carlini and colleagues showed that web-scale datasets built from internet links could be poisoned because the content behind a link can change after the dataset is published, so later users may download something different from what the original curators saw. They estimated that USD 60 would have been enough to poison 0.01 percent of the LAION-400M or COYO-700M image datasets.3
In October 2025 Anthropic, working with the UK AI Security Institute and the Alan Turing Institute, reported that as few as 250 poisoned documents planted a backdoor in models ranging from 600 million to 13 billion parameters. The number needed did not grow with model size. The authors stress the limits: the backdoor they tested made the model produce gibberish on a trigger phrase, and it is unclear whether the result holds for larger models or more harmful behaviours.4
A separate study, “Sleeper Agents”, found that deliberately trained backdoors could survive the standard safety training methods meant to remove unwanted behaviour, including supervised fine-tuning and adversarial training.5 Together these results suggest a cautious stance: you cannot rely on later training to clean up a model that learned from tainted data.
Extraction: Getting The Training Data Back Out
Models memorise some of what they see. A privacy attack tries to recover that material, or to learn whether a particular record was in the training set, which is called membership inference.
In 2021 researchers recovered hundreds of verbatim sequences from GPT-2’s training data, including names, phone numbers and email addresses, some of which appeared only once in the data. They also found that larger models memorised more.6 Two years later a follow-up study showed that aligned production chatbots were not immune: a simple prompt that pushed ChatGPT out of its normal conversational behaviour made it emit memorised training data at roughly 150 times the usual rate.7
For organisations, the lesson is about what goes into training and fine-tuning. OWASP’s 2026 text notes that once data has shaped weights, embeddings or adapters, it can stay extractable after the source records are deleted, which creates tension with data protection erasure rights.8
Model Theft: Copying The Asset
A trained model represents a large investment in compute and data. Model extraction attacks try to copy it, or parts of it, using nothing but the public API, and large language models are not exempt. In 2024 a team recovered one internal layer of OpenAI production models through ordinary API queries, for under USD 20 for the smaller Ada and Babbage models, and estimated that the same approach would cost under USD 2,000 for gpt-3.5-turbo.9 The result did not copy a whole model, but it showed that hidden details can leak through an API that was believed to expose only answers.
Controls By Stage
| Stage | Main Threat | Controls To Consider |
|---|---|---|
| Data collection | Poisoning | Track where data came from, pin dataset versions with hashes, filter and deduplicate, prefer curated sources |
| Training and fine-tuning | Backdoors, memorised secrets | Remove personal data and secrets before training, test for trigger behaviour, keep training pipelines access-controlled |
| Deployment | Tampered model files | Verify model provenance and integrity before loading (see the AI supply chain article) |
| Inference | Extraction and theft | Rate and cost limits, monitor unusual query patterns, limit detail returned by the API, filter outputs for sensitive data |
Data protection is the common thread. In May 2025 the NSA, CISA and the FBI, with partner agencies in Australia, New Zealand and the UK, issued joint guidance on securing the data used to train and operate AI systems. It treats the data supply chain and maliciously modified data as core risks and recommends provenance tracking and digital signatures for data.10 Agentic systems add new poisoning routes, such as tainted long-term memory, which we cover in Agentic AI Security.
Footnotes
-
NIST, AI 100-2 E2025, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations”, March 2025. csrc.nist.gov ↩
-
OWASP GenAI Security Project, “LLM05:2026 Data and Model Poisoning”, OWASP Top 10 for LLM Applications 2026, August 2026. github.com ↩
-
N. Carlini et al., “Poisoning Web-Scale Training Datasets is Practical”, arXiv 2302.10149, February 2023. arxiv.org ↩
-
Anthropic, UK AI Security Institute and The Alan Turing Institute, “A small number of samples can poison LLMs of any size”, 9 October 2025. anthropic.com ↩
-
E. Hubinger et al., “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training”, arXiv 2401.05566, January 2024. arxiv.org ↩
-
N. Carlini et al., “Extracting Training Data from Large Language Models”, USENIX Security 2021, arXiv 2012.07805. arxiv.org ↩
-
M. Nasr, N. Carlini et al., “Scalable Extraction of Training Data from (Production) Language Models”, arXiv 2311.17035, November 2023. arxiv.org ↩
-
OWASP GenAI Security Project, “LLM02:2026 Sensitive Information Disclosure”, OWASP Top 10 for LLM Applications 2026, August 2026. github.com ↩
-
N. Carlini et al., “Stealing Part of a Production Language Model”, arXiv 2403.06634, March 2024. arxiv.org ↩
-
CISA, NSA, FBI and international partners, “AI Data Security: Best Practices for Securing Data Used to Train & Operate AI Systems”, 22 May 2025. cisa.gov ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.