Jailbreaks And Adversarial Examples: Making Models Misbehave On Purpose
Adversarial examples fool a model's judgement; jailbreaks talk a model out of its safety rules. Both exploit the same weakness, and neither has a complete fix. Here is how defenders should think about them.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
A model can be fooled without anyone touching its training data or its code. An attacker only needs to choose the input carefully. For classifiers, such as image, speech or malware detection models, this is called an evasion attack, and the crafted input is an adversarial example: something chosen or altered so that the deployed model reaches the wrong answer. Often the change is tiny. For chat models the best-known version is the jailbreak: an input crafted to make the model produce output its developer tried to forbid.
This article explains both, why they are related, and why the honest answer to “can we stop them?” is “we can make them harder and limit the damage”. It describes the techniques at a conceptual level only.
Adversarial Examples: Crafted Inputs, Wrong Answers
In 2013 a team of researchers including Christian Szegedy and Ian Goodfellow reported a surprising property of neural networks. A change to an image too small for a person to notice could make a well-trained classifier label it as something else entirely, and the same altered image often fooled other networks trained separately.1 The change is not random noise. It is computed to push the model’s decision in the wrong direction.
A year later Goodfellow and two colleagues offered an explanation. They argued that the weakness comes from how close to linear these networks behave, rather than from some quirk of overfitting, and the insight gave them a quick way to generate adversarial examples. They then used those examples as extra training data, an early form of the adversarial training defence discussed below.2
The attack also works off the screen. In 2017 Kevin Eykholt and colleagues fixed black and white stickers to a real stop sign so that a road sign classifier read it as a different sign they had chosen. The trick worked in all of their laboratory images and in 84.8 percent of video frames filmed from a moving vehicle.3 The stickers were plain to see, which shows that an adversarial input does not have to be invisible to work.
The finding generalises well beyond pictures. NIST’s taxonomy describes evasion attacks against audio, video, text and cybersecurity models, such as malware or network traffic classifiers. It also records evasion seen outside the laboratory, for example against face recognition used for identity checks and against phishing detectors.4 In a security product, an evasion attack means a malicious file or connection that the model waves through.
Jailbreaks: Talking A Model Out Of Its Rules
Chat models are trained to refuse some requests. A jailbreak is an input designed to get around those refusals. NIST classes it as a type of direct prompting attack whose goal is to enable misuse.4 It differs from prompt injection mainly in what is targeted: a jailbreak attacks the model developer’s safety rules, while prompt injection usually attacks an application and its users. In practice the same tricks serve both, and OWASP treats jailbreaking as a subset of prompt injection.5
Three lines of research show the range of approaches:
- Hand-written jailbreaks exploit tensions inside the model’s training, for example between being helpful and being safe, or find requests phrased in ways the safety training never covered. NIST calls these competing objectives and mismatched generalisation.4
- Automated adversarial suffixes. In 2023 Andy Zou and colleagues used an automated search to find strings of characters that, appended to a harmful request, made models comply. Suffixes found on open models transferred to commercial chatbots including ChatGPT, Bard and Claude.6 This is the adversarial example idea applied to text.
- Many-shot jailbreaking. In April 2024 Anthropic described filling a long context window with many fabricated dialogues in which an assistant complies with harmful requests. With enough examples, the model followed the pattern. Effectiveness grew with the number of examples in the same way as ordinary in-context learning, which suggests the weakness comes from a useful capability rather than a simple bug.7
| Evasion Of A Classifier | Jailbreak Of A Chat Model | |
|---|---|---|
| Typical target | Image, audio, malware or fraud classifier | Chat assistant with safety training |
| What the attacker changes | A crafted change to a legitimate input, often very small, sometimes physical such as stickers | The wording, framing or context of a request |
| What goes wrong | Wrong label, such as malicious marked as benign | Refused content is produced |
| NIST AI 100-2 class | Evasion (predictive AI) | Direct prompting attack, misuse goal (generative AI) |
| MITRE ATLAS technique (v2026.09) | AML.T0015 Evade AI Model | AML.T0054 LLM Jailbreak |
| Main mitigations | Adversarial training, input checks, defence in depth | Safety training, input and output classifiers, monitoring, limits on what output can trigger |
The ATLAS technique IDs come from MITRE’s ATLAS knowledge base of attacks on AI systems, release 2026.09.8 We look at ATLAS and the other frameworks in the last article of this group.
Why There Is No Complete Fix
Defences exist and help. Adversarial training, which adds attack examples to the training data, makes classifiers more resistant, but NIST notes it is expensive and usually costs some accuracy on normal inputs.4 For chat models, Anthropic reported that classifying and modifying prompts before they reached the model cut the success rate of one many-shot attack from 61 percent to 2 percent.7
None of these closes the door. NIST’s 2025 report points to theoretical results suggesting that systems fully resistant to attack may be mathematically impossible without additional assumptions, and to the difficulty of even measuring progress on defences.4 Each new defence tends to be followed by a new attack that gets around it.
A Practical Defensive Posture
For teams deploying models, the useful question is not “is our model jailbreak-proof?” but “what happens when it is jailbroken?”. The steps below give a workable routine.
- Decide What Failure Would Cost
List the outputs or decisions that would cause real harm if the model were fooled, and the people or systems they reach.
- Add Controls Outside The Model
Use input and output classifiers, allow lists for actions, and human review for high-stakes decisions.
- Red Team Before Release
Test with known jailbreak and evasion techniques, including automated ones, and record what gets through.
- Monitor In Production
Log inputs and outputs, watch for unusual patterns such as very long prompts, and feed findings back into testing.
- Re-Test After Every Model Change
A new model version or provider can behave differently, so earlier test results do not carry over.
Where jailbreaks matter most is when model output can trigger actions. A jailbroken chatbot that writes something offensive is a reputational problem; a jailbroken agent that can send email or move money is a security incident. The second case is covered in Agentic AI Security.
Footnotes
-
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow and R. Fergus, “Intriguing properties of neural networks”, arXiv 1312.6199, December 2013. arxiv.org ↩
-
I. J. Goodfellow, J. Shlens and C. Szegedy, “Explaining and Harnessing Adversarial Examples”, arXiv 1412.6572, December 2014. arxiv.org ↩
-
K. Eykholt et al., “Robust Physical-World Attacks on Deep Learning Models”, arXiv 1707.08945, July 2017. arxiv.org ↩
-
NIST, AI 100-2 E2025, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations”, March 2025, sections 2.2, 3.3 and 4.1. csrc.nist.gov ↩ ↩2 ↩3 ↩4 ↩5
-
OWASP GenAI Security Project, “LLM01:2026 Prompt Injection”, OWASP Top 10 for LLM Applications 2026, August 2026. github.com ↩
-
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter and M. Fredrikson, “Universal and Transferable Adversarial Attacks on Aligned Language Models”, arXiv 2307.15043, July 2023. arxiv.org ↩
-
Anthropic, “Many-shot jailbreaking”, 2 April 2024. anthropic.com ↩ ↩2
-
MITRE, “ATLAS data”, release 2026.09, 15 September 2026, file dist/v6/ATLAS-2026.09.yaml. github.com ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.