AI Red Teaming: How to Stress-Test a Model Before Launch
Table of Contents
Table of Contents
Share

Learn how I would scope AI red teaming before a 2026 launch: the attack classes to test, open-source tools, EU AI Act duties and the release gate to set.
Frequently Asked Questions
- An evaluation measures how often a model behaves well on a fixed test set. Red teaming searches for the inputs and sequences that make it behave badly, including multi-turn conversations, documents it retrieves and tools it can call. Microsoft's AI red team, drawing on more than 100 generative AI products, lists 'AI red teaming is not safety benchmarking' among its eight lessons (arXiv, January 2025). You need both: evals tell you the average case, red teaming tells you the worst case an attacker can reach.
- Probably not as an explicit duty today. Article 55(1)(a) of Regulation (EU) 2024/1689 requires documented adversarial testing from providers of general-purpose AI models with systemic risk, and Chapter V has applied since 2 August 2025. A high-risk system must be resilient against attacks such as data poisoning and adversarial examples under Article 15(5), and Regulation (EU) 2026/1744 moved those obligations to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems (EUR-Lex, July 2026). Most startups building on a hosted model sit outside the first group, so testing is a product decision first.
- Three are worth knowing. NVIDIA's garak is a scanner that checks whether a model can be made to fail through hallucination, data leakage, prompt injection, misinformation, toxicity generation and jailbreaks. Microsoft's PyRIT is an open-source Python framework that helps security engineers identify risks in generative AI systems. Promptfoo is a command-line tool and library for evaluating and red teaming LLM applications, and it fits into a CI pipeline. Scanners give you a fast baseline; they do not replace a person running multi-turn attacks against your own tools and data.
- When a written release gate is met, not when the testing budget runs out. The gate I would sign has five conditions: no open critical finding that an unauthenticated user can reach, a hard limit outside the model on every action that moves money or data, a regression suite built from past findings that passes, a measured over-refusal rate the product owner accepts, and a named owner for post-launch monitoring. If any condition fails, the launch slips or the feature ships with less autonomy.
Don't Miss What's Next
Subscribe to newsletter
AI Red Teaming
LLM Security
Founder Perspectives
Prompt Injection
EU AI Act
AI
Get in Touch
Our team will get back to you within 24 hours.
















