Skip to content
Launchpad Library logo
← All glossary terms

How It Works

AI red teaming

Deliberately attacking an AI system to find ways it can be made to fail, misbehave or cause harm — for example, coaxing a chatbot into giving dangerous instructions, leaking data, or ignoring its rules — so the problems can be fixed before release. Red teams may be internal staff, outside experts or members of the public at organized events.

Origin · no single documented coiner

Borrowed from military and later cybersecurity practice, where a “red team” plays the adversary against the defending “blue team.” AI labs adopted the term for pre-release testing around 2020–2022, and the 2023 U.S. executive order on AI and the DEF CON Generative Red Team event helped make it standard vocabulary.

Read the source →