How It Works
AI red teaming
Deliberately attacking an AI system to find ways it can be made to fail, misbehave or cause harm — for example, coaxing a chatbot into giving dangerous instructions, leaking data, or ignoring its rules — so the problems can be fixed before release. Red teams may be internal staff, outside experts or members of the public at organized events.
Origin · no single documented coiner
Borrowed from military and later cybersecurity practice, where a “red team” plays the adversary against the defending “blue team.” AI labs adopted the term for pre-release testing around 2020–2022, and the 2023 U.S. executive order on AI and the DEF CON Generative Red Team event helped make it standard vocabulary.
