Red Teaming
Red Teaming
The process of testing the safety, security and performance of an Al system through an adversarial lens, typically through the simulation of adversarial attacks on the model to evaluate it against certain benchmarks, jailbreak it and try to make it behave in unintended or inappropriate ways. Red teaming reveals security risks, model flaws, biases, misinformation and other harms, and the results of such testing are passed along to the model developers for evaluation and remediation. Developers use red teaming to improve a model before and after releasing it to the public.