safety
Red Teaming
The practice of deliberately attempting to find vulnerabilities, biases, and failure modes in AI systems by adopting an adversarial perspective. Red teams test models with edge cases, manipulative prompts, and creative attacks to identify weaknesses before deployment.
In practice
A red team might try thousands of prompt variations to find ways to make a model produce harmful or biased content.