Encyclopedia · 176 concepts

Safety, Ethics & Governance · intermediate · concept 170 of 176

Red Teaming AI Systems

Systematically testing AI systems by trying to make them fail, produce harmful outputs, or bypass safety measures. Essential for finding vulnerabilities before deployment. Both manual and automated approaches.

Key terms

JailbreakingAdversarial attacksSafety evaluationPrompt injection

Learn these first

Guides and articles

Courses, papers, and more