Encyclopedia · 187 concepts

Safety, Ethics & Governance · intermediate · concept 179 of 187

Red Teaming AI Systems

Systematically testing AI systems by trying to make them fail, produce harmful outputs, or bypass safety measures. Essential for finding vulnerabilities before deployment. Both manual and automated approaches.

Key terms

JailbreakingAdversarial attacksSafety evaluationPrompt injection

Learn these first

Guides and articles

Courses, papers, and more