Safety, Ethics & Governance · intermediate · concept 170 of 176
Red Teaming AI Systems
Systematically testing AI systems by trying to make them fail, produce harmful outputs, or bypass safety measures. Essential for finding vulnerabilities before deployment. Both manual and automated approaches.
Key terms
JailbreakingAdversarial attacksSafety evaluationPrompt injection
Learn these first
Videos
▶ OWASP's Top 10 Ways to Attack LLMs: AI Vulnerabilities Exposed ↗
IBM Technology · YouTube
Guides and articles
Courses, papers, and more
This unlocks