Red Teaming AI: Finding the Fault Lines Before Attackers Do

A complete course on adversarial testing of AI systems: why models fail, the full attack taxonomy (prompt injection, jailbreaks, poisoning, extraction), the red team process used by Microsoft and the frontier labs, the open-source tooling (PyRIT, garak, ART), real-world incidents from Tay to Grand Theft Claude, and the NIST/EU/ISO governance landscape.

intermediate 90 min 11 lessons

Lessons

  • What Is AI Red Teaming?

    Origins of the practice โ€” from Microsoft's Tay wake-up call to today โ€” and how red teaming differs from penetration testing, QA, and benchmark evaluation.

    8 min
  • How AI Systems Fail

    The AI attack surface: harms taxonomy, capability vs safety vs security failures, and the layers attackers can target โ€” model, data, application, and infrastructure.

    8 min
  • The Red Team Process

    The end-to-end engagement: scoping, threat modeling, the red team charter, playbook, execution, reporting, remediation, and retesting โ€” based on Microsoft's experience across 100 products.

    10 min
  • Prompt Injection

    The top LLM vulnerability: direct and indirect injection, real-world cases from Chevy chatbots to Grand Theft Claude, and the mitigations that actually help.

    9 min
  • Jailbreaks & Why Refusals Fail

    How attackers bypass safety training: DAN and roleplay, many-shot jailbreaking, ASCII art, ciphers, and multi-turn crescendo attacks โ€” and why refusal training alone never holds.

    9 min
  • Beyond Prompts: Data & Model Attacks

    The non-prompt attack surface: adversarial examples, data poisoning and backdoors, model extraction, membership inference, and model inversion.

    9 min
  • The Attack Taxonomy: OWASP & MITRE ATLAS

    The two reference maps for AI security: every category in the OWASP LLM Top 10 and how MITRE ATLAS structures adversarial ML tactics and techniques.

    9 min
  • Standards & Governance: NIST, EU AI Act & ISO

    The regulatory and standards landscape: NIST AI RMF and the GenAI profile, the EU AI Act's systemic-risk obligations, ISO/IEC 42001, and how they shape red teaming duties.

    9 min
  • Tools of the Trade

    Hands-on tooling for red teams: PyRIT, garak, IBM ART, TextAttack, promptfoo, LLM-Attacks, CyberSecEval โ€” plus automated and LLM-vs-LLM red teaming.

    9 min
  • Red Teaming in the Real World

    How OpenAI, Anthropic, Google, Meta, and Microsoft actually run red teams and bounties, the incident timeline from Tay to Grand Theft Claude, and how to build your own program.

    10 min
  • Grand Quiz

    Twenty questions across the whole course โ€” attack taxonomy, process, tools, and governance. Answer all to see your grade.

    10 min