World Model · podcast knowledge graph

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

2026-06-22 · 66 min · 13 entities

Asserted relationships

  • → references Matt Fredrikson person
    0.77
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://linkedin.com/in/matt-fredrikson-7596349
  • → discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • → discusses Technology concept
    0.40
    evidence rules-v4
    Feed category: Technology
  • → references grayswan.ai website
    0.38
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://grayswan.ai/
  • → references zicokolter.com website
    0.38
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://zicokolter.com/
  • → references mattfredrikson.com website
    0.38
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://mattfredrikson.com/

Entities found in this episode

websites 6

  • mentioned grayswan.ai website
    0.45
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://grayswan.ai/
  • mentioned zicokolter.com website
    0.45
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://zicokolter.com/
  • mentioned mattfredrikson.com website
    0.45
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://mattfredrikson.com/
  • references grayswan.ai website
    0.38
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://grayswan.ai/
  • references zicokolter.com website
    0.38
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://zicokolter.com/
  • references mattfredrikson.com website
    0.38
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://mattfredrikson.com/

concepts 5

  • mentioned Science concept
    0.50
    evidence rules-v4
    Feed category: Science
  • mentioned Technology concept
    0.50
    evidence rules-v4
    Feed category: Technology
  • evidence rules-v4
    Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan
  • discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • discusses Technology concept
    0.40
    evidence rules-v4
    Feed category: Technology

persons 2

  • mentioned Matt Fredrikson person
    0.90
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://linkedin.com/in/matt-fredrikson-7596349
  • references Matt Fredrikson person
    0.77
    evidence rules-v4
    Link in episode "Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan": https://linkedin.com/in/matt-fredrikson-7596349
Episode description as stored
AI Engineer World’s Fair regular bird tix will sell out ~today! Join us next week ahead of the Late Bird price hike and get >$40,000 in sponsor credits for attending ! Thanks to the US Government issuing an export control directive on Mythos and Fable , the risks of jailbreaks and (industry term) indirect prompt injection are suddenly the talk of the town, though we have been covering AI security for a few years now, from Hackaprompt to the enigmatic Pliny the Elder . Zico Kolter, member of OpenAI’s board of directors on the Safety & Security Committee , and Matt Fredrikson, CMU professor and CEO of Gray Swan , co-authored the definitive paper on Indirect Prompt Injections , and Gray Swan were cited authorities on the Mythos model card , directly investigating the exact capabilities that are under scrutiny right now: We seized the opportunity to ask them the state of AI Red Teaming, and Shade , the adversarial red teaming tool that Anthropic used to evaluate the robustness of their models against prompt injection attacks in coding environments. Shade is part of their overall toolkit covering Simon Willison’s Lethal Trifecta , including Cygnal , an AI guardrails product, and the world’s largest AI Red Teaming Arena , including AIRT celebrity Wyatt Walls . All of this security tooling, and yet, we’re only staving off the inevitable. The risks of extremely smart AI increasingly feel like gray swan events: an event that everyone can see coming. In this episode, Gray Swan cofounders Zico Kolter and Matt Fredrikson join swyx to explain why AI security is not just “cybersecurity with AI,” why agents introduce a new class of vulnerabilities, and why the next major AI incident may be a gray swan: unlikely, but clearly visible before it happens. We go deep on prompt injection , automated red teaming , model robustness, agent identity, computer-use agents , enterprise guardrails, and the emerging AI insurance/compliance stack. Zico and Matt also explain why frontier models are not automatically safer as they scale, why specialized red-teaming models can now beat humans at breaking AI systems , and why the future of AI security may depend on AI systems attacking, defending, and interpreting other AI systems. We discuss: * Why AI systems need a different security mindset from traditional software * How prompt injection creates a new exploit class for agents like Codex and Claude Code * Gray Swan Arena and the rise of community red teaming * Shade : AI that can outperform humans at breaking models * Why LLMs are an alien form of intelligence that fail differently from humans * Human vs browser-agent robustness and why humans ranked fourth * Why eval awareness and capability elicitation matter * Cygnal : Gray Swan’s guardrail model for policy enforcement * Why bigger models do not automatically become more robust * The lethal trifecta : untrusted data, private data, and exfiltration * Why “just prompt it better” is not enough for enterprise AI security * OpenClaw , computer-use agents, and the agent security nightmare * Agent-native identity , permissions, and enterprise deployment * Why AI security may become part of insurance and compliance * Why the first major AI prompt-injection breach may be inevitable Gray Swan * Website: https://www.grayswan.ai/ Zico Kolter * X: https://x.com/zicokolter * Website: https://zicokolter.com/ * LinkedIn: https://www.linkedin.com/in/zico-kolter-560382a4/ Matt Fredrikson * Website: https://www.mattfredrikson.com/ * LinkedIn: https://www.linkedin.com/in/matt-fredrikson-7596349/ Timestamps 00:00:00 Introduction 00:02:31 Why AI Security Is Different 00:06:38 Testing Claude, Codex, and Prompt Injection 00:07:47 Gray Swan Arena and Automated Red Teaming 00:11:14 AI That Breaks Models Better Than Humans 00:14:00 LLMs as Alien Intelligence 00:19:00 Humans vs AI Agents 00:24:35 Red Teaming, Jailbreaks, and Capability Elicitation 00:26:11 Cygnal: Guardrails for AI Agents 00:34:04 The Lethal Trifecta 00:39:31 Can AI