World Model · podcast knowledge graph

Exploring the Biology of LLMs with Circuit Tracing with Emmanuel Ameisen - #727

2025-04-14 · 94 min · episode 727 · 13 entities

Asserted relationships

  • evidence rules-v4
    Exploring the Biology of LLMs with Circuit Tracing with Emmanuel Ameisen - #727
  • → hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • → discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • → discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • → discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • → discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News
  • → hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

Entities found in this episode

concepts 5

  • mentioned Science concept
    0.50
    evidence rules-v4
    Feed category: Science
  • mentioned Tech News concept
    0.50
    evidence rules-v4
    Feed category: Tech News
  • discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News

companys 4

  • mentioned TWIML company
    0.70
    evidence rules-v4
    Feed author/publisher: TWIML
  • mentioned Technology company
    0.50
    evidence rules-v4
    Feed category: Technology
  • discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

persons 2

  • mentioned Sam Charrington person
    0.70
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington

books 1

podcasts 1

Episode description as stored
In this episode, Emmanuel Ameisen, a research engineer at Anthropic, returns to discuss two recent papers: "Circuit Tracing: Revealing Language Model Computational Graphs" and "On the Biology of a Large Language Model." Emmanuel explains how his team developed mechanistic interpretability methods to understand the internal workings of Claude by replacing dense neural network components with sparse, interpretable alternatives. The conversation explores several fascinating discoveries about large language models, including how they plan ahead when writing poetry (selecting the rhyming word "rabbit" before crafting the sentence leading to it), perform mathematical calculations using unique algorithms, and process concepts across multiple languages using shared neural representations. Emmanuel details how the team can intervene in model behavior by manipulating specific neural pathways, revealing how concepts are distributed throughout the network's MLPs and attention mechanisms. The discussion highlights both capabilities and limitations of LLMs, showing how hallucinations occur through separate recognition and recall circuits, and demonstrates why chain-of-thought explanations aren't always faithful representations of the model's actual reasoning. This research ultimately supports Anthropic's safety strategy by providing a deeper understanding of how these AI systems actually work. The complete show notes for this episode can be found at https://twimlai.com/go/727.