World Model · podcast knowledge graph

Is It Time to Rethink LLM Pre-Training? with Aditi Raghunathan - #747

2025-09-16 · 58 min · episode 747 · 17 entities

Asserted relationships

Entities found in this episode

companys 7

  • mentioned TWIML company
    0.70
    evidence rules-v4
    Feed author/publisher: TWIML
  • mentioned Carnegie Mellon University company
    0.62
    evidence rules-v4
    professor at Carnegie Mellon University
  • mentioned Technology company
    0.50
    evidence rules-v4
    Feed category: Technology
  • works at Carnegie Mellon University company
    0.50
    evidence rules-v4
    professor at Carnegie Mellon University
  • works at Carnegie Mellon University company
    0.50
    evidence rules-v4
    professor at Carnegie Mellon University
  • discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

concepts 6

  • mentioned Science concept
    0.50
    evidence rules-v4
    Feed category: Science
  • mentioned Tech News concept
    0.50
    evidence rules-v4
    Feed category: Tech News
  • discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News
  • mentioned LLM concept
    0.35
    evidence rules-v4
    LLM

persons 3

  • mentioned Aditi Raghunathan person
    0.72
    evidence rules-v4
    Is It Time to Rethink LLM Pre-Training? with Aditi Raghunathan - #747
  • mentioned Sam Charrington person
    0.70
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington

podcasts 1

Episode description as stored
Today, we're joined by Aditi Raghunathan, assistant professor at Carnegie Mellon University, to discuss the limitations of LLMs and how we can build more adaptable and creative models. We dig into her ICML 2025 Outstanding Paper Award winner, “Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction,” which examines why LLMs struggle with generating truly novel ideas. We dig into the "Roll the dice" approach, which encourages structured exploration by injecting randomness at the start of generation, and the "Look before you leap" concept, which trains models to take "leaps of thought" using alternative objectives to create more diverse and structured outputs. We also discuss Aditi’s papers exploring the counterintuitive phenomenon of "catastrophic overtraining," where training models on more data improves benchmark performance but degrades their ability to be fine-tuned for new tasks, and dig into her lab's work on creating more controllable and reliable models, including the concept of "memorization sinks," an architectural approach to isolate and enable the targeted unlearning of specific information. The complete show notes for this episode can be found at https://twimlai.com/go/747.