World Model · podcast knowledge graph

Mamba, Mamba-2 and Post-Transformer Architectures for Generative AI with Albert Gu - #693

2024-07-17 · 58 min · episode 693 · 16 entities

Asserted relationships

Entities found in this episode

companys 7

  • mentioned TWIML company
    0.70
    evidence rules-v4
    Feed author/publisher: TWIML
  • mentioned Carnegie Mellon University company
    0.62
    evidence rules-v4
    professor at Carnegie Mellon University
  • mentioned Technology company
    0.50
    evidence rules-v4
    Feed category: Technology
  • works at Carnegie Mellon University company
    0.50
    evidence rules-v4
    professor at Carnegie Mellon University
  • works at Carnegie Mellon University company
    0.50
    evidence rules-v4
    professor at Carnegie Mellon University
  • discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

concepts 5

  • mentioned Science concept
    0.50
    evidence rules-v4
    Feed category: Science
  • mentioned Tech News concept
    0.50
    evidence rules-v4
    Feed category: Tech News
  • discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News

persons 3

  • mentioned Albert Gu person
    0.72
    evidence rules-v4
    Mamba, Mamba-2 and Post-Transformer Architectures for Generative AI with Albert Gu - #693
  • mentioned Sam Charrington person
    0.70
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington

podcasts 1

Episode description as stored
Today, we're joined by Albert Gu, assistant professor at Carnegie Mellon University, to discuss his research on post-transformer architectures for multi-modal foundation models, with a focus on state-space models in general and Albert’s recent Mamba and Mamba-2 papers in particular. We dig into the efficiency of the attention mechanism and its limitations in handling high-resolution perceptual modalities, and the strengths and weaknesses of transformer architectures relative to alternatives for various tasks. We dig into the role of tokenization and patching in transformer pipelines, emphasizing how abstraction and semantic relationships between tokens underpin the model's effectiveness, and explore how this relates to the debate between handcrafted pipelines versus end-to-end architectures in machine learning. Additionally, we touch on the evolving landscape of hybrid models which incorporate elements of attention and state, the significance of state update mechanisms in model adaptability and learning efficiency, and the contribution and adoption of state-space models like Mamba and Mamba-2 in academia and industry. Lastly, Albert shares his vision for advancing foundation models across diverse modalities and applications. The complete show notes for this episode can be found at https://twimlai.com/go/693.