World Model · podcast knowledge graph

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought with Chengzu Li - #722

2025-03-10 · 42 min · episode 722 · 13 entities

Asserted relationships

  • evidence rules-v4
    Imagine while Reasoning in Space: Multimodal Visualization-of-Thought with Chengzu Li - #722
  • → hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • → discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • → discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • → discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • → discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News
  • → hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

Entities found in this episode

concepts 5

  • mentioned Science concept
    0.50
    evidence rules-v4
    Feed category: Science
  • mentioned Tech News concept
    0.50
    evidence rules-v4
    Feed category: Tech News
  • discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News

companys 4

  • mentioned TWIML company
    0.70
    evidence rules-v4
    Feed author/publisher: TWIML
  • mentioned Technology company
    0.50
    evidence rules-v4
    Feed category: Technology
  • discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

persons 3

  • mentioned Chengzu Li person
    0.72
    evidence rules-v4
    Imagine while Reasoning in Space: Multimodal Visualization-of-Thought with Chengzu Li - #722
  • mentioned Sam Charrington person
    0.70
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington

podcasts 1

Episode description as stored
Today, we're joined by Chengzu Li, PhD student at the University of Cambridge to discuss his recent paper, “Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.” We explore the motivations behind MVoT, its connection to prior work like TopViewRS, and its relation to cognitive science principles such as dual coding theory. We dig into the MVoT framework along with its various task environments—maze, mini-behavior, and frozen lake. We explore token discrepancy loss, a technique designed to align language and visual embeddings, ensuring accurate and meaningful visual representations. Additionally, we cover the data collection and training process, reasoning over relative spatial relations between different entities, and dynamic spatial reasoning. Lastly, Chengzu shares insights from experiments with MVoT, focusing on the lessons learned and the potential for applying these models in real-world scenarios like robotics and architectural design. The complete show notes for this episode can be found at https://twimlai.com/go/722.