World Model · podcast knowledge graph

Why Vision Language Models Ignore What They See with Munawar Hayat - #758

2025-12-09 · 58 min · episode 758 · 13 entities

Asserted relationships

  • evidence rules-v4
    Why Vision Language Models Ignore What They See with Munawar Hayat - #758
  • → hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • → discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • → discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • → discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • → discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News
  • → hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

Entities found in this episode

concepts 5

  • mentioned Science concept
    0.50
    evidence rules-v4
    Feed category: Science
  • mentioned Tech News concept
    0.50
    evidence rules-v4
    Feed category: Tech News
  • discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News

companys 4

  • mentioned TWIML company
    0.70
    evidence rules-v4
    Feed author/publisher: TWIML
  • mentioned Technology company
    0.50
    evidence rules-v4
    Feed category: Technology
  • discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

persons 3

  • mentioned Munawar Hayat person
    0.72
    evidence rules-v4
    Why Vision Language Models Ignore What They See with Munawar Hayat - #758
  • mentioned Sam Charrington person
    0.70
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington

podcasts 1

Episode description as stored
In this episode, we’re joined by Munawar Hayat, researcher at Qualcomm AI Research, to discuss a series of papers presented at NeurIPS 2025 focusing on multimodal and generative AI. We dive into the persistent challenge of object hallucination in Vision-Language Models (VLMs), why models often discard visual information in favor of pre-trained language priors, and how his team used attention-guided alignment to enforce better visual grounding. We also explore a novel approach to generalized contrastive learning designed to solve complex, composed retrieval tasks—such as searching via combined text and image queries—without increasing inference costs. Finally, we cover the difficulties generative models face when rendering multiple human subjects, and the new "MultiHuman Testbench" his team created to measure and mitigate issues like identity leakage and attribute blending. Throughout the discussion, we examine how these innovations align with the need for efficient, on-device AI deployment. The complete show notes for this episode can be found at https://twimlai.com/go/758.