World Model · podcast knowledge graph

Generative Benchmarking with Kelly Hong - #728

2025-04-23 · 54 min · episode 728 · 16 entities

Asserted relationships

  • evidence rules-v4
    Generative Benchmarking with Kelly Hong - #728
  • → hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • → works at Chroma company
    0.50
    evidence rules-v4
    researcher at Chroma
  • → works at Chroma company
    0.50
    evidence rules-v4
    researcher at Chroma
  • → discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • → discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • → discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • → discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News
  • → hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

Entities found in this episode

companys 7

  • mentioned TWIML company
    0.70
    evidence rules-v4
    Feed author/publisher: TWIML
  • mentioned Chroma company
    0.62
    evidence rules-v4
    researcher at Chroma
  • mentioned Technology company
    0.50
    evidence rules-v4
    Feed category: Technology
  • works at Chroma company
    0.50
    evidence rules-v4
    researcher at Chroma
  • works at Chroma company
    0.50
    evidence rules-v4
    researcher at Chroma
  • discusses Technology company
    0.40
    evidence rules-v4
    Feed category: Technology
  • hosted by TWIML company
    0.40
    evidence rules-v4
    Feed author/publisher: TWIML

concepts 5

  • mentioned Science concept
    0.50
    evidence rules-v4
    Feed category: Science
  • mentioned Tech News concept
    0.50
    evidence rules-v4
    Feed category: Tech News
  • discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • discusses News concept
    0.40
    evidence rules-v4
    Feed category: News
  • discusses Tech News concept
    0.40
    evidence rules-v4
    Feed category: Tech News

persons 3

  • mentioned Kelly Hong person
    0.72
    evidence rules-v4
    Generative Benchmarking with Kelly Hong - #728
  • mentioned Sam Charrington person
    0.70
    evidence rules-v4
    Feed author/publisher: Sam Charrington
  • hosted by Sam Charrington person
    0.55
    evidence rules-v4
    Feed author/publisher: Sam Charrington

podcasts 1

Episode description as stored
In this episode, Kelly Hong, a researcher at Chroma, joins us to discuss "Generative Benchmarking," a novel approach to evaluating retrieval systems, like RAG applications, using synthetic data. Kelly explains how traditional benchmarks like MTEB fail to represent real-world query patterns and how embedding models that perform well on public benchmarks often underperform in production. The conversation explores the two-step process of Generative Benchmarking: filtering documents to focus on relevant content and generating queries that mimic actual user behavior. Kelly shares insights from applying this approach to Weights & Biases' technical support bot, revealing how domain-specific evaluation provides more accurate assessments of embedding model performance. We also discuss the importance of aligning LLM judges with human preferences, the impact of chunking strategies on retrieval effectiveness, and how production queries differ from benchmark queries in ambiguity and style. Throughout the episode, Kelly emphasizes the need for systematic evaluation approaches that go beyond "vibe checks" to help developers build more effective RAG applications. The complete show notes for this episode can be found at https://twimlai.com/go/728.