World Model · podcast knowledge graph

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

2026-09-09 · 40 min · episode 1193 · 12 entities

Asserted relationships

  • → hosted by Andreessen Horowitz person
    0.55
    evidence rules-v4
    Feed author/publisher: Andreessen Horowitz
  • → discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • → discusses Technology concept
    0.40
    evidence rules-v4
    Feed category: Technology
  • → discusses Business concept
    0.40
    evidence rules-v4
    Feed category: Business
  • → discusses Entrepreneurship concept
    0.40
    evidence rules-v4
    Feed category: Entrepreneurship
  • → hosted by a16z company
    0.40
    evidence rules-v4
    Feed author/publisher: a16z

Entities found in this episode

concepts 9

  • mentioned Science concept
    0.50
    evidence rules-v4
    Feed category: Science
  • mentioned Technology concept
    0.50
    evidence rules-v4
    Feed category: Technology
  • mentioned Business concept
    0.50
    evidence rules-v4
    Feed category: Business
  • mentioned Entrepreneurship concept
    0.50
    evidence rules-v4
    Feed category: Entrepreneurship
  • mentioned Ben Horowitz & Rayan Krishnan concept
    0.42
    evidence rules-v4
    Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
  • discusses Science concept
    0.40
    evidence rules-v4
    Feed category: Science
  • discusses Technology concept
    0.40
    evidence rules-v4
    Feed category: Technology
  • discusses Business concept
    0.40
    evidence rules-v4
    Feed category: Business
  • discusses Entrepreneurship concept
    0.40
    evidence rules-v4
    Feed category: Entrepreneurship

persons 2

  • mentioned Andreessen Horowitz person
    0.70
    evidence rules-v4
    Feed author/publisher: Andreessen Horowitz
  • hosted by Andreessen Horowitz person
    0.55
    evidence rules-v4
    Feed author/publisher: Andreessen Horowitz

companys 1

  • hosted by a16z company
    0.40
    evidence rules-v4
    Feed author/publisher: a16z
Episode description as stored
a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination. Resources: Follow Rayan Krishnan on X: https://x.com/RayanKrishnan Follow Ben Horowitz on X: https://x.com/bhorowitz Follow Jennifer Li on X: https://x.com/JenniferHli Stay Updated: Find a16z on YouTube: YouTube Find a16z on X Find a16z on LinkedIn Listen to the a16z Show on Spotify Listen to the a16z Show on Apple Podcasts Follow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.