World Model · podcast knowledge graph

The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking

2026-05-01 · 107 min · 21 entities

Asserted relationships

  • → hosted by Erik Torenberg, Nathan Labenz person
    0.55
    evidence rules-v4
    Feed author/publisher: Erik Torenberg, Nathan Labenz
  • → founded OpenPipe company
    0.54
    evidence rules-v4
    founder of OpenPipe
  • → works at OpenPipe company
    0.50
    evidence rules-v4
    founder of OpenPipe
  • → discusses Society & Culture concept
    0.40
    evidence rules-v4
    Feed category: Society & Culture
  • → discusses Technology concept
    0.40
    evidence rules-v4
    Feed category: Technology
  • → discusses Business concept
    0.40
    evidence rules-v4
    Feed category: Business
  • → discusses Entrepreneurship concept
    0.40
    evidence rules-v4
    Feed category: Entrepreneurship
  • → hosted by Turpentine company
    0.40
    evidence rules-v4
    Feed author/publisher: Turpentine
  • → references getvcx.com website
    0.38
    evidence rules-v4
    Link in episode "The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking": https://getvcx.com/
  • → references sequencehq.com website
    0.38
    evidence rules-v4
    Link in episode "The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking": https://sequencehq.com/

Entities found in this episode

concepts 10

  • mentioned Society & Culture concept
    0.50
    evidence rules-v4
    Feed category: Society & Culture
  • mentioned Technology concept
    0.50
    evidence rules-v4
    Feed category: Technology
  • mentioned Business concept
    0.50
    evidence rules-v4
    Feed category: Business
  • mentioned Entrepreneurship concept
    0.50
    evidence rules-v4
    Feed category: Entrepreneurship
  • evidence rules-v4
    The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
  • discusses Society & Culture concept
    0.40
    evidence rules-v4
    Feed category: Society & Culture
  • discusses Technology concept
    0.40
    evidence rules-v4
    Feed category: Technology
  • discusses Business concept
    0.40
    evidence rules-v4
    Feed category: Business
  • discusses Entrepreneurship concept
    0.40
    evidence rules-v4
    Feed category: Entrepreneurship
  • mentioned GRPO concept
    0.35
    evidence rules-v4
    GRPO

companys 5

  • mentioned Turpentine company
    0.70
    evidence rules-v4
    Feed author/publisher: Turpentine
  • mentioned OpenPipe company
    0.68
    evidence rules-v4
    founder of OpenPipe
  • founded OpenPipe company
    0.54
    evidence rules-v4
    founder of OpenPipe
  • works at OpenPipe company
    0.50
    evidence rules-v4
    founder of OpenPipe
  • hosted by Turpentine company
    0.40
    evidence rules-v4
    Feed author/publisher: Turpentine

websites 4

  • mentioned getvcx.com website
    0.45
    evidence rules-v4
    Link in episode "The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking": https://getvcx.com/
  • mentioned sequencehq.com website
    0.45
    evidence rules-v4
    Link in episode "The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking": https://sequencehq.com/
  • references getvcx.com website
    0.38
    evidence rules-v4
    Link in episode "The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking": https://getvcx.com/
  • references sequencehq.com website
    0.38
    evidence rules-v4
    Link in episode "The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking": https://sequencehq.com/

persons 2

Episode description as stored
Kyle Corbitt, founder of OpenPipe, breaks down reinforcement learning and custom fine-tuning for modern AI models. He explains how RL differs from supervised fine-tuning, why GRPO and LLM-as-judge post-training matter, and how these techniques can improve performance, latency, and cost on open source models. The conversation also covers reward hacking, evaluation design, LoRA adapters, and how Chinese labs are using distillation to fast-follow frontier models. Sponsors: Sequence: Sequence handles the full revenue workflow for complex pricing, from quoting and metering to invoicing, revenue recognition, and collections. Book a public demo at https://sequencehq.com and use code Cognizant in the source field to save 20% off year one AvePoint: AvePoint is building the control layer for AI agents so you can securely govern, audit, and recover every action at scale. Design trusted agentic outcomes from day one at https://avpt.co/tcr VCX: VCX, by Fundrise, is the public ticker for private tech, giving everyday investors access to high-growth private companies in AI, space, defense tech, and more. Learn how to invest at https://getvcx.com Claude: Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr