World Model · podcast knowledge graph

Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?

2026-09-22 · 39 min · 18 entities

Asserted relationships

  • → hosted by Claire Vo person
    0.55
    evidence rules-v5
    Feed author/publisher: Claire Vo
  • → discusses Technology company
    0.40
    evidence rules-v5
    Feed category: Technology
  • → references penname.co website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://penname.co/
  • → references Codex website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://openai.com/codex
  • → references chatprd.ai website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://chatprd.ai/
  • → references clairevo.com website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://clairevo.com/
  • → references Claude Opus 5 5 website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://anthropic.com/claude-opus-5-5
  • → references Introducing Gpt 6 Sol And Luna website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://openai.com/index/introducing-gpt-6-sol-and-luna

Entities found in this episode

websites 12

  • mentioned penname.co website
    0.45
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://penname.co/
  • mentioned Codex website
    0.45
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://openai.com/codex
  • mentioned chatprd.ai website
    0.45
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://chatprd.ai/
  • mentioned clairevo.com website
    0.45
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://clairevo.com/
  • mentioned Claude Opus 5 5 website
    0.45
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://anthropic.com/claude-opus-5-5
  • mentioned Introducing Gpt 6 Sol And Luna website
    0.45
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://openai.com/index/introducing-gpt-6-sol-and-luna
  • references penname.co website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://penname.co/
  • references Codex website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://openai.com/codex
  • references chatprd.ai website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://chatprd.ai/
  • references clairevo.com website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://clairevo.com/
  • references Claude Opus 5 5 website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://anthropic.com/claude-opus-5-5
  • references Introducing Gpt 6 Sol And Luna website
    0.38
    evidence rules-v5
    Link in episode "Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?": https://openai.com/index/introducing-gpt-6-sol-and-luna

persons 2

  • mentioned Claire Vo person
    0.70
    evidence rules-v5
    Feed author/publisher: Claire Vo
  • hosted by Claire Vo person
    0.55
    evidence rules-v5
    Feed author/publisher: Claire Vo

companys 2

  • mentioned Technology company
    0.50
    evidence rules-v5
    Feed category: Technology
  • discusses Technology company
    0.40
    evidence rules-v5
    Feed category: Technology

concepts 2

Episode description as stored
I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived. What you’ll learn: How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend Where Sol still wins me over on clear writing, readable PRDs, and price The character SVG results that completely overturned my prediction about Anthropic What happened when I asked these models to edit video, and why I think skills explain part of the disappointment Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t — In this episode, we cover: (00:00) LIVE setup and new model launches (01:30) What’s new in Opus 5.5, Sol, and Luna (04:11) Guardrails, personality, and speed (09:00) The How I AI bench and blind evaluation process (11:31) Email and personal-productivity results (13:50) Frontend prototype vibe checks (24:10) Backend, agent personality, and long-running tasks (28:25) SVG illustration test (29:48) AI video-editing results (30:43) Predictions before the reveal (31:20) Barbie Bench: the 3D fashion-game test (34:17) Results: Astra, Sol, and Opus 5.5 (35:04) Writing clarity and creative surprises (36:51) Why the LLM judge disagreed with me (37:24) What each model is actually best for — Tools referenced: • Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5 • GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/ • Codex (OpenAI): https://openai.com/codex — Where to find Claire Vo: ChatPRD: https://www.chatprd.ai/ Website: https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ X: https://x.com/clairevo — Production and marketing by https://penname.co/ . For inquiries about sponsoring the podcast, email jordan@penname.co.