How to Find the Agent Failures Your Evals Miss with Scott Clark - #767
2026-05-07 · 53 min · episode 767 · 16 entities
Asserted relationships
-
→ appeared on The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence podcast0.68
evidence rules-v4
How to Find the Agent Failures Your Evals Miss with Scott Clark - #767
-
0.55
evidence rules-v4
Feed author/publisher: Sam Charrington
-
0.50
evidence rules-v4
CEO of Distributional
-
0.50
evidence rules-v4
CEO of Distributional
-
0.40
evidence rules-v4
Feed category: Science
-
0.40
evidence rules-v4
Feed category: Technology
-
0.40
evidence rules-v4
Feed category: News
-
0.40
evidence rules-v4
Feed category: Tech News
-
0.40
evidence rules-v4
Feed author/publisher: TWIML
Entities found in this episode
companys 7
-
0.70
evidence rules-v4
Feed author/publisher: TWIML
-
0.62
evidence rules-v4
CEO of Distributional
-
0.50
evidence rules-v4
Feed category: Technology
-
0.50
evidence rules-v4
CEO of Distributional
-
0.50
evidence rules-v4
CEO of Distributional
-
0.40
evidence rules-v4
Feed category: Technology
-
0.40
evidence rules-v4
Feed author/publisher: TWIML
concepts 5
-
0.50
evidence rules-v4
Feed category: Science
-
0.50
evidence rules-v4
Feed category: Tech News
-
0.40
evidence rules-v4
Feed category: Science
-
0.40
evidence rules-v4
Feed category: News
-
0.40
evidence rules-v4
Feed category: Tech News
persons 3
-
0.72
evidence rules-v4
How to Find the Agent Failures Your Evals Miss with Scott Clark - #767
-
0.70
evidence rules-v4
Feed author/publisher: Sam Charrington
-
0.55
evidence rules-v4
Feed author/publisher: Sam Charrington
podcasts 1
-
appeared on The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence podcast0.68
evidence rules-v4
How to Find the Agent Failures Your Evals Miss with Scott Clark - #767
Episode description as stored
In this episode, Scott Clark, co-founder and CEO of Distributional, joins us to explore how teams can reliably operate and improve complex LLM systems and agents in production. Scott introduces a Maslow’s hierarchy of observability: telemetry for logging, monitoring for known signals, and post-production or online analytics to surface unknown unknowns. We dig into examples of real-world failures Scott’s team has seen in production systems, such as “lazy” tool-use hallucinations that standard evals miss, and how mapping traces into vector fingerprints enables clustering and topic discovery to uncover emergent behaviors. Scott explains how analytics can feed the data flywheel by generating evals, guardrails, and training data, and why online, adaptive approaches are essential for non-stationary models. We also touch on practical how-to’s such as instrumentation with OpenTelemetry, the GenAI semantic conventions, and the role of dedicated analytics tools.
The complete show notes for this episode can be found at https://twimlai.com/go/767.