How to Engineer AI Inference Systems with Philip Kiely - #766
2026-04-30 · 55 min · episode 766 · 16 entities
Asserted relationships
-
→ appeared on The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence podcast0.68
evidence rules-v4
How to Engineer AI Inference Systems with Philip Kiely - #766
-
0.55
evidence rules-v4
Feed author/publisher: Sam Charrington
-
0.50
evidence rules-v4
head of AI
-
0.50
evidence rules-v4
head of AI
-
0.40
evidence rules-v4
Feed category: Science
-
0.40
evidence rules-v4
Feed category: Technology
-
0.40
evidence rules-v4
Feed category: News
-
0.40
evidence rules-v4
Feed category: Tech News
-
0.40
evidence rules-v4
Feed author/publisher: TWIML
Entities found in this episode
companys 7
-
0.70
evidence rules-v4
Feed author/publisher: TWIML
-
0.62
evidence rules-v4
head of AI
-
0.50
evidence rules-v4
Feed category: Technology
-
0.50
evidence rules-v4
head of AI
-
0.50
evidence rules-v4
head of AI
-
0.40
evidence rules-v4
Feed category: Technology
-
0.40
evidence rules-v4
Feed author/publisher: TWIML
concepts 5
-
0.50
evidence rules-v4
Feed category: Science
-
0.50
evidence rules-v4
Feed category: Tech News
-
0.40
evidence rules-v4
Feed category: Science
-
0.40
evidence rules-v4
Feed category: News
-
0.40
evidence rules-v4
Feed category: Tech News
persons 3
-
0.72
evidence rules-v4
How to Engineer AI Inference Systems with Philip Kiely - #766
-
0.70
evidence rules-v4
Feed author/publisher: Sam Charrington
-
0.55
evidence rules-v4
Feed author/publisher: Sam Charrington
podcasts 1
-
appeared on The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence podcast0.68
evidence rules-v4
How to Engineer AI Inference Systems with Philip Kiely - #766
Episode description as stored
In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore why inference has become the stickiest and most critical workload in AI, how it blends GPU programming, applied research, and large-scale distributed systems, and where the line sits between inference and model serving. Philip shares how research-to-production can move in hours, not months, and why understanding “the knobs” of inference—batching, quantization, speculation, and KV cache reuse—lets teams design better products and SLAs. We trace the inference maturity journey from closed APIs to dedicated deployments and in-house platforms, discuss GPU lifecycles, and survey today’s runtime landscape, including vLLM, SGLang, and TensorRT LLM. Finally, we look ahead to agents and multimodality, making the case for specialized, workload-specific runtimes when performance and efficiency matter most.
The complete show notes for this episode can be found at https://twimlai.com/go/766.