Distilling Transformers and Diffusion Models for Robust Edge Use Cases with Fatih Porikli - #738
2025-07-09 · 60 min · episode 738 · 13 entities
Asserted relationships
-
→ appeared on The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence podcast0.68
evidence rules-v4
Distilling Transformers and Diffusion Models for Robust Edge Use Cases with Fatih Porikli - #738
-
0.55
evidence rules-v4
Feed author/publisher: Sam Charrington
-
0.40
evidence rules-v4
Feed category: Science
-
0.40
evidence rules-v4
Feed category: Technology
-
0.40
evidence rules-v4
Feed category: News
-
0.40
evidence rules-v4
Feed category: Tech News
-
0.40
evidence rules-v4
Feed author/publisher: TWIML
Entities found in this episode
concepts 5
-
0.50
evidence rules-v4
Feed category: Science
-
0.50
evidence rules-v4
Feed category: Tech News
-
0.40
evidence rules-v4
Feed category: Science
-
0.40
evidence rules-v4
Feed category: News
-
0.40
evidence rules-v4
Feed category: Tech News
companys 4
-
0.70
evidence rules-v4
Feed author/publisher: TWIML
-
0.50
evidence rules-v4
Feed category: Technology
-
0.40
evidence rules-v4
Feed category: Technology
-
0.40
evidence rules-v4
Feed author/publisher: TWIML
persons 3
-
0.72
evidence rules-v4
Distilling Transformers and Diffusion Models for Robust Edge Use Cases with Fatih Porikli - #738
-
0.70
evidence rules-v4
Feed author/publisher: Sam Charrington
-
0.55
evidence rules-v4
Feed author/publisher: Sam Charrington
podcasts 1
-
appeared on The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence podcast0.68
evidence rules-v4
Distilling Transformers and Diffusion Models for Robust Edge Use Cases with Fatih Porikli - #738
Episode description as stored
Today, we're joined by Fatih Porikli, senior director of technology at Qualcomm AI Research for an in-depth look at several of Qualcomm's accepted papers and demos featured at this year’s CVPR conference. We start with “DiMA: Distilling Multi-modal Large Language Models for Autonomous Driving,” an end-to-end autonomous driving system that incorporates distilling large language models for structured scene understanding and safe planning motion in critical "long-tail" scenarios. We explore how DiMA utilizes LLMs' world knowledge and efficient transformer-based models to significantly reduce collision rates and trajectory errors. We then discuss “SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation,” a diffusion-distilled approach that combines generative models with metric depth estimation to produce sharp, accurate monocular depth maps. Additionally, Fatih also shares a look at Qualcomm’s on-device demos, including text-to-3D mesh generation, real-time image-to-video and video-to-video generation, and a multi-modal visual question-answering assistant.
The complete show notes for this episode can be found at https://twimlai.com/go/738.