More The TWIML AI Podcast episodes

How to Engineer AI Inference Systems with Philip Kiely thumbnail

How to Engineer AI Inference Systems with Philip Kiely

Published 30 Apr 2026

Duration: 00:54:47

AI inference deployment is accelerating, emphasizing inference engineering's critical role in optimizing generative models with advanced hardware and complex systems, while addressing challenges like latency, scalability, and modality-specific optimizations amid evolving industry trends and fragmented yet open-source-driven markets.

Episode Description

In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore w...

Overview

The podcast delves into the rapid evolution of AI inference compared to traditional fields like medicine and physics, where model training typically takes weeks or months, while AI inference can occur within hours. It emphasizes the growing importance of inference engineering, which focuses on deploying and optimizing large generative models in real-time, particularly as models scale to billions of parameters. This shift underscores the distinction between inferencecentral to AI-native companiesand earlier ML ops trends, as the complexity of inference increases with hardware demands, distributed systems, and strict latency requirements. Key challenges include managing technical limitations like insufficient compute resources, system orchestration, and the need for interdisciplinary expertise in GPU programming, quantization, and model parallelism. The field is also marked by a fast research-to-implementation cycle, rivaling industries like high-frequency trading in speed.

The discussion highlights how inference engineering has evolved from a niche concern to a critical industry standard, driven by the practical needs of deploying AI at scale. Companies across various sectors increasingly recognize the necessity of inference strategies to balance performance, cost, and reliability, impacting user experience and competitive advantage. The podcast outlines the spectrum of inference control, from limited user customization in closed systems to full flexibility in self-hosted deployments. It also addresses the transition from pay-per-token models to GPU-based infrastructure, influenced by cost, capacity, and scalability. Additionally, the role of specialized hardware, such as Hopper GPUs, and the fragmented yet advancing open-source ecosystem for inference tools (e.g., VLLM, TensorRT) are explored, alongside trends like compute disaggregation and modality-specific optimizations for tasks like vision or text-to-speech.

Finally, the content addresses the growing demand for inference engineers, driven by the complexity of deploying and optimizing AI systems. It emphasizes the need for interdisciplinary expertise to integrate applied research with infrastructure, while also acknowledging the limits of full automation due to hardware-specific optimizations. The podcast further touches on the future of inference, including the specialization of systems for task-specific workloads, the rise of agent-based systems requiring real-time inference, and the challenges of multimodal models. As AI models become more integral to product development, the strategic role of inference engineering in enabling efficient, reliable, and scalable AI applications is underscored, with implications for businesses seeking competitive differentiation through model-level innovation.

Recent Episodes of The TWIML AI Podcast

16 Sept 2026 From Voice Agents to AI Avatars with Alexander Smola

"Voice AI advances toward human-like avatars but faces challenges like latency, emotional authenticity, and hardware limitations, with future systems balancing speed, multimodal training, and ethical considerations."

25 Aug 2026 Why the Next AI Breakthrough May Come from Physics with Max Welling

"AI accelerates scientific research in molecular dynamics and material science through equivariant neural networks, machine learning force fields, and digital twin simulations, revolutionizing fields like carbon capture, semiconductors, and energy storage while integrating physics and machine learning for broader scientific advancements."

27 Jul 2026 Why Models Are AIs Next Training Dataset with Damian Borth

"Explores weights-based learning, treating neural network weights as input data to train new models, improving efficiency, addressing data scarcity, and enabling tasks like model compression and performance prediction, with future directions in scaling, privacy, and cross-domain knowledge transfer."

8 Jul 2026 How AI Learns to Smell with Alex Wiltschko

Digitizing scent using AI involves converting molecules into digital data, creating standardized scent representations, and reproducing odors, addressing biological complexities, leveraging graph neural networks, and exploring applications in fragrance, diagnostics, and emotion while highlighting technical and ethical challenges.

More The TWIML AI Podcast episodes