More MLOps.community episodes

Serving LLMs in Production: Performance, Cost & Scale // CAST AI Roundtable thumbnail

Serving LLMs in Production: Performance, Cost & Scale // CAST AI Roundtable

Published 19 Feb 2026

Duration: 01:05:55

AI model deployment requires careful planning of infrastructure and scalability to ensure smooth transition from experimental to production stages, considering factors like cost, performance, and control.

Episode Description

Roundtable CAST AI episode: Serving LLMs in Production: Performance, Cost & Scale.Join the Community:https://go.mlops.community/YTJoinInGet the newsle...

Overview

The conversation focuses on the difficulties of moving AI and machine learning models from experimental stages into production, emphasizing the importance of infrastructure planning and scalability. Teams often prioritize solving specific problems or proving concepts without considering the complexities of long-term deployment. As AI adoption expands, there's a growing need to shift from experimentation to scaling, which requires robust MLOps practices. The discussion examines different deployment models, such as APIs, managed GPU services, and self-hosting, each with varying trade-offs in cost, performance, and control. Self-hosting provides the most control and flexibility but demands extensive infrastructure setup, including Kubernetes, GPU orchestration, and auto-scaling, presenting significant complexity.

The choice of infrastructure is influenced by the type of workload, like generative, summarization, or chat-like tasks, which have distinct performance and cost requirements. The conversation highlights key performance metricssuch as time to first token, inter-token latency, and goodputas critical for optimizing model serving. Techniques like model quantization, kernel optimizations, and separating pre-fill and decode phases are discussed as ways to improve efficiency. Overall, the discussion stresses the need to align deployment strategies with specific use cases and user expectations to achieve effective and efficient AI model serving.

Recent Episodes of MLOps.community

14 Sept 2026 Why Cost Per Million Tokens Is A Useless KPI?

"AI's exponential growth in cloud services demands new FinOps strategies to manage unbounded costs, unpredictable usage, and real-time tracking, requiring adapted SRE/DevOps principles and use-case-specific financial modeling."

27 Jul 2026 What an Anthropic Engineer Thinks About MCP

"SDKs now see hundreds of millions of downloads annually, with a focus on minimal, extensible designs and a major MCP update shifting to stateless protocols for scalability, balancing simplicity with complexity while prioritizing stability and future-proofing."

More MLOps.community episodes