More MLOps.community episodes

How We Cut LLM Latency 70% With TensorRT in Production thumbnail

How We Cut LLM Latency 70% With TensorRT in Production

Published 10 Apr 2026

Duration: 01:05:20

Optimizing AI systems via TensorRT LLM, efficient GPU use, cold start management with AWS FSX, and model quantization, while addressing challenges in in-house development, scaling strategies, hidden scaling complexities ("AI iceberg"), and balancing technical efficiency with organizational alignment through frameworks like Flywheel and responsible AI practices.

Episode Description

Maher Hanafi is an engineering leader who went from zero AI experience to self-hosting LLMs at enterprise scale managing GPU costs, optimizing inferen...

Overview

The text explores strategies for optimizing AI systems, emphasizing efficiency, cost management, and scalability. Techniques such as TensorRT LLM reduced latency by up to 70% through hardware-specific optimizations, while model quantization and GPU packing maximized throughput and minimized resource usage. Cold start time was addressed via preloaded container images and faster storage solutions like AWS FSX, alongside managing GPU initialization delays. Challenges in in-house AI development included balancing performance, latency, accuracy, and cost, with a focus on GPU selection, model fine-tuning, and architecture design. Scaling strategies like scheduled, dynamic, and proactive GPU allocation were tailored to traffic patterns, particularly in low-usage domains like HR tech. The "AI iceberg" concept highlighted invisible complexities such as cost, latency, and response quality, requiring tailored trade-offs for specific use cases.

Iterative optimization and collaboration across teams were critical, with an emphasis on learning alongside engineers and aligning AI initiatives with business goals. The "flywheel framework" guided planning, building, and refining AI projects to ensure high impact with manageable effort. Cost savings were prioritized through strategic GPU upgrades and dynamic scaling, while tools like an LLM proxy enabled efficient load balancing based on prefilling/decoding needs. Challenges included multilingual support, model hallucination, and ensuring transparency and compliance in AI outputs. The text also underscored the need for responsible AI practices, human oversight, and iterative testing to refine systems and align with user expectations, balancing technical innovation with practical deployment constraints.

Recent Episodes of MLOps.community

27 Jul 2026 What an Anthropic Engineer Thinks About MCP

"SDKs now see hundreds of millions of downloads annually, with a focus on minimal, extensible designs and a major MCP update shifting to stateless protocols for scalability, balancing simplicity with complexity while prioritizing stability and future-proofing."

20 Jul 2026 The Creator of FastMCP Explains the Future of MCP

"Fast MCP streamlined the Multi-Chat Protocol, dominating the market with simplicity and efficiency, while evolving to support interactive UI apps, Python-based token-efficient interfaces, and addressing security and scalability challenges, with AI tools enhancing personal and professional workflows."

13 Jul 2026 What Happens When Every Developer Has 20 AI Agents?

"Modern software development faces bottlenecks from limited human resources and AI-driven shifts, transforming productivity, SaaS models, and workflows while straining infrastructure and open-source ecosystems."

More MLOps.community episodes