More MLOps.community episodes

It's 2026, and We're Still Talking Evals thumbnail

It's 2026, and We're Still Talking Evals

Published 21 Apr 2026

Duration: 00:40:56

Evaluations in AI product development must be integrated early, address real-world complexities, use nuanced metrics beyond accuracy, employ user-centric and iterative testing, leverage post-deployment data, and adapt tailored strategies to balance quality, domain-specific metrics, and system reliability.

Episode Description

Maggie Konstanty is an AI Product Manager at Prosus, one of the world's largest consumer internet companies, where she builds and evaluates AI agents...

Overview

The discussion emphasizes the critical role of evaluations (evals) in AI product development, advocating for their integration from the early stages of ideation to ensure quality before deployment. Teams often delay setting up evals, leading to inefficiencies and confusion. Pre-production evaluations typically rely on simulated tasks, but real-world scenariossuch as unexpected user queriesrequire adaptive methods. Post-deployment evaluations must shift from static tests to dynamic, user-generated data and metrics. Challenges include non-deterministic failure modes in large language models (LLMs), where systems may perform well repeatedly before failing unexpectedly, necessitating nuanced strategies. User-centric approaches involve simulating diverse personas to model real-world interactions, while traditional accuracy metrics are criticized for oversimplification. Metrics like TNR and TPR are preferred for comparing AI responses to human-labeled outcomes. Current practices face limitations, such as LLMs losing coherence or task alignment over time, and the risks of evaluating LLMs with other LLMs. Balancing feature prioritization with iterative testing and failure mode analysis is highlighted as crucial for refining AI agents post-deployment.

Key challenges include the ambiguity of accuracy metrics in real-world applications and the underutilization of error analysis for edge cases. Scenario-based testing, involving predefined interactions and personas, is stressed as a method to measure consistency and identify performance gaps. Pre-release evaluations focus on internal simulations, while post-release analyses rely on user data to uncover hidden failures. Iterative testingsuch as AB testingallows for rapid adjustments based on feedback. Unconventional strategies, like stress-testing with extreme scenarios, help reveal hidden flaws. Continuous monitoring is essential, as real-world unpredictability demands ongoing adaptation. Evaluations are also framed as iterative processes, not one-time tasks, requiring consistent refinement. The discussion underscores the need for tailored evaluation frameworks that align with specific use cases, user needs, and domain-specific metrics (e.g., conversion rates for food ordering vs. satisfaction metrics in automotive services).

The conversation also highlights the cultural and practical challenges of implementing evaluative rigor, such as perceiving error analysis as tedious or resource-intensive, and the tendency to skip it in favor of more immediately gratifying development tasks. Evaluators must be closely aligned with business goals to avoid redundancy, and existing tools often lack support for multi-turn conversations, large datasets, or efficient data export. Custom tools are emphasized for creating pipelines and evaluators that address specific failure modes and regressions. Internal development of evaluators is preferred to maintain control over data privacy and alignment with team-specific needs. Ultimately, the focus remains on defining clear success metrics, understanding user intent ambiguity, and fostering team alignment on evaluation priorities as the foundation for reliable AI systems.

Recent Episodes of MLOps.community

20 Jul 2026 The Creator of FastMCP Explains the Future of MCP

"Fast MCP streamlined the Multi-Chat Protocol, dominating the market with simplicity and efficiency, while evolving to support interactive UI apps, Python-based token-efficient interfaces, and addressing security and scalability challenges, with AI tools enhancing personal and professional workflows."

13 Jul 2026 What Happens When Every Developer Has 20 AI Agents?

"Modern software development faces bottlenecks from limited human resources and AI-driven shifts, transforming productivity, SaaS models, and workflows while straining infrastructure and open-source ecosystems."

6 Jul 2026 AI Agents Should Be Treated Like Hackers

Integrating AI agents with enterprise systems via APIs presents security risks from untrusted access, requiring solutions like the Multi-Cloud Protocol, zero-trust models, and GraphQL to balance innovation with safeguards against data exposure and autonomous decision risks.

6 Jul 2026 Developers May Stop Depending on Libraries

Recommended: There is more than one way to build with AI

Advancements in AI tools like Hugging Face MCP and Fast Agent simplify LLM integration for innovative workflows, emphasizing idea-driven development, Rust's performance, open-source models (e.g., Gemma 4, Quen), and accessible tools for non-experts, while balancing efficiency, transparency challenges, and evolving SDKs.

6 Jul 2026 10 Cities. 4 Countries. One Unexpected MCP Lesson.

The Model Communication Protocol (MCP) enables secure AI-to-tool integration via APIs, with DeepL promoting it through global workshops, hackathons, and practical examples like a Python server, emphasizing security, implementation challenges, and hands-on learning to bridge technical gaps and enhance AI workflows.

More MLOps.community episodes