More The TWIML AI Podcast episodes

How to Find the Agent Failures Your Evals Miss with Scott Clark thumbnail

How to Find the Agent Failures Your Evals Miss with Scott Clark

Published 7 May 2026

Duration: 00:54:02

Distributional employs post-production analytics, unsupervised learning, and LLMs to analyze agent traces, detect patterns and anti-patterns like hallucinations, address distributional shifts, and generate actionable insights for AI system refinement in security and enterprise settings, emphasizing adaptive analytics and domain expertise.

Episode Description

In this episode, Scott Clark, co-founder and CEO of Distributional, joins us to explore how teams can reliably operate and improve complex LLM systems...

Overview

The text outlines Distributional, an AI analytics platform focused on improving agent quality through analysis of production data. It emphasizes a hierarchical approach to observability, starting with foundational telemetry (logging system behavior), progressing to monitoring (real-time tracking of predefined metrics), and culminating in analytics (uncovering hidden patterns in production data to refine agents via unsupervised learning and feedback loops). The platform leverages Bayesian statistics and insights from Scott Clarks work on optimization (e.g., Bayesian methods at Yelp and SIGOpt) to address challenges like overfitting in black-box optimizers and the need for meaningful, domain-specific goals in AI system design. It shifts from pre-production testing to post-production analytics to better align with real-world dynamics, such as detecting agent "hallucinations" (e.g., false tool calls in financial agents) and identifying "unknown unknowns" through statistical anomalies in behavioral patterns.

The text also highlights analytics role in continuous adaptation, using techniques like vector mapping, clustering, and large language model (LLM)-driven analysis to detect subtle deviations from expected behavior (e.g., unusual tool call distributions or emergent risks in non-stationary environments). It contrasts monitoring (ensuring system health) with analytics (identifying optimization opportunities), both being critical for iterative system improvement. Key challenges include parsing unstructured data (e.g., logs), aligning evaluation metrics with business needs, and managing complexity in agentic systems. Security is emphasized as a critical application area, using analytics to uncover hidden signals or anomalies in agent behavior. The platform is framed as a post-production tool, designed for enterprises, with open-source deployment options and a focus on bridging gaps between model performance and real-world reliability. It also addresses the need for new benchmarks to evaluate analytics tools in domains like cybersecurity, where detecting subtle risks is crucial.

Recent Episodes of The TWIML AI Podcast

27 Jul 2026 Why Models Are AIs Next Training Dataset with Damian Borth

"Explores weights-based learning, treating neural network weights as input data to train new models, improving efficiency, addressing data scarcity, and enabling tasks like model compression and performance prediction, with future directions in scaling, privacy, and cross-domain knowledge transfer."

8 Jul 2026 How AI Learns to Smell with Alex Wiltschko

Digitizing scent using AI involves converting molecules into digital data, creating standardized scent representations, and reproducing odors, addressing biological complexities, leveraging graph neural networks, and exploring applications in fragrance, diagnostics, and emotion while highlighting technical and ethical challenges.

9 Jun 2026 Is RAG Dead? Lessons from Building AI for Tax Law with Alex Bowcut

The podcast examines Retrieval-Augmented Generation's evolving role in AI-driven tax compliance, focusing on Spheres AI's TRAM model, challenges in processing fragmented legal data, and the need for accurate citations, taxonomy integration, and real-time compliance automation via a global tax legislation index.

21 May 2026 Relational Foundation Models for Enterprise Data with Jure Leskovec

Relational foundation models and graph-based machine learning, like GNNs, enable accurate predictions on structured data across biomedical research and industries by capturing complex relationships, integrating multi-scale data, and overcoming traditional limitations through automated feature extraction and hybrid modeling.

30 Apr 2026 How to Engineer AI Inference Systems with Philip Kiely

AI inference deployment is accelerating, emphasizing inference engineering's critical role in optimizing generative models with advanced hardware and complex systems, while addressing challenges like latency, scalability, and modality-specific optimizations amid evolving industry trends and fragmented yet open-source-driven markets.

More The TWIML AI Podcast episodes