More Practical AI episodes

AI incidents, audits, and the limits of benchmarks thumbnail

AI incidents, audits, and the limits of benchmarks

Published 13 Feb 2026

Duration: 2572

Experts highlight the need for robust AI safety measures, including developing methods to catalog and prevent AI incidents, and using data and third-party audits to identify and address flaws.

Episode Description

AI is moving fast from research to real-world deployment, and when things go wrong, the consequences are no longer hypothetical. In this episode, Sean...

Overview

The podcast emphasizes the critical need for AI safety, focusing on the difficulties in defining and recording AI-related incidents. It highlights the AI Incident Database, which compiles over 5,000 annotated reports of AI failures to prevent the recurrence of similar issues, inspired by safety practices in other industries. The discussion addresses the shortcomings of current benchmarking methods, the benefits of third-party audits, and risks that arise from improper AI system configurations.

The content also underscores the importance of distinguishing between intentional and unintentional failures in AI systems, and the role of statistical validation in detecting broader systemic weaknesses. It calls for the development of standardized reporting tools and procedures to enhance AI safety. Additionally, insights from the Generative Red Team Challenge at DEF CON are mentioned, where structured testing by hackers exposed significant security flaws in model design and integration processes.

Recent Episodes of Practical AI

17 Sept 2026 How to get discovered in AI search

"Explores AI-driven search's impact on digital marketing, shifting from SEO to strategies like AEO and GEO, and challenges in adapting to AI's unique retrieval methods, brand visibility, and evolving content needs."

10 Sept 2026 Computer-Use Agents and the Future of the Agentic Internet

"AI's real-world impact is explored, covering its role in daily life, work, and creativity, with a focus on evolving AI agents, trust challenges, B2C vs. B2B applications, future agentic interfaces, and ethical concerns like centralized control and rapid advancements."

28 Aug 2026 Building the Foundation for the Agentic AI Era

"AI leader Angie Jones detailed her work at IBM, Twitter, and Block - including training 12,000 employees on AI agents like Goose - while advocating for deeper AI integration, open standards (e.g., MCP), and ethical, human-augmenting AI development."

25 Aug 2026 AI Proficiency: From Users to Builders

"AI's real-world impact in business and daily life, focusing on strategic adoption, workforce transformation, and enhancing human roles through adaptability and measurable value."

More Practical AI episodes