More Scaling DevTools episodes

Lawrence Jones from Incident.io @ AIE Europe: building an AI SRE thumbnail

Lawrence Jones from Incident.io @ AIE Europe: building an AI SRE

Published 14 Apr 2026

Duration: 00:09:26

Advancements in AI-driven SRE include automated root cause analysis tools for managing production system complexity, challenges in scaling AI due to logging and data variability, the need for human expertise in contextual reasoning, handling vast log telemetry, and examples of AI uncovering undocumented system behaviors while emphasizing centralized incident management, ambient analysis tools, and the upcoming AI Incident System (AIS).

Episode Description

Recorded at AI Engineers Europe, Lawrence Jones is an AI engineer at Incident.io and he shares his experiences building an AI SRE.Links:Incident.io ht...

Overview

The podcast discusses the application of AI in Software Reliability Engineering (SRE), focusing on developing tools to manage complexity in production systems through automated root cause analysis and incident investigation. Key challenges include scaling investigations across large numbers of customer accounts and ensuring AI can monitor performance metrics like accuracy and degradation. AIs effectiveness is highlighted with a reported 85-90% accuracy in root cause analysis for well-configured systems, though non-deterministic issues such as logging failures, misconfigurations, and external data inconsistencies pose variability in results. The discussion emphasizes the need for human expertise to complement AI, particularly through runbooks and analysis playbooks that provide contextual reasoning and structure data for meaningful insights, addressing AIs limitations in mathematical analysis and context-based interpretation.

The role of organizational context is critical, with AI tools requiring internal infrastructure and historical data to avoid spurious conclusions, especially compared to generic tools lacking domain-specific knowledge. Log telemetry challenges are also addressed, emphasizing the need for structured summarization to prevent insights from being lost in overwhelming volumes of data. Product integration focuses on enabling direct access to AI-driven investigations via desktop apps linked to tools like Court or Codex, with an emphasis on confidence scoring and alignment with incident management workflows. Examples include AI detecting undocumented system behaviors, such as a 750ms timeout in a telecom providers documentation, and resolving issues faster than manual teams. The potential of ambient analysis tools to identify unpredictable patterns and anomalies is highlighted, alongside plans for the upcoming AI Incident System (AIS), which aims to broadly detect previously undetected issues and streamline collaboration through centralized data sharing during incidents.

Recent Episodes of Scaling DevTools

24 Jun 2026 Robby Russell on Oh My Zsh, Developer Experience, and Open Source

Oh My Zshell, initiated in 2009 by Robbie Russell, simplifies Zsh configurations and Git workflows through modularity and customization, driven by community contributions, educational adoption, and trends like CLI preference over GUI tools, AI's impact on open-source practices, and challenges in sustaining open-source projects amid evolving tech landscapes.

28 May 2026 Joel Griffith from browserless: from GitHub issue to bootstrapped business

Founder of bootstrapped startup Browserless shares their journey from jazz-inspired creativity in Portland to building a profitable tech company, emphasizing structured experimentation, bootstrapping growth, AI-driven innovation, and ethical considerations in democratizing technology through personal relationships and community engagement.

More Scaling DevTools episodes