The discussion centers on the adoption and optimization of AI tools within engineering organizations, with a focus on improving developer workflows through structured evaluation and context management. Key challenges include determining the relevance of contextual information fed into AI models, as outdated or excessive context - referred to as "context rot" - can degrade performance. Experiments showed that removing irrelevant context improved evaluation outcomes, highlighting the need for continuous assessment and pruning of inputs. Organizations like Datadog have established dedicated AI teams to manage tooling, measure impact through metrics such as DORA, and build custom evaluation systems to guide decisions on model and tool adoption.
A major theme is the evolution of evaluation (evals) practices to ensure AI systems are effective and efficient. The team developed agent-based evals to test AI performance on real-world scenarios, using both isolated skill assessments and broader project-level evaluations. Automation is prioritized to avoid human bottlenecks, with nightly eval runs and ad-hoc testing for major changes. Cost efficiency is critical, influencing decisions about evaluation frequency and model selection. Additionally, the conversation explores organizational scaling challenges, such as managing plugin marketplaces, ensuring proper ownership of AI steering documents, and adapting team structures - splitting responsibilities between "signals" (governance, metrics) and "flow" (developer experience) - to support thousands of engineers effectively.