More Practical AI episodes

The Future of AI Infrastructure with CoreWeave thumbnail

The Future of AI Infrastructure with CoreWeave

Published 17 Jul 2026

Duration: 00:50:04

"AI infrastructure demands specialized, application-centric systems for training and inference, addressing challenges like GPU failures and orchestration inefficiencies, while emphasizing observability, cost optimization, and the future of AI-driven workflows and democratized research."

Episode Description

As AI applications become more complex, the infrastructure powering them needs to evolve. Corey Sanders, SVP of Product at CoreWeave, joins Chris to d...

Overview

The podcast discusses the evolving landscape of AI infrastructure and its distinct demands compared to traditional cloud computing. It highlights how AI workloads - particularly training and inference - require specialized, high-performance systems with optimized networking, storage, and orchestration. CoreWeave is presented as a provider focused on delivering AI-centric infrastructure, emphasizing bare-metal Kubernetes, scalability, and efficiency for large-scale deployments. The discussion underscores the limitations of legacy cloud models and the necessity for pre-planned, customized environments to handle the interconnected nature of AI tasks and avoid costly slowdowns.

Further exploration centers on the shift from model-centric to application-centric AI development, where the focus moves beyond building models to integrating them into complex, multi-model workflows. The conversation covers tools and platforms like Weights & Biases and ARIA, which enable experiment tracking, automated analysis, and iterative improvement through an "AI loop." There is a strong emphasis on human-AI collaboration, with AI agents assisting in research and deployment decisions. The future of AI is envisioned as agent-driven, with natural, conversational interactions replacing traditional UIs. Additionally, the podcast touches on democratizing AI by making advanced workflows accessible, supporting multi-cloud and edge environments, and accelerating adoption across industries through open-source integration and talent development.

What If

  • What if you treated your AI app like a lab experiment?

    • Move: Set up automated experiment tracking using Weights & Biases (W&B) for every prompt, model, and inference call; log inputs, outputs, latency, and cost.
    • Why Now?: AI inference loops are evolving fast - manual tracking won't scale, and silent failures (e.g., GPU stragglers, prompt drift) degrade quality without warning.
    • Expected Upside: 30 - 50% faster iteration cycles, early detection of performance drops, and data-backed decisions on which prompts or models to retire, swap, or fine-tune.
  • What if you rebuilt your AI workflow around agent-led execution?

    • Move: Replace static scripts with an agent that runs, evaluates, and suggests next steps (e.g., prompt tweaks, smaller models) using ARIA-style logic - start with a rule-based version in Python.
    • Why Now?: Agent patterns are shifting from sci-fi to production tools; early adopters cut experimentation costs and accelerate learning loops before competitors catch up.
    • Expected Upside: 60% reduction in trial-and-error compute spend, continuous autonomous optimization, and freeing yourself to focus on high-value design vs. grunt debugging.
  • What if you stopped treating infrastructure as generic and customized it for AI?

    • Move: Audit your current provider's GPU utilization, storage latency, and orchestration bottlenecks - then benchmark a move to a bare-metal Kubernetes setup (e.g., via CoreWeave or Sunk on your stack).
    • Why Now?: General cloud providers still treat GPUs as add-ons; AI-native infra like Sunk + InfiniBand networking already cuts training time by 20 - 40% for early adopters.
    • Expected Upside: Faster training/inference, lower cost per run, and preemptive avoidance of GPU stragglers or cascading failures - critical when every hour of downtime burns cash.

Takeaway

  • Prioritize optimizing prompts over immediate model fine-tuning to reduce compute costs and improve AI application performance.
  • Implement experiment tracking tools like Weights & Biases to log, compare, and iterate on AI model results systematically.
  • Design AI workflows with infrastructure awareness by choosing cost-efficient models, caching results, and minimizing inference latency.
  • Adopt agent-assisted development patterns to automate repetitive analysis and accelerate the AI loop of test, evaluate, and improve.
  • Build portable, Kubernetes-based systems that support multi-cloud deployment while optimizing for performance-critical workloads on specialized AI infrastructures.

Recent Episodes of Practical AI

9 Jul 2026 Building Durable AI Agents

The evolution of AI agents from local tools to enterprise systems highlights challenges in scalability, reliability, and infrastructure, emphasizing the need for robust frameworks, open-source innovation, and observability in managing complex, distributed workflows.

2 Jul 2026 Image Generation and Visual Intelligence with Black Forest Labs

The evolution of generative AI progresses from basic outputs to cinematic-quality media via diffusion and autoregressive models, with innovations in noise-removal techniques, preference-based evaluation, multimodal integration, and efficiency-focused research for real-world applications.

25 Jun 2026 AIUC-1: Building trust in AI agents

The development of AI safety and ethics standards, employing a flywheel model of audits, certifications, and red teaming, addresses risks to vulnerable groups and enterprise adoption through frameworks like three-layer structures, probabilistic risk management, and systemic safeguards beyond technical controls.

4 Jun 2026 Breaking down the 2026 Stanford AI Index Report

Recent advancements in AI, highlighted by the Stanford AI Index Report's findings on accelerating capabilities, human-level performance in specialized tasks, impacts on education and work, challenges like flawed benchmarks and the "jagged frontier," robotics limitations, U.S.-China leadership dynamics, governance gaps, and broader implications for labor, creativity, and policy.

28 May 2026 Rebooting Enterprise AI with MCP and Kubernetes

The Multi-Cloud Protocol (MCP) bridges AI systems with enterprise infrastructure, enabling secure, scalable interactions between LLMs and traditional tools via standardized, governance-focused operational frameworks.

More Practical AI episodes