Rethinking Performance Engineering for Agentic AI
Description
Welcome to the Leanpub Launch video for Rethinking Performance Engineering for Agentic AI: A Practitioner's Guide to Performance, Latency Budgets, and Production-Grade Observability for Scalable Enterprise Agentic AI https://leanpub.com/agentic-ai-performance by Kandasamy Selvaraj! 0:00 Introduction: Kandasamy Selvaraj's background in performance engineering and observability 0:50 Unlearning old playbooks and rethinking performance engineering for agentic AI 1:37 Goal of the book: making agentic AI systems work at enterprise scale with millions of interactions 2:23 The core problem illustrated: one agent endpoint doing 1 unit of work vs. 15 depending on the request 3:10 Deterministic vs. probabilistic systems: why old pass/fail metrics break for agentic AI 3:57 Highway analogy: how one misbehaving agent can backlog shared infrastructure for all agents 4:46 Enterprise standards for controlling blast radius and containing failures at the agent level 5:35 Book overview: practical, readable in two sittings, grounded in real enterprise-scale examples 5:35 Target audience: performance engineers onboarding to agentic AI and architects making agent fleets production-grade 5:35 Closing: if the book saves you from one weekend incident, it has done its job About the Book This book is free. Your agent passed every load test. It's timing out in production anyway. The performance playbook that served us for twenty years assumed one thing: the system does roughly the same work for every request. Agentic AI breaks that assumption. An agent decides at runtime how many model calls, tool calls, and reasoning loops each request needs, and every dashboard you own was built for a world where that number never changed. This is a practitioner's book about what replaces the old playbook, written by a performance architect with two decades of production experience: Latency budgets with designed degradation paths, replacing single SLO targets that describe nothing Bounded autonomy: the four config values that turn an unbounded worst case into a fifteen-second one A token SLO per trace, with a real CI gate from the author's own pipeline that blocked a deployment every traditional metric approved Production-grade observability: the LLM metrics that matter (tokens, cache hit rate, context bloat, loop counts), instrumented once via OpenTelemetry and monitored through Langfuse, Splunk, and Arize Evals as the quality corner's monitoring system, AI gateway and orchestrator standards, and a 90-day path from one team's practice to an enterprise standard A first-week plan: five days, one agent, an artifact you keep every day No theory dressed as practice. Every chapter ends with something you can apply this week, and the whole book reads in two sittings. For performance engineers onboarding to agentic AI, and architects making agent fleets production-grade. About the Author Kandasamy Selvaraj is a Principal Architect and performance engineering leader with over two decades of experience making high-volume distributed systems fast, reliable, and observable. For more than a decade of that journey, he has led performance engineering and observability initiatives for large-scale enterprise platforms and systems protecting workloads that process millions of transactions, and today applies that leadership to GenAI and agentic AI systems at enterprise-grade scale. This book is his view and his experience, independently researched, and not the views of any employer. Thank you for watching, please like and leave a comment, we'd love to hear from you! Please Subscribe and Follow! YouTube: https://www.youtube.com/leanpub X: https://x.com/leanpub Instagram: https://www.instagram.com/leanpub Facebook: https://www.facebook.com/leanpub Create Your Own Leanpub Book! You can create your own book anytime here: https://leanpub.com/create/book Here's the tutorial showing how to write and publish a Leanpub book in your browser (it's free!): http://help.leanpub.com/en/articles/2932527-getting-started-writing-a-book-in-the-web-browser-writing-mode If you're a Leanpub author and you'd like to submit your own Launch video for us to publish, or if you'd like to record a Launch video with Len, please go here: https://leanpub.com/launch. #books #leanpublishing #selfpublishing #leanpub #writing #agenticengineering #performancetesting #largelanguagemodels #AgenticAI #PerformanceEngineering #LLMObservability #LatencyBudgets #OpenTelemetry
