Leanpub Header

Skip to main content

Rethinking Performance Engineering for Agentic AI

A Practitioner's Guide to Performance, Latency Budgets, and Production-Grade Observability for Scalable Enterprise Agentic AI

This book is 100% completeLast updated on 2026-07-18

Your agent passed every load test and timed out in production anyway. This free practitioner's guide shows why the traditional performance playbook breaks for agentic AI, and what replaces it: latency budgets, bounded autonomy, token SLOs, and production-grade observability, from a performance architect with two decades in production.

Minimum price

Free!

$1.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
About

About

About the Book

This book is free.

Your agent passed every load test. It's timing out in production anyway.

The performance playbook that served us for twenty years assumed one thing: the system does roughly the same work for every request. Agentic AI breaks that assumption. An agent decides at runtime how many model calls, tool calls, and reasoning loops each request needs, and every dashboard you own was built for a world where that number never changed.

This is a practitioner's book about what replaces the old playbook, written by a performance architect with two decades of production experience:

  • Latency budgets with designed degradation paths, replacing single SLO targets that describe nothing
  • Bounded autonomy: the four config values that turn an unbounded worst case into a fifteen-second one
  • A token SLO per trace, with a real CI gate from the author's own pipeline that blocked a deployment every traditional metric approved
  • Production-grade observability: the LLM metrics that matter (tokens, cache hit rate, context bloat, loop counts), instrumented once via OpenTelemetry and monitored through Langfuse, Splunk, and Arize
  • Evals as the quality corner's monitoring system, AI gateway and orchestrator standards, and a 90-day path from one team's practice to an enterprise standard
  • A first-week plan: five days, one agent, an artifact you keep every day

No theory dressed as practice. Every chapter ends with something you can apply this week, and the whole book reads in two sittings.

For performance engineers onboarding to agentic AI, and architects making agent fleets production-grade.

Author

About the Author

Kandasamy Selvaraj

Kandasamy Selvaraj is a Principal Architect and performance engineering leader with over two decades of experience making high-volume distributed systems fast, reliable, and observable. For more than a decade of that journey, he has led performance engineering and observability initiatives for large-scale enterprise platforms and systems protecting workloads that process millions of transactions, and today applies that leadership to GenAI and agentic AI systems at enterprise-grade scale. This book is his view and his experience, independently researched, and not the views of any employer.

Launch

Launch Video

Subscribe on YouTube

Clips

Clips

Contents

Table of Contents

  • Part I: Performance Engineering for Agentic AI
    • Chapter 1: Why the Old Playbook Breaks Down
    • Chapter 2: The New Anatomy of Latency
    • Chapter 3: From SLOs to Latency Budgets
    • Chapter 4: The Speed-Cost-Quality Triangle
    • Chapter 5: Bounding Autonomy as a Performance Lever
    • Chapter 6: Case Study: Twelve Seconds to Four
    • Chapter 7: Performance Standards at the Edges: Gateways and Orchestrators
  • Part II: Production-Grade Observability for Agentic AI
    • Chapter 8: The Observability Foundation
    • Chapter 9: LLM Metrics That Matter
    • Chapter 10: Evals: Monitoring the Quality Corner
    • Chapter 11: Instrumentation Frameworks: Getting the Data Out
    • Chapter 12: Tool Integrations: Langfuse, Splunk, and Arize
    • Chapter 13: Dashboards and Alerting for Agentic Systems
  • Part III: Applying It All
    • Chapter 14: From Team Practice to Enterprise Standard
    • Chapter 15: Getting Started: Your First Week
    • Chapter 16: Quick Reference
  • Further Reading
  • About the Author

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub