Leanpub Header

Skip to main content

Self-Healing Infrastructure: Building Autonomous Cloud Systems with AI

Self-Healing Infrastructure: Building Autonomous Cloud Systems with AI

Most writing about AI and infrastructure stops at the demo. This book starts on the day the model is confidently wrong at 3 a.m. and an auditor asks who authorised the action it took.

Seven working labs — MCP servers with real identity and audit, closed-loop remediation behind a reversibility gate, and autonomy that can be revoked — for platform engineers in safety-critical and regulated industries.

Free With Membership

With Membership

Free!

$1.99

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
About

About

About the Book

There is no shortage of writing about applying AI to infrastructure. Most of it stops at the demo. You will find a great deal about wiring a language model to an alert manager, and very little about what happens on the day the model is confidently wrong at 3 a.m., while an auditor asks who authorized the action it took.

This book is about the gap between those two things. It is written for platforms operating in safety-critical and regulated industries — aviation, healthcare, finance, public sector — where self-healing is a design discipline rather than a marketing phrase, and where "the model
suggested it" is not an acceptable answer to a regulator.

What you will build. Seven working labs, not toy examples: a production MCP server with OAuth 2.1, per-tool RBAC and an immutable audit trail; alert enrichment that keeps paging when the model provider is down; an incident scribe that drafts a blameless postmortem from a Slack thread; a three-agent investigation system with a human review gate; a closed-loop remediator that acts behind a reversibility gate and rolls back when a probe says the world got worse; and a policy contract with per-action budgets, an approval queue and a kill switch. Every lab runs on a laptop — no corporate cloud account, no paid identity-provider tenant.

What makes it different. Every pattern is anchored to a named AWS or Azure Well-Architected reliability principle, then mapped against the regulation that constrains it: the FAA AI Safety Assurance Roadmap, the NIST AI Risk Management Framework, Singapore's 2026 agentic-AI framework, the EU AI Act and EASA guidance. Appendix A carries the full pattern-by-regulation matrix with a source link for every claim.

The book is also honest about its own limits. Chapter 5 argues against autonomous remediation on the production-alert surface, in a book whose title contains the word autonomous. Chapter 3 concedes that a fifty-line heuristic at one percent of the cost is a perfectly good baseline, and that AI is allowed to lose to a regular expression on the merits.

Who it is for. Platform engineers and cloud architects designing the surfaces AI tooling will run on; senior SREs handed AI tooling and asked to make it reliable; AI/ML infrastructure leads who have the model working and now need it to survive an audit. It assumes you have run something in production and been paged for it. It does not assume you know how a transformer works, and never requires it.

Author

About the Author

Ajin Baby

Ajin Baby is an AI Platform & Cloud Infrastructure Architect. He has spent fifteen years building and operating production infrastructure, the last eight of them in commercial aviation — about as unforgiving an environment as software gets. He works on MCP servers, agentic infrastructure, and the identity, policy, and audit layers that make them safe to run in a regulated industry, currently as a Staff Software Engineer at Jeppesen ForeFlight, Inc. and previously as a Computing Architect at Boeing. Before the architect years he was a two-time founder. He writes at cloudandsre.com.

Contents

Table of Contents

Acknowledgements

About This Book

  1. Who This Book Is For
  2. What You Will Build
  3. How This Book Is Organized
  4. Prerequisites
  5. Conventions Used in This Book
  6. Using the Code Examples
  7. The Companion Blog

About the Author

Foreword

How to Read This Book

  1. 1. The Recoverability Triangle
  2. 2. The Reversibility Gate
  3. 3. Per-Call Economics, Not Per-Demo Economics
  4. 4. The Autonomy Ladder Is a Risk Function, Not a Roadmap
  5. 5. Tools, Not Chat
  6. 6. Context Is the Substrate, Not the Prompt
  7. 7. Identity Travels With the Action
  8. Putting it together
  9. What’s next
  10. Part I. Foundations

Chapter 1. The AI-Native SRE

  1. What this chapter covers
  2. The 3 a.m. problem, in 2026
  3. What actually changed
  4. The AI-native SRE, defined
  5. What this book will and will not do
  6. What you will build
  7. Reliability patterns anchor
  8. Summary
  9. What’s next

Chapter 2. The AI Stack and Its Failure Modes

  1. What this chapter covers
  2. The stack, layer by layer
  3. Hosted vs. self-hosted — the honest tradeoff
  4. Failure modes (the part nobody writes about)
  5. Guardrails and graceful degradation
  6. Regulatory baseline (preview)
  7. Reliability patterns anchor
  8. Summary
  9. What’s next

Chapter 3. Prompt Engineering for Infrastructure

  1. What this chapter covers
  2. Prompts as interface contracts
  3. Structured output is non-negotiable
  4. Testing a non-deterministic interface
  5. Patterns worth their weight
  6. Resilience at the API boundary
  7. Reliability patterns anchor
  8. Summary
  9. What’s next
  10. Part II. The Agentic Platform

Chapter 4. MCP in Production

  1. What this chapter covers
  2. Why MCP deserves its own chapter
  3. MCP, precisely
  4. The 2026 enterprise-readiness shift
  5. Production design patterns
  6. Threat model (the short list)
  7. Lab: Production MCP Server
  8. Regulatory fit
  9. Reliability patterns anchor
  10. Summary
  11. What’s next

Chapter 5. Intelligent On-Call Automation

  1. What this chapter covers
  2. The on-call reality in 2026
  3. The enrichment pipeline
  4. What “good” looks like
  5. Lab: alert-explainer
  6. Reliability patterns anchor
  7. Summary
  8. What’s next

Chapter 6. Incident Response with AI

  1. What this chapter covers
  2. The real incident workflow
  3. The scribe pattern
  4. Blameless postmortems at scale
  5. Lab: incident-scribe
  6. Reliability patterns anchor
  7. Summary
  8. What’s next

Chapter 7. Multi-Agent SRE Systems with MCP

  1. What this chapter covers
  2. Why most multi-agent demos fail in production
  3. The Scheduler-Agent-Supervisor pattern
  4. AgentOps — observability for the AI itself
  5. Failure modes and containment
  6. Lab: sre-agent
  7. Reliability patterns anchor
  8. Summary
  9. What’s next
  10. Part III. Autonomy in Production

Chapter 8. Closed-Loop Remediation

  1. What this chapter covers
  2. The gap Chapter 7 left open
  3. The reversibility-gate, in code
  4. Action taxonomy
  5. Lab: the closed-loop remediator
  6. Failure modes you will create
  7. Reliability patterns anchor
  8. Summary
  9. What’s next

Chapter 9. Bounded Autonomy in Practice

  1. What this chapter covers
  2. The autonomy-ladder, climbed
  3. Policy as the autonomy contract
  4. Blast-radius math
  5. Lab: policy-gated remediation
  6. Reliability patterns anchor
  7. Summary
  8. What’s next

Chapter 10. Building Your AI SRE Platform

  1. What this chapter covers
  2. The reference architecture
  3. Regulatory framing — US first, EU as the strictest case
  4. The GenAIOps operational model
  5. Making the case to leadership
  6. What’s next (and what’s in v2.0)
  7. Reliability patterns anchor
  8. Closing

Appendix A. Reliability Pattern × Regulation Mapping

  1. Table A-1. Patterns mapped to cloud architecture guidance
  2. Table A-2. Patterns mapped to regulatory obligations
  3. Cross-cutting controls (apply to most patterns above)
  4. How to use this appendix
  5. A note on regulatory currency
  6. Beyond the tables: NIST AI RMF and the Singapore agentic framework
  7. Source list

Glossary

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub