- Preface
- Chapter 1 — AI-First vs AI-Added — The First Architecture Decision
- Chapter 2 — The Maturity Model — Architecture Readiness
- Chapter 3 — The Canonical Architecture
- Chapter 4 — LLM-Native Decision & Data Architecture
- Chapter 5 — Design Patterns
- Chapter 6 — Retrieval & Knowledge
- Chapter 7 — Agentic Architectures
- Chapter 8 — Multi-Agent Systems & Orchestration
- Chapter 9 — Prompt & Context Engineering as a Discipline
- Chapter 10 — Evaluation & Quality
- Chapter 11 — The Agent Runtime
- Chapter 12 — Deployment, Scale & Resilience Architecture
- Chapter 13 — Security & Threat Modeling
- Chapter 14 — Cost & FinOps for LLM Systems
- Chapter 15 — Observability & LLMOps
- Chapter 16 — Governance & Compliance, Operationalised
- Chapter 17 — The Engagement Checklist
- Appendix A — Architecture Glossary
- Appendix B — Governance & Regulatory Architecture Mapping
- Appendix C — NovaCred Artifact & Evidence Chain
Architecting Production-Ready Gen AI and Agentic AI Systems
A Practitioner’s Guide to Architecture, Retrieval, Agents, Evaluation, Security, Governance, FinOps, and Production Operations
An AI system becomes an architecture problem when it can influence decisions, invoke tools, carry state or change the outside world. This practitioner playbook shows how to design Gen AI and Agentic AI systems that remain reliable, governable and defensible once they reach production.
Minimum price
$7.99
$9.99
You pay
Author earns
About
About the Book
Architecting Production-Ready Gen AI and Agentic AI Systems - Version 3 - August 2026
An AI system becomes an architecture problem the moment its output can change a workflow, invoke a tool, influence a consequential decision, consume governed data or continue acting after the original request has ended.
At that point, model quality is only one part of the design. The architecture must also determine what the system is allowed to know, which sources it may trust, what authority an agent inherits, where deterministic validation sits, how long-running work survives failure, what happens when a tool call partially succeeds and what evidence will allow the organisation to reconstruct the decision months later.
Architecting Production-Ready Gen AI and Agentic AI Systems is a practitioner playbook for making those decisions deliberately.
The book begins before model selection. It distinguishes AI-added, AI-first, generative, agentic and hybrid systems, then connects those choices to organisational maturity and the level of autonomy the enterprise is actually prepared to operate. From there it develops a canonical production architecture covering decision and data contracts, retrieval and knowledge systems, prompt and context engineering, evaluation, agent runtimes, multi-agent orchestration, deployment, resilience, security, FinOps, observability and operational governance.
Agentic AI receives its own architecture treatment because adding tools, memory, state, retries and delegated action changes the failure model. The book examines bounded authority, execution identity, approval gates, durable state, idempotency, tool contracts, trajectory evidence and the separation between an agent's ability to reason and its authority to execute.
The production runtime is treated as a distributed system rather than a wrapper around an LLM API. Synchronous, streaming, queued and durable execution are considered separately, along with backpressure, admission control and event-driven patterns that keep probabilistic reasoning from becoming an unnecessary dependency of deterministic transaction paths. The Propose → Validate → Commit pattern shows how AI-generated actions can remain inspectable until deterministic policy, authority and state checks permit execution.
Evaluation is treated as release engineering rather than a benchmark exercise. Retrieval quality, model behaviour, tool selection, agent trajectories, operational quality and business outcomes are evaluated independently, tied to thresholds, regression sets and production evidence.
Throughout the book, the fictional regulated enterprise NovaCred forces architecture decisions into concrete form. Retrieval contracts, decision contracts, tool manifests, evaluation gates, approval records, cost baselines, traces, incident evidence and runbooks are connected into a single architecture evidence chain rather than presented as disconnected best practices.
Version 3 is a substantial architectural expansion rather than a cosmetic update. Agentic AI is promoted to a first-class design concern, with dedicated treatment of single-agent and multi-agent architectures, delegated identity, bounded authority, durable execution, state, retries, approval boundaries and trajectory-level evidence. The production runtime is strengthened with synchronous, streaming, queued and durable execution models, backpressure and event-driven decoupling, including the Propose → Validate → Commit pattern for separating probabilistic reasoning from deterministic business execution. Evaluation, security, observability and governance are made explicitly agent-aware, while the NovaCred case, architecture decision records, retrieval and prompt evidence, tool manifests, release gates, cost baselines, runbooks and approval trails are connected into a stronger end-to-end evidence model. The result is a Version 3 that goes much further into what it takes to operate autonomous and semi-autonomous AI safely under real enterprise conditions.
The final engagement checklist turns the material into a practical review method for architects and engineering leaders who need to answer a more demanding question than whether the model works:
Can this system be operated, controlled, explained, recovered and defended when it is placed under real production pressure?
Written for solution architects, AI engineers, platform teams, engineering leaders, product owners, risk professionals and senior practitioners designing enterprise Gen AI and Agentic AI systems.
Feedback
Author
About the Author
Srinivas is a Generative AI Practitioner and Educator specializing in the architectural design and rigorous evaluation of LLM-powered applications. With deep experience in developing multi-agent frameworks and hybrid RAG architectures, he focus on bridging the gap between experimental AI and production-ready systems.
He is the creator of popular technical practice tests on Udemy, including the AWS Certified GenAI Developer - Professional series, and have developed comprehensive frameworks for AI project estimation and compliance. His work frequently involves industry-leading evaluation tools such as RAGAS, Giskard, and Guardrails.ai.
Driven by the mission to help IT professionals navigate the "mindset shift" required for the AI era, Srinivas provides systematic, data-driven methodologies for building AI that is not only innovative but reliable and compliant with emerging standards like the EU AI Act.
Contents
Table of Contents
Get the free Community Edition
You can get the free Community Edition in PDF or EPUB just by sharing your name and email address with the author.
The Leanpub 60 Day 100% Happiness Guarantee
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...
Earn $8 on a $10 Purchase, and $16 on a $20 Purchase
We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.
(Yes, some authors have already earned much more than that on Leanpub.)
In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.
Learn more about writing on Leanpub
Free Updates. DRM Free.
If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).
Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.
Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.
Learn more about Leanpub's ebook formats and where to read them
Write and Publish on Leanpub
You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!
Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.
Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.