Leanpub Header

Skip to main content

Orchestrating AI Agents

Coordinating Claude Code, Codex, Local Models, and MCP with a Persistent Control Plane

Orchestrating AI Agents

A practical guide to operating a fleet of AI coding agents through routing, memory, skills, MCP, guardrails, and a persistent control plane (286 manuscript pages).

Minimum price

$19.99

$29.99

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
About

About

About the Book

Orchestrating AI Agents is a practical, architecture-first guide to operating several AI coding agents as one coordinated system. It has sixteen chapters in four parts, is vendor-neutral by design, and uses a single worked control plane, Hermes, to show how the pieces fit while labeling what is specific to it. It is written for engineers, platform builders, and technical founders who already run more than one agent, and it is not a tutorial for any single tool.

The book starts from a problem and a thesis. Claude Code, Codex, local models, and MCP servers each do something well, but nothing owns persistence, scheduling, cross-project memory, routing, or a capability registry. Because an agent's commands run in a shell and most serious agents offer a headless mode, a persistent control plane can drive any command-line agent as a tool. The book's rule is to orchestrate rather than replace, so the control plane stays a modest coder and becomes an excellent operator.

The middle of the book builds the mechanics. A dated survey reads Claude Code, Codex, Antigravity, Aider, OpenCode, and local models on the axes that matter for routing. The delegation primitive covers briefs, adapters, validated output, write-back, and auth isolation. A routing engine decides by stakes, context size, privacy class, and autonomy among acting directly, delegating, and escalating, under one absolute rule that high stakes reaches a human. Layered memory is shared across harnesses, skills form a capability registry, MCP is the shared tool layer with reads freely granted and writes gated through one seam, and local-first routing is enforced by a model gateway.

The last part is about operating the result. It covers defense-in-depth sandboxing, approval modes and blocklists, scheduling with a lease-based multi-agent work queue, a worked orchestration with a software-release factory and a customer-support fleet as further cases, budgets and spend governance, and a catalogue of failure modes with a discipline of escalation and a clear account of what never to automate. It closes with a maturity arc and an honest ceiling: production, money, and security stay human-reviewed.

Readers come away able to route work by explicit criteria, brief and validate delegated runs, design shared memory that compounds, contain unattended agents, govern their spend, and recover when a standing fleet fails. Tool-specific claims are dated and treated as examples to confirm against current documentation, while the architecture is written to outlast them.

Author

About the Author

Yohan Rodriguez

Yohan is a Senior Full-Stack Software Engineer with extensive experience delivering scalable, end-to-end software solutions across web, enterprise, and cloud-based environments. He specializes in architecting robust platforms, modernizing legacy systems, driving cloud transformation efforts, and building integration-heavy applications that support critical business workflows. He is recognized for translating complex requirements into reliable, maintainable, and high-value solutions across industries such as insurance, cybersecurity, and professional services.

Known for combining strong technical execution with a practical business mindset, he has contributed to projects from concept and design through production delivery and long-term support. His experience includes collaborating with cross-functional teams, improving development workflows, solving complex technical challenges, and helping organizations deliver dependable software products that adapt to changing business needs. He brings a balanced approach to engineering that values quality, efficiency, and continuous improvement.

Contents

Table of Contents

  • Preface i
  • 1 The Multi-Tool Problem 2
    • 1.1 The Year Everyone Got a Fleet 2
    • 1.2 A Tuesday With Four Agents 3
    • 1.3 Five Things Nothing Owns 4
    • 1.4 Why ``Which Agent Is Best?'' Is the Wrong Question 6
    • 1.5 Why the Problem Stayed Invisible 7
    • 1.6 What an Operating Layer Would Do 7
    • 1.7 Hands-On: Audit Your Own Agent Stack 8
    • 1.8 The Operator's Mindset 8
    • 1.9 What This Book Builds 10
  • 2 Orchestrate, Don't Replace 11
    • 2.1 How the Command-Line Agent Got Its Shape 11
    • 2.2 The Limitation That Is the Opening 12
    • 2.3 Agents Have a Back Door: Headless Mode 13
    • 2.4 Two Planes: Control and Execution 15
    • 2.5 A Task, Traced Through Both Planes 16
    • 2.6 Operate, Don't Replace 18
    • 2.7 What the Control Plane Actually Is 18
    • 2.8 Hands-On: The Seed of Delegation 19
    • 2.9 Why This Scales, and Where We Go Next 22
    • 2.10 Implementation Guidance: One Adapter Per Agent 22
    • 2.11 Where the Control Plane Runs 24
    • 2.12 Beyond One Coordinator: Governance and Scale 25
    • 2.13 Mistakes to Avoid 25
  • 3 The Agent Landscape, 2026 29
    • 3.1 How to Read an Agent 29
    • 3.2 Claude Code 31
    • 3.3 Codex CLI 32
    • 3.4 Antigravity CLI 32
    • 3.5 Aider and OpenCode: The Local-First Pair 33
    • 3.6 Local Models as an Execution Tool 33
    • 3.7 The Capability Matrix 34
    • 3.8 Hands-On: The Same Task Through Two Harnesses 35
    • 3.9 A Routing Scenario: One Day, Five Tasks 37
    • 3.10 Onboarding a New Agent 38
    • 3.11 Harness Migration and Version Management 38
    • 3.12 Provider Independence and Exit Strategy 39
    • 3.13 Trade-offs and Routing Implications 41
    • 3.14 Mistakes to Avoid 42
    • 3.15 Where the Landscape Is Going, and Where We Go Next 43
  • 4 Hermes as the Control Plane 45
    • 4.1 Why a Daemon, Not a Command 45
    • 4.2 The Anatomy of a Control Plane 46
    • 4.3 The Agent Loop 47
    • 4.4 Memory as Layers 48
    • 4.5 The Terminal Backend Is a Config Knob 48
    • 4.6 Approvals and the Hardline Blocklist 49
    • 4.7 Gateways and the Scheduler 50
    • 4.8 Configuration: The Knobs That Matter 50
    • 4.9 Hands-On: Reading a Control Plane's State 51
    • 4.10 A Request's Journey Through the Control Plane 52
    • 4.11 Implementation Guidance: Adopt or Build 54
    • 4.12 Availability and Multi-Environment Operation 55
    • 4.13 Change Management and Disaster Recovery 56
    • 4.14 Operational Considerations 57
    • 4.15 Trade-offs 58
    • 4.16 Mistakes to Avoid 58
    • 4.17 Lessons Learned, and Where We Go Next 59
  • 5 Delegation via Headless CLIs 61
    • 5.1 What a Delegation Skill Is 61
    • 5.2 The Brief: Context Assembly 62
    • 5.3 The Delegation Sequence 63
    • 5.4 Capturing and Validating Structured Output 64
    • 5.5 Writing the Outcome Back 65
    • 5.6 Auth Isolation and the Exfiltration Channel 66
    • 5.7 A Delegation, Traced --- and One That Fails 67
    • 5.8 Implementation Guidance and Operational Considerations 69
    • 5.9 Patterns of Delegation 69
    • 5.10 Long-Running and Asynchronous Delegation 70
    • 5.11 Advanced Delegation Patterns 71
    • 5.12 The Limits of the Shell-Out Primitive 72
    • 5.13 Trade-offs 73
    • 5.14 Mistakes to Avoid 74
    • 5.15 Lessons Learned, and Where We Go Next 74
  • 6 The Routing Engine 77
    • 6.1 The Four Axes 77
    • 6.2 The Three Outcomes 78
    • 6.3 The Decision and the Hard Rule 79
    • 6.4 The Router in the Loop 81
    • 6.5 Classifying the Task, Choosing the Agent 81
    • 6.6 Routing Under Uncertainty 83
    • 6.7 Worked Routing: Nine Real Workloads 84
    • 6.8 Routing on Evidence 86
    • 6.9 A Week of Routing, Reviewed 87
    • 6.10 Implementation Guidance and Operational Considerations 87
    • 6.11 Routing at Scale: Tables as Code, and Learned Routing 88
    • 6.12 Trade-offs 89
    • 6.13 Mistakes to Avoid 89
    • 6.14 Lessons Learned, and Where We Go Next 90
  • 7 Memory Across Harnesses 92
    • 7.1 Why Memory Is the Moat 92
    • 7.2 The Layered Model, Revisited in Depth 93
    • 7.3 The Bounded Core: What Earns a Permanent Seat 94
    • 7.4 What a Memory Record Contains 95
    • 7.5 Write Policy: Archive, Propose, Curate 95
    • 7.6 Cross-Harness Write-Back 97
    • 7.7 Relevance Retrieval: The Read Path 98
    • 7.8 Measuring Retrieval Quality 99
    • 7.9 Hands-On: Designing a Memory Schema 100
    • 7.10 Portability Across Harnesses 101
    • 7.11 A Bad Memory, Traced and Repaired 102
    • 7.12 Memory Hygiene: Drift, Poisoning, and Pruning 103
    • 7.13 Memory at Scale: Sharding, Federation, and Conflict 104
    • 7.14 Implementation Guidance and Trade-offs 105
    • 7.15 Mistakes to Avoid 106
    • 7.16 Lessons Learned, and Where We Go Next 106
  • 8 Skills as a Capability Registry 108
    • 8.1 A Skill Is a Procedure on Disk 108
    • 8.2 The Shape of a Complete Skill 109
    • 8.3 The Three Loading Levels 110
    • 8.4 Matching and Debugging the Registry 112
    • 8.5 The Three Classes of Skill 112
    • 8.6 Hands-On: Classify Your Procedures 113
    • 8.7 From Repeated Procedure to Trusted Skill 113
    • 8.8 Testing, Versioning, and Deprecation 115
    • 8.9 Permissions and Environment 115
    • 8.10 Agent-Authored Skills, and the Self-Improvement Claim 116
    • 8.11 Community Skills Are Code Paths 117
    • 8.12 Implementation Guidance and Operational Considerations 117
    • 8.13 Trade-offs and Mistakes to Avoid 118
    • 8.14 Lessons Learned, and Where We Go Next 119
  • 9 MCP: The Shared Tool Layer 121
    • 9.1 What MCP Is 121
    • 9.2 MCP as the Device-Driver Layer 122
    • 9.3 The Wiring 122
    • 9.4 Registration: Making a Server Real 123
    • 9.5 The Protocol Shape: Discover, Invoke, Validate 124
    • 9.6 Designing a Tool Surface: Read Versus Write 125
    • 9.7 Hands-On: Map a Server You Would Build 126
    • 9.8 Authentication and Token Scope 126
    • 9.9 Output Validation and Injection Handling 127
    • 9.10 Routing Writes Through One Seam 128
    • 9.11 Coordinator Access Versus Delegated Access 129
    • 9.12 The Server Landscape: Existing and Proposed 129
    • 9.13 A Memory Server, Made Concrete 130
    • 9.14 A Real Workflow: Build, Verify, Publish 130
    • 9.15 The Security Surface 131
    • 9.16 When the Shared Layer Fails: MCP Outages and Degradation 132
    • 9.17 Protocol Versioning, Transport, and Mixed-Version Fleets 133
    • 9.18 Implementation Guidance, Trade-offs, and Mistakes 135
    • 9.19 Lessons Learned, and Where We Go Next 136
  • 10 Local vs Cloud Routing 138
    • 10.1 Why Local-First Is the Default 138
    • 10.2 The Privacy Classes, Made Operational 139
    • 10.3 The Secrets Boundary 140
    • 10.4 The Three Variables: Cost, Latency, Quality 141
    • 10.5 The Budget and Audit Chokepoint 143
    • 10.6 The Local-vs-Cloud Decision Flow 147
    • 10.7 Hands-On: Classify Your Task Mix 148
    • 10.8 A Week of Local-vs-Cloud Routing, Reviewed 149
    • 10.9 Advanced Routing: Cascades, Fallbacks, and Speculation 149
    • 10.10 Capacity Planning for Local Inference 150
    • 10.11 Deployment Scenarios: Where the Boundary Falls 151
    • 10.12 A Routing Failure, Traced 152
    • 10.13 Operating It: Monitoring, Debugging, Scaling 152
    • 10.14 Implementation Guidance, Trade-offs, and Mistakes 153
    • 10.15 Lessons Learned, and Where We Go Next 154
  • 11 Sandboxing and Guardrails 157
    • 11.1 The Threat Model 157
    • 11.2 The Layered Model, Made Precise 159
    • 11.3 The Isolation Spectrum: Terminal Backends 160
    • 11.4 Approval Modes: Manual, Smart, Off 163
    • 11.5 The Hardline Blocklist 163
    • 11.6 The Pre-Exec Scanner 164
    • 11.7 The Cross-Cutting Controls: Kill-Switch, Caps, Audit, Secrets 166
    • 11.8 The Pre-Flight Checklist for Autonomous Operation 167
    • 11.9 Multi-Agent Considerations 168
    • 11.10 A Contained Failure, Traced 169
    • 11.11 Advanced Isolation Architectures 169
    • 11.12 Guardrails at Fleet Scale: Policy as Code 170
    • 11.13 Identity, Secrets, and Credential Lifecycle 171
    • 11.14 Audit Trails and Compliance for Regulated Operation 173
    • 11.15 A Guardrail Failure, Traced 174
    • 11.16 Operating It: Monitoring, Debugging, Maintenance 174
    • 11.17 Trade-offs and Mistakes 175
    • 11.18 Lessons Learned, and Where We Go Next 176
  • 12 Scheduling and the Work Queue 178
    • 12.1 From One-Shot to Standing Operation 178
    • 12.2 The Work Queue as the Coordination Primitive 179
    • 12.3 Claim and Lock: The Heart of Multi-Agent Safety 180
    • 12.4 Status Reporting and Observability 184
    • 12.5 Observability: Tracing a Request Across the Fleet 184
    • 12.6 Human-Escalation Gates 185
    • 12.7 Deduplication and Idempotency 186
    • 12.8 Wiring the Budget Contract 187
    • 12.9 A Multi-Agent Scheduled Workflow, Traced 187
    • 12.10 Backfill, Catch-Up, and the Cold Start 188
    • 12.11 Advanced Queue Architectures 189
    • 12.12 Hierarchical and Supervisor Orchestration 190
    • 12.13 A Scheduling Failure, Traced 191
    • 12.14 Scaling the Fleet: 10 to 1,000 Agents 192
    • 12.15 Capacity Math: Utilization, Queue Depth, and Little's Law 193
    • 12.16 Operating It: Monitoring, Debugging, Scaling 195
    • 12.17 Trade-offs and Mistakes 195
    • 12.18 Lessons Learned, and Where We Go Next 196
  • 13 A Worked Orchestration 198
    • 13.1 The Factory as an Orchestration 198
    • 13.2 Routing a Feature Through the Pipeline 199
    • 13.3 Where Every Mechanism Appears 201
    • 13.4 The Compile-Verify Loop in Depth 202
    • 13.5 Testing the Orchestration: A Test Pyramid for Agent Systems 204
    • 13.6 The Human Gates 205
    • 13.7 Hands-On: Trace Your Own Pipeline 206
    • 13.8 A Run That Went Wrong, and Recovered 206
    • 13.9 The Same Shape, Other Factories 207
    • 13.10 A Second Worked Case: A Software-Release Factory 208
    • 13.11 A Third Worked Case: A Customer-Support Fleet 209
    • 13.12 Cross-Team and Multi-Repo Orchestration 211
    • 13.13 When Orchestration Is the Wrong Choice 212
    • 13.14 Operating the Factory: Monitoring, Debugging, Scaling 213
    • 13.15 Trade-offs, Mistakes, and Lessons 214
    • 13.16 Lessons Learned, and Where We Go Next 214
  • 14 Budgets and Spend Governance 216
    • 14.1 The Accounting Seam 216
    • 14.2 Per-Task-Class Cost Envelopes 217
    • 14.3 The Cloud-versus-Local Break-Even 219
    • 14.4 Fleet-Wide Caps and the Runaway 221
    • 14.5 Cost Curves and the Shape of Fleet Spend 222
    • 14.6 The Cost of Human Review 223
    • 14.7 A Cost Explosion, Traced 223
    • 14.8 FinOps for Agent Fleets: Chargeback and Incentives at Scale 225
    • 14.9 What the Expensive Model Should Cost 225
    • 14.10 Hands-On: Build a Cost Model for Your Task Mix 226
    • 14.11 Showback: Attributing Cost to Projects and Owners 226
    • 14.12 A Month of Spend, Reviewed 227
    • 14.13 Caching, Token Efficiency, and Performance 228
    • 14.14 Reasoning About the Numbers: Cost and Latency Models 229
    • 14.15 Operating It: Monitoring, Forecasting, Maintenance 231
    • 14.16 Trade-offs and Mistakes 231
    • 14.17 Lessons Learned, and Where We Go Next 232
  • 15 Failure and Escalation 234
    • 15.1 Why Standing Fleets Fail Differently 234
    • 15.2 Failure Mode: Collisions 235
    • 15.3 Failure Mode: Compounding Error 236
    • 15.4 Failure Mode: Poisoned Memory 236
    • 15.5 Failure Mode: Prompt Injection 237
    • 15.6 Failure Mode: Runaway 237
    • 15.7 Advanced Failure Modes: Cascades, Correlation, and Drift 238
    • 15.8 A Cascading Incident, Fully Traced 239
    • 15.9 Evaluation: Measuring Agent Quality Over Time 240
    • 15.10 Chaos Testing for Agent Fleets 242
    • 15.11 The Escalation Discipline 242
    • 15.12 Debugging an Agent in the Wild 244
    • 15.13 Reproducibility and Replay 247
    • 15.14 What to Never Automate 248
    • 15.15 Building a Recovery Posture 248
    • 15.16 Reliability Engineering: SLOs, Error Budgets, and On-Call 249
    • 15.17 An Incident, Traced End to End 250
    • 15.18 A Worked System: A Multi-Agent Operations Center 251
    • 15.19 Trade-offs and Mistakes 253
    • 15.20 Lessons Learned, and Where We Go Next 253
  • 16 The Three-Year Horizon 255
    • 16.1 The Maturity Arc 255
    • 16.2 What Each Phase Unlocks and Gates On 257
    • 16.3 Migration Paths Between Phases 259
    • 16.4 The Honest Ceiling 260
    • 16.5 Hands-On: Locate Yourself and Pick the Next Gate 261
    • 16.6 Operational Lessons for the Long Haul 263
    • 16.7 Deployment Archetypes: The Arc in Four Contexts 263
    • 16.8 A Worked System: An Enterprise Research Organization 264
    • 16.9 How the Team Evolves with the Fleet 266
    • 16.10 How Maturity Regresses: Migration Failure Modes 267
    • 16.11 A Three-Year Trajectory, Traced 267
    • 16.12 Component Lifecycle: Versioning, Compatibility, and Retirement 268
    • 16.13 A Dated Snapshot, So You Can Measure the Drift 269
    • 16.14 Standing Review Triggers 270
    • 16.15 Closing: Orchestrate, Don't Replace 271
  • Conclusion 273

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub