- Preface i
-
1 What an Agentic Coding Harness Is 2
- 1.1 The Moment Coding Tools Stopped Being Just Assistants 2
- 1.2 Why ``AI Coding Tool'' Is Too Vague 3
- 1.3 The Developer Tooling Spectrum 4
- 1.4 Static Autocomplete: Prediction Without Agency 5
- 1.5 Chat Assistants: Reasoning Without Direct Execution 5
- 1.6 Inline Chat and IDE Assistants: Context-Aware Help 7
- 1.7 Agentic Coding Harnesses: The Model Gets Hands 7
- 1.8 A Precise Definition of an Agentic Coding Harness 8
- 1.9 The Four Capabilities of a True Harness 9
- 1.10 Capability 1: Autonomous File-System Read and Write 10
- 1.11 Capability 2: Shell Command Execution 10
- 1.12 Capability 3: Recursive Self-Correction 11
- 1.13 Capability 4: Tool Registration and Tool Schemas 12
- 1.14 Human-in-the-Loop vs. Human-on-the-Loop 13
- 1.15 Why Harnesses Need More Permissions Than Chat 14
- 1.16 What Harnesses Still Cannot Safely Do 15
- 1.17 Tool Calls as the Boundary Between Text and Action 15
- 1.18 Walkthrough: Adding Input Validation with a Harness 16
- 1.19 Comparing the Same Task Across Tool Categories 18
- 1.20 How to Classify Any New AI Coding Tool 19
- 1.21 Common Misconceptions About Coding Agents 20
- 1.22 Hands-On Lab: First Agentic Session 21
- 1.23 Interview Questions 22
- 1.24 Chapter Summary 24
- 1.25 What Comes Next 25
-
2 The Agent Loop 26
- 2.1 The Loop Behind the Illusion 27
- 2.2 Why Harnesses Are Easier to Debug Once You See the Loop 27
- 2.3 A Tiny Tool-Call Round Trip 29
- 2.4 The Five Phases of the Agent Loop 30
- 2.5 Phase 1: Perception 31
- 2.6 Phase 2: Planning 32
- 2.7 Phase 3: Execution 33
- 2.8 Phase 4: Observation 34
- 2.9 Phase 5: Termination 35
- 2.10 Correction: The Loop Learns From Its Own Tool Output 35
- 2.11 Tool Calls as Structured Intent 36
- 2.12 What the Harness, Not the Model, Actually Executes 37
- 2.13 How Context Is Rebuilt on Every Turn 38
- 2.14 Turn Limits, Token Budgets, and Timeouts 38
- 2.15 Permission Denials as Loop Events 39
- 2.16 Failed Commands and Error Re-Injection 40
- 2.17 Clean Termination vs. Stalled Termination 41
- 2.18 The Loop in Pseudocode 42
- 2.19 Walkthrough: A Docstring Task From Start to Finish 43
- 2.20 Reading Agent Logs 46
- 2.21 Diagnosing Common Loop Failures 47
- 2.22 How Different Harnesses Expose the Loop 49
- 2.23 Hands-On Lab: Trace a Real Agent Loop 49
- 2.24 Interview Questions 51
- 2.25 Chapter Summary 53
- 2.26 What Comes Next 54
-
3 Anatomy of a Harness 55
- 3.1 From Loop Behavior to Harness Architecture 55
- 3.2 The Five Components Every Harness Needs 56
- 3.3 Why Anatomy Matters When Agents Misbehave 58
- 3.4 Component 1: The System Prompt 58
- 3.5 System Prompt Layering: Built-In, Global, Project, Session 59
- 3.6 Global Instruction Files 61
- 3.7 Project-Level Instruction Files: AGENTS.md and CLAUDE.md 62
- 3.8 What Belongs in Project Instructions 63
- 3.9 What Does Not Belong in Project Instructions 64
- 3.10 Component 2: Tool Definitions 65
- 3.11 Tool Schemas, Parameters, and Descriptions 66
- 3.12 Built-In Tools vs. Registered Tools 67
- 3.13 MCP as a Tool-Registration Layer 68
- 3.14 Component 3: The Context and State Manager 69
- 3.15 How a Harness Decides What the Model Sees 70
- 3.16 Conversation History, File Contents, Repo Maps, and Summaries 71
- 3.17 Context Drift and Context Pollution 71
- 3.18 Component 4: The Permission Gate 72
- 3.19 Approval, Denial, Confirmation, and Auditability 73
- 3.20 Component 5: The LLM Client 75
- 3.21 Provider APIs, Local Endpoints, and Model Routing 75
- 3.22 Mapping Configuration Files to Harness Components 77
- 3.23 Troubleshooting by Component 79
- 3.24 Hands-On Lab: Build Your First Harness Anatomy Map 80
- 3.25 Interview Questions 81
- 3.26 Chapter Summary 83
- 3.27 What Comes Next 84
-
4 The Model Context Protocol (MCP) 87
- 4.1 Why Harnesses Need External Tools 87
- 4.2 The Problem MCP Solves 88
- 4.3 MCP in One Sentence 89
- 4.4 Clients, Servers, Tools, and Results 90
- 4.5 How MCP Fits Into the Agent Loop 91
- 4.6 Tool Discovery: How the Harness Learns What Exists 92
- 4.7 Tool Schemas: How Capabilities Become Callable 93
- 4.8 Tool Calls: From Model Intent to Server Request 94
- 4.9 Tool Results: Observations Returned to the Loop 95
- 4.10 Transport Option 1: stdio 96
- 4.11 Transport Option 2: HTTP/SSE 97
- 4.12 Choosing stdio vs. HTTP/SSE 97
- 4.13 Registering a Fictional MCP Server 99
- 4.14 Building a Minimal Python MCP Server 100
- 4.15 Validating Tool Inputs 101
- 4.16 Returning Useful Tool Results 103
- 4.17 Error Handling and Recovery 103
- 4.18 Security: Every Tool Is a Capability 104
- 4.19 Authentication and Authorization for Networked MCP 105
- 4.20 Logging and Auditability 106
- 4.21 Common MCP Failure Modes 107
- 4.22 Hands-On Lab: Build a Tiny todo_mcp_server 109
- 4.23 Interview Questions 110
- 4.24 Chapter Summary 112
- 4.25 What Comes Next 113
-
5 Context & Cache Management 114
- 5.1 When Good Agent Sessions Go Stale 114
- 5.2 What ``Context'' Means in a Coding Harness 115
- 5.3 The Context Window Is a Budget, Not a Filing Cabinet 116
- 5.4 What Enters the Context on Each Turn 117
- 5.5 How Context Grows During an Agent Loop 118
- 5.6 Input Tokens, Output Tokens, and Tool Results 119
- 5.7 Why More Context Is Not Always Better 120
- 5.8 Context Pollution: When Observations Become Noise 121
- 5.9 Prompt Caching: Paying Less for Stable Prefixes 122
- 5.10 What Caching Does Not Solve 123
- 5.11 Compaction: Compressing the Session State 124
- 5.12 What Compaction Can Lose 124
- 5.13 Clearing and Restarting Sessions 125
- 5.14 Repo Maps and Symbol-Level Context 126
- 5.15 Full Files vs. Snippets vs. Summaries 128
- 5.16 MCP Results and Context Pressure 129
- 5.17 Task Scoping as the Primary Cost Control 130
- 5.18 Estimating Token Growth Before You Run 131
- 5.19 Cost Formulas Without Hardcoded Prices 132
- 5.20 Local Models and Smaller Context Windows 133
- 5.21 Walkthrough: Rescuing a Drifting Session 134
- 5.22 Context Hygiene Checklist 135
- 5.23 Hands-On Lab: Measure Context Growth 138
- 5.24 Interview Questions 139
- 5.25 Chapter Summary 140
- 5.26 What Comes Next 142
-
6 Tool Design & Safety 143
- 6.1 Tools Turn Text Into Effects 144
- 6.2 Why Tool Design Is a Safety Problem 145
- 6.3 The Standard Built-In Tool Set 146
- 6.4 File Read Tools 147
- 6.5 Preventing Path Traversal 148
- 6.6 File Write and Edit Tools 149
- 6.7 Diffs, Atomic Writes, and Reviewability 151
- 6.8 Search Tools and Their Failure Modes 151
- 6.9 Shell Execution: The Sharpest Tool 152
- 6.10 Dangerous Command Patterns 153
- 6.11 Test Runner Tools as Safer Shell Alternatives 155
- 6.12 Web Fetch and Network-Aware Tools 156
- 6.13 MCP Tools as Capability Surfaces 157
- 6.14 Permission Gates: Allow, Confirm, Deny 158
- 6.15 Pattern-Based Command Policies 159
- 6.16 Path-Based Permission Rules 160
- 6.17 Session Decisions and Audit Logs 161
- 6.18 Sandboxing Beyond Permission Prompts 162
- 6.19 Read-Only Workspaces and Writable Scratch Areas 163
- 6.20 Prompt Injection Through Repository Content 164
- 6.21 Linguistic Defenses vs. Hard Boundaries 165
- 6.22 Least-Privilege Tool Design 166
- 6.23 Returning Safe and Useful Tool Results 167
- 6.24 Walkthrough: Blocking a Risky Cleanup Command 167
- 6.25 Tool Safety Checklist 169
- 6.26 Hands-On Lab: Test a Permission Policy 170
- 6.27 Interview Questions 171
- 6.28 Chapter Summary 173
- 6.29 What Comes Next 174
-
7 Claude Code Deep Dive 177
- 7.1 Why Claude Code Gets the First Deep Dive 177
- 7.2 Claude Code Through the Harness Anatomy 178
- 7.3 Installation and First Run, Without Overfitting to Today's Syntax 180
- 7.4 Working Inside a Repository 181
- 7.5 Instruction Layers: Global, Project, and Session 182
- 7.6 Writing a Safe Project Instruction File 183
- 7.7 The Core Claude Code Workflow 184
- 7.8 File Reads, Diffs, and Reviewable Edits 186
- 7.9 Shell Commands and Permission Prompts 187
- 7.10 Configuring Safe Command Behavior 188
- 7.11 Slash Commands as Repeatable Workflows 188
- 7.12 Custom Commands for Team Conventions 190
- 7.13 Hooks as Quality Gates 191
- 7.14 Hook Design: Useful, Small, and Observable 192
- 7.15 MCP Integration in Claude Code 193
- 7.16 Context Management in Long Sessions 194
- 7.17 Clear, Compact, Restart, and Handoff Notes 195
- 7.18 Workflow 1: Small Bug Fix 196
- 7.19 Workflow 2: Test-First Change 197
- 7.20 Workflow 3: Code Review and Safety Audit 197
- 7.21 Walkthrough: Adding CLI Argument Validation 198
- 7.22 Where Claude Code Fits Best 200
- 7.23 Where Claude Code Is Not the Best Fit 201
- 7.24 Claude Code Safety Checklist 202
- 7.25 Hands-On Lab: Build a Safe Claude Code Workflow 203
- 7.26 Interview Questions 204
- 7.27 Chapter Summary 206
- 7.28 What Comes Next 208
-
8 Aider Deep Dive 209
- 8.1 Why Aider Deserves Its Own Deep Dive 210
- 8.2 Aider Through the Harness Anatomy 210
- 8.3 Installation and First Run, Without Overfitting to Today's Syntax 212
- 8.4 Aider's Center of Gravity: Repo Maps and Git 213
- 8.5 Working Inside a Git Repository 213
- 8.6 The Repo Map: Structural Context Without Full Files 214
- 8.7 What Repo Maps Are Good At 216
- 8.8 What Repo Maps Cannot Replace 217
- 8.9 Selecting Files for the Conversation 218
- 8.10 Keeping the File Set Small Enough to Reason About 219
- 8.11 Git as a Safety Net 219
- 8.12 Auto-Commit, Diff Review, and Undo 221
- 8.13 Edit Formats and Patch Reliability 222
- 8.14 When Edits Fail to Apply 223
- 8.15 Architect Mode: Planning and Editing as Separate Roles 223
- 8.16 Configuration With .aider.conf.yml 225
- 8.17 Model Selection and Provider Flexibility 226
- 8.18 Local Models With Aider 226
- 8.19 Workflow 1: Small Targeted Edit 227
- 8.20 Workflow 2: Repo-Map-Assisted Refactor 228
- 8.21 Workflow 3: Test-Driven Bug Fix 229
- 8.22 Walkthrough: Adding Type Hints and Docstrings 229
- 8.23 Where Aider Fits Best 231
- 8.24 Where Aider Is Not the Best Fit 232
- 8.25 Aider Safety Checklist 233
- 8.26 Hands-On Lab: Compare Repo Map Editing 234
- 8.27 Interview Questions 235
- 8.28 Chapter Summary 237
- 8.29 What Comes Next 238
-
9 OpenCode, Goose, and Codex 240
- 9.1 Opening Scene: Three Agents, Same Bug, Different Loops 241
- 9.2 Why This Chapter Is Not a Product Ranking 242
- 9.3 The Anatomy Lens Returns 243
- 9.4 OpenCode's Center of Gravity 244
- 9.5 OpenCode Through the Five-Component Anatomy 244
- 9.6 A Conceptual OpenCode Workflow 245
- 9.7 Where OpenCode Fits Well 247
- 9.8 Where OpenCode Is a Poor Fit 247
- 9.9 Goose's Center of Gravity 248
- 9.10 Goose Through the Five-Component Anatomy 249
- 9.11 A Conceptual Goose Workflow 250
- 9.12 Goose as Automation Surface, Not Only Coding Surface 252
- 9.13 Where Goose Fits Well 252
- 9.14 Where Goose Is a Poor Fit 253
- 9.15 Codex's Center of Gravity 254
- 9.16 Codex Through the Five-Component Anatomy 255
- 9.17 A Conceptual Codex Workflow 255
- 9.18 Codex and the Delegated-Task Loop 257
- 9.19 Where Codex Fits Well 259
- 9.20 Where Codex Is a Poor Fit 259
- 9.21 Same Bug, Three Harness Loops 260
- 9.22 Comparison Table: Centers of Gravity 261
- 9.23 Comparison Table: Anatomy Mapping 262
- 9.24 Permission and Tool-Surface Differences 263
- 9.25 Runtime Placement: Local, Remote, and Hybrid 264
- 9.26 Configuration and Reproducibility 264
- 9.27 Reviewability as the Deciding Factor 266
- 9.28 Choosing a Harness by Task Shape 266
- 9.29 A Small Field Guide for Mixed-Tool Teams 267
- 9.30 Safety Checklist 268
- 9.31 Lab 9: Compare Three Harness Loops on One Bounded Change 269
- 9.32 Interview Questions 271
- 9.33 Chapter Summary 273
- 9.34 What Comes Next 274
-
10 Running Harnesses on Local Models 276
- 10.1 Opening Scene: The Same Harness, a Different Model Boundary 277
- 10.2 What ``Local Model'' Means in a Harness 278
- 10.3 Why Local Models Matter for Agentic Coding 279
- 10.4 Why Local Models Do Not Solve Everything 280
- 10.5 The Five-Component Anatomy With a Local Model 280
- 10.6 Local Inference Engines: Conceptual Landscape 282
- 10.7 Model Formats, Quantization, and Hardware in Practical Terms 283
- 10.8 A Conceptual Local Endpoint 284
- 10.9 Local Models and Repository Context 285
- 10.10 Local Models and Tool Use 285
- 10.11 Local Models and Structured Output 287
- 10.12 Local Models and Coding Tasks 288
- 10.13 Task Routing: Local, Hosted, or Hybrid 289
- 10.14 The Hybrid Harness Pattern 291
- 10.15 Privacy Boundaries and Mistaken Assumptions 291
- 10.16 Cost Boundaries and Mistaken Assumptions 292
- 10.17 Latency and Developer Experience 293
- 10.18 Evaluation: The Missing Discipline 294
- 10.19 A Small Local-Model Evaluation Suite 295
- 10.20 Failure Modes of Local Models in Coding Harnesses 296
- 10.21 Designing Prompts for Local Models 297
- 10.22 Context Compression and Local Models 298
- 10.23 Local Models in Claude Code, Aider, OpenCode, Goose, and Codex-Like Workflows 299
- 10.24 Configuration Examples 300
- 10.25 Security Checklist for Local Model Use 301
- 10.26 Same Task, Three Model Placements 303
- 10.27 Decision Table: When to Use Local Models 304
- 10.28 Operational Habits 305
- 10.29 Lab 10: Build a Local-Model Routing Plan 306
- 10.30 Interview Questions 307
- 10.31 Chapter Summary 309
- 10.32 What Comes Next 310
-
11 Workspace Configuration & AGENTS.md Files 311
- 11.1 When the Agent Does Not Know the Project 312
- 11.2 Project Instructions as a Repository Contract 312
- 11.3 Global Instructions vs. Project Instructions 313
- 11.4 Session Prompts Are Not Project Configuration 315
- 11.5 The Anatomy of a Good AGENTS.md 316
- 11.6 Project Purpose and Scope 317
- 11.7 Repository Layout Without Oversharing 318
- 11.8 Build, Test, Lint, and Format Commands 319
- 11.9 Coding Conventions That Agents Can Follow 320
- 11.10 Generated Files and Source-of-Truth Rules 320
- 11.11 Dependency and Package-Management Rules 321
- 11.12 Security Constraints and Secret Boundaries 322
- 11.13 Forbidden, Confirmed, and Safe Operations 323
- 11.14 Completion Checklists 324
- 11.15 What Not to Put in Project Instructions 325
- 11.16 Keeping Instructions Short Enough to Stay Useful 326
- 11.17 Configuration Across Different Harnesses 327
- 11.18 Claude-Style Project Instructions 328
- 11.19 Aider Configuration and File Selection 329
- 11.20 OpenCode, Goose, and Codex Configuration Concepts 330
- 11.21 Local-Model Notes in Project Configuration 331
- 11.22 Before and After: Rewriting a Weak AGENTS.md 332
- 11.23 Walkthrough: Teaching sample_cli Its Own Rules 334
- 11.24 Testing Whether the Harness Follows Instructions 335
- 11.25 Maintaining Project Instructions Over Time 337
- 11.26 Workspace Configuration Checklist 337
- 11.27 Hands-On Lab: Write and Test an AGENTS.md 338
- 11.28 Interview Questions 340
- 11.29 Chapter Summary 342
- 11.30 What Comes Next 342
-
12 Choosing a Harness Per Task 344
- 12.1 The Wrong Harness for the Right Task 344
- 12.2 Harness Selection Is Task Selection 345
- 12.3 The Six-Step Selection Model 346
- 12.4 Step 1: Classify the Task 348
- 12.5 Step 2: Estimate Scope and Context Pressure 349
- 12.6 Step 3: Identify Required Capabilities 350
- 12.7 Step 4: Choose Permission Posture 351
- 12.8 Step 5: Choose Local, Cloud, or Hybrid 352
- 12.9 Step 6: Define Verification Before Starting 352
- 12.10 Task Category 1: Small Bug Fix 354
- 12.11 Task Category 2: Test-First Change 355
- 12.12 Task Category 3: Large Refactor 355
- 12.13 Task Category 4: Greenfield Feature 357
- 12.14 Task Category 5: Documentation Update 357
- 12.15 Task Category 6: Code Review 358
- 12.16 Task Category 7: Security Audit 359
- 12.17 Task Category 8: Dependency Upgrade 360
- 12.18 Task Category 9: CI Repair Loop 361
- 12.19 Task Category 10: Low-Cost Local Batch Work 362
- 12.20 Claude Code Fit 363
- 12.21 Aider Fit 364
- 12.22 OpenCode Fit 365
- 12.23 Goose Fit 365
- 12.24 Codex Fit 366
- 12.25 Local Model Fit 366
- 12.26 Cloud Model Fit 367
- 12.27 Permission Postures by Task Risk 368
- 12.28 Decision Tree: Pick the Harness 370
- 12.29 Scenario Walkthroughs 372
- 12.30 Common Selection Mistakes 374
- 12.31 Hands-On Lab: Build Your Harness Decision Matrix 375
- 12.32 Interview Questions 376
- 12.33 Chapter Summary 378
- 12.34 What Comes Next 378
-
13 Cost, Speed, and Capability (Cloud vs. Local) 380
- 13.1 The Cheapest Model Is Not Always the Cheapest Workflow 381
- 13.2 Why Agentic Loops Cost Differently Than Chat 382
- 13.3 The Unit of Analysis: The Whole Task 382
- 13.4 Cloud Cost Components 383
- 13.5 Input Tokens, Output Tokens, and Cached Tokens 384
- 13.6 Turns, Retries, and Tool-Result Growth 385
- 13.7 Prompt Caching and Its Limits 386
- 13.8 Local Cost Components 387
- 13.9 Hardware Amortization Without Hardware Hype 389
- 13.10 Electricity, Maintenance, and Setup Time 390
- 13.11 Human Time as a Real Cost 390
- 13.12 Speed: Latency, Throughput, and Tool Time 391
- 13.13 Why Round Trips Matter in Agent Loops 392
- 13.14 Capability: Reliability Under Tool Use 393
- 13.15 Structured Output and Tool-Calling Failure Rates 394
- 13.16 Context Length and Capability Degradation 395
- 13.17 Privacy and Data Boundaries 396
- 13.18 Why Local Does Not Automatically Mean Safe 396
- 13.19 Cloud Cost Formula 397
- 13.20 Local Cost Formula 398
- 13.21 Break-Even Formula 398
- 13.22 Retry-Adjusted Cost 400
- 13.23 Human-Time-Adjusted Cost 401
- 13.24 Worked Example 1: Documentation Update 401
- 13.25 Worked Example 2: Focused Bug Fix 402
- 13.26 Worked Example 3: Large Refactor 403
- 13.27 Worked Example 4: Read-Only Security Audit 404
- 13.28 Worked Example 5: Repetitive Local Batch Task 405
- 13.29 Hybrid Routing Strategies 406
- 13.30 Routing by Task Shape, Risk, and Verification 408
- 13.31 Common Cost Mistakes 409
- 13.32 Hands-On Lab: Build a Cost and Capability Worksheet 410
- 13.33 Interview Questions 412
- 13.34 Chapter Summary 414
- 13.35 What Comes Next 415
-
14 Multi-Harness Workflows 416
- 14.1 One Harness Does Not Have to Do Everything 417
- 14.2 What a Multi-Harness Workflow Is 418
- 14.3 When Multi-Harness Workflows Help 418
- 14.4 When They Add Unnecessary Complexity 419
- 14.5 The Phase Model: Plan, Edit, Verify, Audit, Review 420
- 14.6 Handoff Artifacts 421
- 14.7 Context Handoff Without Context Pollution 423
- 14.8 Git as Shared Workflow Memory 424
- 14.9 Keeping Review Boundaries Clean 424
- 14.10 Workflow 1: Plan With Claude Code, Edit With Aider 425
- 14.11 Workflow 2: Read-Only Local Audit, Cloud Fix 427
- 14.12 Workflow 3: Aider Refactor, Claude Code Review 428
- 14.13 Workflow 4: Local Documentation Sweep, Cloud Final Review 429
- 14.14 Workflow 5: CI Repair Loop With Strict Limits 430
- 14.15 Workflow 6: Sandboxed Dependency Upgrade 431
- 14.16 Workflow 7: Shared MCP Validation Tool 432
- 14.17 Designing Safe Handoff Notes 433
- 14.18 Designing Plan Files That Another Harness Can Execute 435
- 14.19 Avoiding Cross-Tool Drift 436
- 14.20 Avoiding Duplicate Work 437
- 14.21 Avoiding Permission Escalation 437
- 14.22 Multi-Harness Security Checklist 439
- 14.23 Walkthrough: Refactoring task_tracker With Two Harnesses 440
- 14.24 Common Multi-Harness Mistakes 444
- 14.25 Hands-On Lab: Design a Two-Harness Workflow 445
- 14.26 Interview Questions 446
- 14.27 Chapter Summary 448
- 14.28 What Comes Next 449
-
15 Build Your Own Minimal Harness 451
- 15.1 Why Build a Harness Yourself? 452
- 15.2 The Minimal Harness Feature Set 452
- 15.3 The Target Architecture 453
- 15.4 What This Teaching Harness Will Not Do 455
- 15.5 Representing Tool Definitions 456
- 15.6 Designing a Stable Tool Result Shape 457
- 15.7 Building the LLM Client Boundary 459
- 15.8 Using a Fake Client for Deterministic Tests 460
- 15.9 Safe Path Resolution 462
- 15.10 Implementing read_file 463
- 15.11 Implementing a Conservative Test Runner 464
- 15.12 Why Generic Shell Access Is Dangerous 466
- 15.13 Implementing the Permission Gate 466
- 15.14 Classifying Commands as Allow, Confirm, or Deny 468
- 15.15 Building the Context for Each Turn 469
- 15.16 Dispatching Tool Calls 470
- 15.17 Re-Injecting Tool Results as Observations 471
- 15.18 Termination Conditions 473
- 15.19 The Main Loop in Code 474
- 15.20 Walkthrough: Fixing parse_count in sample_cli 476
- 15.21 Testing the Harness 477
- 15.22 What Commercial Harnesses Add 480
- 15.23 Where This Minimal Harness Is Unsafe 482
- 15.24 How to Extend It Safely 482
- 15.25 Hands-On Lab: Build the Minimal Harness 483
- 15.26 Interview Questions 484
- 15.27 Chapter Summary 487
- 15.28 What Comes Next 487
-
16 Staying Current in a Fast-Moving Ecosystem 489
- 16.1 The Ecosystem Will Keep Moving 490
- 16.2 What Changes Fast and What Changes Slowly 490
- 16.3 Durable Concepts From This Book 491
- 16.4 Version-Sensitive Surfaces 492
- 16.5 Release Notes Are Not Enough 494
- 16.6 The Harness Update Checklist 495
- 16.7 Evaluating a New Harness 497
- 16.8 Evaluating a New Model 499
- 16.9 Evaluating Local-Model Improvements 500
- 16.10 Evaluating MCP and Tooling Changes 502
- 16.11 Evaluating Permission and Sandbox Changes 503
- 16.12 Maintaining Workspace Instructions 504
- 16.13 Maintaining Cost and Routing Rules 505
- 16.14 Maintaining Benchmark Tasks 506
- 16.15 Building a Small Evaluation Suite 508
- 16.16 Upgrade, Wait, or Roll Back 508
- 16.17 Solo Developer Workflow 510
- 16.18 Team Workflow 511
- 16.19 Security Review Cadence 512
- 16.20 Decision Records for Tooling Changes 513
- 16.21 Common Ways Teams Drift 514
- 16.22 A Practical Adoption Scorecard 516
- 16.23 Hands-On Lab: Evaluate a Harness Update 518
- 16.24 Interview Questions 519
- 16.25 Final Synthesis 521
- Conclusion 523
Agentic Coding Harnesses, Compared & Explained
Master Claude Code, Aider, OpenCode, Goose, and Codex — Then Run Them Locally
A vendor-neutral, mechanism-level field guide to operating and extending agentic coding harnesses (542 manuscript pages).
Minimum price
$19.99
$29.99
You pay
Author earns
About
About the Book
Coding agents now run autonomously on your behalf, searching a repository, editing files, running commands and tests, and correcting their own errors. Tools such as Claude Code, Aider, OpenCode, Goose, and Codex look very different on the surface, yet all are described with the same few words.
Agentic Coding Harnesses, Compared & Explained looks underneath. Its premise is that the intelligence lives in the model but the agency lives in the harness, the program that turns the model's requests into real actions and decides when to stop. The book follows the agent loop turn by turn, takes a harness apart into its components, and covers the infrastructure they share: the Model Context Protocol, context and cache management, and tool design and safety.
It then applies that lens to specific tools, with deep dives on Claude Code and Aider and a survey of OpenCode, Goose, and Codex, and moves on to practice: running harnesses on local models, writing project instruction files, choosing a harness per task, weighing cost against capability, and combining several harnesses in one workflow.
The book closes by building a minimal harness in Python against a scripted fake model, and by showing how to evaluate new tools as the ecosystem keeps moving. The chapters are vendor-neutral, use fictional projects, and treat specific commands and settings as conceptual sketches to verify against current documentation.
Author
About the Author
Yohan is a Senior Full-Stack Software Engineer with extensive experience delivering scalable, end-to-end software solutions across web, enterprise, and cloud-based environments. He specializes in architecting robust platforms, modernizing legacy systems, driving cloud transformation efforts, and building integration-heavy applications that support critical business workflows. He is recognized for translating complex requirements into reliable, maintainable, and high-value solutions across industries such as insurance, cybersecurity, and professional services.
Known for combining strong technical execution with a practical business mindset, he has contributed to projects from concept and design through production delivery and long-term support. His experience includes collaborating with cross-functional teams, improving development workflows, solving complex technical challenges, and helping organizations deliver dependable software products that adapt to changing business needs. He brings a balanced approach to engineering that values quality, efficiency, and continuous improvement.
Contents
Table of Contents
Get the free sample chapters
Click the buttons to get the free sample in PDF or EPUB, or read the sample online here
The Leanpub 60 Day 100% Happiness Guarantee
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...
Earn $8 on a $10 Purchase, and $16 on a $20 Purchase
We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.
(Yes, some authors have already earned much more than that on Leanpub.)
In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.
Learn more about writing on Leanpub
Free Updates. DRM Free.
If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).
Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.
Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.
Learn more about Leanpub's ebook formats and where to read them
Write and Publish on Leanpub
You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!
Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.
Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.