Agent Specification Engineering
- Building Reliable, Testable, and Governable AI Systems with DSLs, Contracts, and Executable Specifications
Copyright
- Code License
Disclaimer
About the Author
Preface
How to Use This Book
Terminology and Notation
The Missing Engineering Layer Between Intent and Execution
- The Production Problem Is Larger Than Answer Quality
- Prompting and Specification Solve Different Problems
- Goals Are Not Instructions, and Autonomy Is Not Authority
- A First AgentSpec
- Executable Does Not Mean Fully Formal
- The Reader Promise
Part I — Why Agents Need Specifications
Chapter 1 — The Agent Reliability Problem
- The Demonstration Hides the State Space
- Eight Reliability Failures Behind a Fluent Answer
- Why a Better Prompt Cannot Close These Gaps
- Unsafe Design: The Prompt as Control Plane
- Corrected Design: Proposals Inside an Enforcement Shell
- Failure Semantics Must Be Part of the Product
- Security and Privacy Are Execution Properties
- Reliability and Operational Implications
- When Not to Build an Agent
- Design Review: A Fluent but Unreliable Support Agent
- Production Checklist
- Key Takeaways
- Exercises
Chapter 2 — From Prompts to Behavioral Contracts
- A Ladder of Explicitness
- Design by Contract, Adapted Carefully
- Preconditions Are Not Suggestions to Prepare Better
- Postconditions Describe Observable Guarantees
- Invariants Protect the Execution Across Steps
- An Insufficient Structured Specification
- A Contract-Oriented Refactoring
- What Belongs in Prose, Data, Code, or Policy
- Guardrails Are Not More Instructions
- Security, Reliability, and Operations
- Trade-offs and When Not to Formalize Further
- Design Review: Refactoring an Operations Prompt
- Production Checklist
- Key Takeaways
- Exercises
Chapter 3 — Anatomy of an Agent Specification
- AgentSpec’s Design Boundary
- The Minimal Specification
- Metadata and Portability Make Context Reproducible
- Vocabulary Prevents Domain Drift
- Inputs and Outputs Define Ports, Not Merely Formats
- Objectives Separate Direction from Termination
- Capabilities and Tools Separate Ability from Adapters
- The Contract Section Defines Cross-Cutting Obligations
- State Turns a Loop into a Workflow
- Policies Express Contextual Authority
- Memory, Context, and Data Are Three Policies
- Budgets and Resilience Define the Outer Failure Envelope
- Observability and Audit Serve Different Consumers
- Evaluation Makes Release Expectations Executable
- The Insufficient Specification Revisited
- Implementation and Deployment Considerations
- When AgentSpec Is Not the Right Representation
- Design Review: Finding the Missing Section
- Production Checklist
- Key Takeaways
- Exercises
Part II — Designing Agent Contracts
Chapter 4 — Goals, Outcomes, and Completion Conditions
- A Vocabulary of Ending
- Goals Need Bounds
- Output Is Not Outcome
- Acceptance Criteria Belong to Different Owners
- Completion Must Be a Runtime Decision
- Terminal States Are Semantically Different
- Business Success Is Often Delayed
- Model Confidence Is One Signal
- An Unsafe Completion Design
- Cancellation and Changing Intent
- Security, Reliability, and Operational Implications
- When Not to Use a Complex Outcome Model
- Design Review: When “Resolved” Is Three Different Things
- Production Checklist
- Key Takeaways
- Exercises
Chapter 5 — Tool Contracts and Side Effects
- A Tool Is More Than a Function Signature
- Typed Inputs and Outputs
- Read, Reversible Write, Irreversible Write
- Preconditions and Postconditions Around Effects
- Timeouts Create Knowledge Problems
- Designing Idempotency Keys
- Retry Is a Classified Decision
- Compensation Is Forward Recovery
- Human Approval Belongs Before the Effect
- Tool-Result Provenance
- An Unsafe and Corrected Tool Design
- Security, Privacy, Reliability, and Operations
- When Not to Expose a Tool
- Design Review: The Timeout After Commit
- Production Checklist
- Key Takeaways
- Exercises
Chapter 6 — State, Memory, and Context
- Four Information Planes
- Workflow State Is Not a Transcript
- Domain State Has an Owner
- Context Is a Purpose-Built Projection
- Memory Needs a Retention Contract
- Staleness Is a Semantic Property
- Conflicting Context Must Remain Visible
- Context Minimization and Privacy
- An Unsafe and Corrected Design
- Reliability and Operational Implications
- When Not to Use Persistent Memory
- Design Review: The Convenient Summary That Became Truth
- Production Checklist
- Key Takeaways
- Exercises
Chapter 7 — Decisions and State Transitions
- The Model Proposes; the State Machine Commits
- States Should Represent Operational Meaning
- Valid and Invalid Transitions
- Guards Are Current Predicates
- Transition Effects Need Their Own Safety
- Terminal, Suspended, and Unknown States
- Loops Need Progress Measures
- Temporal Rules
- Recovery and Versioning
- An Unsafe and Corrected Workflow
- Security and Operational Implications
- When Not to Use a Rich State Machine
- Design Review: Two Workers, One Approval
- Production Checklist
- Key Takeaways
- Exercises
Chapter 8 — Policies, Permissions, and Authority
- Authority Is a Tuple, Not a Boolean
- Least Privilege Starts With Architecture
- Capabilities Are Delegable Units of Authority
- Roles and Attributes Answer Different Questions
- Deny by Default and Explicit Prohibition
- Policy as Code
- Tenant and Regional Boundaries
- Approval Thresholds and Separation of Duties
- Prompt Injection as an Authorization Problem
- Confused Deputies
- An Unsafe and Corrected Engineering Agent
- Reliability and Operational Implications
- When Not to Delegate Authority
- Design Review: The Confused Deputy in a Repository Agent
- Production Checklist
- Key Takeaways
- Exercises
Part III — Execution Architecture
Chapter 9 — Compiling Specifications into Workflows
- Compilation as a Pipeline
- Parsing and Normalization
- Structural Validation
- Semantic Validation
- Binding to Deployment Reality
- From State Model to Workflow Graph
- Planning Is Constrained Graph Construction
- Model-Facing Projections
- Generated Runtime Artifacts
- Runtime Execution
- An Unsafe and Corrected Compiler
- Security and Supply-Chain Considerations
- Operational Implications
- When Not to Generate
- Design Review: The Specification Passed but the Deployment Drifted
- Production Checklist
- Key Takeaways
- Exercises
Chapter 10 — Deterministic and Probabilistic Boundaries
- Five Responsibility Zones
- What the Model Should Interpret
- What Deterministic Code Should Validate
- What Policy Engines Should Authorize
- What Transaction Services Should Execute
- What Humans Should Approve
- The Deterministic Shell
- Constrained Decoding and Schema Enforcement
- Confidence Thresholds
- False Determinism
- An Unsafe and Corrected Decision Path
- Latency and Cost Trade-offs
- Security and Operational Implications
- When Not to Use a Model
- Design Review: JSON That Is Valid and Wrong
- Production Checklist
- Key Takeaways
- Exercises
Chapter 11 — Orchestration, Delegation, and Multi-Agent Systems
- Multi-Agent Is an Organizational Choice
- Supervisor and Specialists
- Delegation Contracts
- Task Ownership
- Authority Propagation
- Shared State Without Shared Confusion
- Handoff Schemas
- Result Validation and Trust
- Coordination Versus Decomposition
- Circular Delegation and Termination
- Unsafe and Corrected Multi-Agent Design
- Security and Privacy
- Reliability and Operations
- When Not to Use Multiple Agents
- Design Review: When More Agents Create Less Accountability
- Production Checklist
- Key Takeaways
- Exercises
Chapter 12 — Failure Semantics and Recovery
- Failure Is a Contract Outcome
- Transient and Permanent Failures
- Invalid Model Output
- Tool Timeout and Unknown Outcome
- Partial Execution
- Duplicate and Out-of-Order Results
- Dependency Failure
- Approval Timeout and Denial
- Conflicting Decisions
- Backoff and Jitter
- Dead-Letter Handling
- Compensation and Escalation
- Retry Versus Safe Replay
- Failure Budgets
- Unsafe and Corrected Recovery
- Security and Privacy
- Operational Implications
- When Not to Recover Automatically
- Design Review: Recovery Without Wishful Rollback
- Production Checklist
- Key Takeaways
- Exercises
Part IV — Verification and Quality
Chapter 13 — Testing Agent Specifications
- A Layered Test Strategy
- Schema Tests
- Parser and Normalization Tests
- Semantic Tests
- Contract Tests for Tools
- State-Transition Tests
- Invariant Tests
- Property-Based Testing
- Model-Based Testing
- Simulation and Failure Injection
- Adversarial Tests
- Golden Datasets and Regression Suites
- Mutation Testing
- Test Doubles and Deterministic Models
- Stochastic Test Strategy
- Security and Operations
- When Not to Use an End-to-End Model Test
- Design Review: A Test Suite That Passed the Wrong System
- Production Checklist
- Key Takeaways
- Exercises
Chapter 14 — Evaluating Behavior, Not Just Answers
- Evaluation Begins With a Behavioral Record
- Separate Hard Constraints From Quality
- Response Quality
- Task Completion
- Tool Selection and Argument Correctness
- Policy Compliance and Execution Order
- Escalation Quality
- Evidence Use and Source Attribution
- Latency, Cost, and Reliability
- Behavioral Scorecards
- Offline Evaluation
- Online Evaluation
- Evaluation Leakage
- Evaluator Disagreement
- Unsafe and Corrected Evaluation
- Security and Privacy
- When Not to Use a Single Metric
- Design Review: The 92 Percent Agent
- Production Checklist
- Key Takeaways
- Exercises
Chapter 15 — Guardrails and Runtime Enforcement
- A Guardrail Must Control a Boundary
- Input Validation
- Output Validation
- Domain Validation
- Policy and Authorization
- Rate Limits and Budgets
- Sandboxing
- Content Controls
- Data-Loss Prevention
- Approval Gates
- Kill Switches and Capability Rollback
- Layered Guardrails
- Unsafe and Corrected Engineering Guardrails
- Guardrail Composition and Failure
- Security and Privacy
- When Not to Add Another Guardrail
- Design Review: The Guardrail That Lived in the Prompt
- Production Checklist
- Key Takeaways
- Exercises
Chapter 16 — Observability, Auditability, and Replay
- Three Evidence Products
- Traces and Spans
- Execution Timelines
- Decision Records
- Version Identity
- Tool Calls and Policy Decisions
- Provenance
- Token and Cost Accounting
- Audit Logs
- Redaction
- Correlation Identifiers
- Safe Replay
- Forensic Investigation
- Unsafe and Corrected Observability
- Privacy Implications
- When Not to Store a Transcript
- Design Review: Reconstructing a Duplicate Effect
- Production Checklist
- Key Takeaways
- Exercises
Part V — Production Patterns
Chapter 17 — Customer-Support Triage Agent
- Scope and Architecture
- AgentSpec Boundary
- Data Model
- State Machine
- Tool Contracts
- Policies and Privacy
- Classification and Deterministic Rules
- Testing Strategy
- Walkthrough: The Suspicious-Login Ticket
- Decision Table
- Data-Threat Review
- Observability and Audit
- Failure Analysis
- Unsafe Design
- Corrected Design and Trade-offs
- When Not to Use This Agent
- Design Review: Mixed Intent Under Degraded Dependencies
- Production Checklist
- Key Takeaways
- Exercises
Chapter 18 — Software-Engineering Agent
- Scope and Architecture
- AgentSpec Boundary
- Repository Permissions
- Sandboxing and Command Contracts
- Planning Contract
- Patch and Artifact Provenance
- Test Gates
- Pull-Request Preparation and Human Review
- Walkthrough: Protected Workflow Drift
- State and Failure Paths
- Evaluation and Tests
- Observability
- Trade-offs
- When Not to Use This Agent
- Design Review: A Correct Patch from an Unsafe Process
- Production Checklist
- Key Takeaways
- Exercises
Chapter 19 — Financial-Operations Agent
- Scope and Architecture
- AgentSpec Boundary
- Authoritative Data and Immutable Evidence
- Threshold Policy
- Separation of Duties
- Review Package
- State and Human Authorization
- Idempotency and Duplicate Containment
- Compensation and Reversal
- Data Privacy and Explainability
- Walkthrough: Threshold Drift
- Tests and Evaluation
- Observability and Audit
- Failure Containment
- Trade-offs and When Not to Use
- Design Review: The Threshold at the Currency Cutoff
- Production Checklist
- Key Takeaways
- Exercises
Chapter 20 — Multi-Agent Incident Response
- Roles and Authority
- Architecture
- Supervisor AgentSpec
- Immutable Incident Timeline
- Evidence Collection
- Conflicting Hypotheses
- Time Bounds and Budgets
- Blast-Radius Control
- Remediation State Machine
- Communications Contract
- Walkthrough: The Wrong-Region Hypothesis
- Failure and Recovery
- Testing and Evaluation
- Observability and Post-Incident Replay
- Trade-offs and When Not to Use
- Design Review: Approval Invalidated by Late Telemetry
- Production Checklist
- Key Takeaways
- Exercises
Chapter 21 — Portability, Evolution, and Governance
- Model Portability
- Framework Portability
- Specification Versioning
- Schema Evolution
- Migration
- Deprecation
- Governance as an Engineering System
- Review Processes
- Agent Catalogs
- Organizational Maturity
- Production-Readiness Gates
- Unsafe and Corrected Evolution
- Security and Operational Implications
- When Not to Pursue Portability
- Design Review: A Portable Specification on an Incomplete Framework
- Production Checklist
- Key Takeaways
- Exercises
From Agent Demos to Engineered Systems
- The Discipline of Agent Specification Engineering
- The Central Separations
- A Practical Adoption Sequence
- A Ninety-Day Adoption Plan
- What AgentSpec Contributes
- The Standard for “Reliable”
- The Engineering Model to Carry Forward
Appendix A — AgentSpec 0.1 Language Reference
- Document and Serialization
- Conformance Levels
- Top-Level Inventory
metadataagentpurposeandvocabulary- Data Ports:
inputsandoutputs objectivescapabilitiestools- Predicates
contractstatepoliciesmemorycontextdatabudgetsresilienceobservabilityauditevaluationportability- Validation Order
- Minimum Runtime Events
- Compatibility and Migration
Appendix B — AgentSpec JSON Schema
- Dialect and Identity
- Closed Top-Level Structure
- Reusable Definitions
- Required Versus Optional Fields
- Enumerations
- Bounds
- Arrays and Uniqueness
- Embedded Port and Tool Schemas
- Schema Validation in the Reference Engine
- Versioning the Schema
- Using the Schema in CI
- Complete AgentSpec 0.1 Schema
Appendix C — Threat-Modeling Checklist
- 1. System and Purpose
- 2. Assets
- 3. Entry Points and Trust Boundaries
- 4. Identity and Delegation
- 5. Capability and Tool Authority
- 6. Prompt Injection and Untrusted Content
- 7. Data Privacy and Isolation
- 8. Context and Memory Integrity
- 9. Model and Provider Boundary
- 10. State and Workflow Integrity
- 11. Side Effects and Distributed Failure
- 12. Sandbox and Execution Environment
- 13. Human Approval and Social Engineering
- 14. Output and External Communication
- 15. Observability, Audit, and Replay
- 16. Supply Chain and Control Plane
- 17. Availability and Abuse
- 18. Evaluation and Detection
- 19. Residual Risk and Decision
- Threat-Model Review Record
Appendix D — Production-Readiness Checklist
- Release Header
- Gate 1 — Purpose, Ownership, and Scope
- Gate 2 — Specification and Build Integrity
- Gate 3 — Identity, Authorization, and Secrets
- Gate 4 — Data Governance and Privacy
- Gate 5 — Tooling and Side-Effect Safety
- Gate 6 — State, Concurrency, and Recovery
- Gate 7 — Model and Context Behavior
- Gate 8 — Sandbox, Network, and Supply Chain
- Gate 9 — Verification and Evaluation
- Gate 10 — Observability and Audit
- Gate 11 — Operations and Incident Response
- Gate 12 — Rollout and Change Control
- Release Decision
- Fast Red-Flag Review
Appendix E — Evaluation Scorecards
- Scorecard Header
- Universal Hard Gates
- Behavioral Rating Scale
- Core Behavioral Dimensions
- Weighting Template
- Risk Slices
- Customer-Support Triage Scorecard
- Software-Engineering Agent Scorecard
- Financial-Operations Agent Scorecard
- Multi-Agent Incident Scorecard
- Online Evaluation and Stop Conditions
- Adjudication Protocol
- Completed Scorecard Record
Appendix F — Exercise Solutions and Discussion Notes
- Chapter 1 — The Agent Reliability Problem
- Chapter 2 — From Prompts to Behavioral Contracts
- Chapter 3 — Anatomy of an Agent Specification
- Chapter 4 — Goals, Outcomes, and Completion Conditions
- Chapter 5 — Tool Contracts and Side Effects
- Chapter 6 — State, Memory, and Context
- Chapter 7 — Decisions and State Transitions
- Chapter 8 — Policies, Permissions, and Authority
- Chapter 9 — Compiling Specifications into Workflows
- Chapter 10 — Deterministic and Probabilistic Boundaries
- Chapter 11 — Orchestration, Delegation, and Multi-Agent Systems
- Chapter 12 — Failure Semantics and Recovery
- Chapter 13 — Testing Agent Specifications
- Chapter 14 — Evaluating Behavior, Not Just Answers
- Chapter 15 — Guardrails and Runtime Enforcement
- Chapter 16 — Observability, Auditability, and Replay
- Chapter 17 — Customer-Support Triage Agent
- Chapter 18 — Software-Engineering Agent
- Chapter 19 — Financial-Operations Agent
- Chapter 20 — Multi-Agent Incident Response
- Chapter 21 — Portability, Evolution, and Governance
Appendix G — Glossary
Appendix H — Bibliography and Further Reading
- Contracts, Formal Reasoning, and State
- Data Formats, APIs, and Events
- Security, Privacy, and Risk
- Observability, Reliability, and Supply Chain
- Language Models, Reasoning, and Tools
- Suggested Reading Paths
- Source Maintenance Note