Leanpub Header

Skip to main content

AI Agent Evals

AI Agent Evals
This book is 91% completeLast updated on 2026-10-09

An agent's answer can look right while its actions are wrong. Learn to evaluate tasks, traces, tool calls and changed state with reproducible harnesses, tested graders and regression gates.

Minimum price

$14.99

$39.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
About

About

About the Book

A convincing answer is not proof that an AI agent did the right thing. To evaluate an agent, you need to inspect its actions, the state it changed, the constraints it respected and the evidence it left behind.

AI Agent Evals is a practical guide for engineers and architects building evaluation systems they can inspect and challenge. Starting with a small support system, it follows the work from testable task definitions and representative datasets through reproducible runs, traces, code graders, human review and model judges.

Learn how to evaluate retrieval and evidence use, tool calls and long-running state, conversations and adversarial behaviour. Test the evaluators themselves, make comparisons under uncertainty, and turn failures into regression tests and release gates rather than a reassuring dashboard.

The book connects evaluation design to CI and operational decisions: what should block a change, what needs human review, and which observations justify another experiment? Its supplied labs use synthetic tasks and offline execution. A recorded live-model case study examines actual questions, actions and refund ledgers; it is a small experiment, not a model ranking or proof of production performance.

For practitioners who want to move beyond scoring final answers and build evidence for decisions about AI agents.

Author

About the Author

Contents

Table of Contents

AI Agent Evals

  1. Dedication
  2. Preface
  3. Author’s note
  4. Reading contents

How to use this book

  1. Who this is for
  2. Get the companion code
  3. Start with the local example
  4. Prepare a working copy
  5. Choose a reading route
  6. What counts as an agent evaluation?
  7. Ask four questions
  8. Read the evidence labels
  9. A note on sources

Part 1: Foundations

Chapter 1: Watch an agent fail

  1. For decision-makers
  2. Decide what the customer should receive
  3. Run the first trial
  4. Follow the evidence into the grader
  5. Make the promised change
  6. Check the evaluator before trusting it
  7. Distinguish failure from missing measurement
  8. Repeat without inheriting success
  9. When doing nothing is correct
  10. Exercise: catch an unrelated side effect
  11. Takeaways
  12. Sources

Chapter 2: Turn success into a testable task

  1. For decision-makers
  2. Write the rule before the check
  3. Use a small, checked card
  4. Review the card without running the candidate
  5. Keep the earlier checkpoint intact
  6. Watch a lucky guess fail
  7. Complete the exchange
  8. Accept outcomes without prescribing a transcript
  9. Exercise: expose an unjustified rule
  10. Exercise: change the reply and reconcile the reviewers
  11. Takeaways
  12. Sources

Chapter 3: Build a dataset that represents the work

  1. For decision-makers
  2. Choose families before writing variations
  3. Wrap the existing task rather than invent another runner
  4. Make origin inspectable
  5. Exercise: find the hole before opening the repair
  6. Curate a golden dataset
  7. Keep ordinary work beside challenges
  8. Execute the cards without claiming model quality
  9. Define a scorecard before adding a headline
  10. Expand scope only when you can support the expectation
  11. Grow cases from decisions, then from traffic
  12. Review the coverage policy as well as the rows
  13. Leave a usable next checkpoint
  14. Transfer the contract to a second domain
  15. Takeaways
  16. Sources

Chapter 4: Create labels and protect comparisons

  1. For decision-makers
  2. Decide what a label means
  3. Preserve the first judgement
  4. Separate meaning from shared origin
  5. Freeze permitted uses before tuning
  6. Break the split before trusting it
  7. Know what duplicate checks cannot prove
  8. Make the fresh comparison claim narrow
  9. Exercise: change the rubric without recycling the comparison
  10. Keep golden status separate from holdout status
  11. Test the protection, then carry the packet forward
  12. Takeaways
  13. Sources

Part 2: Evaluation machinery

Chapter 5: Build the evaluation harness

  1. For decision-makers
  2. Start with the existing boundaries
  3. Follow one refund through the runner
  4. Give each trial its own state
  5. Cross the provider boundary explicitly
  6. Test the protocol without measuring a model
  7. Keep errors in the result
  8. Bound the work and name the timeout
  9. Compare direct and bounded retry fairly
  10. Run a crash and read the complete report
  11. Inspect the boundary that failed
  12. Keep live execution behind a separate decision
  13. Takeaways
  14. Sources

Chapter 6: Prepare the harness for real models

  1. For decision-makers
  2. The provider-neutral adapter
  3. Record and replay
  4. Reading real traces
  5. Capturing cost and latency honestly
  6. The first live run
  7. Lab: grade native tasks and retain failed evidence
  8. Lab: replay the cassettes through the Chapter 5 harness
  9. Plan your first live run
  10. Takeaways
  11. References

Chapter 7: Make runs reproducible and inspectable

  1. For decision-makers
  2. Keep the failed evidence before repairing anything
  3. Give the report a complete identity
  4. Read the failed refund in execution order
  5. Replay stored evidence without executing the agent
  6. Derive a new identity when the grader changes
  7. Redact the display, retain the grading boundary
  8. Exercise: verify the change between manifests
  9. Takeaways
  10. Sources

Chapter 8: Write and test code graders

  1. For decision-makers
  2. Choose evidence that can answer the question
  3. State the requirement before implementing the predicate
  4. Exactness, tolerance and malformed observations
  5. Compare deltas without losing invariants
  6. Preserve multiplicity and accept valid alternatives
  7. Write fixtures that disagree with the implementation
  8. Verification exercise: kill a duplicate-blind mutant
  9. Project trials without erasing errors
  10. Implement a noncompensatory scorecard
  11. Read the denominators before comparing agents
  12. Takeaways
  13. Sources

A live run: correct refunds, unnecessary questions

  1. Freeze the claim before the run
  2. Read the result with its denominators
  3. Failure one: the refund succeeds after a needless question
  4. Failure two: a correct refusal has the same problem
  5. What changed before Astra ran
  6. Carry the evidence into the next experiment

Part 3: Human and model judgement

Chapter 9: Design human review that teaches you something

  1. For decision-makers
  2. Decide what the reviewer is judging
  3. Write anchors before collecting ratings
  4. Build a packet with separate evidence roles
  5. Break the blinding, then repair it
  6. Preserve uncertainty without calling it bad quality
  7. Record independence before adjudication
  8. Bring in expertise for the right question
  9. Estimate the effort before scheduling it
  10. Exercise: replace a vague criterion
  11. Takeaways
  12. Sources

Chapter 10: Build a model judge without trusting it

  1. For decision-makers
  2. Begin with one question
  3. Give the judge relevant evidence
  4. Separate authority from content
  5. Pointwise scores and pairwise preferences
  6. Make invalid answers unusable
  7. Exercise: preserve refusal and unknown
  8. Build the provider path without claiming a live run
  9. Bind a judgement to its evaluator
  10. What this checkpoint establishes
  11. Choose a judge design, then test its errors
  12. Takeaways
  13. Sources

Chapter 11: Calibrate and challenge your judges

  1. For decision-makers
  2. Start with a label you can defend
  3. Read the matrix before the headline
  4. Exercise: choose a threshold before seeing the answer
  5. Check your choice against the report
  6. Count the cases the judge did not rate
  7. A skewed set rewards an always-pass judge
  8. Freeze before opening the audit
  9. Compare the threshold costs
  10. Challenge order and unnecessary length
  11. Diagnose a disagreement before changing the judge
  12. Keep the report strict and the authority narrow
  13. Takeaways
  14. Sources

Part 4: Measuring under uncertainty

Chapter 12: Measure nondeterministic agents

  1. For decision-makers
  2. Exercise: decide what the numerator means
  3. Preserve tasks while adding trials
  4. Read across a row before reading down a column
  5. Break the headline, then repair it
  6. Separate observed events from probability estimates
  7. Why the global mean cannot do this work
  8. Errors must remain in the picture
  9. Spend the next fixed budget deliberately
  10. Keep the checkpoint testable
  11. Takeaways
  12. Sources

Chapter 13: Compare changes with statistical care

  1. For decision-makers
  2. Exercise: discover what was counted
  3. Specify the quantity before choosing an interval
  4. Pair versions before aggregating
  5. Keep lineage separate from taxonomy
  6. Freeze the comparison card before opening outcomes
  7. Run the cumulative comparison
  8. Resample whole bundles
  9. Execute the false precision
  10. Make absence visible before interpreting success
  11. Spend effort on the unresolved uncertainty
  12. Read uncertainty against the decision threshold
  13. Retain a comparison another engineer can challenge
  14. Sources
  15. Takeaways

Part 5: Agent-specific evaluation

Chapter 14: Evaluate retrieval and evidence use

  1. For decision-makers
  2. Exercise: repair the evidence path first
  3. Define the task before labelling passages
  4. Keep effective dates separate from freshness
  5. Bound the retriever and freeze its inputs
  6. Measure relevance, coverage and order separately
  7. A citation establishes an address, not support
  8. Execute the false pass and its repair
  9. Handle conflict, incompleteness and absence honestly
  10. Why reference-free scores remain proxies
  11. Check the exercise without moving the target
  12. Extend relevance judgements carefully
  13. Takeaways
  14. Sources

Chapter 15: Evaluate tools, trajectories and long-running state

  1. For decision-makers
  2. Exercise: choose the dependencies before the transcript
  3. Evaluate selection, arguments and authority separately
  4. Specify a partial order, not a favourite route
  5. Give the business operation a durable identity
  6. Run the failure before repairing it
  7. Unknown outcome is not failed operation
  8. Build the tool fault matrix around effects
  9. Treat memory as a claim about shared state
  10. Attribute a handoff without inventing collaboration evidence
  11. Compare the trace with state in both directions
  12. Check the exercise and retain the counterexamples
  13. Takeaways
  14. Sources

Chapter 16: Simulate users and evaluate conversations

  1. For decision-makers
  2. First predict the disagreement
  3. Treat the user as an experimental component
  4. Build a loop without discarding the earlier lab
  5. Run the cumulative lab
  6. Repair the question, not the evidence
  7. Measure the whole conversation
  8. Replay recorded users without inventing them
  9. Validate the simulator against people
  10. Keep failures attributable
  11. When the user changes the environment
  12. Worked solution and next experiment
  13. Takeaways
  14. Sources

Chapter 17: Test adversarial behaviour without losing utility

  1. For decision-makers
  2. Exercise: decide what an incomplete comparison supports
  3. Define what the attacker can change
  4. Match the task, not just its label
  5. Observe the executed comparison
  6. Make denominators survive a failure
  7. Reproduce the misleading security result
  8. Repair behaviour without changing the ruler
  9. Attack the evaluator separately
  10. Verify claims as well as hashes
  11. Read a zero without overstating it
  12. Check the exercise and choose a new experiment
  13. Follow the instruction across the trust boundary
  14. Takeaways
  15. Sources

Chapter 18: Evaluate coding agents

  1. For decision-makers
  2. Exercise: specify an alternative before seeing the repair
  3. Give the issue a starting state
  4. Separate repair tests from preservation tests
  5. Put the evaluator outside the patch
  6. Observe what actually executes
  7. Reproduce the false green result
  8. Repair behaviour and retain the counterexamples
  9. Keep environment errors out of the success column
  10. Bind the report without mistaking it for authentication
  11. Inspect the evidence before making a delivery decision
  12. Check the alternative and decide what remains untested
  13. Move beyond a single-function patch
  14. Extend the protected runner to a real issue
  15. Takeaways
  16. Sources

Part 6: Production and operations

Chapter 19: Gate releases on complete evidence

  1. For decision-makers
  2. Exercise: decide what an incomplete run permits
  3. Give each evaluation a place in delivery
  4. Retain the scheduled work before execution
  5. Reproduce a successful skip
  6. Turn statuses into explicit decisions
  7. Build the local delivery configuration
  8. Validate evidence at consumption
  9. Keep runtime authority on a separate path
  10. Read the matrix as a decision aid
  11. Reuse an observation, then decide again
  12. Deliver the evidence with the change
  13. Takeaways
  14. Sources

Chapter 20: Instrument agents with traces you can evaluate

  1. For decision-makers
  2. Guided application exercise: decide what the record permits
  3. Choose identities before drawing a tree
  4. Record fields according to what they mean
  5. Use OpenTelemetry names without inventing an exporter
  6. Capture at the point of observation
  7. Repair the missing handoff
  8. Convert into the actual task contract
  9. Decide what survives ingestion
  10. Keep collection loss visible in monitoring
  11. Test the collector as evaluation code
  12. Evaluate delegated and tool-server boundaries
  13. Takeaways
  14. Sources

Chapter 21: Learn from production without fooling yourself

  1. For decision-makers
  2. Exercise: decide what the queue can tell you
  3. Start with the question, then collect the trace
  4. Preserve the selected schedule
  5. Repair the claim without hiding the sample
  6. Delayed outcomes need their own clock
  7. Monitor signals, investigate changes
  8. Keep the financial failure visible
  9. Turn an incident into a reviewed regression
  10. Privacy is part of collection design
  11. Check the exercise, then choose the next action
  12. Takeaways
  13. Sources

Chapter 22: Detect drift and catch regressions in production

  1. For decision-makers
  2. Exercise: choose the next action before seeing the plots
  3. Separate the things that can change
  4. Freeze identities before freezing thresholds
  5. Put fixed windows on a calendar
  6. Measure categorical input change
  7. A p-chart for abrupt proportion changes
  8. A CUSUM for persistent excess
  9. Run the offline experiment
  10. Read the evidence before choosing severity
  11. Many slices create an alerting policy
  12. Keep online experiments separate
  13. Stage the change and rehearse the fallback
  14. Give the runbook an owner and an evidence packet
  15. Worked check
  16. Takeaways
  17. Sources

Interlude: Roll out and roll back

  1. Exercise: separate a stop from a rollback
  2. Carry the complete decision into rollout
  3. Increase exposure without increasing authority by accident
  4. Check the connected rehearsal
  5. Takeaways
  6. Sources

Chapter 23: Diagnose failures at scale and close the loop

  1. For decision-makers
  2. Read the evidence before naming the defect
  3. Exercise: choose the next intervention
  4. Turn observations into falsifiable hypotheses
  5. Classify only after locating the boundary
  6. Reproduce the false improvement
  7. Compare one component under one evaluator
  8. Keep budgets and attempts visible
  9. Check the retrieval exercise
  10. Decide what evidence to collect next
  11. Record a root cause that can survive challenge
  12. Work through a queue, not a favourite incident
  13. Exercise: choose what deserves investigation
  14. Group notes without laundering judgement
  15. Reject a persuasive but inconsistent report
  16. Close the engineering loop without claiming approval
  17. Transfer the diagnosis to identity review
  18. Takeaways
  19. Sources

Chapter 24: Operate an evaluation programme that earns its cost

  1. For decision-makers
  2. Exercise: spend less without claiming the same coverage
  3. Start with the existing accounting boundary
  4. Price a scenario without manufacturing a bill
  5. Compare the simplest system that meets the constraints
  6. Reduce frequency before deleting evidence
  7. Hand over an obligation, not a dashboard
  8. Read the decision across the cumulative lab
  9. Takeaways
  10. Sources

Appendix A: Metrics and denominators

  1. Scheduled trials and assessable outcomes
  2. Equal trials and equal tasks
  3. Effects and communication have separate denominators
  4. Judge errors use the reference class
  5. Selected queues and population estimates
  6. Unknown values and cost ratios
  7. Statistical quantities

Appendix B: Exercise checks and debugging

  1. Check the decision, not just the number
  2. Diagnose a failed local check

Appendix C: Reusable templates

  1. Choose the right kind of file
  2. Task: state the acceptable effect
  3. Dataset: retain lineage before counting coverage
  4. Rubric: separate anchors from ratings
  5. Comparison: freeze the question before opening outcomes
  6. Release: collect evidence rather than fill in a pass
  7. Run and retain the check

Appendix D: Defensive evaluation engineering for Chapters 1-4 and 7

  1. Keep the expectation outside the candidate’s ownership
  2. Validate exact types and relationships
  3. Reject duplicate JSON members at the reader boundary
  4. Read exits in the context of the command
  5. Bind versions without confusing identity with truth
  6. Chapter 7: canonical JSON and manifest validation
  7. Checks to retain when extending a lab

Appendix E: Statistics toolkit

  1. For decision-makers
  2. Exercise: write the analysis before opening the answers
  3. Wilson intervals: uncertainty about a pass probability
  4. Plan precision and power separately
  5. Design effect: repeated rows buy less independent information
  6. McNemar: retain the repairs and regressions
  7. Kappa: agreement adjusted for the raters’ margins
  8. Alpha: count pairable values, not all possible ratings
  9. Holm: define the family before testing slices
  10. Repeated looks: an inspectable always-valid example
  11. CUSUM: surveillance is a different decision
  12. Numerical representation
  13. Takeaways
  14. Sources

Appendix F: Choose an evaluation framework and benchmark

  1. F.1 Start with the evidence you need
  2. F.2 Translate the book’s objects
  3. F.3 Verify portability before choosing convenience
  4. F.4 Select the workload, then the benchmark
  5. F.5 Interpret contamination and saturation honestly
  6. F.6 A worked selection decision
  7. F.7 Exercise
  8. Source keys

Appendix G: Delivery workflow and operational records

  1. G.1 Run the workflow locally
  2. G.2 Preserve missing evidence when using a cache
  3. G.3 Choose the record before filling it
  4. G.4 Work through a synthetic release and a deliberate refusal
  5. G.5 Write a monitoring specification that can refuse a decision
  6. G.6 Carry an incident into a reviewed evaluation
  7. G.7 Migrate before the fallback disappears

Appendix H: Governance questions for an evaluation owner

  1. H.1 Start with scope, not a familiar document number
  2. H.2 Match artifacts to questions
  3. H.3 Record the unresolved work

Glossary

  1. Terms to return to
  2. Additional glossary terms

Concept index

  1. A
  2. B
  3. C
  4. D
  5. E
  6. F
  7. G
  8. H
  9. I
  10. J
  11. K
  12. L
  13. M
  14. N
  15. O
  16. P
  17. R
  18. S
  19. T
  20. U
  21. W

Bibliography

  1. Companion files

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub