AI Agent Evals
- Dedication
- Preface
- Author’s note
- Reading contents
How to use this book
- Who this is for
- Get the companion code
- Start with the local example
- Prepare a working copy
- Choose a reading route
- What counts as an agent evaluation?
- Ask four questions
- Read the evidence labels
- A note on sources
Part 1: Foundations
Chapter 1: Watch an agent fail
- For decision-makers
- Decide what the customer should receive
- Run the first trial
- Follow the evidence into the grader
- Make the promised change
- Check the evaluator before trusting it
- Distinguish failure from missing measurement
- Repeat without inheriting success
- When doing nothing is correct
- Exercise: catch an unrelated side effect
- Takeaways
- Sources
Chapter 2: Turn success into a testable task
- For decision-makers
- Write the rule before the check
- Use a small, checked card
- Review the card without running the candidate
- Keep the earlier checkpoint intact
- Watch a lucky guess fail
- Complete the exchange
- Accept outcomes without prescribing a transcript
- Exercise: expose an unjustified rule
- Exercise: change the reply and reconcile the reviewers
- Takeaways
- Sources
Chapter 3: Build a dataset that represents the work
- For decision-makers
- Choose families before writing variations
- Wrap the existing task rather than invent another runner
- Make origin inspectable
- Exercise: find the hole before opening the repair
- Curate a golden dataset
- Keep ordinary work beside challenges
- Execute the cards without claiming model quality
- Define a scorecard before adding a headline
- Expand scope only when you can support the expectation
- Grow cases from decisions, then from traffic
- Review the coverage policy as well as the rows
- Leave a usable next checkpoint
- Transfer the contract to a second domain
- Takeaways
- Sources
Chapter 4: Create labels and protect comparisons
- For decision-makers
- Decide what a label means
- Preserve the first judgement
- Separate meaning from shared origin
- Freeze permitted uses before tuning
- Break the split before trusting it
- Know what duplicate checks cannot prove
- Make the fresh comparison claim narrow
- Exercise: change the rubric without recycling the comparison
- Keep golden status separate from holdout status
- Test the protection, then carry the packet forward
- Takeaways
- Sources
Part 2: Evaluation machinery
Chapter 5: Build the evaluation harness
- For decision-makers
- Start with the existing boundaries
- Follow one refund through the runner
- Give each trial its own state
- Cross the provider boundary explicitly
- Test the protocol without measuring a model
- Keep errors in the result
- Bound the work and name the timeout
- Compare direct and bounded retry fairly
- Run a crash and read the complete report
- Inspect the boundary that failed
- Keep live execution behind a separate decision
- Takeaways
- Sources
Chapter 6: Prepare the harness for real models
- For decision-makers
- The provider-neutral adapter
- Record and replay
- Reading real traces
- Capturing cost and latency honestly
- The first live run
- Lab: grade native tasks and retain failed evidence
- Lab: replay the cassettes through the Chapter 5 harness
- Plan your first live run
- Takeaways
- References
Chapter 7: Make runs reproducible and inspectable
- For decision-makers
- Keep the failed evidence before repairing anything
- Give the report a complete identity
- Read the failed refund in execution order
- Replay stored evidence without executing the agent
- Derive a new identity when the grader changes
- Redact the display, retain the grading boundary
- Exercise: verify the change between manifests
- Takeaways
- Sources
Chapter 8: Write and test code graders
- For decision-makers
- Choose evidence that can answer the question
- State the requirement before implementing the predicate
- Exactness, tolerance and malformed observations
- Compare deltas without losing invariants
- Preserve multiplicity and accept valid alternatives
- Write fixtures that disagree with the implementation
- Verification exercise: kill a duplicate-blind mutant
- Project trials without erasing errors
- Implement a noncompensatory scorecard
- Read the denominators before comparing agents
- Takeaways
- Sources
A live run: correct refunds, unnecessary questions
- Freeze the claim before the run
- Read the result with its denominators
- Failure one: the refund succeeds after a needless question
- Failure two: a correct refusal has the same problem
- What changed before Astra ran
- Carry the evidence into the next experiment
Part 3: Human and model judgement
Chapter 9: Design human review that teaches you something
- For decision-makers
- Decide what the reviewer is judging
- Write anchors before collecting ratings
- Build a packet with separate evidence roles
- Break the blinding, then repair it
- Preserve uncertainty without calling it bad quality
- Record independence before adjudication
- Bring in expertise for the right question
- Estimate the effort before scheduling it
- Exercise: replace a vague criterion
- Takeaways
- Sources
Chapter 10: Build a model judge without trusting it
- For decision-makers
- Begin with one question
- Give the judge relevant evidence
- Separate authority from content
- Pointwise scores and pairwise preferences
- Make invalid answers unusable
- Exercise: preserve refusal and unknown
- Build the provider path without claiming a live run
- Bind a judgement to its evaluator
- What this checkpoint establishes
- Choose a judge design, then test its errors
- Takeaways
- Sources
Chapter 11: Calibrate and challenge your judges
- For decision-makers
- Start with a label you can defend
- Read the matrix before the headline
- Exercise: choose a threshold before seeing the answer
- Check your choice against the report
- Count the cases the judge did not rate
- A skewed set rewards an always-pass judge
- Freeze before opening the audit
- Compare the threshold costs
- Challenge order and unnecessary length
- Diagnose a disagreement before changing the judge
- Keep the report strict and the authority narrow
- Takeaways
- Sources
Part 4: Measuring under uncertainty
Chapter 12: Measure nondeterministic agents
- For decision-makers
- Exercise: decide what the numerator means
- Preserve tasks while adding trials
- Read across a row before reading down a column
- Break the headline, then repair it
- Separate observed events from probability estimates
- Why the global mean cannot do this work
- Errors must remain in the picture
- Spend the next fixed budget deliberately
- Keep the checkpoint testable
- Takeaways
- Sources
Chapter 13: Compare changes with statistical care
- For decision-makers
- Exercise: discover what was counted
- Specify the quantity before choosing an interval
- Pair versions before aggregating
- Keep lineage separate from taxonomy
- Freeze the comparison card before opening outcomes
- Run the cumulative comparison
- Resample whole bundles
- Execute the false precision
- Make absence visible before interpreting success
- Spend effort on the unresolved uncertainty
- Read uncertainty against the decision threshold
- Retain a comparison another engineer can challenge
- Sources
- Takeaways
Part 5: Agent-specific evaluation
Chapter 14: Evaluate retrieval and evidence use
- For decision-makers
- Exercise: repair the evidence path first
- Define the task before labelling passages
- Keep effective dates separate from freshness
- Bound the retriever and freeze its inputs
- Measure relevance, coverage and order separately
- A citation establishes an address, not support
- Execute the false pass and its repair
- Handle conflict, incompleteness and absence honestly
- Why reference-free scores remain proxies
- Check the exercise without moving the target
- Extend relevance judgements carefully
- Takeaways
- Sources
Chapter 15: Evaluate tools, trajectories and long-running state
- For decision-makers
- Exercise: choose the dependencies before the transcript
- Evaluate selection, arguments and authority separately
- Specify a partial order, not a favourite route
- Give the business operation a durable identity
- Run the failure before repairing it
- Unknown outcome is not failed operation
- Build the tool fault matrix around effects
- Treat memory as a claim about shared state
- Attribute a handoff without inventing collaboration evidence
- Compare the trace with state in both directions
- Check the exercise and retain the counterexamples
- Takeaways
- Sources
Chapter 16: Simulate users and evaluate conversations
- For decision-makers
- First predict the disagreement
- Treat the user as an experimental component
- Build a loop without discarding the earlier lab
- Run the cumulative lab
- Repair the question, not the evidence
- Measure the whole conversation
- Replay recorded users without inventing them
- Validate the simulator against people
- Keep failures attributable
- When the user changes the environment
- Worked solution and next experiment
- Takeaways
- Sources
Chapter 17: Test adversarial behaviour without losing utility
- For decision-makers
- Exercise: decide what an incomplete comparison supports
- Define what the attacker can change
- Match the task, not just its label
- Observe the executed comparison
- Make denominators survive a failure
- Reproduce the misleading security result
- Repair behaviour without changing the ruler
- Attack the evaluator separately
- Verify claims as well as hashes
- Read a zero without overstating it
- Check the exercise and choose a new experiment
- Follow the instruction across the trust boundary
- Takeaways
- Sources
Chapter 18: Evaluate coding agents
- For decision-makers
- Exercise: specify an alternative before seeing the repair
- Give the issue a starting state
- Separate repair tests from preservation tests
- Put the evaluator outside the patch
- Observe what actually executes
- Reproduce the false green result
- Repair behaviour and retain the counterexamples
- Keep environment errors out of the success column
- Bind the report without mistaking it for authentication
- Inspect the evidence before making a delivery decision
- Check the alternative and decide what remains untested
- Move beyond a single-function patch
- Extend the protected runner to a real issue
- Takeaways
- Sources
Part 6: Production and operations
Chapter 19: Gate releases on complete evidence
- For decision-makers
- Exercise: decide what an incomplete run permits
- Give each evaluation a place in delivery
- Retain the scheduled work before execution
- Reproduce a successful skip
- Turn statuses into explicit decisions
- Build the local delivery configuration
- Validate evidence at consumption
- Keep runtime authority on a separate path
- Read the matrix as a decision aid
- Reuse an observation, then decide again
- Deliver the evidence with the change
- Takeaways
- Sources
Chapter 20: Instrument agents with traces you can evaluate
- For decision-makers
- Guided application exercise: decide what the record permits
- Choose identities before drawing a tree
- Record fields according to what they mean
- Use OpenTelemetry names without inventing an exporter
- Capture at the point of observation
- Repair the missing handoff
- Convert into the actual task contract
- Decide what survives ingestion
- Keep collection loss visible in monitoring
- Test the collector as evaluation code
- Evaluate delegated and tool-server boundaries
- Takeaways
- Sources
Chapter 21: Learn from production without fooling yourself
- For decision-makers
- Exercise: decide what the queue can tell you
- Start with the question, then collect the trace
- Preserve the selected schedule
- Repair the claim without hiding the sample
- Delayed outcomes need their own clock
- Monitor signals, investigate changes
- Keep the financial failure visible
- Turn an incident into a reviewed regression
- Privacy is part of collection design
- Check the exercise, then choose the next action
- Takeaways
- Sources
Chapter 22: Detect drift and catch regressions in production
- For decision-makers
- Exercise: choose the next action before seeing the plots
- Separate the things that can change
- Freeze identities before freezing thresholds
- Put fixed windows on a calendar
- Measure categorical input change
- A p-chart for abrupt proportion changes
- A CUSUM for persistent excess
- Run the offline experiment
- Read the evidence before choosing severity
- Many slices create an alerting policy
- Keep online experiments separate
- Stage the change and rehearse the fallback
- Give the runbook an owner and an evidence packet
- Worked check
- Takeaways
- Sources
Interlude: Roll out and roll back
- Exercise: separate a stop from a rollback
- Carry the complete decision into rollout
- Increase exposure without increasing authority by accident
- Check the connected rehearsal
- Takeaways
- Sources
Chapter 23: Diagnose failures at scale and close the loop
- For decision-makers
- Read the evidence before naming the defect
- Exercise: choose the next intervention
- Turn observations into falsifiable hypotheses
- Classify only after locating the boundary
- Reproduce the false improvement
- Compare one component under one evaluator
- Keep budgets and attempts visible
- Check the retrieval exercise
- Decide what evidence to collect next
- Record a root cause that can survive challenge
- Work through a queue, not a favourite incident
- Exercise: choose what deserves investigation
- Group notes without laundering judgement
- Reject a persuasive but inconsistent report
- Close the engineering loop without claiming approval
- Transfer the diagnosis to identity review
- Takeaways
- Sources
Chapter 24: Operate an evaluation programme that earns its cost
- For decision-makers
- Exercise: spend less without claiming the same coverage
- Start with the existing accounting boundary
- Price a scenario without manufacturing a bill
- Compare the simplest system that meets the constraints
- Reduce frequency before deleting evidence
- Hand over an obligation, not a dashboard
- Read the decision across the cumulative lab
- Takeaways
- Sources
Appendix A: Metrics and denominators
- Scheduled trials and assessable outcomes
- Equal trials and equal tasks
- Effects and communication have separate denominators
- Judge errors use the reference class
- Selected queues and population estimates
- Unknown values and cost ratios
- Statistical quantities
Appendix B: Exercise checks and debugging
- Check the decision, not just the number
- Diagnose a failed local check
Appendix C: Reusable templates
- Choose the right kind of file
- Task: state the acceptable effect
- Dataset: retain lineage before counting coverage
- Rubric: separate anchors from ratings
- Comparison: freeze the question before opening outcomes
- Release: collect evidence rather than fill in a pass
- Run and retain the check
Appendix D: Defensive evaluation engineering for Chapters 1-4 and 7
- Keep the expectation outside the candidate’s ownership
- Validate exact types and relationships
- Reject duplicate JSON members at the reader boundary
- Read exits in the context of the command
- Bind versions without confusing identity with truth
- Chapter 7: canonical JSON and manifest validation
- Checks to retain when extending a lab
Appendix E: Statistics toolkit
- For decision-makers
- Exercise: write the analysis before opening the answers
- Wilson intervals: uncertainty about a pass probability
- Plan precision and power separately
- Design effect: repeated rows buy less independent information
- McNemar: retain the repairs and regressions
- Kappa: agreement adjusted for the raters’ margins
- Alpha: count pairable values, not all possible ratings
- Holm: define the family before testing slices
- Repeated looks: an inspectable always-valid example
- CUSUM: surveillance is a different decision
- Numerical representation
- Takeaways
- Sources
Appendix F: Choose an evaluation framework and benchmark
- F.1 Start with the evidence you need
- F.2 Translate the book’s objects
- F.3 Verify portability before choosing convenience
- F.4 Select the workload, then the benchmark
- F.5 Interpret contamination and saturation honestly
- F.6 A worked selection decision
- F.7 Exercise
- Source keys
Appendix G: Delivery workflow and operational records
- G.1 Run the workflow locally
- G.2 Preserve missing evidence when using a cache
- G.3 Choose the record before filling it
- G.4 Work through a synthetic release and a deliberate refusal
- G.5 Write a monitoring specification that can refuse a decision
- G.6 Carry an incident into a reviewed evaluation
- G.7 Migrate before the fallback disappears
Appendix H: Governance questions for an evaluation owner
- H.1 Start with scope, not a familiar document number
- H.2 Match artifacts to questions
- H.3 Record the unresolved work
Glossary
- Terms to return to
- Additional glossary terms
Concept index
- A
- B
- C
- D
- E
- F
- G
- H
- I
- J
- K
- L
- M
- N
- O
- P
- R
- S
- T
- U
- W
Bibliography
- Companion files