- Chapter 1 — Why AI Systems Break Classic Testing
- Chapter 2 — The Four-Layer Playbook
- Chapter 3 — Golden Datasets
- Chapter 4 — LLM-as-Judge, and Judging the Judge
- Chapter 5 — RAG Evaluation: The Triad
- Chapter 6 — Adversarial and Multi-Turn Evals
- Chapter 7 — Running Evals: Gates, Economics, and Production
- Chapter 8 — Agentic Evals: When the System Under Test Takes Actions
- Chapter 9 — When Evals Become Training Data
- About the Author
Evals
Quality Engineering for AI Systems
Every test you have ever written assumed the same input gives the same output. Large language models broke that assumption, and with it most of what our profession knows about verification. This is the discipline that replaces it — golden datasets, property assertions, LLM judges you can actually trust, retrieval evaluation, adversarial suites, and CI gates that let a pipeline say no to a model.
Minimum price
$12.99
$24.99
You pay
Author earns
About
About the Book
Somebody has shipped an AI feature, and somebody has asked you whether it is good.
If you came up through test automation, you already know the uncomfortable part: none of your instruments fit. You cannot diff the output. You cannot snapshot it. You cannot run it twice and expect the same thing — and the vocabulary you would normally reach for, flaky and regression and expected value, describes a defect rather than the thing in front of you, which is the product working as designed.
This book is the discipline that replaces those instruments.
WHAT'S INSIDE
Nine chapters, each one dense, none padded to reach a page count.
Why AI systems break classic testing — including why "just set the temperature to zero" fails, and why it would not help even if it worked.
The four-layer playbook — property families sorted by cost, from deterministic checks that run in milliseconds to judges that bill per token.
Golden datasets — the regression suite nobody hands you: anatomy of an entry, sizing and slicing, cold-start from the spec, the circularity trap, and why structured facts get exact matches back.
LLM-as-judge, and judging the judge — rubrics that measure quality instead of fluency, the four biases that corrupt your metrics, and how to calibrate a judge against human labels before you let it gate a release.
RAG evaluation — the triad, evaluating retrieval on its own with recall@k and MRR, why chunking is an eval parameter rather than a preprocessing detail, and a fourth metric the triad misses entirely: did the answer leave something out?
Adversarial and multi-turn evals — context retention, corrections that must supersede, escalation that carries its context, direct and indirect prompt injection, and the PII masking order that catches people out.
Running evals — per-intent gates, tiered economics, triggers beyond code, drift monitoring, and an honest maturity model most teams will not enjoy reading.
Agentic evals — scoring trajectories rather than answers, environments that must be rebuilt rather than reset, and the ways an agent games a verifier.
And a closing argument: the engineer who writes the eval is now, functionally, part of the training team.
HOW IT IS WRITTEN
Every concept arrives with its failure story — how it breaks in production, and how you catch it. The carrot cake that scored perfectly. The judge that gave a fabricated number nine out of ten. The escalation metric that stayed green while every handoff arrived empty. The concepts here are not difficult; recognizing them in the wild, at eleven at night, in someone else's codebase, is the skill, and that is learned from cases rather than definitions.
You will receive every future update free, and there will be updates — this field moves faster than print.
Feedback
Author
About the Author
Pudur Ramaswamy is a quality and performance engineering leader with 20+ years at Apple, Netflix, PayPal, and eBay, where he built automation frameworks, ran the performance practice for a 56-application portfolio, and tested distributed systems at 5,000 transactions per second and 40,000 concurrent users.
He is co-inventor on U.S. Patent 6,510,402 — a distributed, integrated test-environment architecture granted in 2003 that anticipated today's parallel and sharded test execution.
He now works mostly on the evaluation of AI systems: LLM-as-judge harnesses, RAG and agentic evals, and the question of how you verify something that never returns the same answer twice. He is the creator of QAAutoPilot, an AI-powered QA product for on-premise enterprise deployment, and the author of The Pragmatic Comprehensive Senior SDET.
He writes at aitestguru.com.
Contents
Table of Contents
The Leanpub 60 Day 100% Happiness Guarantee
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...
Earn $8 on a $10 Purchase, and $16 on a $20 Purchase
We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.
(Yes, some authors have already earned much more than that on Leanpub.)
In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.
Learn more about writing on Leanpub
Free Updates. DRM Free.
If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).
Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.
Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.
Learn more about Leanpub's ebook formats and where to read them
Write and Publish on Leanpub
You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!
Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.
Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.