Leanpub Header

Skip to main content

Beyond "Ship and Pray"

Testing Agentic Systems with Geometric Ground Truth

This book is 100% completeLast updated on 2026-07-23

An agent can score well on average and still fail exactly where it matters. Beyond “Ship and Pray” shows how to replace benchmark averages with designed experiments, geometric ground truth and failure attribution—so teams can discover when an agent breaks, identify the responsible component and test whether it fails safely under tool faults.

Minimum price

$9.95

$29.95

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
About

About

About the Book

Most AI evaluation ends with a single number: average accuracy. That number may look reassuring while hiding concentrated failures on adversarially framed requests, numerically dense passages, distracting context or broken tools. Beyond “Ship and Pray” begins with a more useful set of questions: Under which conditions does an agent fail? How badly? And why? It reframes evaluation as a designed experiment rather than a benchmark and treats the agent’s full trajectory—not merely its final answer—as the object to be tested.

Agus Sudjianto and Wing Yan Lau develop a reproducible methodology built on geometric ground truth. A verified knowledge graph and Exact Numerical Memory provide answers that are correct by construction. Base questions establish what is true, while ground-truth-invariant enrichment factors vary wording, persona, context, conversation, adversarial pressure, noise and operating constraints without changing the correct answer. Space-filling experimental designs then use a limited testing budget efficiently, producing balanced scenarios for both supervised training data and systematic evaluation.

The result is not another leaderboard. Readers learn how to test individual components, final outcomes, trajectory structure, governance, provenance, process health and resilience under tool faults. Logistic attribution turns a scalar score into a diagnosis of which conditions drive failure and which component is responsible. A full capstone applies the framework to a governed banking complaint agent end to end, supported by companion notebooks that make the methods reproducible.

Written for AI engineers, validation teams, model risk practitioners and governance leaders, Beyond “Ship and Pray” provides a systematic way to determine whether an agent is ready for deployment—not merely whether it usually works.

Author

About the Author

Agus Sudjianto and Wing Yan Lau

Agus Sudjianto

Agus Sudjianto is the Chief Scientist at KnowlytiX. He has spent more than two decades building, governing and validating quantitative models inside major financial institutions. He was Executive Vice President and Head of Model Risk at Wells Fargo, where he served on the Management Committee and led enterprise model risk management. Earlier in his career he held senior quantitative risk roles at Lloyds Banking Group and Bank of America. Since leaving corporate industry, he has continued this work as an advisor, builder and researcher across banking, fintech and AI.

Agus's work sits at the intersection of machine learning, model risk and governed AI systems. He created PiML and MoDeVa, toolkits for interpretable model development and validation, and his more recent work extends that same discipline into agentic AI, graph-grounded retrieval and geometric memory. Across these projects, the through-line is consistent: high-stakes AI should be built with the same rigor expected of high-stakes statistical models.

He is also co-author of Design and Modeling for Computer Experiments, holds several U.S. patents and has long worked across engineering, quantitative finance and applied machine learning. His current research centers on learning as geometry discovery in both predictive machine learning and generative AI.

In this series, Agus brings the perspective of someone who has spent a career asking not only whether a model works, but whether it can be governed, defended and trusted in practice.

WingYan Lau

Wing Yan Lau is the Chief Technology Officer at KnowlytiX. Her work centers on the systems layer that makes GMS usable in practice: document ingestion, knowledge-store construction, query infrastructure, verification pathways and the interfaces that connect governed AI to real enterprise data. She is a co-author of KnowlytiX's research on graph-verified evaluation and structured financial-document retrieval, including work reflected in FinStructBench and in the company's broader knowledge and testing stack.

Wing brings more than two decades of database and data-platform engineering experience to that work. She has contributed to core systems at IBM, SAP and Workday, with technical work spanning query optimization, storage systems and execution infrastructure. That background is visible throughout the KnowlytiX platform, where the challenge is not only to generate answers, but to connect models to structured knowledge in ways that remain exact, inspectable and operationally reliable.

In this series, Wing brings implementation discipline to every layer of the system: how documents become structured stores, how numeric facts remain exact, how graph-backed retrieval is made usable and how governed workflows are turned into code rather than left as intentions in prose. Her contribution is what turns the ideas in the architecture into systems an engineer can actually build, test and run.

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub