Leanpub Header

Skip to main content

Auditable LLM Pipelines

A case study in tracing, fallbacks, and honest evaluation

Auditable LLM Pipelines

Your LLM pipeline passed every check, and a fallback model still served 50 of its 71 successful calls. This case study follows one real multi-stage pipeline through its failures and shows how to make each stage traceable, attributable and honestly measured.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
About

About

About the Book

A multi-stage LLM pipeline can fail and still look green. A provider hits its quota, the fallback chain quietly hands the work to a weaker model, and the critique stage scores the result 85 and accepts it. Nothing in the logs tells you which model actually wrote what, which source backs which claim, or whether last week's "improvement" changed anything at all.

This book follows one real pipeline from start to finish: a harness that researches, plans, writes, critiques and revises whole technical books in stages. It shows what broke, how each failure was found, and the mechanism that now makes it visible. Every number comes from that pipeline's own runs, traces and evaluation sets, including the results that came back null.

  • Files as the source of truth, so every stage can be re-run on its own and a crash costs one chapter, not the run.
  • Typed outputs and human approval gates at each stage boundary.
  • Fallback chains that record which model actually served each call.
  • OpenTelemetry tracing with Phoenix as an audit trail you can query.
  • Hybrid retrieval with citation ids that survive from search to the finished text.
  • Measuring retrieval honestly: graded gold sets, a paired bootstrap, and the reranker that looked like a clear win and moved the score by 0.001.
  • Critique and revise loops that are typed, gated and replayable.
  • Context digests, cost control with cheap smoke runs, and catching machine prose habits.

It is written for engineers who build LLM systems with more than one step and need to answer "how do you know it did what it claims?" with evidence rather than a demo.

Researched and drafted with AI assistance by Ground Truth Books, using the pipeline the book describes, then reviewed and corrected by hand against that pipeline's code, run logs and traces. The pipeline itself is not published; its commands are shown as illustrations. Product names are used descriptively, and this book is not affiliated with or endorsed by their owners.

Author

About the Author

Ground Truth Books

Ground Truth Books publishes practical technical books on computer science, IT tools, data science, machine learning, software engineering and AI.

The name comes from machine learning, where "ground truth" means the real, verified answers you check a model against. That's how these books are made. They're researched and drafted with AI assistance, and then checked against reality. Examples are run against real captures and data in a lab, and those runs are cited like any other source, so you can see which output came straight from the tool. Every claim is cited to its source, with official documentation, specifications and source code preferred over blog posts.

Each book is built around doing the work. Chapters end with exercises, and an appendix gives worked answers you can reproduce on your own machine, using the same freely available data and captures.

Books are updated when the tools change. If you find an error, please report it: a corrected edition is free for every reader, which is one of the best things about Leanpub.

Contents

Table of Contents

About This Book

Introduction

  1. Trusting multi-stage pipelines
  2. Who this book is for
  3. The bookgen case study
  4. The route through the book
  5. The auditable pipeline

Files as the Source of Truth

  1. Why files beat state
  2. Per-project artifact layout
  3. JSON authoritative, corpus.db derived
  4. Content-keyed caches
  5. The chapters_done resume, read from files
  6. Fork and select versions
  7. Where determinism breaks

Typed Outputs and Human Gates

  1. Pydantic models at every boundary
  2. Validation before you trust
  3. Stage gate reviews
  4. The human gate checkpoint
  5. What the gates caught
  6. Failing fast at boundaries

Fallbacks and Model Provenance

  1. Why fallbacks exist
  2. pydantic-ai FallbackModel
  3. OpenRouter as the router
  4. Free-tier degraded fallbacks
  5. request.model versus response.model
  6. The degraded free-tier run
  7. Provenance through the stack

Tracing as Audit

  1. OpenTelemetry GenAI semantics
  2. Spans for every model call
  3. Tool invocations on the trace
  4. Phoenix as the audit pane
  5. The first full trace
  6. Tool errors hidden from the run
  7. pinned_only on two of thirty
  8. Retrieved but never cited
  9. Evidence you can trust

Hybrid Retrieval and Citation IDs

  1. Chunking the corpus
  2. BM25 and vector branches
  3. Reciprocal rank fusion
  4. Pinned boost and diversity decay
  5. Citation ids tie claims to sources
  6. Context assembly and limits
  7. Early retrieval failures

Measuring Retrieval Honestly

  1. Building graded gold sets
  2. Hash matching lost 518 judgements
  3. Containment matching fix
  4. NDCG and recall at K
  5. The paired bootstrap test
  6. The reranker null result
  7. Pooling bias and lower bounds
  8. Removing the reranker
  9. Reading numbers without lying
  10. Closing the loop

Critique and Revise Loops

  1. The critique stage
  2. Self-refine feedback loops
  3. Generated books as stress test
  4. Fabrications under valid citations
  5. The deterministic citation validator
  6. Resolving only known ids
  7. When revise makes it worse
  8. Replaying critiques for evaluation
  9. When a deterministic check is wrong

Context Rot and Digests

  1. Lost in the middle
  2. How context rots
  3. The measured rot trigger
  4. Chapter digests at 187 tokens
  5. Continuity, not mitigation
  6. Part intros from digests
  7. Stale digests after re-approval
  8. Measure before you mitigate

Cost Control and Smoke Runs

  1. Where the money goes
  2. Cheap end-to-end smoke runs
  3. Bounding spend per run
  4. Free tier as cost lever
  5. Real costs of bookgen

Stripping Machine Prose Habits

  1. The style lint pass
  2. Machine prose tells
  3. Blind voice evaluation
  4. AUC 0.42 and 12 of 20
  5. The habit tell list
  6. Linting without flattening voice

Conclusion

  1. What you can now do
  2. The mechanisms that transfer
  3. Where bookgen goes next
  4. The audit reflex

Appendix A: bookgen CLI and stage command reference

  1. Command reference
  2. Stage pipeline
  3. Configuration
  4. Project layout
  5. Measuring the harness
  6. Tracing
  7. The critique and revision loop
  8. Publishing

Appendix B: GenAI semantic convention attribute table

  1. Inference span attributes
  2. Tool span attributes
  3. Operation names
  4. Metrics and events
  5. Renamed or removed
  6. What bookgen reads

Appendix C: Gold set and scoring schema reference

  1. Files
  2. Relevance scale
  3. Building a gold
  4. Covering rules
  5. Score.py reference
  6. Comparing systems
  7. The two golds
  8. Limits

Glossary

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub