A multi-stage LLM pipeline can fail and still look green. A provider hits its quota, the fallback chain quietly hands the work to a weaker model, and the critique stage scores the result 85 and accepts it. Nothing in the logs tells you which model actually wrote what, which source backs which claim, or whether last week's "improvement" changed anything at all.
This book follows one real pipeline from start to finish: a harness that researches, plans, writes, critiques and revises whole technical books in stages. It shows what broke, how each failure was found, and the mechanism that now makes it visible. Every number comes from that pipeline's own runs, traces and evaluation sets, including the results that came back null.
- Files as the source of truth, so every stage can be re-run on its own and a crash costs one chapter, not the run.
- Typed outputs and human approval gates at each stage boundary.
- Fallback chains that record which model actually served each call.
- OpenTelemetry tracing with Phoenix as an audit trail you can query.
- Hybrid retrieval with citation ids that survive from search to the finished text.
- Measuring retrieval honestly: graded gold sets, a paired bootstrap, and the reranker that looked like a clear win and moved the score by 0.001.
- Critique and revise loops that are typed, gated and replayable.
- Context digests, cost control with cheap smoke runs, and catching machine prose habits.
It is written for engineers who build LLM systems with more than one step and need to answer "how do you know it did what it claims?" with evidence rather than a demo.
Researched and drafted with AI assistance by Ground Truth Books, using the pipeline the book describes, then reviewed and corrected by hand against that pipeline's code, run logs and traces. The pipeline itself is not published; its commands are shown as illustrations. Product names are used descriptively, and this book is not affiliated with or endorsed by their owners.