Leanpub Header

Skip to main content

100 LLM Autopsies

What broke, why nobody noticed, and how it was found

100 LLM Autopsies
This book is 100% completeLast updated on 2026-08-29

A model that crashes is a good day. The dangerous failures return answers — plausible, fluent, and wrong. 100 failures. 63 diagnostic instruments. One rule: inspect what actually happened.

Minimum price

$32.00

$32.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
307
Pages
About

About

About the Book


A model that crashes is a good day.

Crashes have stack traces. They have a line number, a repro, an owner, and a fix that either works or does not. Nobody debates for three weeks whether a segmentation fault is real.

The failures in this book do not crash. They return an answer — fluent, well-formed, confident, and wrong by four percent. They pass the test suite. They ship. They sit in production for months while a team optimises around them, and they are found eventually by accident, or by an outsider, or by someone who finally opened an artifact that had been sitting there since the first deployment.
This book was created through a process that combines careful human planning, content direction, and advanced AI technology, followed by thorough refinement and review to ensure a high-quality final work.

This is a book of one hundred such failures, opened up and examined.

Not a book about models being unreliable

Of the hundred cases here, the number caused by the model's weights being wrong is very small. The rest live in the machinery around it: a tokenizer, a template, a mask, a cache key, a config default, a retry, a metric. Those are ordinary software components, and they fail in ordinary software ways.

What makes them hard is that the system's output stays plausible while they fail — so nothing alerts, nothing throws, and the only signal is a number that is slightly worse than it should be.

  • A duplicated beginning-of-sequence token that cost seven points of instruction adherence.
  • An evaluation whose score changed by four points depending on the order the questions were asked in.
  • A prefix cache bounded by entry count, and therefore unbounded in memory.
  • A safety mask that, when retrieval returned nothing, made the model answer from a uniform blend of everything.
  • Two identical support tickets, classified differently, every time, correctly.

Four levels, defined by how much you must hold in your head

The cases are not organised by topic. They are organised by how much of the system you must understand at once to find the bug.

  • Level I — Beginner (01–30). One model, one machine, one request. Every bug is visible in an artifact somebody could have opened on the first day: a token array, a request body, a startup log, a cache key.
  • Level II — Intermediate (31–60). The bug does not exist for a single request. It is created by grouping, batching, adapting, indexing or training. You now need to reproduce a situation, not an input.
  • Level III — Advanced (61–85). Load, time, kernels, replicas. These bugs do not exist on your laptop, and several do not exist until the system has been running for six hours.
  • Level IV — Expert (86–100). The deceptive cases. The evidence is complete, the reasoning is sound, and the conclusion is wrong.

Every case follows the same rhythm

Case card, report, evidence, a failed suspect or two, the turn, the verdict, the fix — and where relevant, what the fix broke. Then the lesson, generalised past the specific bug, and a diagnostic instrument you keep.

The failed suspects are not padding. Learning which plausible explanations to discard is most of the skill, and a book that shows only correct reasoning teaches nothing about how to reason before you know the answer.

What you actually keep

If you remember a hundred bugs, you have learned a hundred bugs, and the hundred and first will be new.

So the cases are not the payload. Running through them is a set of 63 instruments — small, cheap, reusable diagnostic procedures, each introduced inside the case that makes you want it. Most take under twenty minutes to build; several are a single assertion. They share one property:

Every one of them answers a question about what is actually happening, rather than what the code says should be happening.

On evidence, and on trust

Every case card carries an Evidence basis line, with one of four values — documented incident, documented behaviour, reconstructed, or composite — because you are entitled to know how much of what you are reading actually happened. Most cases are reconstructed, and the book says so on the page rather than implying a hundred first-hand war stories.

What is not reconstructed is the mechanism. Every root cause is a real failure mode of real software. Where a case reports a measurement, the arithmetic in the table reconciles. Where a case gives you code as a fix, the code runs.

Who this is for

Anyone who has shipped a system with a language model inside it and has had the experience of not being able to explain what it just did. You do not need to have trained a model. You do need to be comfortable reading code — Python where models, training and evaluation live; C++ and configuration where engines and serving live.

What it is not

It is not about prompt engineering. It is not an API reference. And it is not a book about models being unreliable — almost every model in these hundred cases behaved exactly as it was built to behave. It was given two beginning-of-sequence tokens, or a mask that let it attend to padding, or a prompt in a format it had never been trained on, and it did the only thing it could do with what it was handed.

The model is rarely the patient. It is usually the witness.

Bundles

Bundles that include this book

Author

About the Author

Hatem M.

Hatem M. is a programmer and technical author whose work focuses on modern C++, large language models, and AI systems.

His books combine first-principles explanations with complete implementations and reproducible experiments. They include C++ Algorithmic Mastery, an eight-volume series on algorithms and problem solving; Build an LLM Inference Engine in C++, which constructs a GPT-style inference engine from scratch; LLM Quantization: From the Bits Up, which develops the theory and practice of neural network quantization from the bit level upward; and C++ Autopsy, a forensic investigation of ten subtle C++ bugs that compiled successfully, ran correctly, and still produced the wrong answers.

Contents

Table of Contents

Level I — Beginner

One model, one machine, one request.

  1. The Model That Got Worse in Production
  2. The Conversation the Server Never Had
  3. The GPU That Wasn't
  4. Correct, and Scored Wrong
  5. The Generation That Would Not Stop
  6. The Seed That Only Worked Alone
  7. The dtype That Ate the Memory
  8. The Persona That Faded at Turn Nine
  9. The Token That Printed Itself
  10. Graded on What Was Cut Off
  11. Two Models Behind One Endpoint
  12. The Parameter That Was Never Applied
  13. The Arabic Prompt That Cost Triple
  14. The Context Window That Was a Setting
  15. The Adapter That Never Loaded
  16. The Few-Shot That Gave Away the Answer
  17. The Fine-Tune That Forgot Its Manners
  18. The Repetition Penalty That Broke the Code
  19. The Last Chunk That Never Arrived
  20. Truncated From the Wrong End
  21. The Tokenizer That Aged
  22. The Space That Cost Eleven Points
  23. The Cache That Answered Yesterday's Question
  24. Temperature Zero Was Not Zero
  25. The User Who Typed a System Prompt
  26. Characters, Not Tokens
  27. The Benchmark It Had Already Seen
  28. The A/B Test That Compared a Model to Itself
  29. The Stop Sequence That Never Fired
  30. The Retry That Asked Twice
Level II — Intermediate

The bug does not exist for a single request.

  1. The Eval That Changed Its Mind When You Sorted It
  2. The Answer Meant for Another User
  3. Per-Tensor Where Per-Channel Was Required
  4. Training on the Question
  5. Two Embedding Models, One Index
  6. Padded on the Wrong Side
  7. Gibberish After Token 8192
  8. The LoRA on the Wrong Modules
  9. The Prefix That Belonged to Someone Else
  10. The Answer That Fell Between Two Chunks
  11. Off By One Position
  12. Calibrated on the Wrong Thousand Sentences
  13. The Collator That Dropped the Mask
  14. The Theta That Disagreed
  15. Merged Twice
  16. Cosine on Unnormalised Vectors
  17. The Eval That Remembered the Previous Question
  18. The Layer You Must Never Quantize
  19. Where One Document Ended
  20. Position IDs After Padding
  21. Every Citation Was One Document Off
  22. Full Attention on a Sliding-Window Model
  23. The Effective Batch Nobody Computed
  24. The Draft Model That Spoke a Different Language
  25. The Tie That Was Cut
  26. The Reranker That Read Half the Passage
  27. The Resume That Forgot
  28. Faster, and Different
  29. The Argument That Was a String
  30. The Same Shuffle Every Epoch
Level III — Advanced

Load, time, kernels, replicas.

  1. The Throughput That Decayed for Six Hours
  2. The Overflow in the Attention Scores
  3. The Request That Waited Forever
  4. The Conversation That Changed Replicas
  5. The Tail Nobody Logged
  6. Flash vs Eager
  7. The Leak With a Confident Hit Rate
  8. Backpressure Measured as Latency
  9. Epsilon on the Wrong Side
  10. The Locale in the Sidecar
  11. The Stampede After Every Deploy
  12. The Bill That Disagreed With the Logs
  13. A Softmax Without Its Maximum
  14. The Timeout That Looked Like a Model Failure
  15. The Pointer the Graph Captured
  16. All-Reduce Order
  17. The Positions That Stopped Being Distinct
  18. The Guardrail That Rewrote the Prompt
  19. Chunked Prefill Changed the Numbers
  20. The Quantized Cache That Only Failed at Depth
  21. The Atomic That Raced
  22. The Fallback Nobody Announced
  23. The Bubble Blamed on the Model
  24. The Experiment That Moved Its Own Baseline
  25. Two Optimisations That Were Correct Alone
Level IV — Expert

The evidence is complete and the conclusion is wrong.

  1. The Sink That Was Evicted
  2. One Layer, One Channel
  3. The Metric That Improved Because of the Bug
  4. One Node's Clock
  5. The Grammar That Allowed What It Forbade
  6. The Reference That Was Also Wrong
  7. The Ablation That Removed Nothing
  8. The NaN That Softmax Hid
  9. Everyone's Model Got Worse on the Same Day
  10. Leakage Through the Merge Table
  11. It Learned Position, Not Content
  12. Reward Hacked by a Preprocessing Artifact
  13. Better Benchmarks, Broken Calibration
  14. Fixed for the Wrong Reason
  15. The Bug That Was Never a Bug
Appendix

The Instruments — 63 diagnostic procedures, indexed by case and by symptom.

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub