- 1. What a Prompt Actually Is
- The only thing the model does
- A prompt is billed, budgeted and processed in tokens
- Why prompting works at all
- 2. In-Context Learning
- Designing a test that cannot be faked
- The measurement
- What the demonstrations are actually doing
- 3. Chain of Thought
- The depth limit
- Designing a task that isolates depth
- The measurement
- What this predicts about real systems
- 4. Structured Output
- Why asking nicely has a floor
- The experiment
- Validity against temperature
- The result that settles the argument
- What the constraint cannot do
- Choosing an approach
- 5. Hallucination and Grounding
- Why there is no ``I don't know'' in the mechanism
- Measuring it
- What grounding does, mechanically
- What follows for real systems
- 6. Prompt Injection
- The setup
- The measurement
- What this does and does not show
- What actually reduces the risk
- 7. What a Prompt Costs
- Why length is the wrong unit
- The cost of an encoding
- The cost of the parts of a prompt
- Costing a technique before adopting it
- 8. Evaluating Prompts
- The first noise floor: the run
- The second noise floor: the test set
- The two floors combine
- What to measure, beyond accuracy
- 9. What Prompting Can and Cannot Do
- The limits, and what causes each
- What prompting does reliably
- What this book did not establish
- The reverse map: your certification to these chapters
- How to use the questions
- Closing
Prompting and Controlled Output
Why prompting works when it works — built, broken and measured from scratch
Prompting has more advice than evidence. This book measures it instead: a model reading a
pattern it has never seen, the same model failing at a task in one pass and solving it
perfectly in four, a decoder that guarantees a format instead of improving the odds, and a
model that is 2.4% accurate while 78.6% confident.
Six central claims are built from scratch and measured on a single CPU core. Everything
taken from published research sits in a grey box marked "Not measured here" — so you never
have to guess which is which.
Minimum price
$16.00
$16.00
You pay
Author earns
About
About the Book
Prompting is the part of this field with the worst ratio of advice to evidence. Almost
everyone has a list of techniques that work; almost nobody can tell you why a technique
works, when it stops working, or what it costs. So the same practices circulate as folklore
— say "think step by step", assign the model a role, offer a tip — and when one of them
fails there is nothing to reason from.
The certifications ask about this material in the way folklore cannot answer. They describe
a system that behaves oddly and ask what to change. They give you two prompt designs and ask
which costs more. They describe an output that fails to parse and ask for the fix that
guarantees the format rather than merely improving the odds.
This book is organised by topic rather than by exam. Each chapter explains one mechanism
from the ground up, then shows exactly where each certification stops and what it will ask
you. The final chapter maps every certification back to the chapters it needs, in the order
to read them.
And it measures rather than asserts. Prompting is a behaviour of large models, and the build
environment for this book has no access to one — so instead of assembling other people's
results, it finds out how much of the subject can be built small enough to measure. The
answer turned out to be most of it:
- **In-context learning.** A model reproduces a pattern from the prompt alone at 0.90
accuracy against a chance floor of 0.03 — on a task whose content is regenerated at random
for every example, so the answer provably cannot be in the weights. A model one layer
shallower manages 0.26, which locates the ability in the architecture rather than the data.
- **Chain of thought.** Two identical models, the same task, the same training budget,
differing only in whether they may write intermediate steps: 1.000 against 0.089. The one
that must answer directly never leaves chance at any difficulty.
- **Structured output.** The same trained model decoded two ways. Free decoding never
reaches perfect validity at any training budget tested — twenty-four times the training
moved it from 0.933 to 0.987 and stopped. Constraining the decoder reaches 1.000 always,
discarding about one per cent of the model's probability mass.
- **Hallucination.** A model that answers known questions perfectly answers questions about
subjects it has never seen correctly 2.4% of the time — while reporting 0.786 confidence
in those answers, and 0.781 confidence specifically on the wrong ones. Putting the fact in
the context, with no change to the model, moves it to 1.000.
- **Prompt injection.** One conflicting instruction placed among the data takes a model from
1.000 to as low as 0.06 — and the model that recognises instructions by content, the more
capable one, is the more damaged. Capability and exposure turn out to be the same property.
- **What prompts cost.** The same twenty records in six formats, tokenised with a tokeniser
trained here: a 1.70x spread between the cheapest and most expensive way of writing
identical data. Instructions are nearly free; demonstrations dominate.
There is also a chapter on evaluation, which measures the thing that invalidates most claims
about prompting: eight identical configurations differing only in random seed produced
accuracies from 0.821 to 1.000, and one unchanged model measured on 20 test items returned
anything from 0.650 to 1.000.
Where a result comes from published research rather than this book's experiments, it appears
in a grey box marked "Not measured here", and no chapter rests its argument on one. Two
experiments that failed are reported as failures rather than quietly dropped.
What you get:
- 9 chapters covering prompts as context, in-context learning, chain of thought, structured
output and constrained decoding, hallucination and grounding, prompt injection, prompt
cost, evaluation, and a final chapter of limits.
- 48 original practice questions, tagged by certification and by difficulty, where every
option is explained — not just why the right answer is right, but why each wrong answer is
wrong, because on these exams the distractors are where the teaching is.
- A reverse map from each certification to the chapters that serve it, in reading order.
- All the code and experiment scripts, so every table and figure can be regenerated.
Written to the published objectives of NVIDIA NCA-GENL, NVIDIA NCP-GENL, NVIDIA NCP Agentic
AI, AWS Certified Generative AI Developer – Professional, and Databricks Certified
Generative AI Engineer. Objectives were checked in August 2026; confirm the current
blueprint with the certifying body before you sit.
You need to be able to read Python. That is all — and nothing in this volume assumes you
have read any other book.
Every question is original, written from published exam objectives. Nothing is reproduced
from, or based on recollection of, any live examination.
This book was created through a process that combines careful human planning, content direction, and advanced AI technology, followed by thorough refinement and review to ensure a high-quality final work.
Feedback
Author
About the Author
Hatem M. is a programmer and technical author whose work focuses on modern C++, large language models, and AI systems.
His books combine first-principles explanations with complete implementations and reproducible experiments. They include C++ Algorithmic Mastery, an eight-volume series on algorithms and problem solving; Build an LLM Inference Engine in C++, which constructs a GPT-style inference engine from scratch; LLM Quantization: From the Bits Up, which develops the theory and practice of neural network quantization from the bit level upward; and C++ Autopsy, a forensic investigation of ten subtle C++ bugs that compiled successfully, ran correctly, and still produced the wrong answers.
Contents
Table of Contents
The Leanpub 60 Day 100% Happiness Guarantee
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...
Earn $8 on a $10 Purchase, and $16 on a $20 Purchase
We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.
(Yes, some authors have already earned much more than that on Leanpub.)
In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.
Learn more about writing on Leanpub
Free Updates. DRM Free.
If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).
Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.
Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.
Learn more about Leanpub's ebook formats and where to read them
Write and Publish on Leanpub
You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!
Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.
Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.