The Leanpub 60 Day 100% Happiness Guarantee
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...

What really happens between a prompt and the next generated token? This collection goes behind the abstractions to explore the engineering of modern LLM inference, from transformer fundamentals and memory management to quantization, KV caching, GPU optimization and production serving. Across C, Rust, vLLM, GGUF and local models, these books show how inference systems are built, optimized and operated when performance, cost and reliability actually matter.
Bought separately
$163
Minimum price
$99.00
$119
About the Bundle
LLM inference is where AI meets real engineering constraints. Models may be impressive, but getting them to run fast, efficiently and reliably on actual hardware is a very different problem.
This bundle brings together five practical books that explore that problem from different angles. From inference architecture and production systems to building complete engines in C and Rust, working with vLLM and understanding quantization, GGUF and local models, the books focus on what happens beneath the surface.
You will learn how inference works, where performance is gained or lost, how memory and compute shape system design and how to build and operate inference systems without relying on black boxes. The emphasis throughout is on fundamentals, source code, measurements and real engineering trade-offs.
Whether you are building an inference engine from scratch, optimizing an existing serving stack or simply want to understand what happens between a prompt and the generated token, this collection gives you a practical path from the fundamentals to production-grade systems.
About the Books
Modern AI inference is no longer a side task tacked on to training pipelines. It is a distinct engineering discipline with its own performance constraints, reliability requirements, cost dynamics and failure modes. This book explains how production inference systems are designed, optimized, deployed and operated. It is written for software engineers, ML engineers, platform engineers and technical leads who need to build inference platforms that are fast, reliable, cost-efficient and understandable. You will learn not only what inference engineering techniques exist but why they work, how they work internally, when to use them, what trade-offs they introduce and how to measure their impact. The book progresses from first principles through sophisticated production-grade systems, always grounding guidance in quantitative reasoning, reproducible measurement and trade-off analysis rather than marketing claims or prescriptive recipes.
This book teaches experienced engineers how to build, measure, and optimize production LLM inference systems. Using vLLM as its primary case study, it walks you from the first principles of transformer inference and GPU architecture through vLLM's internal design, advanced performance techniques, distributed deployment, and operational reliability. You will learn not just how to configure vLLM, but how to reason about memory bottlenecks, scheduling trade-offs, parallelism strategies, and cost-performance optimization so you can architect inference infrastructure that scales. The material is grounded in primary sources, source code, and measured benchmarks; version-dependent details are flagged explicitly so the underlying engineering principles endure beyond any single release.
This book takes you from the mathematical foundations of transformer inference all the way to a complete, working large language model inference engine written in portable C. You will understand exactly how transformers generate text one token at a time, how weights are stored and loaded from SafeTensors format, how KV caching makes inference practical, and how to optimize every component for performance. By the final chapter you will have a genuinely executable program that loads a real TinyLlama-1.1B model from HuggingFace, tokenizes your prompt, runs the full transformer forward pass on your CPU, maintains the KV cache, and streams generated text to your terminal. No pseudocode, no placeholders, no exercises left for you to implement on your own: every line of code is provided, explained, and tested as a complete end-to-end system.
This book is a complete guide to understanding and building large language model inference engines in Rust. It takes you from the mathematical foundations of the transformer architecture through tokenization, attention mechanisms, KV caching, quantization, batching strategies, GPU and CPU optimization, distributed serving, and production deployment. Along the way, you will see idiomatic Rust code examples, performance benchmarks, comparisons with leading frameworks like vLLM, TGI, llama.cpp, Candle, Burn, and mistral.rs, and practical insights for shipping inference systems that rival the best open-source and commercial offerings. Whether you are a systems programmer interested in AI infrastructure or an ML engineer curious about the Rust ecosystem, this book gives you the depth to build, optimize, and understand.
This book teaches systems engineers how large language model inference actually works on real hardware, from transformer mechanics and quantization mathematics through GGUF internals and llama.cpp deployment to production-grade local inference services. You will learn to calculate exact memory requirements, diagnose performance bottlenecks systematically, configure GPU offloading correctly, benchmark rigorously, and build reliable multi-user inference systems that run efficiently on commodity CPUs and GPUs. No hand-waving about AI magic, just the engineering details that matter when you are responsible for making it work.
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...
We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.
(Yes, some authors have already earned much more than that on Leanpub.)
In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.
Learn more about writing on Leanpub
If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).
Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.
Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.
Learn more about Leanpub's ebook formats and where to read them
You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!
Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.
Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.