The Leanpub 60 Day 100% Happiness Guarantee
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...

Build three substantial software engines in C++: a compiler, a SQL database engine, and an LLM inference engine.
Bought separately
$79.00
$49.00
About the Bundle
Build C++: Compilers, Databases & LLM Engines
Three in-depth books for developers who want to go beyond using software and understand how complex software engines are built.
This bundle brings together three hands-on C++ projects:
Together, the books cover three very different domains while sharing the same engineering approach: understand the internals, implement the components, test them, and build the complete engine in C++.
This is not a collection of introductory C++ books. It is a collection for developers who want to work at the implementation level and see how substantial software is constructed from its core components.
3 books. 3 engines. One C++ engineering journey.
About the Books
Build an LLM Inference Engine in C++ — Through Challenges
You don't truly understand how large language models run
until you've built the engine yourself.
This book takes you from a blank C++ project to a complete,
working inference engine that loads a real Llama-family model
and generates text — one challenge at a time.
What you'll build:
- A strided tensor system with zero-copy views and arena allocation
- Math kernels: RMSNorm, SwiGLU, softmax, GEMM with SIMD
- A byte-level BPE tokenizer
- A full Transformer: RoPE, GQA, Flash Attention, KV Cache
- int8/int4 quantization with direct block multiplication
- GGUF model loading with mmap
- Sampling, streaming, speculative decoding, and continuous batching
- An optional CUDA capstone for the heaviest kernels
By Unit 14, the engine runs a real model on CPU. Every concept
earns its place right after you've built the thing it improves.
Companion Source Code
The complete source code for the book is available on GitHub:
https://github.com/Hatem-M-lab/llm-inference-engine
Development Methodology
This book was created through a process that combines careful human planning, content direction, and advanced AI technology, followed by thorough refinement and review to ensure a high-quality final work.
Who this is for:
C++20 developers comfortable with algorithms
and memory layout who want to understand what actually happens
inside an LLM runtime — not by reading, but by building.
Most compiler books start with a lexer, and nothing runs for chapters. This one starts with twelve bytes of machine code.
In the first hour you encode an x86-64 program by hand, wrap it in an ELF executable by hand, and run it. No assembler and no linker touches it. Then you climb, one layer per unit, until you have a complete compiler for a language called Slate: an assembler, relocatable object files the system linker accepts, a machine IR, two register allocators, SSA, an optimiser, a lexer, a parser, a type checker, a one-command driver, and DWARF line tables that let `gdb` step through Slate source one line at a time.
The order is the point. Because the back end comes first, every stage you write later has a real machine underneath it, and when the front end arrives in Unit 13 it already has somewhere to send its output.
This book was created through a process that combines careful human planning, content direction, and advanced AI technology, followed by thorough refinement and review to ensure a high-quality final work.
**Nothing is taken on trust.** The byte counts, addresses, sizes and exit statuses in the book come from running the code, and a five-layer verification suite stands behind them: hand-written encodings assembled independently by `nasm`, nearly every transcript re-executed, every "break it and watch it fail" experiment actually performed, every listing compared byte for byte with the file it came from, and random-program oracles that compile what nobody wrote. The repository's fault log lists 171 real mistakes found while the book was written, and what caught each one.
The last unit is the proof. A 270-line ray tracer is written twice, in Slate and in C. Your compiler builds one, `gcc` builds the other, and forty-one scenes must come out identical: the same picture to the byte, and the floating-point numbers behind every pixel the same to the bit.
What you get
- 19 units in five parts, more than 400 pages laid out for the screen
- More than 130 code listings, each with its real file path and line numbers, and more than 200 terminal transcripts
- About 22,000 lines of C++23 you can build, run and break, with the verification suite that checks them
- Linux on x86-64 and the usual toolchain (g++ 13+, nasm, binutils, gdb); no framework, no package manager, no dependencies
Build a working SQL database engine in C++20 -- from an empty directory to
a query processor that runs real SQL against data on disk -- through a
single relentless method: nothing is asserted; everything is demonstrated.
Every data structure in this book is built, compiled, and run. Every
performance claim is a table printed by a benchmark whose source code is on
the page in front of you. Every design decision is followed by the
measurement that justifies it -- and, where the design has a cost, by the
measurement that exposes that cost too. And once per unit, something fails
in front of you: a real bug, reproduced deterministically, diagnosed from
the evidence, fixed, and locked shut with a regression test.
This is not a survey of database theory. It is a lab manual. You will not
find hand-waving about how B+Trees are "generally logarithmic" -- you will
find the fan-out arithmetic, the page-count math, and a benchmark that
walks a tree of a million keys and prints the real number.
This book takes the engine across two complete parts and eight units:
PART I -- STORAGE
1. Slotted Pages & the Pager -- self-describing 4 KiB pages, records with
stable slot ids, a pager with a free list
2. B+Tree: Insert & Search -- logarithmic lookup, proven against a
million-key tree
3. B+Tree: Delete & Range Scans -- rebalancing, merges, ordered range
queries
4. The Buffer Pool -- a real cache with clock eviction, measured 2.5x
faster with identical logical work
PART II -- FROM BYTES TO A QUERY
5. The Record Layer -- typed rows: a schema-aware codec and a table heap,
reached by key through the index
6. The Catalog -- persistent schemas and named tables: the engine's
self-knowledge
7. The Front End -- a SQL tokenizer and recursive-descent parser, with
compiler-quality caret-pointed errors
8. The Executor -- the milestone: a real executor that runs CREATE TABLE,
INSERT, and SELECT end to end, against a database of hundreds of
thousands of rows
By the last page, the engine answers a SQL query it parsed from text,
against rows it stored on disk, through an index it built and a cache it
manages itself -- and every number in the book came from actually running
that code.
Reference machine: g++ 13.3.0, Ubuntu 24.04, C++20, stdlib + POSIX only.
Every benchmark ships with the exact command that produced it, so you can
run it yourself and get the same structural numbers.
Source Code:
https://github.com/Hatem-M-lab/labdb
Development Methodology
This book was created through a process that combines careful human planning, content direction, and advanced AI technology, followed by thorough refinement and review to ensure a high-quality final work.
A companion volume, "Build a SQL Database Engine in C++, Book 2: Parts III
& IV," continues the engine into durability (crash recovery via a
write-ahead log), concurrency, and performance -- available separately, and
as a discounted bundle with this book.
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...
We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.
(Yes, some authors have already earned much more than that on Leanpub.)
In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.
Learn more about writing on Leanpub
If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).
Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.
Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.
Learn more about Leanpub's ebook formats and where to read them
You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!
Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.
Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.