Leanpub Header

Skip to main content

Building High-Performance LLM Inference Engines in C

From Transformer Fundamentals to a Complete SafeTensors Inference Runtime

Building High-Performance LLM Inference Engines in C
This book is 100% completeLast updated on 2026-09-29

Build a real LLM inference engine from the ground up in portable C. Learn how transformers work, load SafeTensors weights, implement KV caching and optimize inference on the CPU. By the end, you will run TinyLlama locally and generate text with a complete working runtime.

Minimum price

$25.00

$35.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
188
Pages
About

About

About the Book

This book takes you from the mathematical foundations of transformer inference all the way to a complete, working large language model inference engine written in portable C. You will understand exactly how transformers generate text one token at a time, how weights are stored and loaded from SafeTensors format, how KV caching makes inference practical, and how to optimize every component for performance. By the final chapter you will have a genuinely executable program that loads a real TinyLlama-1.1B model from HuggingFace, tokenizes your prompt, runs the full transformer forward pass on your CPU, maintains the KV cache, and streams generated text to your terminal. No pseudocode, no placeholders, no exercises left for you to implement on your own: every line of code is provided, explained, and tested as a complete end-to-end system.

Author

About the Author

Steve Publications

Steve is a technology professional with more than 20 years of experience in software development, server infrastructure, cybersecurity, vulnerability research and reverse engineering. Throughout his career, he has designed, secured, analyzed and tested complex software and infrastructure, with a particular focus on understanding how systems fail and how they can be made more secure.

Outside of work, Steve enjoys sharing knowledge with the technology community. He collaborates with researchers, industry experts and technology professionals to write practical books covering software development, cybersecurity, cloud computing, networking, DevOps, artificial intelligence and enterprise technologies. His books focus on practical learning through clear explanations, real-world examples and hands-on exercises. With more than two decades of industry experience, his goal is to help IT professionals, students and technology enthusiasts build useful skills and stay current in a rapidly changing industry.

We believe readers deserve to know how our books are created. Most of our authors are not native English speakers, so we use AI to help translate, proofread manuscripts, fix grammar, improve sentence structure and make technical explanations easier to read. AI is used as an editing tool only. It does not replace the research, technical knowledge or hands-on experience behind our books. Some of our authors also prefer to remain anonymous for privacy or professional reasons. In those cases, we publish their work under a different name. The author's name may be different, but the quality of the content and our review process remain the same.

Every book is written, reviewed and maintained by experienced technology professionals, with contributions from our private technical community of more than 420 engineers and researchers. We spend far more time validating technical accuracy and keeping our content up to date than generating text. We are always interested in working with experienced professionals who have deep expertise in a particular technology or domain. If you would like to publish a book with us or help review an existing manuscript, we'd love to hear from you. Send us a message describing your area of expertise. We are especially interested in niche technologies, specialized skills and emerging topics that are underrepresented in existing technical literature.

If you look through the contents of our books, you'll see practical examples, detailed explanations and material that is regularly updated. Our goal is to publish books that professionals can actually rely on, not low-effort AI-generated content. If you ever feel that one of our books does not meet that standard, Leanpub offers a 60-day money-back guarantee. Feel free to request a refund if you are not satisfied with your purchase.

Contents

Table of Contents

From Transformer Fundamentals to a Complete SafeTensors Inference Runtime

Introduction: How LLMs Actually Generate Text

Chapter 1: How LLMs Think: The Transformer Inference Journey

  1. The Inference Journey from Prompt to Response
  2. Why Inference Is Different from Training
  3. The Transformer Blueprint at a Glance
  4. What This Book Builds
  5. A First Look: Loading, Running, and Generating

Chapter 2: Tokens, Vocabularies, and the Language of Models

  1. What a Token Really Is
  2. Byte Pair Encoding and Subword Tokenization
  3. Loading and Interpreting Tokenizer Files
  4. Implementing Token Encoding in C
  5. Implementing Token Decoding and Special Tokens

Chapter 3: Tensors, Shapes, and Memory Layouts

  1. The Tensor as a Multi-Dimensional Array
  2. Row-Major vs Column-Major and Why It Matters
  3. Shapes, Strides, and Offsets
  4. Float16, BFloat16, and Float32 Representations
  5. Contiguous vs Strided vs Packed Tensors

Chapter 4: Building the Tensor Core Library

  1. Tensor Allocation and Lifetime Management
  2. Views, Reshaping, and Slicing Without Copies
  3. Element-Wise Operations and Broadcasting
  4. Matrix Multiplication from First Principles
  5. Reductions, Sums, and Normalization Primitives

Chapter 5: The Transformer Layer in Depth

  1. Layer Normalization and RMSNorm
  2. Multi-Head Self-Attention Mechanism
  3. Feed-Forward Networks and Activation Functions
  4. Residual Connections and Gate Linear Units
  5. Putting It All Together: One Transformer Block

Chapter 6: Positional Encodings and RoPE

  1. Why Positional Information Matters
  2. Learnable vs Absolute vs Rotary Encodings
  3. The Mathematics of RoPE
  4. Implementing RoPE in C
  5. RoPE Scaling and Long-Context Strategies

Chapter 7: Softmax, Logits, and Sampling

  1. From Hidden States to Logits
  2. Numerically Stable Softmax
  3. Temperature Scaling and Exploration
  4. Top-K Filtering and Nucleus Sampling
  5. Deterministic and Reproducible Generation

Chapter 8: The KV Cache and Autoregressive Generation

  1. The Autoregressive Generation Loop
  2. Key-Value Caching: The Core Optimization
  3. KV Cache Layouts and Memory Planning
  4. Prefill vs Decode Phases
  5. Managing Cache Growth and Limits

Chapter 9: The SafeTensors File Format

  1. Why SafeTensors and What Problem It Solves
  2. Header Structure and Serialization
  3. Tensor Descriptors, Data Types, and Offsets
  4. Endianness, Alignment, and File Boundaries
  5. Supported Data Types and Shape Encoding

Chapter 10: Implementing the SafeTensors Loader in C

  1. Header Reading and Length Prefix Parsing
  2. JSON Metadata Parsing in C
  3. Tensor Descriptor Extraction and Validation
  4. Mapping and Reading Tensor Data Blocks
  5. Error Handling, Boundaries, and Integrity Checks

Chapter 11: Model Configuration and Weight Mapping

  1. The config.json: Architecture Parameters
  2. Weight Naming Conventions and Patterns
  3. Mapping Transformer Layer Weights
  4. Handling Embeddings and Output Projections
  5. Validation: Checking Loaded Weight Shapes

Chapter 12: Building the Model Forward Pass

  1. Token Embedding Lookup
  2. The Layer Loop: Attention, FFN, and Residuals
  3. Positional Encoding at Each Position
  4. Final Layer Normalization
  5. Computing Output Logits
  6. The Complete Attention Implementation
  7. Putting It All Together: The Forward Pass

Chapter 13: End-to-End Inference: From Prompt to Text

  1. Runtime State and Context Management
  2. The Generation Loop Implementation
  3. Integration: Tokenizer, Model, and Sampler
  4. Streaming Output and Callbacks
  5. Command-Line Interface and First Run

Chapter 14: Verifying Correctness Before Optimizing

  1. Reference Implementations and Ground Truth
  2. Comparing Embeddings and Intermediate Activations
  3. Logit Comparison and Numerical Tolerance
  4. Generating Test Cases from Known Outputs
  5. Assertions, Sanitizers, and Debugging Techniques

Chapter 15: Performance Profiling and Bottleneck Analysis

  1. Latency, Throughput, and Tokens Per Second
  2. Profiling Tools: Perf, Valgrind, and Flame Graphs
  3. Compute-Bound vs Memory-Bound Analysis
  4. Cache Misses, Branch Mispredictions, and Latency
  5. Identifying the Critical Path

Chapter 16: SIMD and Vectorization

  1. Understanding SIMD Architectures: AVX2, AVX-512, NEON
  2. Autovectorization vs Hand-Written Intrinsics
  3. Vectorized Element-Wise Operations
  4. Vectorized Matrix Multiplication
  5. Cross-Platform SIMD Portability

Chapter 17: Memory and Threading Optimizations

  1. Cache Locality and Data Layout Optimization
  2. Loop Tiling and Blocking for MatMul
  3. Multithreading with OpenMP and Native Threads
  4. False Sharing, Synchronization, and Contention
  5. NUMA Awareness and Large Memory Models
  6. Memory Bandwidth, Compute-Bound vs Memory-Bound Workloads

Chapter 18: Advanced Optimization Techniques

  1. Weight Quantization and Low-Bit Representations
  2. Packed Weights and Efficient Memory Use
  3. Batching and Continuous Batching
  4. Speculative Decoding and Draft Models
  5. Asynchronous Execution and Pipeline Parallelism
  6. Model-Size-Dependent Optimization Strategies

Conclusion: What We Built and Why It Matters

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub