Leanpub Header

Skip to main content

Rust for LLM Inference

Building High-Performance LLM Inference Engine from Scratch

This book is 100% completeLast updated on 2026-07-16

Learn how modern LLM inference engines work by building one from scratch in Rust. From transformers and tokenization to KV caching, quantization, batching, and GPU optimization, this book combines theory, hands-on code, and performance engineering to help you create fast, production-ready AI systems.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
182
Pages
About

About

About the Book

This book is a complete guide to understanding and building large language model inference engines in Rust. It takes you from the mathematical foundations of the transformer architecture through tokenization, attention mechanisms, KV caching, quantization, batching strategies, GPU and CPU optimization, distributed serving, and production deployment. Along the way, you will see idiomatic Rust code examples, performance benchmarks, comparisons with leading frameworks like vLLM, TGI, llama.cpp, Candle, Burn, and mistral.rs, and practical insights for shipping inference systems that rival the best open-source and commercial offerings. Whether you are a systems programmer interested in AI infrastructure or an ML engineer curious about the Rust ecosystem, this book gives you the depth to build, optimize, and understand.

Bundles

Bundles that include this book

Author

About the Author

Steve Publications

Steve is a technology professional with more than 20 years of experience in software development, server infrastructure, cybersecurity, vulnerability research and reverse engineering. Throughout his career, he has designed, secured, analyzed and tested complex software and infrastructure, with a particular focus on understanding how systems fail and how they can be made more secure.

Outside of work, Steve enjoys sharing knowledge with the technology community. He collaborates with researchers, industry experts and technology professionals to write practical books covering software development, cybersecurity, cloud computing, networking, DevOps, artificial intelligence and enterprise technologies. His books focus on practical learning through clear explanations, real-world examples and hands-on exercises. With more than two decades of industry experience, his goal is to help IT professionals, students and technology enthusiasts build useful skills and stay current in a rapidly changing industry.

We believe readers deserve to know how our books are created. Most of our authors are not native English speakers, so we use AI to help translate, proofread manuscripts, fix grammar, improve sentence structure and make technical explanations easier to read. AI is used as an editing tool only. It does not replace the research, technical knowledge or hands-on experience behind our books. Some of our authors also prefer to remain anonymous for privacy or professional reasons. In those cases, we publish their work under a different name. The author's name may be different, but the quality of the content and our review process remain the same.

Every book is written, reviewed and maintained by experienced technology professionals, with contributions from our private technical community of more than 400 engineers and researchers from Ukraine, Belarus and Russia. We spend far more time validating technical accuracy and keeping our content up to date than generating text. We are always interested in working with experienced professionals who have deep expertise in a particular technology or domain. If you would like to publish a book with us or help review an existing manuscript, we'd love to hear from you. Send us a message describing your area of expertise. We are especially interested in niche technologies, specialized skills and emerging topics that are underrepresented in existing technical literature.

If you look through the contents of our books, you'll see practical examples, detailed explanations and material that is regularly updated. Our goal is to publish books that professionals can actually rely on, not low-effort AI-generated content. If you ever feel that one of our books does not meet that standard, Leanpub offers a 60-day money-back guarantee. Feel free to request a refund if you are not satisfied with your purchase.

Contents

Table of Contents

Building High-Performance LLM Inference Engine from Scratch

  1. About This Book

Introduction: Why Rust for LLM Inference

  1. The Inference Engine We Will Build: llm-rs
  2. The LLM Serving Landscape: A Competitive Ecosystem
  3. Why Rust? The Systems Argument
  4. How to Read This Book
  5. Framework Trade-offs: When Rust Makes Sense (and When It Does Not)
  6. What This Book Is Not

Chapter 1: The Transformer Architecture

  1. What a Transformer Is and Why It Works
  2. Encoder-Decoder vs Decoder-Only
  3. Layer-by-Layer Anatomy
  4. The Forward Pass as Computation Graph
  5. Model Configuration in Rust
  6. Try It Yourself: Building a Model Config

Chapter 2: Tokenization and the Vocabulary

  1. Byte-Pair Encoding and Its Variants
  2. Building and Using Tokenizers in Rust
  3. Prompt Assembly and Token Counting
  4. Try It Yourself: Implementing a Simple BPE Tokenizer
  5. Try It Yourself: Building a Prompt Assembler

Chapter 3: Attention Mechanisms Deep Dive

  1. Scaled Dot-Product Attention Derivation
  2. Multi-Head Attention and Head Pruning
  3. FlashAttention and I/O-Aware Computation
  4. Causal (Masked) Attention for Autoregressive Generation
  5. Rotary Position Embeddings (RoPE) Detail
  6. Try It Yourself: Implementing Scaled Dot-Product Attention
  7. Try It Yourself: Implementing RoPE

Chapter 4: Building the Core Inference Loop

  1. Setting Up a Minimal Rust Project
  2. Implementing Linear Layers, RMSNorm, and Activations
  3. Composing a Complete Transformer Layer
  4. The Autoregressive Generation Loop
  5. Benchmarking a Single-Layer Forward Pass
  6. Try It Yourself: Building a Complete Inference Loop
  7. Try It Yourself: Sampling Strategies

Chapter 5: KV Caching and Autoregressive Optimization

  1. Why Naive Autoregressive Inference Is O(n²) in Memory and Compute
  2. Key-Value Cache Design and Tensor Shapes
  3. PagedAttention: Memory Management as an Operating System Problem
  4. Implementing PagedAttention in Rust
  5. Cache Eviction Strategies
  6. Try It Yourself: Implementing a Contiguous KV Cache
  7. Try It Yourself: Implementing PagedAttention

Chapter 6: Quantization for Memory and Speed

  1. Floating-Point Formats: FP32, FP16, BF16, FP8
  2. INT8 and INT4 Quantization Schemes
  3. GGUF/GGML Format and llama.cpp’s Approach
  4. Implementing Dequantize-and-Multiply in Rust
  5. Quantization Schemes Compared
  6. Try It Yourself: Implementing INT4 Quantization
  7. Try It Yourself: Parsing GGUF Files

Chapter 7: Batching Strategies

  1. Static vs Dynamic Batching
  2. The Continuous Batching Scheduler
  3. Continuous Batching and PagedAttention
  4. Request Scheduling and Priority
  5. Throughput-Latency Tradeoffs
  6. Try It Yourself: Implementing a Scheduler
  7. Try It Yourself: Prefix Caching with RadixAttention

Chapter 8: CPU Optimization Techniques

  1. SIMD Vectorization with Rust
  2. Cache-Conscious Memory Layout
  3. Parallelism with Rayon and Work-Stealing
  4. BLAS Integration for Matmul
  5. Benchmarking Methodology and Results
  6. Try It Yourself: SIMD Matrix Multiplication
  7. Try It Yourself: Rayon Parallel Matmul
  8. Try It Yourself: CPU Profiling with perf

Chapter 9: GPU Acceleration with CUDA and ROCm

  1. Rust CUDA Bindings: Burn, Candle, llama.cpp
  2. Building a CUDA Backend for llm-rs
  3. Writing Custom CUDA Kernels
  4. Memory Allocation Patterns for GPU Tensors
  5. Profiling with Nsight
  6. torch.compile-Style Kernel Fusion
  7. Try It Yourself: Writing a CUDA Kernel
  8. Try It Yourself: Implementing a Dequantize-and-Matmul Kernel
  9. Try It Yourself: GPU Memory Pool

Chapter 10: Model Formats and Serialization

  1. PyTorch .bin and Safetensors Formats
  2. GGUF/GGML, ONNX, MLX, and TensorRT
  3. Serialization in Rust
  4. Weight Loading and Layout Transformation
  5. Loading GGUF Models in Rust
  6. Try It Yourself: Building a Model Loader
  7. Try It Yourself: GGUF Parser

Chapter 11: Distributed Inference and Serving

  1. Tensor Parallelism vs Pipeline Parallelism
  2. Communication Patterns: All-Reduce and All-Gather
  3. Serving Frameworks: vLLM, TGI, Ollama, Mistral.rs
  4. Building a Production Rust Inference Server
  5. Try It Yourself: Building a Multi-GPU Engine
  6. Try It Yourself: Streaming API

Chapter 12: Production Deployment and Monitoring

  1. Containerization and Orchestration
  2. Observability: Metrics, Tracing, Logging
  3. Fault Tolerance and Recovery
  4. A/B Testing New Models and Quantizations
  5. Cost Analysis: Dollars per Million Tokens
  6. Try It Yourself: Deploying with Docker
  7. Try It Yourself: Setting Up Prometheus Metrics
  8. Try It Yourself: A/B Testing Pipeline

Appendix A: Benchmarking Methodology

  1. Hardware Specifications
  2. Software Versions
  3. Benchmarking Procedure
  4. Obtaining the source code
  5. Quick start
  6. Understanding Benchmark Numbers
  7. Statistical Rigor

Conclusion: Where Rust Fits in the LLM Ecosystem

  1. The Trade-offs We Made
  2. Where Rust Shines (and Where It Does Not)
  3. The Pragmatic Stack
  4. The Future of Rust in LLM Inference
  5. What Comes Next

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub