Leanpub Header

Skip to main content

Rust for LLM Inference

Building High-Performance LLM Inference Engine from Scratch

This book is 100% completeLast updated on 2026-07-13

Learn how modern LLM inference engines work by building one from scratch in Rust. From transformers and tokenization to KV caching, quantization, batching, and GPU optimization, this book combines theory, hands-on code, and performance engineering to help you create fast, production-ready AI systems.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
About

About

About the Book

This book is a complete guide to understanding and building large language model inference engines in Rust. It takes you from the mathematical foundations of the transformer architecture through tokenization, attention mechanisms, KV caching, quantization, batching strategies, GPU and CPU optimization, distributed serving, and production deployment. Along the way, you will see idiomatic Rust code examples, performance benchmarks, comparisons with leading frameworks like vLLM, TGI, llama.cpp, Candle, Burn, and mistral.rs, and practical insights for shipping inference systems that rival the best open-source and commercial offerings. Whether you are a systems programmer interested in AI infrastructure or an ML engineer curious about the Rust ecosystem, this book gives you the depth to build, optimize, and understand.

Bundles

Bundles that include this book

Author

About the Author

Steve T. Publications

Steve T. Publications is a specialized book publishing company dedicated to delivering high-quality technical resources for IT professionals, students, educators, and technology enthusiasts. Our mission is to make complex technology concepts accessible through well-structured, practical, and industry-relevant publications.

We focus on publishing books across a wide range of information technology disciplines, including software development, cloud computing, cybersecurity, artificial intelligence, data science, networking, DevOps, databases, and enterprise technologies. Every publication is designed to bridge the gap between theory and real-world application, helping readers build the skills needed to succeed in today's rapidly evolving digital landscape.

At Steve T. Publications, we collaborate with experienced industry experts, educators, and technology professionals to produce accurate, up-to-date, and engaging content. We are committed to maintaining the highest editorial standards while empowering learners and professionals with trusted technical knowledge.

Whether you're beginning your IT journey, preparing for professional certifications, or advancing your expertise in emerging technologies, Steve T. Publications is your trusted source for authoritative and practical technical books.

Contents

Table of Contents

Building High-Performance LLM Inference Engine from Scratch

  1. About This Book

Introduction: Why Rust for LLM Inference

  1. The Inference Engine We Will Build: llm-rs
  2. The LLM Serving Landscape: A Competitive Ecosystem
  3. Why Rust? The Systems Argument
  4. How to Read This Book
  5. Framework Trade-offs: When Rust Makes Sense (and When It Does Not)
  6. What This Book Is Not

Chapter 1: The Transformer Architecture

  1. What a Transformer Is and Why It Works
  2. Encoder-Decoder vs Decoder-Only
  3. Layer-by-Layer Anatomy
  4. The Forward Pass as Computation Graph
  5. Model Configuration in Rust
  6. Try It Yourself: Building a Model Config

Chapter 2: Tokenization and the Vocabulary

  1. Byte-Pair Encoding and Its Variants
  2. Building and Using Tokenizers in Rust
  3. Prompt Assembly and Token Counting
  4. Try It Yourself: Implementing a Simple BPE Tokenizer
  5. Try It Yourself: Building a Prompt Assembler

Chapter 3: Attention Mechanisms Deep Dive

  1. Scaled Dot-Product Attention Derivation
  2. Multi-Head Attention and Head Pruning
  3. FlashAttention and I/O-Aware Computation
  4. Causal (Masked) Attention for Autoregressive Generation
  5. Rotary Position Embeddings (RoPE) Detail
  6. Try It Yourself: Implementing Scaled Dot-Product Attention
  7. Try It Yourself: Implementing RoPE

Chapter 4: Building the Core Inference Loop

  1. Setting Up a Minimal Rust Project
  2. Implementing Linear Layers, RMSNorm, and Activations
  3. Composing a Complete Transformer Layer
  4. The Autoregressive Generation Loop
  5. Benchmarking a Single-Layer Forward Pass
  6. Try It Yourself: Building a Complete Inference Loop
  7. Try It Yourself: Sampling Strategies

Chapter 5: KV Caching and Autoregressive Optimization

  1. Why Naive Autoregressive Inference Is O(n²) in Memory and Compute
  2. Key-Value Cache Design and Tensor Shapes
  3. PagedAttention: Memory Management as an Operating System Problem
  4. Implementing PagedAttention in Rust
  5. Cache Eviction Strategies
  6. Try It Yourself: Implementing a Contiguous KV Cache
  7. Try It Yourself: Implementing PagedAttention

Chapter 6: Quantization for Memory and Speed

  1. Floating-Point Formats: FP32, FP16, BF16, FP8
  2. INT8 and INT4 Quantization Schemes
  3. GGUF/GGML Format and llama.cpp’s Approach
  4. Implementing Dequantize-and-Multiply in Rust
  5. Quantization Schemes Compared
  6. Try It Yourself: Implementing INT4 Quantization
  7. Try It Yourself: Parsing GGUF Files

Chapter 7: Batching Strategies

  1. Static vs Dynamic Batching
  2. The Continuous Batching Scheduler
  3. Continuous Batching and PagedAttention
  4. Request Scheduling and Priority
  5. Throughput-Latency Tradeoffs
  6. Try It Yourself: Implementing a Scheduler
  7. Try It Yourself: Prefix Caching with RadixAttention

Chapter 8: CPU Optimization Techniques

  1. SIMD Vectorization with Rust
  2. Cache-Conscious Memory Layout
  3. Parallelism with Rayon and Work-Stealing
  4. BLAS Integration for Matmul
  5. Benchmarking Methodology and Results
  6. Try It Yourself: SIMD Matrix Multiplication
  7. Try It Yourself: Rayon Parallel Matmul
  8. Try It Yourself: CPU Profiling with perf

Chapter 9: GPU Acceleration with CUDA and ROCm

  1. Rust CUDA Bindings: Burn, Candle, llama.cpp
  2. Building a CUDA Backend for llm-rs
  3. Writing Custom CUDA Kernels
  4. Memory Allocation Patterns for GPU Tensors
  5. Profiling with Nsight
  6. torch.compile-Style Kernel Fusion
  7. Try It Yourself: Writing a CUDA Kernel
  8. Try It Yourself: Implementing a Dequantize-and-Matmul Kernel
  9. Try It Yourself: GPU Memory Pool

Chapter 10: Model Formats and Serialization

  1. PyTorch .bin and Safetensors Formats
  2. GGUF/GGML, ONNX, MLX, and TensorRT
  3. Serialization in Rust
  4. Weight Loading and Layout Transformation
  5. Loading GGUF Models in Rust
  6. Try It Yourself: Building a Model Loader
  7. Try It Yourself: GGUF Parser

Chapter 11: Distributed Inference and Serving

  1. Tensor Parallelism vs Pipeline Parallelism
  2. Communication Patterns: All-Reduce and All-Gather
  3. Serving Frameworks: vLLM, TGI, Ollama, Mistral.rs
  4. Building a Production Rust Inference Server
  5. Try It Yourself: Building a Multi-GPU Engine
  6. Try It Yourself: Streaming API

Chapter 12: Production Deployment and Monitoring

  1. Containerization and Orchestration
  2. Observability: Metrics, Tracing, Logging
  3. Fault Tolerance and Recovery
  4. A/B Testing New Models and Quantizations
  5. Cost Analysis: Dollars per Million Tokens
  6. Try It Yourself: Deploying with Docker
  7. Try It Yourself: Setting Up Prometheus Metrics
  8. Try It Yourself: A/B Testing Pipeline

Appendix A: Benchmarking Methodology

  1. Hardware Specifications
  2. Software Versions
  3. Benchmarking Procedure
  4. Reproducing the Results
  5. Understanding Benchmark Numbers
  6. Statistical Rigor

Conclusion: Where Rust Fits in the LLM Ecosystem

  1. The Trade-offs We Made
  2. Where Rust Shines (and Where It Does Not)
  3. The Pragmatic Stack
  4. The Future of Rust in LLM Inference
  5. What Comes Next

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub