Leanpub Header

Skip to main content

Running LLMs on Tiny Hardware

A Systems Engineer's Guide to Edge Inference on Resource-Constrained Devices

Running LLMs on Tiny Hardware
This book is 100% completeLast updated on 2026-09-29

What does it take to run an LLM on hardware that barely has enough memory to breathe? This practical guide takes you from transformer math and model compression to inference runtimes, benchmarking and real-world deployment, with hands-on case studies for building fully offline LLM systems on tiny devices.

Minimum price

$25.00

$35.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
179
Pages
About

About

About the Book

This book teaches you how to make modern language models run on extremely constrained devices: microcontrollers with kilobytes of RAM, single-board computers with gigabytes, and mobile chips with NPUs. It is a practical engineering handbook for systems programmers who want to deploy LLM inference entirely offline on real hardware. You will learn the physics of why LLM inference is hard on small devices, the mathematics of transformer computation, the techniques for shrinking and accelerating models, the architecture of inference runtimes, and the skills needed to benchmark, debug, and deploy working systems. Every chapter builds on the previous one, culminating in complete case studies that show you how to design and implement an on-device LLM system from scratch.

Author

About the Author

Steve Publications

Steve is a technology professional with more than 20 years of experience in software development, server infrastructure, cybersecurity, vulnerability research and reverse engineering. Throughout his career, he has designed, secured, analyzed and tested complex software and infrastructure, with a particular focus on understanding how systems fail and how they can be made more secure.

Outside of work, Steve enjoys sharing knowledge with the technology community. He collaborates with researchers, industry experts and technology professionals to write practical books covering software development, cybersecurity, cloud computing, networking, DevOps, artificial intelligence and enterprise technologies. His books focus on practical learning through clear explanations, real-world examples and hands-on exercises. With more than two decades of industry experience, his goal is to help IT professionals, students and technology enthusiasts build useful skills and stay current in a rapidly changing industry.

We believe readers deserve to know how our books are created. Most of our authors are not native English speakers, so we use AI to help translate, proofread manuscripts, fix grammar, improve sentence structure and make technical explanations easier to read. AI is used as an editing tool only. It does not replace the research, technical knowledge or hands-on experience behind our books. Some of our authors also prefer to remain anonymous for privacy or professional reasons. In those cases, we publish their work under a different name. The author's name may be different, but the quality of the content and our review process remain the same.

Every book is written, reviewed and maintained by experienced technology professionals, with contributions from our private technical community of more than 420 engineers and researchers. We spend far more time validating technical accuracy and keeping our content up to date than generating text. We are always interested in working with experienced professionals who have deep expertise in a particular technology or domain. If you would like to publish a book with us or help review an existing manuscript, we'd love to hear from you. Send us a message describing your area of expertise. We are especially interested in niche technologies, specialized skills and emerging topics that are underrepresented in existing technical literature.

If you look through the contents of our books, you'll see practical examples, detailed explanations and material that is regularly updated. Our goal is to publish books that professionals can actually rely on, not low-effort AI-generated content. If you ever feel that one of our books does not meet that standard, Leanpub offers a 60-day money-back guarantee. Feel free to request a refund if you are not satisfied with your purchase.

Contents

Table of Contents

A Systems Engineer’s Guide to Edge Inference on Resource-Constrained Devices

Introduction

  1. What This Book Covers
  2. How to Read This Book
  3. Technology Context
  4. The Central Thesis

Chapter 1: The Physics of Tiny

  1. Why Memory Rules Everything
  2. SRAM, DRAM, and Flash: Capacity, Speed, and Cost
  3. Memory Hierarchy and Cache Behavior
  4. Compute vs. Bandwidth: Arithmetic Intensity
  5. Power Budgets and Thermal Walls
  6. The Numbers: How Much Memory and Bandwidth Does an LLM Actually Need

Chapter 2: Inside an LLM Inference Step

  1. Tokenization: Strings to Integers
  2. The Embedding Layer and Positional Encoding
  3. Transformer Blocks: Self-Attention and Feed-Forward
  4. Multi-Head and Grouped-Query Attention
  5. Layer Normalization and Residual Connections
  6. Autoregressive Decoding and the KV Cache
  7. The Complete Inference Dataflow

Chapter 3: Measuring Feasibility

  1. Parameter Storage: Bits, Bytes, and Formats
  2. Activation Memory: What You Cannot See Coming
  3. KV Cache: The Hidden RAM Killer
  4. Arithmetic Intensity and Compute Requirements
  5. Bandwidth Calculations: The True Bottleneck
  6. The Feasibility Checklist

Chapter 4: Model Architectures for Constrained Inference

  1. Dense Decoder-Only Transformers: The Baseline
  2. Mixture of Experts: Selective Activation
  3. Recurrent and Hybrid Architectures
  4. Compact Language Models: Design for Small
  5. Architecture Choices That Matter for Embedded: Context Window, Heads, FFN Width
  6. Selecting a Model Family for Your Target

Chapter 5: Quantization Fundamentals

  1. Why Quantization Works: Redundancy in Floating-Point Weights
  2. Symmetric vs. Asymmetric Quantization: Scales and Zero-Points
  3. Per-Tensor vs. Per-Channel vs. Group-wise
  4. Calibration Strategies
  5. The Quantization-Error-Accuracy Tradeoff
  6. Quantization Formats: FP16, BF16, INT8, and Lower

Chapter 6: Weight-Only Quantization

  1. How Weight-Only Quantization Works
  2. Dequantization on the Fly: The Compute Cost
  3. Packing Schemes: 4-bit, 5-bit, 6-bit Weights
  4. Weight-Only Quantization and Activation Outliers
  5. Implementing Weight-Only Matrix Multiplication
  6. When Weight-Only Is Enough (and When It Is Not)

Chapter 7: Activation Quantization and Mixed Precision

  1. Why Activations Are Harder to Quantize
  2. Activation Outliers and Their Impact
  3. Mixed Precision Strategies
  4. Quantization-Aware Training
  5. Post-Training Quantization with Activation Calibration
  6. INT4/INT3 Activation Quantization: The Cutting Edge

Chapter 8: Sparsity, Pruning and Structured Compression

  1. Unstructured vs. Structured Sparsity
  2. Pruning Methods and Magnitude Thresholding
  3. Compressed Sparse Row and Other Formats
  4. The Problem of Sparsity on Small CPUs
  5. Block-wise and Channel-wise Pruning
  6. Low-Rank Factorization
  7. When to Use Pruning and Sparsity

Chapter 9: Knowledge Distillation and Architecture Modification

  1. Distillation: Making Small Models Smarter
  2. Architecture Surgery: Head Merging, Layer Thinning
  3. Vocabulary Reduction and Tokenizer Optimization
  4. Context Window Management and Compression
  5. Speculative Decoding: Drafting and Verifying
  6. When to Modify the Model vs. When to Optimize the Runtime

Chapter 10: KV Cache Engineering

  1. KV Cache Layout: Memory Mapping the Context
  2. KV Cache Quantization and Compression
  3. Eviction Policies and Sliding Windows
  4. Activation Checkpointing and Recomputation
  5. Streaming and Weight-Reuse Strategies
  6. KV Cache for Long Context on Tiny Memory

Chapter 11: CPU, SIMD and Vector Optimization

  1. Vector Units: NEON, SSE, AVX, and RISC-V V-Extensions
  2. Writing Vectorized Matmul Kernels
  3. Packing and Unpacking for SIMD
  4. Instruction-Level Parallelism and Loop Unrolling
  5. Memory Access Patterns and Cache Locality
  6. Tuning for Specific CPU Microarchitectures

Chapter 12: GPU, NPU and DSP Acceleration

  1. Mobile GPUs: Mali, Adreno, PowerVR
  2. Neural Processing Units: NPU Architectures and Constraints
  3. DSP Acceleration for Embedded Systems
  4. Tensor Cores and Matrix Multiply Accelerators
  5. Accelerator APIs and Vendor SDKs
  6. When Accelerators Help (and When They Get in the Way)

Chapter 13: Inference Runtimes and Software Architecture

  1. The Runtime Stack: From Model Files to Generated Tokens
  2. Graph Representations and IR Formats
  3. Lightweight Transformer Runtimes
  4. Vendor-Specific Runtimes vs. Portable Solutions
  5. Kernel Selection and Code Generation
  6. Designing Your Own Minimal Inference Runtime

Chapter 14: Model Conversion and Deployment Pipelines

  1. Exporting Models: ONNX, GGUF, safetensors, and Others
  2. Tokenizer Packaging and Conversion
  3. Quantization Pipelines: Tools and Techniques
  4. Cross-Compilation and Build Systems
  5. Binary Size, Flash Constraints, and OTA Updates
  6. Model Versioning and Update Strategies

Chapter 15: Performance Engineering and Benchmarking

  1. Metrics That Matter: TTFT, ITL, Throughput, and Energy
  2. Benchmarking Methodology: Warm-Up, Sampling, Reproducibility
  3. Profiling: Finding the Real Bottleneck
  4. Compute-Bound vs. Memory-Bound vs. Synchronization-Bound
  5. Compiler Optimization and Its Surprises
  6. Performance Regression and Change Management

Chapter 16: Power, Energy and Thermal Management

  1. Measuring Power and Energy in Real Deployments
  2. Energy per Token: A Practical Metric
  3. Dynamic Voltage and Frequency Scaling
  4. Thermal Throttling and Sustained Performance
  5. Power-Optimized Scheduling and Batching

Chapter 17: Debugging and Failure Analysis

  1. Incorrect Outputs: Numerical Error vs. Model Behavior
  2. Memory Exhaustion and Fragmentation
  3. Stack Overflow and Heap Corruption
  4. Quantization Bugs: Zero-Point and Scale Errors
  5. Hardware-Specific Issues: Alignment, Endianness, FPU Modes
  6. Silent Accuracy Degradation and Regression Testing

Chapter 18: Security and Privacy

  1. The Security Model of Local Inference
  2. Model Extraction and Protection
  3. Supply Chain Risks: Malicious Models and Tooling
  4. Input and Output Security: Prompt and Data Handling
  5. Secure OTA Updates and Integrity Verification
  6. Privacy Advantages and Their Limits

Chapter 19: Case Study : An LLM in 256 MB of RAM

  1. The Hardware and Its Constraints
  2. Choosing and Preparing the Model
  3. Memory Budget: Weights, KV Cache, Runtime Overhead
  4. Implementation Decisions and Trade-offs
  5. Benchmarking Results
  6. Lessons Learned

Chapter 20: Case Study : Offline Edge Assistant on an ARM SBC

  1. System Requirements and Architecture
  2. Model and Tokenizer Selection
  3. Building the Inference Engine
  4. Application Integration: Input, Output, Streaming
  5. Robustness: Error Handling, Watchdogs, Recovery
  6. Production Readiness Checklist

Chapter 21: Case Study : Pushing the Limits on MCU-Class Hardware

  1. MCU Constraints: Kilobytes of RAM, Megahertz Clocks
  2. Micro-Models: Can Transformers Run on an MCU?
  3. Extreme Quantization and Architecture Reduction
  4. Alternative Approaches: Recurrent and State-Space Models
  5. Measured Performance and Practical Viability
  6. The Real Limit: Where LLMs Cannot Go

Chapter 22: Production Deployment Patterns

  1. Firmware Integration and Boot-Time Constraints
  2. Multi-Model and Fallback Strategies
  3. Handling Concurrent Requests on Constrained Systems
  4. Observability: Logging, Metrics, Telemetry
  5. A/B Testing and Model Rollout at the Edge
  6. Maintenance and Lifecycle Management

Chapter 23: Future Directions

  1. Architectural Innovations: Mamba, RWKV, and State-Space Models
  2. Hardware Roadmaps: What Next-Gen Edge Silicon Will Bring
  3. Algorithmic Advances: Better Quantization, Sparsity, Distillation
  4. Compiler and Runtime Innovations
  5. The Economics of Edge Inference
  6. The Next Frontier: What Comes After Tiny LLMs

Conclusion: The Engineer’s Playbook

  1. The Decision Tree: Choosing Your Optimization Strategy
  2. A Checklist for Edge LLM Deployment
  3. Accepting Trade-offs: No Solution Is Free
  4. Final Thoughts: The Promise and Limits of Edge Intelligence

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub